Articles tagged "moe"

03-Sept-2026
Qwen3.6-35B-A3B on a 16 GB RTX 5060 Ti: Why Context Beats Tok/s for Coding
Qwen3.6-35B-A3B on a 16 GB RTX 5060 Ti with FreeToken: long-context tuning, FP8 vs NVFP4 throughput and correctness, and real memory requirements.

29-Aug-2026
Running Qwen3.6-35B-A3B at 262K Context with FreeToken MoE Offload
Qwen3.6-35B-A3B-FP8 on RTX 4090 with FreeToken: near-262K context verification, 51/51 statistical rerun, and a 50-task correctness eval (46/50).