Benchmarking Qwen 3.6 35B MoE (3B active) on RTX 3090

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 GPU. The test evaluates performance on a single consumer-grade GPU. It focuses on

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 GPU. The test evaluates performance on a single consumer-grade GPU. It focuses on the 35B parameter variant with 3B active. Results include throughput and latency metrics for inference. The study highlights the feasibility of running large models on a single RTX 3090. It provides insights for developers seeking cost-effective inference solutions. The benchmark informs hardware selection for AI workloads. Further details are available on the author’s website.