Laguna XS 2.1 33B Achieves 296 Tokens/sec on RTX 3090, 152 Tokens/sec with 256K Context
The Laguna XS 2.1 model, featuring 33 billion parameters, was benchmarked on an RTX 3090 GPU. Peak performance recorded 296 tokens per second. When operating with a 256 k token
The Laguna XS 2.1 model, featuring 33 billion parameters, was benchmarked on an RTX 3090
GPU. Peak performance recorded 296 tokens per second. When operating with a 256 k token
context window, throughput dropped to 152 tokens per second. The tests were documented on
the Lucebox blog. Results illustrate the trade‑off between model size and context length
on high‑end hardware. The findings help developers gauge realistic inference speeds for
large language models. The RTX 3090 remains a popular choice for experimental AI workloads
despite newer GPUs. Observers may compare these numbers with other models to assess
efficiency.