What happened
AWS Machine Learning published a benchmark comparing two 30B Mixture-of-Experts models—Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B—across G5, G6, G6e, and G7 GPU instances on Amazon SageMaker AI.
The benchmark measured throughput, latency, and cost-per-token for real-time LLM inference, highlighting how G7's NVIDIA Blackwell GPUs deliver measurable price-performance improvements.
Why it matters
As organizations deploy large language models, choosing the right GPU instance is critical for balancing performance and cost. This benchmark provides data-driven guidance for optimizing real-time inference workloads.
The results suggest that newer Blackwell-based instances (G7) can offer better price-performance, potentially lowering the total cost of ownership for high-throughput LLM applications.
Key facts
Benchmark compared Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B, both 30B Mixture-of-Experts models.
Instances tested: G5, G6, G6e, and G7 on Amazon SageMaker AI.
Metrics: throughput, latency, and cost-per-token.
G7 instances use NVIDIA Blackwell GPUs and show measurable price-performance gains.
What to watch next
Future benchmarks may expand to other model sizes and architectures, offering broader guidance for instance selection.
As Blackwell GPUs become more widely available, expect more comparisons across different AI workloads and deployment scenarios.
