Different chips, different speeds, different costs
Current market prices and workload economics to help you choose the right hardware.
Current Hardware Planning Comparison
Compare planning throughput with an all-in cost that includes full-system CAPEX, power and maintenance.
Swipe to compare CAPEX, OPEX and total cost.
| System | Benchmark TPS | System Power | Planning CAPEX | Annual OPEX | 3-Year TCO / 1M Tokens |
|---|---|---|---|---|---|
NVIDIA HGX B300Blackwell Ultra 8-GPU server | 101,608 | 15,443W | $501,853 | $71,289 | $0.11 |
NVIDIA HGX B200Blackwell 8-GPU server | 99,159 | 14,147W | $393,608 | $58,694 | $0.09 |
NVIDIA HGX H200 8-GPU server | 32,955 | 6,845W | $326,084 | $41,963 | $0.21 |
NVIDIA HGX H100Hopper 8-GPU server | 30,576 | 6,377W | $297,220 | $38,437 | $0.20 |
NVIDIA HGX H800Hopper CN 8-GPU server | 26,000 | 6,377W | $282,400 | $36,955 | $0.23 |
Three-year TCO equals installed CAPEX plus three years of OPEX, divided by productive token output over the same period.
Planning assumptions: three-year straight-line hardware life, 70% productive utilization, $0.12/kWh electricity, 1.3 PUE and annual maintenance equal to 10% of installed CAPEX.
Prices were checked on September 16, 2026. CAPEX (capital expenditure) includes the listed complete server purchase price. B300 through H100 use public configured-system prices; H800 uses a reported 8-GPU market reference of approximately $282,400. OPEX (operating expenditure) includes facility-adjusted power and annual maintenance, but excludes staffing, networking, tax and financing.
Benchmark throughput uses the 2026 Lenovo Press MLPerf comparison for Llama 3.1 405B on matched eight-GPU configurations. H800 is a planning estimate because the same comparison does not include that model. Revalidate every row with your model, quantization, batch size and latency target before purchase.
Public Market Anchors
Five comparable NVIDIA generations shown at a consistent 8-GPU system scope.
Matched System Scope
Every row compares a complete eight-GPU server, not a loose accelerator card.
Benchmark Your Workload
Model, quantization, batch size and latency targets can materially change cost per token.
Not sure which to pick?
Run a quick diagnosis on the homepage — I'll recommend the best fit.