What it takes to run the biggest open models
The weights are free to download. The hardware is not. Kimi-K2-Instruct-0905 has 1026B parameters and needs about 632 GB of memory even squeezed down to Q4 — roughly 9 H100 cards, or $243,000 of silicon before you have powered any of it on.
Nobody is running these at a desk — that isn't the point. It's worth seeing the scale of what “open” actually means at the top end, and how far it sits from the models on the rest of this section. What your own computer can run →
For scale: the largest computer you can simply order — a Mac Studio M3 Ultra, 512GB at about $9,500 — holds 512GB. That is still 1.2× too small for Kimi-K2-Instruct-0905.
The models, by size
| Model | Parameters | Memory @ Q4 | H100s | Rough cost |
|---|---|---|---|---|
| Kimi-K2-Instruct-0905Moonshot AI | 1026B | 632 GB | 9 | $243,000 |
| MiMo-V2.5-ProXiaomi | 1023B | 635 GB | 9 | $243,000 |
| DeepSeek-V3.1DeepSeek | 685B | 416 GB | 6 | $162,000 |
| RakutenAI-3.0Rakuten | 671B | 415 GB | 6 | $162,000 |
| ERNIE-4.5-VL-424B-A47B-PTBaidu AI · 47B active | 424B | 260 GB | 4 | $108,000 |
| Qwen3.5-397B-A17BAlibaba / Qwen · 17B active | 403B | 244 GB | 4 | $108,000 |
| Ornith-1.0-397BDeepreinforce · 79B active | 397B | 246 GB | 4 | $108,000 |
| ERNIE-4.5-300B-A47B-PTBaidu AI · 47B active | 300B | 184 GB | 3 | $81,000 |
| MiniMax-M2.7MiniMax · 46B active | 229B | 144 GB | 2 | $54,000 |
| MiniMax-M2.5MiniMax · 46B active | 229B | 143 GB | 2 | $54,000 |
Memory is at Q4_K_M with 4K context — the smallest anyone sensibly runs these at; at full precision they are four times larger again. Card counts divide memory by the H100's 80GB with a 10% allowance, so they are a floor: real deployments lose more to activations, long-context caches and sharding. Prices are indicative street prices for the silicon alone, excluding the server, networking, power and cooling around it.
Kimi-K2-Instruct-0905 on different hardware
The same model, and what it would take on each generation of accelerator.
| Accelerator | Memory each | Cards needed | Rough cost | Est. speed |
|---|---|---|---|---|
| A100 80GBNVIDIA published specification: 80 GB HBM2e, 2,039 GB/s | 80 GB | 9 | $135,000 | ~2.1 tok/s |
| H100 80GBNVIDIA published specification: 80 GB HBM3, 3.35 TB/s | 80 GB | 9 | $243,000 | ~3.5 tok/s |
| H200 141GBNVIDIA published specification: 141 GB HBM3e, 4.8 TB/s | 141 GB | 5 | $160,000 | ~5.0 tok/s |
| B200 192GBNVIDIA published specification: 192 GB HBM3e, 8 TB/s | 192 GB | 4 | $160,000 | ~8.4 tok/s |
| Mac Studio M3 Ultra, 512GBApple's largest unified-memory configuration: 512 GB at 819 GB/s | 512 GB | Does not fit | $9,500 | — |
Speeds assume a single card's bandwidth and are optimistic: splitting a model across several cards adds communication cost that a single-card figure ignores.
What about GPT, Claude and Gemini?
We can't tell you. OpenAI, Anthropic and Google don't publish parameter counts or release weights for their frontier models, so any figure for them would be a guess dressed up as a fact. Every number on this page comes from a published file size or an official specification, and we'd rather leave a gap than fill it with invention.
That difference is itself the point of the page. With an open-weight model you can measure exactly what it costs to run, because the weights are right there. With a closed one you get a price per token and no way to check what sits behind it.