Best Cheap LLM in 2026

Models released in the last 12 months, ranked by blended price per million tokens (3:1 input-to-output mix).

Updated automatically as pricing changes. Full model database →

#ModelInput /1M
1Nex-N2-MiniNex Agi$0.025
2Qwen3.7 FlashAlibaba / Qwen$0.03
3Granite 4.1 8BIbm$0.05
4Laguna XS 2.1Poolside$0.06
5DeepSeek V4 Flash LatestDeepSeek$0.09
6Qwen3.5-9BAlibaba / Qwen$0.1
7Qwen3.5-FlashAlibaba / Qwen$0.065
8Hy3 PreviewTencent$0.066
9ERNIE 4.5 21B A3B ThinkingBaidu AI$0.07
10ERNIE 4.5 21B A3BBaidu AI$0.07
11ERNIE 4.5 VL 28B A3BBaidu AI$0.07
12DeepSeek V4 FlashDeepSeek$0.098
13Laguna S 2.1Poolside$0.1
14Step-3.5-FlashStepFun$0.1
15Step-3.5-Flash-BaseStepFun$0.1
16Seed-2.0-MiniByteDance$0.1
17MiMo-V2.5Xiaomi$0.14
18DeepSeek V4 Flash 0731DeepSeek$0.14
19Reka EdgeReka$0.2
20NVIDIA Nemotron-3-Super-120B-A12BNVIDIA$0.2
21Nemotron 3 SuperNVIDIA$0.1
22Ring-2.6-1TInclusionai$0.075
23Laguna M.1Poolside$0.2
24KAT-Coder-Air V2.5Kwaipilot$0.15
25Grok 4.1 ThinkingxAI$0.2
26Mistral SabaMistral AI$0.2
27Qwen3.5-35B-A3BAlibaba / Qwen$0.14
28Qwen3.6 35B A3BAlibaba / Qwen$0.161
29Mercury 2Inception$0.25
30Step-3.7-FlashStepFun$0.2

Finding the best value LLM

Price alone doesn't tell the whole story. A model that costs twice as much but solves problems in half the calls is actually cheaper. When evaluating cost, consider:

  • Input vs output pricing — for chat and RAG, input is usually 80%+ of your tokens. For generation-heavy tasks (writing, summarization), output price matters more.
  • Context window — larger contexts let you process more in a single call, reducing round trips and total token usage.
  • Open-weight models — if you can self-host, models like Mistral and LLaMA can cost near zero at scale. Check the “open weights” column on the model database.
  • Quality vs cost — compare benchmark scores on the benchmark leaderboard to find the best performance per dollar.

Also see: Best Coding LLM, Best Reasoning LLM, Compare any two models.