Best Cheap LLM in 2026

Models released in the last 12 months, ranked by blended price per million tokens (3:1 input-to-output mix).

Updated automatically as pricing changes. Full model database →

#ModelInput /1M
1Nex-N2-MiniNex Agi$0.025
2Solar Pro 4Upstage$0.03
3Qwen3.7 FlashAlibaba / Qwen$0.03
4Granite 4.1 8BIbm$0.05
5Mercury 2.5 PreviewInception$0.04
6Laguna XS 2.1Poolside$0.06
7Hy-MT2-1.8BTencent$0.044
8Ling 3.0 Flash FinInclusionai$0.06
9DeepSeek V4 Flash LatestDeepSeek$0.09
10Qwen3.5-9BAlibaba / Qwen$0.1
11Granite 4.2 8BIbm$0.1
12Qwen3.5-FlashAlibaba / Qwen$0.065
13Hy3 PreviewTencent$0.066
14GLM-5.3-FlashZhipu AI$0.075
15ERNIE 4.5 21B A3B ThinkingBaidu AI$0.07
16DeepSeek V4 FlashDeepSeek$0.098
17Laguna S 2.1Poolside$0.1
18Meta: Muse Spark 1.3 ContributorMeta AI$0.1
19Hy-MT2-7BTencent$0.074
20Hy-MT2-30B-A3BTencent$0.074
21Nemotron 3.5 LightningNVIDIA$0.1
22Step-3.5-FlashStepFun$0.1
23Step-3.5-Flash-BaseStepFun$0.1
24Seed-2.0-MiniByteDance$0.1
25MiMo-V2.5Xiaomi$0.14
26DeepSeek V4 Flash 0731DeepSeek$0.14
27Reka EdgeReka$0.2
28Nemotron 3 SuperNVIDIA$0.1
29NVIDIA Nemotron-3-Super-120B-A12BNVIDIA$0.2
30Ring-2.6-1TInclusionai$0.075

Finding the best value LLM

Price alone doesn't tell the whole story. A model that costs twice as much but solves problems in half the calls is actually cheaper. When evaluating cost, consider:

  • Input vs output pricing — for chat and RAG, input is usually 80%+ of your tokens. For generation-heavy tasks (writing, summarization), output price matters more.
  • Context window — larger contexts let you process more in a single call, reducing round trips and total token usage.
  • Open-weight models — if you can self-host, models like Mistral and LLaMA can cost near zero at scale. Check the “open weights” column on the model database.
  • Quality vs cost — compare benchmark scores on the benchmark leaderboard to find the best performance per dollar.

Also see: Best Coding LLM, Best Reasoning LLM, Compare any two models.