What can you run on an 8GB graphics card?

With 8GB VRAM you can run 57 of the 129 open-source models we track — up to about 16B parameters at Q4 quantization. The most popular that fits is Qwen3.5-9B, needing 6.8 GB and generating around ~31 tok/s on typical hardware of this size.

Typical machines: RTX 4060 · RTX 3070 · RTX 3060 Ti · RX 7600. After the operating system takes its share, about 7.7 GB is free for a model.

Having the memory is not the same as having the speed

8GB VRAM decides whether a model loads. How fast it answers is decided by memory bandwidth — how quickly your machine can read the model, which it must do once for every word it writes. Machines with 8GB VRAM span 272–448 GB/s, so the fastest is about 1.6× quicker than the slowest with identical capacity. The table quotes a RTX 4060; pick yours below to see the difference. Why bandwidth and not the processor →

272 GB/sGeForce RTX 4060 (8GB)· NVIDIA published specification: 272 GB/s GDDR6, 128-bit
288 GB/sRadeon RX 7600 (8GB)· AMD published specification: 288 GB/s GDDR6, 128-bit
448 GB/sGeForce RTX 3070 or 3060 Ti (8GB)· NVIDIA published specification: 448 GB/s GDDR6, 256-bit

Models that run on 8GB VRAM in 2026

ModelSizeMemoryEst. speedDownloadsFeels like
granite-embedding-97m-multilingual-r2Ibm97M603 MBRuns well~3005 tok/s97,079Fast
LFM2.5-Encoder-230MLiquid Ai230M793 MBRuns well~1275 tok/s13,432Fast
LFM2.5-230MLiquid Ai230M808 MBRuns well~1152 tok/s55,999Fast
granite-embedding-311m-multilingual-r2Ibm312M821 MBRuns well~698 tok/s64,530Fast
granite-speech-5.0-470m-turboctcIbm473M857 MBRuns well~619 tok/s11,294Fast
harrier-oss-v1-270mMicrosoft268M865 MBRuns well~698 tok/s727,222Fast
LFM2.5-Encoder-350M-Policy-LinterLiquid Ai355M885 MBRuns well~825 tok/s9,957Fast
LFM2.5-VL-450MLiquid Ai449M900 MBRuns well~771 tok/s67,396Fast
nemotron-3.5-asr-streaming-0.6bNVIDIA638M1.1 GBRuns well~357 tok/s917,259Fast
Qwen3.5-0.8BAlibaba / Qwen873M1.3 GBRuns well~332 tok/s2,881,964Fast
privacy-filterOpenAI · MoE1.4B1.4 GBRuns well~1046 tok/s280M active456,933Fast
LFM2.5-1.2B-InstructLiquid Ai1.2B1.4 GBRuns well~242 tok/s366,847Fast
LFM2.5-VL-1.6BLiquid Ai1.6B1.4 GBRuns well~242 tok/s248,110Fast
Nemotron-3-Embed-1B-BF16NVIDIA1.1B1.5 GBRuns well~257 tok/s386,914Fast
Hy-MT2-1.8BTencent2B1.9 GBRuns well~156 tok/s24,475Fast
MiniMax-Music3MiniMax2.4B2.1 GBRuns well~127 tok/s23,217Fast
Qwen3.5-2BAlibaba / Qwen2.3B2.1 GBRuns well~127 tok/s3,097,484Fast
cohere-transcribe-arabic-07-2026Cohere2.1B2.3 GBRuns well~113 tok/s51,893Fast
cohere-transcribe-03-2026Cohere2.1B2.3 GBRuns well~113 tok/s607,964Fast
granite-speech-4.1-2bIbm2.3B2.4 GBRuns well~113 tok/s237,181Fast
LFM2.5-2.6BLiquid Ai2.7B2.5 GBRuns well~106 tok/s133,677Fast
LFM2.5-VL-3BLiquid Ai3.1B2.5 GBRuns well~106 tok/s24,282Fast
granite-4.0-1b-speechIbm2.3B2.5 GBRuns well~110 tok/s74,644Fast
LFM2.5-Audio-1.5BLiquid Ai1.5B2.6 GBRuns well~97 tok/s1,702Fast
Unlimited-OCRBaidu AI3.3B2.7 GBRuns well~91 tok/s3,008,635Fast
Nemotron-Labs-Diffusion-3BNVIDIA3B2.8 GBRuns well~98 tok/s35,414Fast
LocateAnything-3BNVIDIA3.8B2.8 GBRuns well~84 tok/s95,054Fast
granite-4.1-3bIbm3.4B2.9 GBRuns well~86 tok/s122,076Fast
granite-4.2-3bIbm3.7B3.1 GBRuns well~79 tok/s14,073Fast
granite-4.0-3b-visionIbm4B3.1 GBRuns well~79 tok/s1,069Fast
Shieldstral-1.0-3BMistral AI3.8B3.1 GBRuns well~82 tok/s20,895Fast
Ministral-3-3B-Instruct-2512Mistral AI3.8B3.1 GBRuns well~82 tok/s544,419Fast
Nemotron-Labs-Diffusion-3B-BaseNVIDIA3.8B3.3 GBRuns well~76 tok/s19,213Fast
Cosmos3-EdgeNVIDIA3.9B3.3 GBRuns well~76 tok/s1,325,763Fast
Nemotron-3.5-Content-SafetyNVIDIA4.3B3.6 GBRuns well~71 tok/s134,547Fast
gemma-4-E2B-itGoogle DeepMind5.1B3.8 GBRuns well~57 tok/s3,235,641Fast
Qwen3.5-4BAlibaba / Qwen4.7B3.8 GBRuns well~65 tok/s7,681,585Fast
Fara1.5-4BMicrosoft4.5B4.0 GBRuns well~61 tok/s2,786Fast
MolmoWeb-4BAi24.9B4.1 GBRuns well~60 tok/s2,152Fast
NVIDIA-Nemotron-3-Nano-4B-BF16NVIDIA4B4.1 GBRuns well~62 tok/s712,701Fast
Fara-7BMicrosoft8.3B5.5 GBRuns well~38 tok/s1,593Fast
Hy-MT2-7BTencent8B5.7 GBRuns well~38 tok/s12,004Fast
LFM2.5-8B-A1BLiquid Ai · MoE8.5B5.9 GBRuns well~290 tok/s1B active96,675Fast
Nemotron-3-Embed-8B-BF16NVIDIA8B5.9 GBRuns well~37 tok/s81,804Fast
Nemotron-Labs-Diffusion-8BNVIDIA8B5.9 GBRuns well~37 tok/s163,233Fast
Cosmos3-NanoNVIDIA16B6.2 GBRuns well~35 tok/s244,209Fast
gemma-4-E4B-itGoogle DeepMind8B6.2 GBRuns well~33 tok/s4,740,694Fast
Nemotron-Labs-Diffusion-8B-BaseNVIDIA8.5B6.2 GBRuns well~34 tok/s280,417Fast
Ministral-3-8B-Instruct-2512Mistral AI8.9B6.3 GBRuns well~34 tok/s119,379Fast
granite-4.1-8bIbm8.8B6.3 GBRuns well~35 tok/s1,441,995Fast
MolmoWeb-8BAi28.7B6.4 GBRuns well~34 tok/s1,679Fast
granite-4.2-8bIbm8.8B6.6 GBRuns well~33 tok/s12,050Fast
Ornith-1.0-9BDeepreinforce9B6.7 GBTight~31 tok/s2,250,422Fast
Qwen3.5-9BAlibaba / Qwen9.7B6.8 GBTight~31 tok/s13,660,443Fast
GLM-4.6V-FlashZhipu AI10B7.0 GBTight~29 tok/s282,562Fast
Fara1.5-9BMicrosoft9.4B7.0 GBTight~30 tok/s4,461Fast
Olmo-3-7B-InstructAi27.3B7.2 GBTight~40 tok/s406,776Fast

Sorted by memory, smallest first — click any column to change it.

Upgrade

Including gemma-4-12B-it (12B), Ministral-3-14B-Instruct-2512 (14B) and Nemotron-Labs-Diffusion-14B (14B) — around ~45 tok/s on a typical machine that size.

See everything 12GB VRAM runs →

All figures at Q4_K_M quantization, 4K context. “Memory” includes the model, its context cache and runtime overhead. Speed is calculated, not benchmarked: 272 GB/s × 0.65 ÷ bytes read per token, assuming a RTX 4060. Scale it by your own machine's bandwidth from the list above. Mixture-of-experts models read only their active parameters per token, which is why some large ones outrun smaller dense models.

Image generators that also fit

These make pictures rather than text. Their memory use scales with the resolution you generate at, not conversation length, so treat these figures as a floor rather than a ceiling.

What you'd need more memory for

The next models up, and what they ask for.

Other hardware

Verified against HuggingFace on 2026-09-04New to this? Start here →Runs on a laptop and up · last 12 months only