research

Gemma 4, DeepSeek V4, and ZAYA1 Deploy KV Cache Compression to Cut Long-Context Memory Costs

TL;DR

Recent open-weight LLM releases from Google, DeepSeek, and others are adopting architectural techniques that reduce KV cache size by approximately 50% at long contexts. These include cross-layer KV sharing in Gemma 4, which saves 2.7 GB at 128K context for the E2B model, and compressed convolutional attention in ZAYA1-8B.

3 min read
0

Gemma 4, DeepSeek V4, and ZAYA1 Deploy KV Cache Compression to Cut Long-Context Memory Costs

Multiple open-weight LLM releases in April and May 2026 have adopted architectural techniques specifically designed to reduce KV cache size and memory traffic at long contexts, according to a technical analysis by Sebastian Raschka.

Cross-Layer KV Sharing in Gemma 4

Google's Gemma 4 suite, released in early April, implements cross-layer KV sharing in its E2B and E4B variants. Instead of computing separate key-value projections in each transformer layer, later layers reuse KV tensors from earlier layers while still computing their own query projections.

The Gemma 4 E2B model has 35 transformer layers but only 15 compute their own KV projections—the final 20 layers reuse KV tensors from previous layers. According to Raschka's calculations, this saves approximately 2.7 GB of memory at 128K context length (bfloat16 precision) for the E2B model. The E4B variant, with 42 layers (24 computing KV, 18 sharing), saves approximately 6 GB at 128K context.

The technique is not exclusive to Gemma 4. The cross-layer attention approach was described in Brandon et al.'s "Reducing Transformer Key-Value Cache Size with Cross-Layer Attention" (NeurIPS 2024), but Gemma 4 represents the first major open-weight implementation.

Additional Architecture Techniques

Beyond Gemma 4, other recent releases have implemented complementary approaches:

  • ZAYA1-8B: Uses compressed convolutional attention to reduce memory footprint
  • Laguna XS.2: Implements layer-wise attention budgeting to allocate attention compute selectively
  • DeepSeek V4: Combines multi-head compressed (mHC) attention with additional compression techniques

All three models also use Grouped Query Attention (GQA), which shares key-value heads across multiple query heads—a now-standard technique for KV cache reduction.

Model Variants and Target Use Cases

The Gemma 4 release includes three categories:

  1. E2B and E4B models: Optimized for mobile and embedded devices (IoT)
  2. 26B mixture-of-experts (MoE): Designed for efficient local inference
  3. 31B dense model: Optimized for maximum quality and easier fine-tuning

The E2B and E4B variants combine cross-layer KV sharing with a 4:1 pattern of regular GQA and sliding window attention. Specifically, E2B uses MQA (the single-KV-head special case of GQA) rather than full GQA.

Technical Implementation Details

In cross-layer KV sharing, sliding-window attention layers share KV with previous sliding-window layers, while full-attention layers share with previous full-attention layers. Each layer still computes its own query projections, allowing distinct attention patterns while eliminating redundant KV cache storage.

The memory savings scale with context length. At very long contexts (128K+), the KV cache becomes the dominant memory consumer, making these optimizations critical for reasoning models and agent workflows that maintain extended conversation history.

What This Means

The convergence on KV cache reduction techniques across multiple independent releases signals that long-context efficiency has become a primary architectural constraint. With reasoning models and agent workflows keeping more tokens active for longer periods, memory traffic and attention costs now dominate over pure compute. The ~50% KV cache reduction achieved through cross-layer sharing makes 128K+ context windows practical on consumer hardware. Expect these techniques to become standard in future model releases, particularly for edge deployment and long-context applications.

Related Articles

research

Tavus says 48% of testers mistook its Griffin video AI for a real person on a one-minute call

Tavus has introduced Griffin, which it calls the first 'Human Interaction Model' for real-time face-to-face video conversation. In a Tavus study, 48% of participants believed Griffin was a real person after a one-minute call, versus a 2% maximum for earlier systems. A limited research preview, Griffin-Lite, is open only to select testers.

research

Graphite: Opus 5.5 uses 'this matters' 116x more than humans as AI writing tells persist

Marketing firm Graphite identified 13,000 phrases that appear at least twice as often in AI-generated writing as in human writing. Claude Opus 5.5 uses "this matters" 116 times more than humans, while OpenAI's Astra favors "corrective framing" more than 100 times as often. Em-dash use has collapsed across frontier models, but total tells are holding steady, according to Graphite.

research

Ai2 releases Olmo-core 3, an open MoE training stack benchmarked at 1.2T parameters

Ai2 released Olmo-core 3, an open training framework for large mixture-of-experts models. The company reports about 2.7× the throughput of its earlier FSDP-based implementation and benchmarks up to 1.2 trillion total parameters on 512 NVIDIA B300 GPUs. It is the infrastructure for Ai2's next MoE-based Olmo, not a model release.

research

Anthropic Red Team: GLM-5.3 Matches Claude on Binary Exploitation for First Time

Anthropic's Frontier Red Team reports that Zhipu AI's GLM-5.3 achieved full control flow hijacks in 4% of binary exploitation trials, versus 6% for Claude Mythos Preview. Predecessor models Claude Opus 4.6 and GLM-5.2 scored zero, marking what Anthropic calls a crossed threshold in offensive cyber capability.

Comments

Loading...