OpenAI's GPT 5.4 ties Gemini 3.1 Pro at 72.4% on Google's Android coding benchmark
Google's Android Bench—a benchmark measuring AI model performance for Android app development—shows OpenAI's GPT 5.4 and Google's Gemini 3.1 Pro Preview tied at 72.4% in the latest April 2026 update. OpenAI's GPT 5.3-Codex ranks third at 67.7%, while Anthropic's Claude Opus 4.6 scores 66.6%.
OpenAI's GPT 5.4 Ties Gemini 3.1 Pro at Top of Google's Android Coding Benchmark
Google's Android Bench—introduced in March 2026 as a resource for evaluating AI models in Android app development—released its first update today, showing OpenAI's latest model matching Google's flagship offering.
Benchmark Scores
GPT 5.4 and Gemini 3.1 Pro Preview both scored 72.4%, the highest on the list. OpenAI's GPT 5.3-Codex follows at 67.7%. Anthropic's Claude Opus 4.6 ranks fourth at 66.6%.
Complete April 2026 rankings:
- GPT 5.4: 72.4% (new)
- Gemini 3.1 Pro Preview: 72.4%
- GPT 5.3-Codex: 67.7% (new)
- Claude Opus 4.6: 66.6%
- GPT-5.2 Codex: 62.5%
- Claude Opus 4.5: 61.9%
- Gemini 3 Pro Preview: 60.4%
- Claude Sonnet 4.6: 58.4%
- Claude Sonnet 4.5: 54.2%
- Gemini 3 Flash Preview: 42%
- Gemini 2.5 Flash: 16.1%
Methodology
Google's evaluation framework assesses model capabilities across Android development essentials: Jetpack Compose for UI construction, Coroutines and Flows for asynchronous programming, Room for data persistence, and Hilt for dependency injection. Testing of OpenAI's models occurred in mid-March 2026, prior to their public release this week.
Important Caveats
Google explicitly noted that benchmark results should not be treated as definitive. Real-world performance varies significantly based on workflow, pricing, integration ease, and specific use cases. The methodology measures controlled scenarios that may not reflect production development conditions.
The remainder of the ranking remained unchanged from the initial late-February test run, with no new models added beyond OpenAI's two entries.
Context
Google positioned Android Bench as a tool to help developers "be more productive" and ultimately "deliver higher quality apps across the Android ecosystem." The benchmark represents Google's effort to provide transparent model comparisons for a specific, high-impact use case.
What This Means
The tight tie between GPT 5.4 and Gemini 3.1 Pro suggests convergence at the frontier for code-generation tasks. Developers choosing between them will likely base decisions on factors beyond benchmark scores—pricing, API latency, context window size, and ecosystem integration. The substantial performance gap between top-ranked models (72.4%) and older versions (Gemini 2.5 Flash at 16.1%) indicates rapid model improvement in specialized coding tasks. For teams standardized on either Google or OpenAI infrastructure, this benchmark provides useful calibration without necessarily indicating a decisive winner.
Related Articles
OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness
OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.
Gemini 3.5 Flash ranks 6th in Android coding benchmark at 3x cost of Gemini 3.1 Pro
Google's latest Android Bench results show Gemini 3.5 Flash ranking 6th with a 63.7% success rate, despite averaging $147.10 per benchmark run compared to Gemini 3.1 Pro Preview's $47.90. The newer model used 355.9 tokens per run versus 73.3 for its predecessor, while GPT 5.5 leads the benchmark at 74% success rate.
Nvidia Claims Groq 3 LPX Hits 3,400 Tokens/Sec, 4x Cerebras — But Needs 64 Chips to Do It
Nvidia's new Groq 3 LPX inference accelerator hit 3,400 tokens per second on Gemma 4 31B, a figure the company says is four times faster than Cerebras. Experts note the benchmark uses at least 64 LPX chips versus Cerebras' one or two accelerators, making the comparison far less clean than it appears.
Artificial Analysis Launches Search Index Benchmark for AI Agent Search APIs
Artificial Analysis has released the Search Index, a benchmark measuring how search API providers perform for AI agents across quality, cost, and speed. Parallel, Exa, and Firecrawl lead the initial rankings, with search access boosting model scores from 33 to as high as 75 points.
Comments
Loading...