benchmark

OpenAI's GPT 5.4 ties Gemini 3.1 Pro at 72.4% on Google's Android coding benchmark

TL;DR

Google's Android Bench—a benchmark measuring AI model performance for Android app development—shows OpenAI's GPT 5.4 and Google's Gemini 3.1 Pro Preview tied at 72.4% in the latest April 2026 update. OpenAI's GPT 5.3-Codex ranks third at 67.7%, while Anthropic's Claude Opus 4.6 scores 66.6%.

2 min read
0

OpenAI's GPT 5.4 Ties Gemini 3.1 Pro at Top of Google's Android Coding Benchmark

Google's Android Bench—introduced in March 2026 as a resource for evaluating AI models in Android app development—released its first update today, showing OpenAI's latest model matching Google's flagship offering.

Benchmark Scores

GPT 5.4 and Gemini 3.1 Pro Preview both scored 72.4%, the highest on the list. OpenAI's GPT 5.3-Codex follows at 67.7%. Anthropic's Claude Opus 4.6 ranks fourth at 66.6%.

Complete April 2026 rankings:

  • GPT 5.4: 72.4% (new)
  • Gemini 3.1 Pro Preview: 72.4%
  • GPT 5.3-Codex: 67.7% (new)
  • Claude Opus 4.6: 66.6%
  • GPT-5.2 Codex: 62.5%
  • Claude Opus 4.5: 61.9%
  • Gemini 3 Pro Preview: 60.4%
  • Claude Sonnet 4.6: 58.4%
  • Claude Sonnet 4.5: 54.2%
  • Gemini 3 Flash Preview: 42%
  • Gemini 2.5 Flash: 16.1%

Methodology

Google's evaluation framework assesses model capabilities across Android development essentials: Jetpack Compose for UI construction, Coroutines and Flows for asynchronous programming, Room for data persistence, and Hilt for dependency injection. Testing of OpenAI's models occurred in mid-March 2026, prior to their public release this week.

Important Caveats

Google explicitly noted that benchmark results should not be treated as definitive. Real-world performance varies significantly based on workflow, pricing, integration ease, and specific use cases. The methodology measures controlled scenarios that may not reflect production development conditions.

The remainder of the ranking remained unchanged from the initial late-February test run, with no new models added beyond OpenAI's two entries.

Context

Google positioned Android Bench as a tool to help developers "be more productive" and ultimately "deliver higher quality apps across the Android ecosystem." The benchmark represents Google's effort to provide transparent model comparisons for a specific, high-impact use case.

What This Means

The tight tie between GPT 5.4 and Gemini 3.1 Pro suggests convergence at the frontier for code-generation tasks. Developers choosing between them will likely base decisions on factors beyond benchmark scores—pricing, API latency, context window size, and ecosystem integration. The substantial performance gap between top-ranked models (72.4%) and older versions (Gemini 2.5 Flash at 16.1%) indicates rapid model improvement in specialized coding tasks. For teams standardized on either Google or OpenAI infrastructure, this benchmark provides useful calibration without necessarily indicating a decisive winner.

Related Articles

benchmark

OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness

OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.

benchmark

Gemini 3.5 Flash ranks 6th in Android coding benchmark at 3x cost of Gemini 3.1 Pro

Google's latest Android Bench results show Gemini 3.5 Flash ranking 6th with a 63.7% success rate, despite averaging $147.10 per benchmark run compared to Gemini 3.1 Pro Preview's $47.90. The newer model used 355.9 tokens per run versus 73.3 for its predecessor, while GPT 5.5 leads the benchmark at 74% success rate.

benchmark

Nvidia Claims Groq 3 LPX Hits 3,400 Tokens/Sec, 4x Cerebras — But Needs 64 Chips to Do It

Nvidia's new Groq 3 LPX inference accelerator hit 3,400 tokens per second on Gemma 4 31B, a figure the company says is four times faster than Cerebras. Experts note the benchmark uses at least 64 LPX chips versus Cerebras' one or two accelerators, making the comparison far less clean than it appears.

benchmark

Artificial Analysis Launches Search Index Benchmark for AI Agent Search APIs

Artificial Analysis has released the Search Index, a benchmark measuring how search API providers perform for AI agents across quality, cost, and speed. Parallel, Exa, and Firecrawl lead the initial rankings, with search access boosting model scores from 33 to as high as 75 points.

Comments

Loading...