Baidu releases ERNIE-Image, an 8B parameter text-to-image model with strong text rendering capabilities
Baidu has released ERNIE-Image, an 8B parameter text-to-image generation model built on a single-stream Diffusion Transformer architecture. The model is designed for complex instruction following, text rendering, and structured image generation, and can run on consumer GPUs with 24GB VRAM.
Baidu releases ERNIE-Image, an 8B parameter text-to-image model with strong text rendering capabilities
Baidu has released ERNIE-Image, an 8B parameter text-to-image generation model that claims state-of-the-art performance among open-weight models in its size class. The model is built on a single-stream Diffusion Transformer (DiT) architecture and includes a lightweight Prompt Enhancer component that expands user inputs into structured descriptions.
Technical specifications
ERNIE-Image uses 8 billion DiT parameters and can generate images at multiple resolutions including 1024x1024, 848x1264, and 1264x848 pixels. The model requires 50 inference steps with a guidance scale of 4.0 for the base version. According to Baidu, the model can run on consumer GPUs with 24GB VRAM.
Baidu has also released ERNIE-Image-Turbo, a faster variant optimized with Distribution Matching Distillation (DMD) and reinforcement learning that generates images in 8 inference steps.
Benchmark performance
On the GENEval benchmark, ERNIE-Image with Prompt Enhancer scored 0.8728 overall, outperforming FLUX.2-klein-9B (0.8481) and Z-Image (0.8400). The model scored particularly well on single object generation (0.9906) and two object generation (0.9596).
For text rendering specifically, ERNIE-Image achieved 0.9733 average score on LongTextBench across English and Chinese, trailing only Seedream 4.5 (0.9882) but ahead of GLM-Image (0.9656) and Nano Banana 2.0 (0.9650).
On the OneIG-EN benchmark measuring alignment, text, reasoning, style, and diversity, ERNIE-Image with Prompt Enhancer scored 0.5750 overall, ranking third behind Nano Banana 2.0 (0.5780) and Seedream 4.5 (0.5760).
Intended use cases
Baidu positions ERNIE-Image for commercial applications requiring precise control over generated content, including posters, comics, multi-panel layouts, infographics, and UI mockups. The model supports multiple visual styles including realistic photography, design-oriented imagery, and stylized aesthetic outputs.
The model is available on Hugging Face with both Diffusers and SGLang inference support. Baidu has not disclosed pricing for commercial API access.
What this means
ERNIE-Image represents a strategic release from Baidu targeting practical commercial applications rather than purely aesthetic generation. The 8B parameter count makes it computationally accessible while the benchmark scores suggest competitive performance with larger models. The emphasis on text rendering and instruction following addresses specific pain points in text-to-image generation where models often struggle with accurate text and complex layouts. The availability of a turbo variant with 8-step inference indicates Baidu's focus on deployment efficiency alongside quality.
Related Articles
DeepSeek Releases V4-Flash-Vision-Exp, First Multimodal Model in V4 Family
DeepSeek has released DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the V4 family, adding visual understanding to the V4-Flash architecture. The 305B-parameter model shows substantial gains on multimodal agent benchmarks while holding steady on text-only tasks.
Tencent Open-Sources Hy4 Preview: 770B-Parameter MoE Model with 1M-Token Context
Tencent's Hy Team has open-sourced Hy4 preview, a 770-billion-parameter Mixture-of-Experts model with 49 billion activated parameters and a 1-million-token context window. The model is available under Apache 2.0 alongside an FP8-quantized variant, with Tencent claiming it beats GLM 5.3 and Kimi K3 on internal engineering evaluations.
IBM Releases Granite 4.2 8B, a Dense Reasoning Model with 131K Context and Three Thinking Modes
IBM has released Granite 4.2 8B, a dense reasoning model built for math, code generation, and agentic workflows. The model supports 131K context, 12 languages, and three switchable reasoning modes, priced at $0.10 per 1M input tokens and $0.15 per 1M output tokens.
Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents
Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.
Comments
Loading...