Microsoft Releases Lens-Turbo: 3.8B-Parameter Text-to-Image Model Trained on 800M GPT-4.1-Captioned Images
Microsoft has released Lens-Turbo, a 3.8B-parameter foundational text-to-image model designed for efficient training and fast generation. The model was trained on Lens-800M, an 800 million image-text corpus with GPT-4.1 captions, and supports resolutions up to 1440×1440 with 4-step distilled inference.
Microsoft Releases Lens-Turbo: 3.8B-Parameter Text-to-Image Model
Microsoft has released Lens-Turbo, a 3.8 billion-parameter text-to-image model trained on 800 million GPT-4.1-captioned images. The model uses a 48-block MMDiT (multi-modal diffusion transformer) architecture and supports generation at resolutions up to 1440×1440 pixels.
Technical Architecture
Lens combines several technical approaches:
- Training corpus: Lens-800M dataset containing 800 million image-text pairs with long-form GPT-4.1 captions
- Architecture: 48-block MMDiT denoiser with 3.8B parameters
- Latent encoding: Uses FLUX.2 semantic VAE for image encoding
- Text encoding: Concatenated multi-layer GPT-OSS features for prompt following and multilingual support
- Resolution handling: Mixed-resolution training enables aspect ratios from 1:2 to 2:1
Inference Speed
The distilled Lens-Turbo variant supports 4-step generation, according to Microsoft. The base model went through reinforcement learning post-training for improved visual quality and artifact suppression before distillation.
Resolution and Aspect Ratio Support
The model supports flexible output resolutions:
- Maximum resolution: 1440×1440 pixels
- Aspect ratio range: 1:2 to 2:1
- Multiple resolution presets: 1248×1664, 1664×1248, and square formats
Microsoft states the mixed-resolution training approach enables inference across different aspect ratios without quality degradation.
Training Efficiency Claims
Microsoft claims Lens reaches "competitive quality with substantially less training compute than larger T2I models" through dense-caption pre-training that maximizes information density per training batch. The company has not disclosed specific benchmark scores, training compute requirements, or comparisons to specific competing models.
Model Availability
The model is available on Hugging Face under the repository microsoft/Lens-Turbo. Microsoft has released minimal inference code for generating images from Lens DiT checkpoints. Pricing information for API access has not been disclosed.
What This Means
Lens-Turbo represents Microsoft's entry into the sub-4B parameter text-to-image model category, emphasizing training efficiency through high-quality captions rather than dataset scale. The 4-step distilled inference and flexible resolution support position it for applications requiring fast generation across varied aspect ratios. The reliance on GPT-4.1 for caption generation suggests Microsoft is leveraging its existing LLM infrastructure to improve training data quality, though the actual performance relative to models like Stable Diffusion 3 or FLUX.1 remains unverified without published benchmarks.
Related Articles
Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October
Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.
Microsoft's ThinkingBox: Claude Opus 5.5 passes all 20 runs on just 241 of 507 stateful agent tasks
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the database state they leave behind rather than their responses. Across 507 workflows run 20 times each, Claude Opus 5.5 leads at 67.16% pass@1 but passes all 20 attempts on only 241 tasks.
Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages
Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.
TII releases 1.6B Falcon-ASR, claims 20.92% Arabic WER against best listed 23.17%
The Technology Innovation Institute (TII) released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect. TII claims a 20.92% average word error rate across six Arabic test sets, versus 23.17% for the next-best system on the leaderboard snapshot it used. A demo is live on Hugging Face. Pricing and API availability have not been disclosed.
Comments
Loading...