model releaseMicrosoft

Microsoft Releases Lens-Turbo: 3.8B-Parameter Text-to-Image Model Trained on 800M GPT-4.1-Captioned Images

TL;DR

Microsoft has released Lens-Turbo, a 3.8B-parameter foundational text-to-image model designed for efficient training and fast generation. The model was trained on Lens-800M, an 800 million image-text corpus with GPT-4.1 captions, and supports resolutions up to 1440×1440 with 4-step distilled inference.

2 min read
0

Microsoft Releases Lens-Turbo: 3.8B-Parameter Text-to-Image Model

Microsoft has released Lens-Turbo, a 3.8 billion-parameter text-to-image model trained on 800 million GPT-4.1-captioned images. The model uses a 48-block MMDiT (multi-modal diffusion transformer) architecture and supports generation at resolutions up to 1440×1440 pixels.

Technical Architecture

Lens combines several technical approaches:

  • Training corpus: Lens-800M dataset containing 800 million image-text pairs with long-form GPT-4.1 captions
  • Architecture: 48-block MMDiT denoiser with 3.8B parameters
  • Latent encoding: Uses FLUX.2 semantic VAE for image encoding
  • Text encoding: Concatenated multi-layer GPT-OSS features for prompt following and multilingual support
  • Resolution handling: Mixed-resolution training enables aspect ratios from 1:2 to 2:1

Inference Speed

The distilled Lens-Turbo variant supports 4-step generation, according to Microsoft. The base model went through reinforcement learning post-training for improved visual quality and artifact suppression before distillation.

Resolution and Aspect Ratio Support

The model supports flexible output resolutions:

  • Maximum resolution: 1440×1440 pixels
  • Aspect ratio range: 1:2 to 2:1
  • Multiple resolution presets: 1248×1664, 1664×1248, and square formats

Microsoft states the mixed-resolution training approach enables inference across different aspect ratios without quality degradation.

Training Efficiency Claims

Microsoft claims Lens reaches "competitive quality with substantially less training compute than larger T2I models" through dense-caption pre-training that maximizes information density per training batch. The company has not disclosed specific benchmark scores, training compute requirements, or comparisons to specific competing models.

Model Availability

The model is available on Hugging Face under the repository microsoft/Lens-Turbo. Microsoft has released minimal inference code for generating images from Lens DiT checkpoints. Pricing information for API access has not been disclosed.

What This Means

Lens-Turbo represents Microsoft's entry into the sub-4B parameter text-to-image model category, emphasizing training efficiency through high-quality captions rather than dataset scale. The 4-step distilled inference and flexible resolution support position it for applications requiring fast generation across varied aspect ratios. The reliance on GPT-4.1 for caption generation suggests Microsoft is leveraging its existing LLM infrastructure to improve training data quality, though the actual performance relative to models like Stable Diffusion 3 or FLUX.1 remains unverified without published benchmarks.

Related Articles

model release

Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October

Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.

benchmark

Microsoft's ThinkingBox: Claude Opus 5.5 passes all 20 runs on just 241 of 507 stateful agent tasks

Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the database state they leave behind rather than their responses. Across 507 workflows run 20 times each, Claude Opus 5.5 leads at 67.16% pass@1 but passes all 20 attempts on only 241 tasks.

model release

Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages

Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.

model release

TII releases 1.6B Falcon-ASR, claims 20.92% Arabic WER against best listed 23.17%

The Technology Innovation Institute (TII) released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect. TII claims a 20.92% average word error rate across six Arabic test sets, versus 23.17% for the next-best system on the leaderboard snapshot it used. A demo is live on Hugging Face. Pricing and API availability have not been disclosed.

Comments

Loading...