researchMistral AI

Mistral AI fine-tunes Pixtral-12B on satellite imagery, boosting classification accuracy from 56% to 91%

TL;DR

Mistral AI reports that fine-tuning its Pixtral-12B vision model on satellite imagery increased classification accuracy from 56% to 91% on the Aerial Image Dataset. The company used LoRA (Low-Rank Adaptation) to train on 8,000 samples for under $10, reducing hallucinations from 5% to 0.1%.

2 min read
0

Mistral AI fine-tunes Pixtral-12B on satellite imagery, boosting classification accuracy from 56% to 91%

Mistral AI reports that fine-tuning its Pixtral-12B vision model on satellite imagery increased classification accuracy from 56% to 91% on the Aerial Image Dataset, demonstrating how domain-specific adaptation can dramatically improve model performance on specialized tasks.

The experiment

Mistral used the publicly available Aerial Image Dataset (AID), splitting it into 8,000 training samples and 2,000 test samples across 30 scene categories. The categories included challenging distinctions like dense residential versus medium residential areas, and ambiguous terms like "center."

The base Pixtral-12B model, using only prompt engineering with a system prompt listing all target classes, achieved 56% accuracy. The model also hallucinated invalid class names 5% of the time, producing labels not in the original set.

Fine-tuning approach

Mistral applied Low-Rank Adaptation (LoRA), which injects small trainable matrices into the model's weights rather than retraining the entire model. According to Mistral, the fine-tuning required no extensive hyperparameter tuning and cost under $10 to complete.

The company used its fine-tuning API, training the model by providing correct labels for system prompts and input images. Mistral recommends starting with a single epoch and monitoring for overfitting risk.

Results

After fine-tuning on the 8,000 training samples:

  • Overall accuracy: 91% (up from 56%)
  • Hallucination rate: 0.1% (down from 5%)
  • Performance became more consistent across all 30 classes

Mistral highlighted a specific example where the base model confused "Playground" and "Stadium" categories, both classified incorrectly as "Stadium." The fine-tuned model learned to distinguish them based on features like the presence of seating surrounding sports fields.

Technical details

Mistral's LoRA implementation allows developers to adapt models without modifying full model weights. The technique proves particularly useful when prompt engineering produces inconsistent results or when dealing with complex, domain-specific visual patterns.

The company offers two fine-tuning options: direct API calls for granular hyperparameter control, or the LaPlateforme UI, which automatically computes optimal batch size based on dataset size.

What this means

This research demonstrates that vision-language models can achieve significant performance gains on specialized visual domains through relatively inexpensive fine-tuning. The 1.6x accuracy improvement on satellite imagery with just 8,000 samples suggests similar approaches could work for other underrepresented visual domains in VLM training sets, such as medical imaging, surveillance analysis, or manuscript transcription.

The under-$10 cost and single-epoch training make this approach accessible for organizations with limited budgets working on specialized computer vision tasks. Mistral has published a cookbook with full implementation details at github.com/mistralai/cookbook.

Related Articles

research

Stanford, Caltech Researchers Wire GPT-6 Astra Directly Into a Robot to Clean an Unfamiliar Kitchen

Researchers built HomeBody, a system that connects GPT-6 Astra directly to a Unitree G1 robot's skill library, letting it explore, map, and tidy an unfamiliar kitchen without a trained control layer in between. The team reports latency, overheating servos, and compute cost as current limitations.

research

Anthropic's Claude Science builds first complete all-sky ultraviolet map, filling gaps with AI inpainting

Anthropic says its Claude Science system produced the first complete ultraviolet map of the sky. AI agents downloaded, calibrated and merged data from multiple space missions, then used inpainting to fill gaps. In tests, the filled-in predictions deviated about 10% from actual measurements on average.

research

OpenAI publishes 372 AI-generated math results on GitHub, claims they solve or advance open problems

OpenAI has published 372 mathematical results generated by an unnamed internal frontier model, hosted on GitHub instead of in peer-reviewed journals. The company claims each result solves or substantially advances an open problem, at an average of about three hours of ChatGPT Pro Thinking compute per result. Many include Lean formalizations, but independent validation of significance is still pending.

research

Google's RRSI cuts overfitting in self-improving agents: up to 4.7-point gains on unseen tasks with ~30% fewer tokens

Google Cloud AI Research and several universities introduced RRSI, a method that stops self-optimizing agent harnesses from memorizing their test tasks. According to the paper, it gains up to 4.7 points on five unseen benchmarks and uses about 30% fewer runtime tokens than the unregularized version, with the underlying model frozen.

Comments

Loading...