Amazon Bedrock now supports fine-tuning for Nova models with three customization approaches
Amazon Bedrock now enables fine-tuning of Amazon Nova models using supervised fine-tuning (SFT), reinforcement fine-tuning (RFT), and model distillation. The service automates infrastructure provisioning and training orchestration, requiring only data upload to S3 and a single API call. Fine-tuned models run on-demand at standard inference pricing without provisioned capacity requirements.
Amazon Nova 2 Lite — Quick Specs
Amazon Bedrock Adds Fine-tuning for Nova Models
Amazon has announced fine-tuning capabilities for Amazon Nova models through Amazon Bedrock, enabling customers to customize models for domain-specific tasks without deep machine learning expertise.
Three Customization Approaches
Bedrock supports three fine-tuning techniques:
Supervised Fine-tuning (SFT): Trains models on labeled input-output examples, embedding domain knowledge directly into model weights.
Reinforcement Fine-tuning (RFT): Uses reward functions—either custom code or an LLM acting as judge—to guide learning toward target behaviors.
Model Distillation: Transfers knowledge from larger teacher models into smaller, faster student models for resource-constrained environments.
All three approaches use parameter-efficient fine-tuning (PEFT), reducing memory requirements and training time while maintaining model quality compared to full fine-tuning.
Supported Models
Amazon Nova 2 Lite and Nova Micro support fine-tuning. Nova 2 Lite is a multimodal model with a 1-million token context window, processing text, images, and video for document processing, video understanding, and code generation. Nova Micro, the smallest in the lineup, targets low-cost inference for pipeline processing tasks like data extraction and address fixing.
Implementation and Pricing
Amazon Bedrock automates the entire training pipeline. Users upload training data to Amazon S3 and initiate the job via AWS Management Console, CLI, or API. The service manages infrastructure provisioning, compute allocation, and training orchestration—no cluster configuration required.
Fine-tuned models run on-demand at the same inference pricing as non-customized versions, with no provisioned capacity requirement. This contrasts with traditional approaches requiring expensive Provisioned Throughput.
Performance Gains
Amazon's internal testing demonstrated measurable improvements. Amazon Customer Service customized Nova Micro for specialized support, improving accuracy by 5.4% on domain-specific issues and 7.3% on general issues while reducing latency.
Fine-tuning eliminates token consumption overhead compared to prompt engineering and Retrieval-Augmented Generation (RAG), which supply context at inference time. While context-based techniques offer immediate deployment and dynamic updates, fine-tuning embeds knowledge directly, reducing cumulative token costs and improving generalization to novel phrasings and edge cases.
When to Fine-tune
Amazon recommends fine-tuning for high-volume, well-defined tasks with quality labeled examples—such as intent classification, brand voice consistency, or replacing traditional ML classifiers. The upfront investment in data labeling and training pays off through reduced per-request inference costs for applications with sustained traffic.
Fine-tuned small LLMs like Nova Micro increasingly replace traditional classifiers for tasks requiring flexibility with natural language variation without retraining.
Training Visibility
Bedrock provides sensible hyperparameter defaults (epochCount, learningRateMultiplier) and real-time training monitoring through loss curves. Clear documentation covers data preparation, format specifications, and schema requirements.
What this means
Bedrock's fine-tuning removes infrastructure barriers for model customization, making it accessible to teams without ML ops expertise. The on-demand pricing model—eliminating provisioned capacity costs—alters economics for domain-specific deployments. This positions Nova models as viable replacements for traditional classifiers in production pipelines, particularly where cost and latency matter more than raw capability. The focus on parameter-efficient approaches preserves inference speed, critical for high-volume applications.
Related Articles
Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
AWS adds managed Web Search to Claude Desktop via Bedrock AgentCore Gateway in three Regions
AWS published a walkthrough for connecting Claude Desktop on Amazon Bedrock to a managed, MCP-compatible Web Search capability through Amazon Bedrock AgentCore Gateway. According to AWS, the search is backed by an Amazon web index spanning tens of billions of documents, and query traffic stays within AWS infrastructure. Web Search is available in three AWS Regions; pricing is not disclosed in the post.
Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support
Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.
AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints
Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.
Comments
Loading...