product updateAmazon Web Services

uniopen lifts Amazon Nova 2 Lite moderation F1 from 0.585 to 0.855 using LoRA fine-tuning on SageMaker AI

TL;DR

Taiwan retail platform uniopen adapted Amazon Nova 2 Lite to its two-axis moderation policy using LoRA supervised fine-tuning in Amazon SageMaker AI, plus a prompt-format change. According to AWS, Per Behavior Macro F1 rose from 0.5852 to 0.8550 and Subject Type Macro F1 from 0.4162 to 0.8491, both above production targets.

3 min read
0

uniopen, the digital communication and membership platform from Taiwan's Uni-President Enterprises Group, raised Amazon Nova 2 Lite's Per Behavior Macro F1 from 0.5852 to 0.8550 and its Subject Type Macro F1 from 0.4162 to 0.8491 using supervised fine-tuning and a prompt change, according to an AWS Machine Learning blog post. Both scores now exceed the company's production targets of 0.8500 and 0.8200. The figures are self-reported by AWS and uniopen and have not been independently verified.

The moderation task

uniopen applies one moderation policy across web, tablet, and mobile. Each interaction is classified on two axes:

  • Behavior: what occurred, across nine categories.
  • Subject: what the behavior refers to: brand, other, or forbidden.

Both labels must be correct for a decision to be useful. The taxonomy is specific to uniopen's business.

Method

The team ran supervised fine-tuning with Low-Rank Adaptation (LoRA) on Amazon Nova 2 Lite in Amazon SageMaker AI. The training set held 3,391 conversation windows, each a bounded segment of a customer conversation. All configurations were evaluated on the same held-out set of 737 windows.

After fine-tuning, the team changed the output format from JSON to a line-based format and clarified how multiple behaviors should be returned. This required no additional training.

Results

Metric Baseline Fine-tuned (JSON) Prompt-optimized Target
Per Behavior Macro F1 0.5852 0.8364 0.8550 ≥ 0.8500
Subject Type Macro F1 0.4162 0.8302 0.8491 ≥ 0.8200

Fine-tuning produced most of the gain: +0.2512 on behavior and +0.4140 on subject type. The prompt change added +0.0186 and +0.0189. After fine-tuning alone, subject type already cleared its target, while behavior (0.8364) remained below 0.8500. The prompt-format change closed that gap.

Macro F1 weights each class equally, so strong results on frequent categories cannot mask weak results on rarer ones.

Production architecture

Amazon Nova 2 Lite handles live moderation requests. Amazon Nova 2 Pro generates candidate corrections for user-reported errors, but a human reviewer must verify each one before it enters the training set. Verified corrections are stored in Amazon S3.

The remaining components:

  • Amazon DynamoDB tracks active and candidate model configurations.
  • Argo Workflows on Amazon EKS orchestrates prompt optimization, evaluation, and deployment.
  • Argo CD applies approved configurations to production.
  • Amazon SNS and CloudWatch alert operators when a gate fails or a candidate needs attention.

Promotion uses two gate types. Hard gates are must-pass regression tests, and a failure stops the workflow and sends an alert. Soft gates are warning signals such as low confidence or a per-class performance drop. A candidate that passes the hard gates but triggers a soft gate waits for administrator approval. Otherwise it is promoted automatically. The post also mentions optional Amazon Bedrock Guardrails content filters on inputs and outputs.

Not disclosed

The post does not give Nova 2 Lite's context window, per-token pricing, or training cutoff in the portion reviewed. It also does not give the LoRA hyperparameters, training run time, or inference cost after customization. Model availability varies by AWS Region.

What this means

This is a customer case study, not a model release. It still shows that a small, low-cost-tier model can reach a narrow production target with a few thousand labeled examples. The roughly 0.25 to 0.41 F1 gains from LoRA show how far general-purpose models sit from business-specific taxonomies, especially for the subject-type axis, where the baseline was 0.4162.

The workflow design matters as much as the scores. Human-verified labels, fixed test sets, and hard-gate regression checks keep model-generated corrections from becoming ground truth, which is a common failure mode in feedback loops. The cheap prompt-format change also delivered nearly two points of F1 on each metric, so output formatting is worth testing before another training round.

Caveats: the 737-window test set is small, and the results come from one deployment with no third-party replication. Teams should validate against their own policies before treating these gains as typical.

Related Articles

model release

Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base

Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.

product update

AWS details ambient agent pattern on Bedrock AgentCore: S3 events trigger jobs, one ask_human tool pauses for approval

AWS published a reference implementation for ambient agents on Amazon Bedrock AgentCore. S3 uploads or scheduled events create jobs that an agent runs, pausing for human input through a single ask_human tool. Each agent turn is capped at the 15-minute Lambda timeout.

product update

AWS Publishes Reference Architecture for Contract Intelligence Using Bedrock AgentCore and Dual Claude Models

AWS published a reference architecture showing how to combine structured data extraction with Bedrock AgentCore, dual Claude models, and Amazon Quick to answer portfolio-wide questions that standard RAG systems get wrong. The design uses Claude Sonnet 4.6 for extraction and Claude Haiku 4.5 for independent verification, with Amazon Textract as a deterministic tiebreaker.

product update

OpenAI launches Dots agent on GPT-6 Astra, limited to $20/month Pro tier and above

OpenAI unveiled Dots at DevDay 2026, a personal agent powered by GPT-6 Astra, available for now only on its $20/month Pro plan and above. Meta's rival Muse agent is free for anyone with a Meta account. OpenAI is positioning Dots as a long-horizon, knowledge-work tool with stronger privacy controls.

Comments

Loading...