Microsoft expands Copilot Cowork with AI model critique feature and cross-model comparison
Microsoft is expanding Copilot Cowork availability and introducing a Critique function that enables one AI model to review another's output. The update also includes a new Researcher agent claiming best-in-class deep research performance, outperforming Perplexity by 7 points, and a Model Council feature for direct model comparison.
Microsoft Expands Copilot Cowork With AI Models Reviewing Each Other's Work
Microsoft is broadening access to Copilot Cowork and introducing a new Critique function that lets AI models evaluate each other's outputs, part of Wave 3 of Microsoft 365 Copilot.
The expanded Copilot Cowork feature builds on the previously announced Claude Cowork capability, enabling the system to handle multi-step tasks using tools, accessing and outputting files, calendar planning, and daily briefings. The feature is now available through Microsoft's Frontier program.
AI Models Checking Each Other's Work
The new Critique function represents a shift toward ensemble model validation. In this workflow, one AI model generates a draft response while a second model reviews and critiques the output. Microsoft draws from both Anthropic and OpenAI models for this capability, allowing different model combinations to work in tandem.
This approach addresses a persistent challenge in AI deployment: single models can propagate errors or miss nuances without external validation. By enabling cross-model review, Microsoft is attempting to improve output quality through algorithmic consensus.
Researcher Agent Performance Claims
Microsoft introduced a new Researcher tool featuring the Critique function and claims it achieves "best-in-class deep research performance." According to Microsoft's benchmark, the Researcher agent outperforms Perplexity with Claude Opus 4.6 by 7 points.
However, the benchmark notably excludes comparison with OpenAI's GPT-5-based Deep Research, limiting assessment of competitive positioning in this capability area.
Model Council for Side-by-Side Comparison
A new Model Council feature allows users to compare answers from different AI models simultaneously, displaying where models agree or diverge. This provides transparency into model behavior and reasoning differences, helping users identify which model performs better for specific tasks.
The feature addresses a practical pain point for organizations deploying multiple models: without direct comparison tools, determining model strengths for different use cases requires manual testing.
What This Means
Microsoft's emphasis on AI-to-AI validation and explicit model comparison reflects industry movement toward collaborative and competitive model architectures. Rather than optimizing single models in isolation, these updates suggest a strategy of leveraging multiple models as checks on each other—reducing hallucination, improving reasoning accuracy, and giving enterprise users visible control over model selection.
The Critique function's reliance on both Anthropic and OpenAI models demonstrates Microsoft's hedging strategy in the multimodel ecosystem. However, the absence of OpenAI's latest deep research tool from benchmarks raises questions about how these capabilities stack up against competitors' newest offerings. The limited claim (7-point margin over one competing tool) suggests marginal rather than substantial advantage.
Related Articles
OpenAI to watermark ChatGPT and Codex text in the EU under AI Act; API opt-in available worldwide
OpenAI will add an invisible watermark to text generated by ChatGPT and Codex in the European Union to comply with the EU AI Act's transparency rules. Developers anywhere can enable it on select API models starting today, but it is off by default. OpenAI's own tests show detection falling from about 92% to 66% after 10% of words are replaced with synonyms.
OpenAI to watermark ChatGPT and Codex text in the EU with textGrain; API watermarking is opt-in worldwide
OpenAI will switch on invisible text watermarks called textGrain for ChatGPT and Codex users in the EU over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory approach for Claude. OpenAI's own data shows detection drops sharply when text is edited.
Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support
Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.
Microsoft's ThinkingBox: Claude Opus 5.5 passes all 20 runs on just 241 of 507 stateful agent tasks
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the database state they leave behind rather than their responses. Across 507 workflows run 20 times each, Claude Opus 5.5 leads at 67.16% pass@1 but passes all 20 attempts on only 241 tasks.
Comments
Loading...