product updateMicrosoft

Microsoft Copilot Researcher adds multi-model features using GPT and Claude

TL;DR

Microsoft has enabled its Copilot Researcher tool to simultaneously leverage OpenAI's GPT and Anthropic's Claude through two new features: Critique, which uses GPT responses refined by Claude, and Model Council, which displays side-by-side outputs with agreement/disagreement analysis. Both features are rolling out in the Microsoft 365 Copilot Frontier early access program.

2 min read
0

Microsoft Copilot Researcher Adds Multi-Model Architecture Using GPT and Claude

Microsoft has expanded its Copilot Researcher tool with dual new features that combine OpenAI's GPT and Anthropic's Claude simultaneously for research tasks, according to a blog post announcing the Copilot Cowork platform.

Critique Feature: Sequential Model Refinement

The Critique feature generates initial responses using GPT, then refines them through Claude. Microsoft claims this architecture "creates a powerful feedback loop that delivers higher-quality results across factual accuracy, analytical breadth, and presentation."

The company states that Researcher's process mirrors "academic and professional research settings." According to Microsoft, the upgrade scores higher than Perplexity's Deep Research models on the Deep Research Accuracy, Completeness, and Objectivity benchmark—though specific benchmark numbers were not disclosed.

Model Council: Comparative Analysis

Alternatively, users can select Model Council to receive side-by-side responses from both Anthropic and OpenAI models. The feature includes a report highlighting where the two models agree and disagree, allowing researchers to evaluate multiple perspectives on complex queries.

Microsoft designed Researcher specifically for multi-step research tasks, distinguishing it from standard Copilot. Anthropic independently operates a Research feature using multiple Claude agents for similar purposes.

Availability and Context

Both Critique and Model Council features are currently available exclusively in Microsoft 365 Copilot's Frontier program, which functions as an early access testing ground for the company's AI innovations. No general availability date has been announced.

The move reflects a broader industry trend of combining complementary AI models to offset individual model weaknesses. Neither OpenAI nor Anthropic has publicly objected to this architecture, suggesting potential partnership or licensing arrangements, though details remain undisclosed.

What this means

Microsoft is positioning multi-model research workflows as enterprise standard practice rather than single-model reliance. This approach benefits organizations requiring high-confidence outputs on complex research tasks. The lack of published benchmark numbers limits independent verification of claims versus Perplexity's comparable offering. Availability remains gated to early access users, suggesting Microsoft is gathering feedback before broader rollout.

Related Articles

product update

OpenAI to watermark ChatGPT and Codex text in the EU under AI Act; API opt-in available worldwide

OpenAI will add an invisible watermark to text generated by ChatGPT and Codex in the European Union to comply with the EU AI Act's transparency rules. Developers anywhere can enable it on select API models starting today, but it is off by default. OpenAI's own tests show detection falling from about 92% to 66% after 10% of words are replaced with synonyms.

product update

OpenAI to watermark ChatGPT and Codex text in the EU with textGrain; API watermarking is opt-in worldwide

OpenAI will switch on invisible text watermarks called textGrain for ChatGPT and Codex users in the EU over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory approach for Claude. OpenAI's own data shows detection drops sharply when text is edited.

product update

Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support

Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.

benchmark

Microsoft's ThinkingBox: Claude Opus 5.5 passes all 20 runs on just 241 of 507 stateful agent tasks

Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the database state they leave behind rather than their responses. Across 507 workflows run 20 times each, Claude Opus 5.5 leads at 67.16% pass@1 but passes all 20 attempts on only 241 tasks.

Comments

Loading...