changelogAnthropic

Anthropic Python SDK v0.104.0 adds thinking token count estimates for streaming responses

TL;DR

Anthropic released version 0.104.0 of its Python SDK on May 21, 2026. The update adds support for a thinking-token-count beta feature that provides estimated token counts in thinking block deltas when streaming responses from reasoning models.

1 min read
0

Anthropic Python SDK v0.104.0 adds thinking token count estimates for streaming responses

Anthropic released version 0.104.0 of its Python SDK on May 21, 2026, adding support for tracking token usage in reasoning model thought processes.

What's new

The update introduces a thinking-token-count beta feature that provides estimated token counts within thinking block deltas during streaming responses. This allows developers to monitor token consumption in real-time as Claude's reasoning models process extended chains of thought.

The feature specifically targets streaming scenarios where Claude models with thinking capabilities—such as Claude 3.5 Sonnet with extended thinking mode—generate internal reasoning before producing final outputs.

Technical details

The implementation provides token count estimates as part of the delta stream, enabling developers to:

  • Track thinking token usage during active streaming
  • Estimate costs for reasoning operations in real-time
  • Monitor and debug extended thinking processes
  • Optimize prompts based on thinking token consumption

The feature is marked as beta, indicating the API may change in future releases.

Version information

  • Version: 0.104.0
  • Release date: May 21, 2026
  • Type: Minor version update
  • Full changelog: Available at github.com/anthropics/anthropic-sdk-python/compare/v0.103.1...v0.104.0

What this means

This update addresses a key observability gap for developers using Claude's reasoning capabilities. Previously, tracking token usage in thinking blocks required waiting for complete responses. With streaming token counts, developers can now monitor costs and performance in real-time, particularly important for applications using extended thinking modes where reasoning token counts can significantly exceed output tokens. The beta designation suggests Anthropic is still refining how thinking token metrics are calculated and surfaced to developers.

Related Articles

changelog

OpenAI Python SDK v3.21.0 Adds GPT-6.1 Sol Model Identifier

OpenAI released v3.21.0 of its official Python SDK on September 29, 2026, adding a constant for a new 'GPT-6.1 Sol' model identifier. No pricing, benchmarks, or capability details have been disclosed.

changelog

Anthropic Python SDK 1.9.0 Adds Reference to Unreleased 'claude-sonnet-5-5' Model ID

Anthropic's anthropic-sdk-python v1.9.0 release adds a reference to an unannounced 'claude-sonnet-5-5' model ID, a new between_tools thinking type, and the ability to run tool calls while a reply streams. No pricing, context window, or benchmark data for the model has been disclosed.

product update

Anthropic adds Mods to Claude Code, a plugin system that hooks into tool calls, prompts and UI rendering

Anthropic released Mods for Claude Code, a plugin system built on JavaScript and TypeScript functions that hook into events such as tool calls, user prompts and UI rendering. Mods are not sandboxed and run with the user's permissions. They work in the CLI, the desktop app and, partly, the VS Code extension.

research

Graphite: Opus 5.5 uses 'this matters' 116x more than humans as AI writing tells persist

Marketing firm Graphite identified 13,000 phrases that appear at least twice as often in AI-generated writing as in human writing. Claude Opus 5.5 uses "this matters" 116 times more than humans, while OpenAI's Astra favors "corrective framing" more than 100 times as often. Em-dash use has collapsed across frontier models, but total tells are holding steady, according to Graphite.

Comments

Loading...