Anthropic Python SDK v0.104.0 adds thinking token count estimates for streaming responses
Anthropic released version 0.104.0 of its Python SDK on May 21, 2026. The update adds support for a thinking-token-count beta feature that provides estimated token counts in thinking block deltas when streaming responses from reasoning models.
Anthropic Python SDK v0.104.0 adds thinking token count estimates for streaming responses
Anthropic released version 0.104.0 of its Python SDK on May 21, 2026, adding support for tracking token usage in reasoning model thought processes.
What's new
The update introduces a thinking-token-count beta feature that provides estimated token counts within thinking block deltas during streaming responses. This allows developers to monitor token consumption in real-time as Claude's reasoning models process extended chains of thought.
The feature specifically targets streaming scenarios where Claude models with thinking capabilities—such as Claude 3.5 Sonnet with extended thinking mode—generate internal reasoning before producing final outputs.
Technical details
The implementation provides token count estimates as part of the delta stream, enabling developers to:
- Track thinking token usage during active streaming
- Estimate costs for reasoning operations in real-time
- Monitor and debug extended thinking processes
- Optimize prompts based on thinking token consumption
The feature is marked as beta, indicating the API may change in future releases.
Version information
- Version: 0.104.0
- Release date: May 21, 2026
- Type: Minor version update
- Full changelog: Available at github.com/anthropics/anthropic-sdk-python/compare/v0.103.1...v0.104.0
What this means
This update addresses a key observability gap for developers using Claude's reasoning capabilities. Previously, tracking token usage in thinking blocks required waiting for complete responses. With streaming token counts, developers can now monitor costs and performance in real-time, particularly important for applications using extended thinking modes where reasoning token counts can significantly exceed output tokens. The beta designation suggests Anthropic is still refining how thinking token metrics are calculated and surfaced to developers.
Related Articles
OpenAI Python SDK v3.21.0 Adds GPT-6.1 Sol Model Identifier
OpenAI released v3.21.0 of its official Python SDK on September 29, 2026, adding a constant for a new 'GPT-6.1 Sol' model identifier. No pricing, benchmarks, or capability details have been disclosed.
Anthropic Python SDK 1.9.0 Adds Reference to Unreleased 'claude-sonnet-5-5' Model ID
Anthropic's anthropic-sdk-python v1.9.0 release adds a reference to an unannounced 'claude-sonnet-5-5' model ID, a new between_tools thinking type, and the ability to run tool calls while a reply streams. No pricing, context window, or benchmark data for the model has been disclosed.
Anthropic adds Mods to Claude Code, a plugin system that hooks into tool calls, prompts and UI rendering
Anthropic released Mods for Claude Code, a plugin system built on JavaScript and TypeScript functions that hook into events such as tool calls, user prompts and UI rendering. Mods are not sandboxed and run with the user's permissions. They work in the CLI, the desktop app and, partly, the VS Code extension.
Graphite: Opus 5.5 uses 'this matters' 116x more than humans as AI writing tells persist
Marketing firm Graphite identified 13,000 phrases that appear at least twice as often in AI-generated writing as in human writing. Claude Opus 5.5 uses "this matters" 116 times more than humans, while OpenAI's Astra favors "corrective framing" more than 100 times as often. Em-dash use has collapsed across frontier models, but total tells are holding steady, according to Graphite.
Comments
Loading...