Amazon Bedrock adds programmatic tool calling to reduce latency and token usage in multi-step workflows
Amazon Bedrock now supports programmatic tool calling (PTC), a technique that allows LLMs to generate Python code for multi-step tool orchestration rather than making sequential API calls. AWS offers three implementation paths: self-hosted Docker sandboxes on ECS, managed execution via Amazon Bedrock AgentCore Code Interpreter, and Anthropic SDK-compatible proxy integration.
Amazon Bedrock adds programmatic tool calling to reduce latency and token usage in multi-step workflows
Amazon Web Services has introduced programmatic tool calling (PTC) for Amazon Bedrock, enabling language models to generate Python code that orchestrates multiple tool invocations in a single inference cycle rather than requiring sequential round trips.
How it works
Traditional tool calling requires models to invoke tools one at a time, with each call requiring a full inference round trip. For a query like "Which engineering team members exceeded their Q3 travel budget?", a model using standard tool calling would need to make 20+ separate API calls to retrieve team members and expense records, passing thousands of intermediate data points through its context window.
With PTC, the model generates Python code once that handles all tool calls, data processing, and filtering within a sandboxed execution environment. Using asyncio.gather(), the code can execute tool calls in parallel. Only the final processed result returns to the model's context.
According to AWS, this approach reduces both latency and token consumption for workflows involving multiple tool calls, data aggregation, or numerical calculations.
Three implementation options
AWS provides three ways to implement PTC on Bedrock:
Self-hosted Docker sandbox on Amazon ECS: Offers maximum control with model-agnostic support for Claude, Qwen, MiniMax, Llama, Nova, and other Bedrock models. Developers can customize the sandbox environment, install domain-specific Python packages, and keep code execution within their AWS account. The architecture uses an orchestrator (ECS task or Lambda) that calls the InvokeModel API via Boto3 and manages Docker sandbox lifecycle.
Managed solution via Amazon Bedrock AgentCore Code Interpreter: A fully managed sandbox environment that handles execution without requiring custom infrastructure.
Anthropic SDK-compatible proxy: Designed for teams already using Anthropic's SDK who want PTC functionality while maintaining their existing developer workflow.
System prompt engineering
The self-hosted implementation relies on injecting tool definitions into the system prompt rather than using the standard tool_config parameter. The prompt instructs models to write Python code with specific rules: each execute_code call runs in a fresh stateless environment, tool calls must use await, and all operations must complete in a single code block.
The orchestrator intercepts tool calls through IPC over stdin/stderr, executes them externally, and injects results back into the sandbox.
What this means
PTC addresses a legitimate bottleneck in agentic workflows where sequential tool calling creates compounding latency. The model-agnostic self-hosted option is particularly significant—it extends a pattern originally introduced by specific providers to any model available on Bedrock. This matters for enterprises already committed to AWS infrastructure who want to avoid vendor lock-in at the model level.
The technique works best for workflows involving data aggregation, filtering operations, or scenarios where intermediate data shouldn't enter the model's context for privacy reasons. However, it requires models capable of generating correct async Python code and adds complexity around sandbox security and resource management.
Related Articles
Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
Google may expand Gemini 'Call for Me' to personal calls, per Android Authority APK teardown
Android Authority reports that an APK teardown shows a 'Gemini Calling' intro screen, suggesting Google may extend its 'Call for Me' AI feature from business calls to personal ones. Google has not announced the feature, and the report notes it may never ship publicly.
OpenAI to watermark ChatGPT and Codex text in the EU under AI Act; API opt-in available worldwide
OpenAI will add an invisible watermark to text generated by ChatGPT and Codex in the European Union to comply with the EU AI Act's transparency rules. Developers anywhere can enable it on select API models starting today, but it is off by default. OpenAI's own tests show detection falling from about 92% to 66% after 10% of words are replaced with synonyms.
OpenAI to watermark ChatGPT and Codex text in the EU with textGrain; API watermarking is opt-in worldwide
OpenAI will switch on invisible text watermarks called textGrain for ChatGPT and Codex users in the EU over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory approach for Claude. OpenAI's own data shows detection drops sharply when text is edited.
Comments
Loading...