AWS launches AgentCore Code Interpreter to process documents beyond context window limits using recursive LLM architectu
Amazon Web Services released AgentCore Code Interpreter, a sandboxed Python environment that enables recursive language models to process documents of unlimited length by treating context as an external environment rather than loading it into the model's context window. The system orchestrates sub-LLM calls from within the sandbox, maintaining intermediate results as Python variables across a persistent session.
AWS launches AgentCore Code Interpreter to process documents beyond context window limits using recursive LLM architecture
Amazon Web Services released AgentCore Code Interpreter, a sandboxed Python runtime that implements recursive language models (RLMs) to analyze documents of unlimited length without context window constraints.
How it works
The system treats input documents as an external environment rather than loading them directly into a model's context window. A root LLM agent writes Python code to search and slice documents iteratively, delegating semantic analysis to sub-LLM calls that keep results in working memory as Python variables.
The architecture has three components:
- A root LLM agent built with the Strands Agents SDK that receives queries and generates code
- An AgentCore Code Interpreter session running in PUBLIC network mode with the full document loaded as a Python variable
- An
llm_query()function injected into the sandbox that calls Amazon Bedrock directly, keeping sub-LLM results in Python variables instead of the root LLM's context window
The persistent session state accumulates variables and intermediate results across multiple code executions, providing working memory throughout the analysis.
Technical specifications
According to AWS, the system:
- Supports documents of varying lengths with no upper bound on context size
- Maintains persistent state across executions in a sandboxed Python 3.10+ environment
- Requires IAM permissions for
bedrock:InvokeModel,bedrock-agentcore:StartCodeInterpreterSession,bedrock-agentcore:InvokeCodeInterpreter, andbedrock-agentcore:StopCodeInterpreterSession - Sets maximum session timeout at 3,600 seconds (1 hour)
- Uses PUBLIC network mode to enable outbound API calls to Amazon Bedrock
Evaluation results
AWS tested the system on the Financial Multi-Document QA subset of LongBench v2, a benchmark with 15 multiple-choice questions requiring analysis across multiple financial reports with context lengths up to approximately 2 million characters.
The company compared RLM against two baselines:
- Base approach: Sending the full document directly to the model with a 200K token context window
- Long Context approach: Using Claude's 1 million token context window
AWS measured success rate (percentage of questions processed without errors) and accuracy (percentage of correct answers). Specific numeric results were not disclosed in the announcement.
Implementation requirements
Developers need:
- AWS account with access to Amazon Bedrock foundation models
- Python 3.10 or later
- AWS CLI configured with appropriate credentials
- AgentCore Code Interpreter configured with PUBLIC network mode
- Strands Agents SDK for orchestration
The implementation involves starting a Code Interpreter session, loading documents into the sandbox, defining the llm_query() helper function, and creating a Strands Agent with an execute_python tool.
What this means
This approach addresses the fundamental limitation of fixed context windows by changing the interaction model between LLMs and long documents. Instead of expanding context windows—which remain bounded and suffer from attention degradation in long inputs—the recursive architecture treats documents as queryable environments. The practical impact is that document length becomes decoupled from model limitations, enabling analysis of arbitrarily long financial reports, legal documents, or technical specifications without preprocessing or chunking strategies. However, the system adds complexity through code generation and orchestration overhead, and the actual performance gains depend on the quality of the root LLM's code generation and the effectiveness of its document exploration strategy.
Related Articles
Anthropic launches Claude for Google Workspace add-on in public beta, adding sidebars to Docs, Sheets and Slides
Anthropic has released the Claude for Google Workspace add-on in public beta, placing a Claude sidebar inside Google Docs, Sheets, and Slides. It is available to all paid Claude users through the Google Workspace Marketplace, and includes an "ask before edits" preview mode.
AWS publishes reference build for voice airline concierge using Nova 2.5 Sonic and Bedrock AgentCore
AWS has published a reference architecture for a voice travel concierge that pairs Amazon Nova 2.5 Sonic with Bedrock AgentCore runtime, AgentCore Gateway over MCP, and Bedrock Knowledge Bases. It deploys with a single AWS CDK script and runs against a sample airline backend with synthetic data. The post does not disclose pricing, context window, or benchmark figures for Nova 2.5 Sonic.
Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
Google may expand Gemini 'Call for Me' to personal calls, per Android Authority APK teardown
Android Authority reports that an APK teardown shows a 'Gemini Calling' intro screen, suggesting Google may extend its 'Call for Me' AI feature from business calls to personal ones. Google has not announced the feature, and the report notes it may never ship publicly.
Comments
Loading...