AWS expanded Amazon Bedrock, Amazon Bedrock AgentCore and Strands in August 2026, giving AI builders million-token model context, cross-Region inference, agent sessions lasting up to 14 days and workflows that extend to physical robots. The changes let developers tackle larger and longer-running workloads while applying geographic, behavioral, access and spending controls.
More context and control in Bedrock
OpenAI’s GPT-5.6 Sol, Terra and Luna now support million-token context windows on Amazon Bedrock, alongside prompt caching intended to reduce latency and cost when applications reuse context. Bedrock Web Search can retrieve current information beyond model training data and return citations, while direct retrieval from public websites supports cases involving changing material such as live prices or newly published documentation. AWS says the combination can analyze a large working set, compare it with current public information and produce a cited answer through one API call.
Cross-Region inference makes the three GPT-5.6 models accessible from more than 25 AWS Regions. Global profiles prioritize broad capacity and lower per-token prices, while Geo profiles constrain inference processing to a selected geography. The distinction gives teams a way to balance throughput during demand spikes against data-processing boundaries.
AWS also added IAM-principal cost allocation for assigning inference spending to users, teams, projects, applications or cost centers. Cost Anomaly Detection now monitors third-party foundation-model spending in Bedrock and supplies root-cause breakdowns for unexpected changes. OpenAI separately announced lower Bedrock pricing for Sol, Terra and Luna, although the recap does not state the revised rates.
Long-running agents with policy and budget boundaries
AgentCore runtime instances run agents on dedicated Amazon EC2 compute, with GPU-accelerated, memory-optimized and compute-optimized configurations and sessions of up to 14 days. The service also expanded to US West (N. California) and Asia Pacific (Hyderabad), supporting long-running research, coding and monitoring closer to relevant users or systems.
New temporal policies assess an action in light of an agent’s previous actions. They can enforce ordering, prerequisites, approval gates, consistent values between calls and freshness requirements. Rate limits can govern requests, inference tokens and concurrent connections by user or group. AgentCore payments adds controlled access to paid APIs, MCP resources and content, backed by infrastructure-enforced spending limits and end-to-end observability.
AgentCore Web Search can restrict or exclude domains and filter results by publication date. Its memory system can derive long-term memories from structured JSON sources—including activity, behavioral and system events—rather than relying only on conversations. Fine-grained access controls isolate memories by user or tenant. Separately, AWS Agent Registry provides a governed, searchable catalog for agents, MCP servers, skills and custom resources, with organization-wide discovery across connected AWS accounts.
Security and regulated workloads
OpenAI’s Daybreak Blue and Daybreak Red are available to eligible Bedrock customers. AWS positions Blue for defensive tasks including vulnerability discovery, detection engineering and incident response, while Red targets authorized vulnerability research, exploit reproduction and mitigation development. Access is paired with identity verification, monitoring, access controls and zero-operator-access infrastructure.
AWS GovCloud (US) Regions gained Claude Opus 5 with zero data retention enabled by default, plus GPT-5.6 Terra and Luna with million-token context and prompt caching. Amazon Nova Multimodal Embeddings supports retrieval across text, documents, images, video and audio in AWS GovCloud (US-West). AgentCore memory, policy and its managed harness also bring managed context, controls and orchestration to GovCloud workloads.
From demonstrations to physical robots
Strands Robots connects Strands Agents, LeRobot and Hugging Face Storage Buckets in a workflow for recording demonstrations, streaming datasets in the LeRobot format, training policies and deploying them to simulated or physical hardware without converting the underlying data. It also supports mesh-based discovery and coordination across devices: Zenoh connects robots on a local network, while AWS IoT Core supports geographically distributed fleets.
Strands Robots and AWS are participating in a limited research preview of the Model Hardware Standard. That preview status is an important qualification: the source supports experimentation with standardized physical-equipment controls, not general availability of the standard.
Source: AWS Machine Learning Blog.
Definition. The August 2026 AWS AI update is a set of Bedrock, AgentCore and Strands capabilities for larger-context models, long-running governed agents, regulated workloads and robotics deployment.
| Inference profile | Primary tradeoff |
|---|---|
| Global profile | Prioritizes broad capacity and lower per-token prices. |
| Geo profile | Constrains inference processing to a selected geography. |
Key takeaways
- GPT-5.6 Sol, Terra and Luna support million-token context windows and prompt caching on Amazon Bedrock.
- Cross-Region inference spans more than 25 AWS Regions, with Global profiles for capacity and Geo profiles for geographic processing constraints.
- AgentCore runtime instances support dedicated EC2 configurations and sessions lasting up to 14 days.
- Temporal policies, rate limits and infrastructure-enforced spending limits add behavioral and budget boundaries for agents.
- GovCloud additions include Claude Opus 5, GPT-5.6 Terra and Luna, multimodal embeddings and AgentCore capabilities.
- Strands Robots supports a workflow from recorded demonstrations and policy training to simulated or physical deployment.
FAQ
Which OpenAI models gained million-token context on Amazon Bedrock?
GPT-5.6 Sol, Terra and Luna gained million-token context windows alongside prompt caching.
How long can AgentCore sessions run?
AgentCore runtime instances support sessions lasting up to 14 days.
What is the difference between Global and Geo inference profiles?
Global profiles prioritize broad capacity and lower per-token prices, while Geo profiles constrain inference processing to a selected geography.
What controls were added for long-running agents?
AWS added temporal policies, request and token rate limits, concurrent-connection controls, fine-grained memory access and infrastructure-enforced spending limits.
What did AWS add for regulated workloads?
AWS GovCloud gained additional models, million-token context and prompt caching, multimodal embeddings, and AgentCore memory, policy and managed orchestration capabilities.
Is the Model Hardware Standard generally available?
No. Strands Robots and AWS are participating in a limited research preview, which supports experimentation rather than general availability.