The systems behind useful intelligence.

Compute, inference, data, observability, and the technical foundations that make reliable AI possible.

Compute  /  Reliability  /  Scale

A soccer player takes a penalty kick before a goalkeeper, framed by film reels, strips of match footage, photographs, and glowing audio-wave forms in a stadium archive.
AI Infrastructure3 min read

Marengo Embed 3.0 brings managed multimodal search to Amazon Bedrock Knowledge Bases

AWS customers can search video, audio and images by meaning, while paying separately for storage, retrieval and Marengo embedding generation.

Sep 10, 2026
3 min read
Server racks with GPU accelerators and illuminated local storage arrays beside a distant data center connected by glowing network paths.
AI Infrastructure4 min read

AWS adds model caching to cut SageMaker HyperPod inference cold starts

The generally available feature can make cached pods available in seconds, although the first download, per-node storage cost and stale-cache risks remain.

Sep 10, 2026
4 min read
Glowing braided data conduits pass through a routing junction and branch toward illuminated memory modules in several physical server racks inside a data center.
AI Infrastructure4 min read

Amazon SageMaker adds prefix-aware routing to cut LLM response latency

The largest AWS-reported gains came from long-context workloads with substantial shared prefixes, while effective cache reuse still depends on workload shape and prefix caching in the…

Sep 10, 2026
4 min read
Robotics engineer working beside an industrial robot arm in a facility with autonomous ground robots, a quadruped robot, and a drone.
AI Infrastructure3 min read

AWS expands Bedrock and AgentCore with million-token context, 14-day agent sessions

The August update combines longer-context OpenAI models, geographically controlled inference, persistent agent infrastructure, GovCloud expansion and a path from robot training to physical deployment.

Sep 10, 2026
3 min read
Open computer hardware with CPU cooling, memory modules, graphics accelerators, and connected processing units arranged on a workbench overlooking mountains at sunset.
AI Infrastructure4 min read

ONNX Runtime 1.30 expands generative AI inference across CUDA, WebGPU and CPUs

The release broadens attention, decoding and quantization support, but several CUDA paths remain hardware-specific or opt-in and CPU FP16 execution now depends on acceleration.

Sep 10, 2026
4 min read
AI-generated editorial illustration: An industrial data center at dusk, a robotic arm carefully rearranging plain unmarked server modules.
AI Infrastructure10 min read

HyperPod InstantStart Turns Complex SageMaker Operations Into Guarded Agent Workflows

The open-source control plane gives infrastructure teams a web interface, REST APIs and an AI agent backed by the same validation, reconciliation and persisted state for…

Sep 06, 2026
10 min read
AI-generated editorial illustration: Macro photography of a silver processor with many colored light paths across a dark silicon surface.
AI Infrastructure9 min read

SGLang v0.5.19 expands model support, inference performance and hardware reach

The release combines 786 pull requests from 214 contributors, adding nine model entries, beam search, DeepEP v2, broader speculative decoding, unified caching, diffusion improvements and extensive…

Sep 06, 2026
9 min read
AI-generated editorial illustration: A robotic hand learning to grasp a red cube in an industrial studio, with a mirrored physical twin.
AI Infrastructure13 min read

How NVIDIA Cosmos 3 and SageMaker HyperPod Power a Physical AI Model Factory

AWS outlines a persistent, shared GPU architecture for synthetic data generation, distributed post-training, and closed-loop evaluation, with end-to-end GPU goodput—not isolated job throughput—as the central operating…

Sep 06, 2026
13 min read