The systems behind useful intelligence.
Compute / Reliability / Scale
Marengo Embed 3.0 brings managed multimodal search to Amazon Bedrock Knowledge Bases
AWS customers can search video, audio and images by meaning, while paying separately for storage, retrieval and Marengo embedding generation.
3 min read
AWS adds model caching to cut SageMaker HyperPod inference cold starts
The generally available feature can make cached pods available in seconds, although the first download, per-node storage cost and stale-cache risks remain.
4 min read
Amazon SageMaker adds prefix-aware routing to cut LLM response latency
The largest AWS-reported gains came from long-context workloads with substantial shared prefixes, while effective cache reuse still depends on workload shape and prefix caching in the…
4 min read
AWS expands Bedrock and AgentCore with million-token context, 14-day agent sessions
The August update combines longer-context OpenAI models, geographically controlled inference, persistent agent infrastructure, GovCloud expansion and a path from robot training to physical deployment.
3 min read
ONNX Runtime 1.30 expands generative AI inference across CUDA, WebGPU and CPUs
The release broadens attention, decoding and quantization support, but several CUDA paths remain hardware-specific or opt-in and CPU FP16 execution now depends on acceleration.
4 min read
HyperPod InstantStart Turns Complex SageMaker Operations Into Guarded Agent Workflows
The open-source control plane gives infrastructure teams a web interface, REST APIs and an AI agent backed by the same validation, reconciliation and persisted state for…
10 min read
SGLang v0.5.19 expands model support, inference performance and hardware reach
The release combines 786 pull requests from 214 contributors, adding nine model entries, beam search, DeepEP v2, broader speculative decoding, unified caching, diffusion improvements and extensive…
9 min read
How NVIDIA Cosmos 3 and SageMaker HyperPod Power a Physical AI Model Factory
AWS outlines a persistent, shared GPU architecture for synthetic data generation, distributed post-training, and closed-loop evaluation, with end-to-end GPU goodput—not isolated job throughput—as the central operating…
13 min read