The systems
behind useful
intelligence.
Compute, inference, data, observability, and the technical foundations that make reliable AI possible.
Compute / Reliability / Scale
ONNX Runtime 1.30 expands generative AI inference across CUDA, WebGPU and CPUs
The release broadens attention, decoding and quantization support, but several CUDA paths remain hardware-specific or opt-in and CPU FP16 execution now depends on acceleration.
4 min read
HyperPod InstantStart Turns Complex SageMaker Operations Into Guarded Agent Workflows
The open-source control plane gives infrastructure teams a web interface, REST APIs and an AI agent backed by the same validation, reconciliation and persisted state for…
10 min read
SGLang v0.5.19 expands model support, inference performance and hardware reach
The release combines 786 pull requests from 214 contributors, adding nine model entries, beam search, DeepEP v2, broader speculative decoding, unified caching, diffusion improvements and extensive…
9 min read
How NVIDIA Cosmos 3 and SageMaker HyperPod Power a Physical AI Model Factory
AWS outlines a persistent, shared GPU architecture for synthetic data generation, distributed post-training, and closed-loop evaluation, with end-to-end GPU goodput—not isolated job throughput—as the central operating…
13 min read
OpenAI GPT-5.6 reaches Australian Amazon Bedrock Regions through global inference
AWS has opened access to GPT-5.6 Sol, Terra, and Luna through Amazon Bedrock endpoints in Sydney and Melbourne, with three API paths, prompt caching, federated Codex…
8 min read
Hugging Face Releases 207 WebGPU Kernels for Local AI in the Browser
The new @huggingface/kernels JavaScript library loads versioned GPU operations from the Hugging Face Hub, while Fleet gathers browser-based performance and correctness evidence across real-world hardware.
7 min read