The systems
behind useful
intelligence.
Compute, inference, data, observability, and the technical foundations that make reliable AI possible.
Compute / Reliability / Scale
NVIDIA Dynamo-Triton Adds Multi-GPU TensorRT Serving Through One Model Endpoint
NVIDIA’s eight-GPU Cosmos 3 Nano test cut mean generation latency from 156.595 seconds to 34.183 seconds, but it measured neither throughput nor deployment economics.
3 min read
xAI’s Grok 4.6 reaches Amazon Bedrock with two deployment paths
Developers gain a 500K-token model with Converse and cross-Region inference, but API features, residency options and pricing differ substantially by endpoint.
5 min read
TensorRT Edge-LLM Cuts Jetson Agentic Benchmark Run to 24 Minutes
NVIDIA attributes the 6.4x completion-time advantage over the published llama.cpp reference to NVFP4 quantization, cache reuse and tree-based multi-token prediction, though the two runs used different…
3 min read
AWS documents six Amazon Bedrock prompt-caching patterns for lower inference costs
Cached input can cost up to 90% less on a hit, but write premiums, minimum token thresholds and expiration determine the savings developers actually realize.
6 min read
Marengo Embed 3.0 brings managed multimodal search to Amazon Bedrock Knowledge Bases
AWS customers can search video, audio and images by meaning, while paying separately for storage, retrieval and Marengo embedding generation.
3 min read
AWS adds model caching to cut SageMaker HyperPod inference cold starts
The generally available feature can make cached pods available in seconds, although the first download, per-node storage cost and stale-cache risks remain.
4 min read
Amazon SageMaker adds prefix-aware routing to cut LLM response latency
The largest AWS-reported gains came from long-context workloads with substantial shared prefixes, while effective cache reuse still depends on workload shape and prefix caching in the…
4 min read
AWS expands Bedrock and AgentCore with million-token context, 14-day agent sessions
The August update combines longer-context OpenAI models, geographically controlled inference, persistent agent infrastructure, GovCloud expansion and a path from robot training to physical deployment.
3 min read