The systems behind useful intelligence.

Compute, inference, data, observability, and the technical foundations that make reliable AI possible.

Compute  /  Reliability  /  Scale

A physical scene representing ordered language-token components and merging pieces.
AI Infrastructure1 min read

Ollama v0.40.0-rc1 updates MLX tokenizer handling to match publisher semantics

The release candidate adds tokenizer compatibility fixes and shared reference cases intended to catch differences between Ollama’s MLX path and published tokenizers.

Oct 04, 2026
1 min read
A physical high-performance computing assembly with interconnected accelerator modules and translucent streams converging into a central processing chamber.
AI Infrastructure1 min read

PyTorch adds per-parameter mixed precision policy to FSDP2

The change supports mixed parameter dtypes in distributed training while retaining a single collective where dtypes allow it, according to PyTorch’s reported benchmarks.

Oct 03, 2026
1 min read
Physical GPU computing hardware and interconnected components representing optimized high-performance inference tuning.
AI Infrastructure2 min read

PyTorch narrows NVFP4 autotuning for Blackwell decode workloads

The Inductor update keeps shape-specific GEMM tactics while cutting the default ranked candidate pool, which PyTorch says reduced tuning time in its tests.

Oct 03, 2026
2 min read
Accelerator hardware with a central compute die and illuminated pathways splitting into three routed channels.
AI Infrastructure2 min read

PyTorch folds NVFP4 output scaling into NVGEMM candidates for decode workloads

The Inductor change can remove a separate scale launch before QKV fan-out while retaining fallback and protected tensor-parallel paths.

Oct 03, 2026
2 min read
An open server chassis with eight interconnected GPU accelerator modules beside a robotic arm cleaning a ceramic plate with a sponge in a server room.
AI Infrastructure3 min read

NVIDIA Dynamo-Triton Adds Multi-GPU TensorRT Serving Through One Model Endpoint

NVIDIA’s eight-GPU Cosmos 3 Nano test cut mean generation latency from 156.595 seconds to 34.183 seconds, but it measured neither throughput nor deployment economics.

Sep 21, 2026
3 min read
A hiker and a dog stand on a rocky mountain overlook above a winding lake at sunrise.
AI Infrastructure5 min read

xAI’s Grok 4.6 reaches Amazon Bedrock with two deployment paths

Developers gain a 500K-token model with Converse and cross-Region inference, but API features, residency options and pricing differ substantially by endpoint.

Sep 21, 2026
5 min read
A compact edge-computing unit beside a rugged autonomous robot, with illuminated branching paths connecting small processor-like modules on an outdoor worktable.
AI Infrastructure3 min read

TensorRT Edge-LLM Cuts Jetson Agentic Benchmark Run to 24 Minutes

NVIDIA attributes the 6.4x completion-time advantage over the published llama.cpp reference to NVFP4 quantization, cache reuse and tree-based multi-token prediction, though the two runs used different…

Sep 16, 2026
3 min read
A worker feeds bundled documents into a central caching machine that distributes the stored material through glowing conduits to three readers.
AI Infrastructure6 min read

AWS documents six Amazon Bedrock prompt-caching patterns for lower inference costs

Cached input can cost up to 90% less on a hit, but write premiums, minimum token thresholds and expiration determine the savings developers actually realize.

Sep 15, 2026
6 min read