The systems behind useful intelligence.
Compute / Reliability / Scale
Ollama v0.40.0-rc1 updates MLX tokenizer handling to match publisher semantics
The release candidate adds tokenizer compatibility fixes and shared reference cases intended to catch differences between Ollama’s MLX path and published tokenizers.
1 min read
PyTorch adds per-parameter mixed precision policy to FSDP2
The change supports mixed parameter dtypes in distributed training while retaining a single collective where dtypes allow it, according to PyTorch’s reported benchmarks.
1 min read
PyTorch narrows NVFP4 autotuning for Blackwell decode workloads
The Inductor update keeps shape-specific GEMM tactics while cutting the default ranked candidate pool, which PyTorch says reduced tuning time in its tests.
2 min read
PyTorch folds NVFP4 output scaling into NVGEMM candidates for decode workloads
The Inductor change can remove a separate scale launch before QKV fan-out while retaining fallback and protected tensor-parallel paths.
2 min read
NVIDIA Dynamo-Triton Adds Multi-GPU TensorRT Serving Through One Model Endpoint
NVIDIA’s eight-GPU Cosmos 3 Nano test cut mean generation latency from 156.595 seconds to 34.183 seconds, but it measured neither throughput nor deployment economics.
3 min read
xAI’s Grok 4.6 reaches Amazon Bedrock with two deployment paths
Developers gain a 500K-token model with Converse and cross-Region inference, but API features, residency options and pricing differ substantially by endpoint.
5 min read
TensorRT Edge-LLM Cuts Jetson Agentic Benchmark Run to 24 Minutes
NVIDIA attributes the 6.4x completion-time advantage over the published llama.cpp reference to NVFP4 quantization, cache reuse and tree-based multi-token prediction, though the two runs used different…
3 min read
AWS documents six Amazon Bedrock prompt-caching patterns for lower inference costs
Cached input can cost up to 90% less on a hit, but write premiums, minimum token thresholds and expiration determine the savings developers actually realize.
6 min read