AI Evaluation
A growing collection of articles on AI: big questions, useful tools, and ideas that stay with you.
Real ideas / Practical perspectives / A brighter next
PyTorch adds per-parameter mixed precision policy to FSDP2
The change supports mixed parameter dtypes in distributed training while retaining a single collective where dtypes allow it, according to PyTorch’s reported benchmarks.
1 min read
PyTorch folds NVFP4 output scaling into NVGEMM candidates for decode workloads
The Inductor change can remove a separate scale launch before QKV fan-out while retaining fallback and protected tensor-parallel paths.
2 min read
Microsoft open-sources RetroChimera for chemist-aligned synthesis planning
Expert reviewers accepted its routes for nine of ten challenging targets, though faster laboratory synthesis remains a projected benefit rather than a demonstrated outcome.
3 min read
TensorRT Edge-LLM Cuts Jetson Agentic Benchmark Run to 24 Minutes
NVIDIA attributes the 6.4x completion-time advantage over the published llama.cpp reference to NVFP4 quantization, cache reuse and tree-based multi-token prediction, though the two runs used different…
3 min read
Nemotron Post-Training Pipeline Reaches IMO 2026 Gold Threshold
The authors report a 30-of-42 score from a natural-language proof system and are releasing specialist checkpoints, code, data and a 200-problem benchmark.
2 min read