Topics
Generative AI
A growing collection of articles on AI: big questions, useful tools, and ideas that stay with you.
Real ideas / Practical perspectives / A brighter next
ONNX Runtime 1.30 expands generative AI inference across CUDA, WebGPU and CPUs
The release broadens attention, decoding and quantization support, but several CUDA paths remain hardware-specific or opt-in and CPU FP16 execution now depends on acceleration.
Sep 10, 2026
4 min read
4 min read