PyTorch has added a bounded, shape-aware NVFP4 autotuning policy to Inductor for Blackwell decode workloads, reducing the default set of GEMM candidates users must compile and profile while retaining selected tactics for measured shapes.

The change centralizes NVFP4 shape decisions in torch/_inductor/heuristics/template/nv_universal_gemm.py. That module identifies the NVFP4 1×16 recipe, determines when swap-AB is available, sets CUDA-graph replay and cold-weight behavior, applies the profiling cap, and returns primary and prefetch candidate keys. Code generation passes problem metadata to the module rather than maintaining independent shape cutoffs.

Ranked nvMatmulHeuristics results remain the primary pool, but successful rankings are capped by default at three candidates in both ordinary and epilogue-fusion-capable pools. Fallback, supplemental and wide-tile routes can add candidates. The policy deduplicates candidate keys, retains primary non-prefetch variants for unranked fallbacks, and excludes narrow-N tactics when N is dynamic; explicit controls can request fewer candidates or a broader fixed supplemental table.

For NVGEMM-only decode requests, the policy uses 16 CUDA-graph replays and rotates weight, scale and output buffers to avoid selecting against permanently L2-hot data. It adds measured transposed 32/64-wide tactics for logical M=32, N=4096 projections missed by the existing heuristics, while native M=128 projections with N at or below 16,384 can admit 128×64 and 128×128 prefetch variants. FP8, MXFP4 and BF16 keep their ordinary policy.

PyTorch reported that changing only the ranked-family cap from five to three preserved or improved selected-kernel latency across all 15 tested shapes on a later complete stack with final selective-PDL policy. Aggregate candidate time fell from 189.03 seconds to 164.93 seconds, a 12.75% reduction. A Qwen3-14B batch-32 run reduced NVGEMM tuning from 49.684 seconds to 40.738 seconds, or 18.0%.

The Qwen run’s reported end-to-end change was 0.10%, which PyTorch said was below unlocked-clock run spread. Focused validation covered recipe detection, replay and cold-cache boundaries, profiling caps, candidate and fallback behavior, dynamic N, MXFP4 exclusion and PDL-independent integration. PyTorch describes the change as having no user-facing backward-compatibility break.

Source: PyTorch release notes

Definition. NVFP4 autotuning is Inductor’s process for compiling and profiling candidate GEMM tactics to select kernels for NVFP4 workloads.

Reported measureBefore and after
Aggregate candidate time across 15 tested shapes189.03 seconds to 164.93 seconds (12.75% reduction)
Qwen3-14B batch-32 NVGEMM tuning49.684 seconds to 40.738 seconds (18.0% reduction)
Qwen3-14B batch-32 end-to-end change0.10%

Key takeaways

  • The shape-aware policy is centralized in Inductor’s NVFP4 universal GEMM heuristics module.
  • Successful ranked nvMatmulHeuristics results are capped at three candidates by default.
  • Fallback, supplemental, and wide-tile routes may add candidates beyond the ranked pool.
  • The policy excludes narrow-N tactics when N is dynamic.
  • PyTorch reported aggregate candidate time fell from 189.03 seconds to 164.93 seconds across 15 tested shapes.
  • A reported Qwen3-14B batch-32 run reduced NVGEMM tuning time from 49.684 to 40.738 seconds.

FAQ

What changed in PyTorch’s NVFP4 autotuning policy?

Inductor now uses a bounded, shape-aware policy that limits successful ranked GEMM candidates to three by default while retaining selected shape-specific and fallback tactics.

How much tuning-time reduction did PyTorch report?

PyTorch reported aggregate candidate time fell 12.75%, from 189.03 seconds to 164.93 seconds, across 15 tested shapes.

Does the policy apply to FP8, MXFP4, and BF16 workloads?

No. The post says FP8, MXFP4, and BF16 keep their ordinary policy.

Is there a user-facing backward-compatibility break?

PyTorch describes the change as having no user-facing backward-compatibility break.