Researchers Radha Gulhane, Quentin Anthony and Beren Millidge have published a paper describing correlated expert placement and token shuffling for mixture-of-experts pretraining. For teams running expert-parallel MoE training across GPUs, the techniques aim to keep more token-to-expert work local and reduce the all-to-all communication that can constrain training steps.

Why communication is the target

In an MoE layer, tokens are routed to selected expert networks while expert parallelism distributes those experts across GPUs. Each layer therefore uses all-to-all collectives to dispatch tokens and combine results in both the forward and backward passes. The authors report that, on a cluster with eight AMD Instinct MI300X GPUs per node, those collectives accounted for 45% of a training step at expert-parallel degree 32 with top-2 routing and 60% with top-6 routing.

Correlations in router choices

The paper says routers develop correlated expert-selection patterns early in pretraining, both within a layer and between successive layers. At top-2 routing, 0.8% of expert pairs were jointly selected by 42% of tokens in the authors’ analysis. The experts selected for a token in one layer also predicted its selections in the next layer, providing the signal used by the two placement methods.

Two locality techniques

Correlated expert placement groups experts that are often selected together on the same GPU. Paired with a dispatcher that sends each token to a GPU once, the authors report that it removes up to 58% of dispatched rows.

Token shuffling applies when sequence parallelism shards tokens across the expert-parallel group. During the reduce-scatter after attention, it moves a token to the GPU predicted to contain its next-layer experts. In a one-node result, the share of token-to-expert assignments served on the token’s own GPU rose from 12.5% to 59%.

Reported evaluation and scope

In Megatron-LM tests spanning expert-parallel degrees 8 through 64 and top-2 and top-6 routing, the authors report all-to-all time reductions of 1.16–2.63X and end-to-end step-time improvement of up to 1.41X. These are experimental results for the stated configurations, not evidence that every MoE training deployment will see the same gains. The methods alter expert and token placement rather than the models’ underlying routing decisions or expert parameters.

Source: arXiv:2610.09372

Definition. Correlated expert placement groups experts frequently selected together to keep more MoE token-to-expert work local to a GPU.

TechniqueReported effect
Correlated expert placementGroups frequently co-selected experts on the same GPU; paired with a single-send dispatcher, it reportedly removed up to 58% of dispatched rows.
Token shufflingMoves tokens toward predicted next-layer experts; local assignments rose from 12.5% to 59% in a one-node result.

Key takeaways

  • All-to-all collectives accounted for 45% of a training step with top-2 routing and 60% with top-6 routing in the reported expert-parallel degree 32 setup.
  • Correlated expert placement groups experts often selected together on the same GPU.
  • With a dispatcher that sends each token to a GPU once, correlated expert placement reportedly removed up to 58% of dispatched rows.
  • In a one-node token-shuffling result, local token-to-expert assignments rose from 12.5% to 59%.
  • The methods change expert and token placement, not routing decisions or expert parameters.

FAQ

What problem do the techniques address?

They target all-to-all communication used to dispatch tokens and combine results in expert-parallel MoE layers.

What is correlated expert placement?

It places experts that routers often select together on the same GPU, aiming to make more work local.

What does token shuffling do?

During the reduce-scatter after attention, it moves a token to the GPU predicted to contain its next-layer experts.

Do these methods change MoE routing decisions?

No. The post states that they alter expert and token placement rather than underlying routing decisions or expert parameters.

Are the reported gains guaranteed for every MoE deployment?

No. The results are experimental findings for the stated Megatron-LM configurations.

Sources