Hybrid Mamba-Transformer MoEs Hide Their Stalls in Places Dashboards Do Not Look

A trace of a hybrid Mamba-Transformer MoE inference run, broken down by layer type. The MoE all-to-all collective stalls dominate the tail. The dashboards saw 96% GPU utilization the entire window.

TL;DR

Hybrid Mamba-Transformer architectures (Nemotron 3 Nano Omni, Jamba and friends) shipped at speed in late April. These models break the assumptions vLLM and SGLang dashboards make about prefill/decode shape: Mamba state-space layers have one runtime profile, Transformer attention has another, MoE router blocks have a third (with all-to-all collective comm). The aggregate looks fine on a duty-cycle counter; the per-layer tail is full of hybrid MoE stalls nobody is decomposing. We trace one and decompose it.

What changed in late April

NVIDIA Nemotron 3 Nano Omni (Apr 28, open multimodal MoE) is the most prominent recent shipment, but it is one of several. The shape is consistent: a hybrid Mamba-Transformer backbone with mixture-of-experts routing, tuned to claim higher throughput than pure-Transformer baselines at comparable parameter counts.

A trace of a hybrid Mamba-Transformer MoE inference run, broken down by layer type. The MoE all-to-all collective stalls dominate the tail. The dashboards saw 96% GPU utilization the entire window.

TL;DR

What changed in late April

Hybrid Mamba-Transformer MoEs Hide Their Stalls in Places Dashboards Do Not Look

Hybrid Mamba-Transformer MoEs Hide Their Stalls in Places Dashboards Do Not Look

Other newsrooms on this story

Related reading

KTransformers的5个隐藏用法，17K Star的MoE推理框架背后没写在README里的能力

KTransformers: 5 Hidden Uses of the 17K-Star MoE Inference Stack from Tsinghua…

Nvidia's Nemotron 3 swaps pure Transformers for a Mamba hybrid to run AI agents…

NVIDIA AI Releases Nemotron 3 Ultra: An Open 550B Mixture-of-Experts Hybrid…

Running Mixtral 8x7B at 21+ TPS on Pure CPU via io_uring and Predictive Caching

Mixture of Experts (MoE): what it actually does under the hood, and when it…

Other newsrooms on this story

Related reading

KTransformers的5个隐藏用法，17K Star的MoE推理框架背后没写在README里的能力

KTransformers: 5 Hidden Uses of the 17K-Star MoE Inference Stack from Tsinghua…

Nvidia's Nemotron 3 swaps pure Transformers for a Mamba hybrid to run AI agents…

NVIDIA AI Releases Nemotron 3 Ultra: An Open 550B Mixture-of-Experts Hybrid…

Running Mixtral 8x7B at 21+ TPS on Pure CPU via io_uring and Predictive Caching

Mixture of Experts (MoE): what it actually does under the hood, and when it…