Mixture-of-Kittens: Cursor’s Open-Source Kernel Trains AI Up to 2.37x Faster

Cursor open-sourced Mixture-of-Kittens, a kernel that fuses all the work of training mixture-of-experts models into a single operation. On NVIDIA NVL72 racks it trains up to 2.37x faster per layer and 41% faster end to end.
artificial-intelligence
software-engineering
Author

Kabui, Charles

Published

2026-08-26

Keywords

mixture-of-kittens, moe-training, megakernel, nvidia-nvl72, training-efficiency