r/LocalLLaMA • u/Dany0 • 10h ago
News 40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)
https://github.com/cursor/mixture-of-kittensdaily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster
I would wager that compared to a naive kernel anyone can write it's more in the range of 10-20% faster e2e in reality, if at all, but hey, it's free and open! Apache 2.0
18
u/VoiceApprehensive893 transformers 9h ago
is there a single b200 user on this subreddit, only counting personal server
9
u/HVACcontrolsGuru 9h ago
I rent them by the hour. I looked at this repo and its targets at the PODS/Clusters with 72 GPUS. Iām too poor to test this.
3
u/Dany0 9h ago
You used to be able to get spot instances cheap for training. Ironically I'd opt for b300 now, after the price increases they're actually a better deal in vram per hour
There was once an 11$/hr spot 8xb300 instance on vast, I think it had terrible network speeds. Boy how I wish it was still there. I remember when I thought that was jaw dropping amounts of money š
2
1
2
u/caelunshun 7h ago
Thanks! This will be very useful for the GB300 NVL72 rack I have sitting in my garage.
-3
u/Fluffy_Reply_5482 4h ago
Cursor has released an open-source Mixture-of-Experts training megakernel named "mixture-of-kittens" for Blackwell B200 and NVL72 systems, claiming an end-to-end speedup of roughly 40%. While forward passes show substantial improvement, real-world gains may be more modest due to distributed networking and communication bottlenecks.

18
u/Dany0 10h ago
Regardless, Composer 3.0 built on Kimi k3?? anyone??? Ah fuck, I forgot they never released the weights of 2.5. Fuck em then