moe
5 stories and discussions about moe, aggregated from every source we track.
2.
A worklog documenting the journey of scaling expert parallelism to achieve high-throughput pretraining.
3.
Which experts fire, and how sure is the router? Token-by-layer routing maps for Qwen3.5 / Qwen3.6 MoE models.
4.
Fork off LLama cpp that newly scales GPU 2x 4x etc also on big models not fully fitting in vram - neurall/llama.cpp
5.
Xing4.0-29B-A4B is China Telecom's new 29B MoE model trained on Ascend NPUs, benchmarked against Gemma4-26B-A4B and Qwen3.6-35B-A3B.