MT-Bench prompts answered by Qwen3.5 and Qwen3.6 (35B-A3B), with every MoE routing decision captured: which of the 256 experts fire for each token in each of the 40 layers, and how sure the router is.
1 comment
gslaller2 days ago
Did some work on visualising the firing pattern of MoE (Mixture of Experts) modules for Qwen models and other statistics. You'll see temporal consistency in the router, as one would expect.
Read the full thread on Hacker News →
Related stories
- Static Analysis vs. Taint Analysis: Which One Secures Your Python Code?nocomplexity.substack.comHacker News · 2 points · 1 day ago
- Binary Mutation Analysis of Tests Using Reassembleable Disassemblywww-users.cs.umn.eduLobsters · 5 points · over 6 years ago
- Hacker News · 2 points · 5 days ago
- Ars Technica · 0 points · 12 days ago
- Qwen Image 2.1qwen.aiHacker News · 728 points · 11 days ago
- Hacker News · 2 points · 11 days ago