Moe inference optimizations: 15% lower expert load by request reordering(blog.doubleword.ai)3 points by mezark 106 days ago | 0 commentsNo comments yet