Back to papers
March 19, 2026cs.LGAdvanced

DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge

AI-Generated Summary

This paper presents DyMoE, a technique to make large AI models with multiple expert components run efficiently on edge devices (like mobile phones or IoT devices) by intelligently compressing less important experts while keeping critical ones intact. The method uses dynamic compression strategies that adapt based on which experts matter most and where in the model they're located, achieving 3-22x faster inference speeds compared to existing approaches while maintaining accuracy.

Difficulty
Advanced
Categories

cs.LG

AI Tags
model compressionquantizationmixture of expertsedge inferenceoptimizationefficient AI