OMP-MoE prunes MoE LLM experts without training, keeping 93.3%

Mixture-of-Experts (MoE) models let large language models scale efficiently, but they need a lot of memory to deploy. The paper's authors say existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. Their answer is OMP-MoE, a training-free compression framework that reduces expert redundancy in MoE-based LLMs.

The method has three parts. First, based on observations of how experts contribute, the authors recast pruning as a sparse signal reconstruction task and solve it with Orthogonal Matching Pursuit. Each expert's contribution is treated as a dictionary atom, and the method greedily selects the experts that minimise reconstruction error, with linear computational complexity. Second, the number of experts kept in each layer is set by a water-filling strategy that accounts for both reconstruction quality and routing stability. Third, the authors introduce OMP-MoE-dagger, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction.

The experiments cover Qwen, DeepSeek-V2, GPT-OSS and Mixtral MoE models. The authors report consistent improvements over existing methods at pruning ratios of 25-50%. The headline result is for Qwen3-30B-A3B at 50% compression: the pruned model retains 93.3% of original performance, with 33x faster search and a 1.55x inference speedup. The baselines behind the search and inference speedups are not named in the source, and the 93.3% figure is given only for this one model and setting. The authors say code will be available after acceptance.

Key facts

  • OMP-MoE is a training-free framework for pruning redundant experts from Mixture-of-Experts LLMs, casting pruning as sparse signal reconstruction solved with Orthogonal Matching Pursuit.
  • Experts are picked greedily to minimise reconstruction error with linear computational complexity; per-layer allocation uses a water-filling strategy that weighs reconstruction quality and routing stability.
  • OMP-MoE-dagger is an adaptive inference mechanism that adjusts expert activation based on energy prediction.
  • On Qwen3-30B-A3B at 50% compression, the authors report 93.3% of original performance, 33x faster search and 1.55x inference speedup.
  • Tests span Qwen, DeepSeek-V2, GPT-OSS and Mixtral at 25-50% pruning ratios; code is promised after acceptance.

Why it matters

MoE models scale well, but their memory needs make them hard to deploy. Pruning experts is one way to cut that cost, and the authors argue current methods are either too expensive to search or ignore how experts depend on each other. A training-free method that also searches much faster (33x, against baselines the source does not name) would make pruning cheaper to try on large models. The reported gains are practical rather than a change in approach to MoE design.

Who it affects

Teams that deploy or compress MoE-based LLMs are the direct audience. The tested families are Qwen, DeepSeek-V2, GPT-OSS and Mixtral, so users of those models are the most relevant. Researchers working on model compression and expert routing will also find the reformulation of pruning as sparse reconstruction of interest.

How to use it

Not yet, in practice. The authors say code will be available after acceptance, and no repository exists yet. Release is conditional on acceptance and no venue is named. Until then, the paper's description is the only guide: select experts greedily by reconstruction error, allocate experts across layers with water-filling, and optionally use the adaptive inference variant.

How solid is it

The claims come from the authors' own abstract, and the source is the paper's summary text. They report consistent improvements over existing methods across four model families at 25-50% pruning. Many details are not stated: the baseline methods are not named, the benchmarks and metrics behind 'original performance' are not named, and the 93.3% figure is given only for Qwen3-30B-A3B at 50% compression. Results for other models are not quantified.

Risks and caveats

The baselines for the 33x search speedup and the 1.55x inference speedup are not stated. It is also not stated whether the 1.55x speedup comes from OMP-MoE itself or from the adaptive inference variant. At 50% compression the pruned model keeps 93.3% of original performance, so roughly 6.7% is lost on the authors' measure. Without released code, none of this can be reproduced yet.

“Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts.”

— OMP-MoE paper abstract