MoE router telemetry leaks whether data was used in fine-tuning
Mixture-of-Experts (MoE) language models produce routing information during inference: which experts handled which input. That information may be logged or exposed for monitoring, debugging, load analysis and safety auditing. The paper asks a privacy question about it. Unlike ordinary model outputs, this telemetry gives a view of the model's internal computation, so can it reveal whether a particular example was used to fine-tune the deployed model?
The authors introduce a router-augmented membership inference attack. It combines conventional output-side signals with aggregated routing features. A membership classifier is learned from independently fine-tuned shadow models and then applied to the target model.
The evaluation covers three MoE architectures and three data domains, nine settings in all. In every one of them, router telemetry improves membership inference over a strong ensemble of output signals. The gain in true positive rate (TPR) at a 1% false positive rate (FPR) ranges from 2.7 to 9.4 percentage points.
The leakage is not tied to one training recipe. It persists across full fine-tuning, frozen-router training, LoRA and instruction tuning. It also remains observable when the attacker has only discrete expert selections, restricted telemetry, or a single shadow model.
A mechanistic analysis explains why. The leakage does not require router-specific memorization. Fine-tuning introduces membership information into the model's hidden representations, and the router exposes a projection of that signal even when its parameters are frozen. Perturbing the telemetry reduces the extra leakage only as its fidelity degrades. The authors conclude that router telemetry can turn an operational signal into an additional privacy surface for fine-tuned MoE models.
Key facts
- The paper presents a router-augmented membership inference attack on fine-tuned Mixture-of-Experts models, combining output-side signals with aggregated routing features.
- Across three MoE architectures and three data domains, adding router telemetry raises TPR at 1% FPR by 2.7 to 9.4 percentage points over a strong output-signal ensemble, in all nine settings.
- The leakage persists under full fine-tuning, frozen-router training, LoRA and instruction tuning, and with only discrete expert selections, restricted telemetry or a single shadow model.
- The mechanism: fine-tuning writes membership information into hidden representations, and the router exposes a projection of it even with frozen router parameters.
- Perturbing the telemetry cuts the extra leakage only as its fidelity degrades.
Why it matters
Routing data is usually treated as an operational detail. Expert selections may be logged or exposed for monitoring, debugging, load analysis and safety auditing. This paper argues that the same signal is also a privacy channel: it can help an observer tell whether a specific example was part of a model's fine-tuning data, beyond what the model's outputs already give away. The extra leakage is measured against a strong output-signal baseline, not a weak one, and it shows up in every one of the nine tested settings.
Who it affects
The paper's concern is fine-tuned MoE language models whose routing information is logged or exposed. That points to teams that fine-tune MoE models on data they want to keep private and also collect or share expert-selection telemetry for monitoring, debugging, load analysis or safety auditing. The source does not say that any specific real-world deployment or provider is vulnerable or exposes such telemetry.
How to use it
This is an attack and analysis paper, not a tool. The practical reading for anyone running a fine-tuned MoE model is that routing logs are not neutral: they carry some information about the fine-tuning data. The one mitigation the paper tests is perturbing the telemetry, and it helps only as the telemetry's fidelity degrades, so protection comes at the cost of usefulness for monitoring. The source does not recommend a specific defense beyond that, and it does not mention a code or dataset release.
How solid is it
The evidence is a single arXiv preprint, and the claims here rest on its abstract. The headline result is consistent across all nine settings (three architectures by three data domains), and the authors test several training regimes and restricted-access conditions, which makes the finding more than a one-off. Against that, the source does not name the three architectures or the three domains, and it gives only the gain in percentage points, not absolute TPR values or the baseline TPR at 1% FPR. A 2.7 to 9.4 point gain is therefore hard to put in context from the abstract alone.
Risks and caveats
The attack assumes a membership classifier learned from independently fine-tuned shadow models, although the authors report that a single shadow model is enough to see the leakage. The gains are measured at a strict 1% false positive rate and vary by a factor of about 3.5 between the weakest and strongest setting (2.7 versus 9.4 points). The perturbation result is a trade-off: noise that meaningfully reduces leakage also degrades the telemetry. The source names no authors or institutions, no submission date or venue, and no concrete deployment, so claims about real systems would go beyond what it shows.
“Our results show that router telemetry can turn an operational signal into an additional privacy surface for fine-tuned MoE models.”
— From the paper's abstract