Multilingual GSM-Symbolic maps what drives cross-language transfer

A paper listed on Hugging Face introduces Multilingual GSM-Symbolic, a math dataset built to test how capabilities learned in one language carry over to another. The authors say little is understood about this transfer, because existing evaluations rely on incomparable, saturation-prone datasets and rarely examine the determinants of transfer jointly. They argue that knowing what predicts transfer would let researchers avoid exhaustive evaluation across all language pairs, and let developers target the factors that limit performance in low-resource languages.
The dataset is described as extensible. It holds 30,000 item-matched question-answer pairs across 15 languages. It uses symbolic templates, which the authors say prevents overfitting and ensures generalisation, since a single sample can generate millions of high-quality variations.
Using the dataset, the authors estimate the largest determinants of capability. Model size comes first (β= 1.77), followed by language resource level (β= 0.77), reasoning (β= 0.67) and typological distance (β= -0.25), which is negative. Because the estimation is joint, the determinants can be expressed in terms of one another. The headline example: a 32B model evaluated in Marathi performs like a 10B model in English.
The authors say the findings matter for model developers. Model size and reasoning narrow the performance gap between low- and high-resource languages (β= -0.27 and β= -0.20, respectively). Similar levers have little or no effect on typologically distant languages.
Overall, the analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation. It predicts a model's performance on an unseen language within 6.0pp (r=.96). Adding measurements from just 10 templates in the target language cuts that error to 4.19pp, which the authors say allows reasonable performance estimates with little or no downstream dataset.
Key facts
- Multilingual GSM-Symbolic has 30,000 item-matched question-answer pairs across 15 languages, built from symbolic templates that can generate millions of variations from one sample.
- Estimated determinants of capability: model size (β= 1.77), language resource level (β= 0.77), reasoning (β= 0.67), typological distance (β= -0.25).
- Joint estimation yields an equivalence: a 32B model evaluated in Marathi performs like a 10B model in English.
- Model size and reasoning narrow the gap between low- and high-resource languages (β= -0.27 and β= -0.20), but have little or no effect for typologically distant languages.
- The framework explains 92% of between-language variation but only 23% of model-by-language variation, and predicts performance on an unseen language within 6.0pp (r=.96), or 4.19pp with 10 target-language templates.
Why it matters
Multilingual evaluation is expensive, and the authors say current benchmarks are hard to compare and prone to saturation. A framework that predicts how a model will do in a language it has not been tested on could replace exhaustive testing across all language pairs. The Marathi example makes the trade-off concrete: for this setup, a 32B model in Marathi performs like a 10B model in English, so the language penalty can be read in model-size terms.
Who it affects
The authors point first to model developers, who could target the factors that limit performance in low-resource languages. Researchers who evaluate multilingual models are the other audience, since the approach is meant to cut the number of language pairs that must be tested. Speakers of low-resource languages are the ones who bear the performance gap the study measures.
How to use it
The practical takeaway from the abstract is the 10-template shortcut: measurements from just 10 templates in a target language bring the prediction error for that language down from 6.0pp to 4.19pp, so a team could estimate performance with little or no downstream dataset. The abstract does not say that the dataset or code is publicly released, nor give a licence.
How solid is it
The numbers come from the abstract alone, and it names no authors or institutions. It gives no submission date, venue or peer-review status. It does not name the models tested or say how many there were, and it does not state the units or scale of the beta coefficients. The 32B versus 10B equivalence is stated only for Marathi versus English. The fit is strong between languages (92% explained, r=.96 on unseen languages) but weak for individual model-by-language combinations (23%).
Risks and caveats
The framework explains only 23% of the variation across model-language combinations, so it describes language-level trends far better than it predicts how one specific model will behave in one specific language. Model size and reasoning help low-resource languages, but the authors say similar levers have little or no effect on typologically distant languages. The dataset is mathematical, and the abstract makes no claim about other kinds of capability. A prediction error of 6.0pp, or 4.19pp with 10 templates, is the figure the authors report for an unseen language.
“a 32B model evaluated in Marathi performs like a 10B model in English”
— Abstract of the Multilingual GSM-Symbolic paper