Falcon-Emirati-7B adapts Falcon-H1-Arabic to Emirati dialect, scores 84.83% on Alyah

Falcon-Emirati-7B adapts Falcon-H1-Arabic to Emirati dialect, scores 84.83% on Alyah

Falcon-Emirati-7B is a dialect-specialized model for Emirati Arabic, built on the 7B variant of Falcon-H1-Arabic, the authors' Arabic model family. The blog post opens with the problem: Arabic is really a family of languages. Modern Standard Arabic (MSA) is what people read in the news, but everyday talk, humor, negotiation and storytelling in the UAE happen in Emirati Arabic, a Gulf dialect with its own vocabulary, rhythm and culture. Nabati poetry, proverbs and short anecdotes carry meaning that does not survive a word-for-word reading, so a model that knows only MSA can translate every word and still miss the point.

The base family, Falcon-H1-Arabic, uses the Falcon-H1 hybrid architecture, with Mamba state space layers and Transformer attention running in parallel inside every block. It comes in 3B, 7B and 34B sizes with context windows up to 128K and 256K tokens, and was already trained on MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) plus English and multilingual data. The authors picked the 7B variant as the best balance: the 34B model would likely push quality a bit further but at a training and serving cost they say does not make sense for a dialect chat model, while 3B leaves too little headroom for cultural and linguistic depth.

The authors list three reasons dialect adaptation is hard. Emirati is mostly spoken, so there is far less written text online than for MSA or other Gulf and Levantine dialects. Meaning is often non-literal. And there is no established recipe for how much dialect data is enough, how to mix it with MSA, or which training stage (continued pre-training, SFT or preference optimization) matters most. They say much of the work came down to trial and error, with ablations on how much dialect data to inject and when, how to balance crawled and synthetic data without overfitting to synthetic patterns, and how much MSA cultural context was needed.

The data pipeline has three sources. First, authentic dialect web data crawled from Emirati websites and forums, written natively rather than translated or transliterated from MSA. Second, MSA material about Emirati culture, heritage, language, customs, values, history and social norms, which does not teach dialect writing but teaches the model what it is talking about. Third, a large amount of synthetic Emirati-dialect data, generated under strict rules plus glossaries and dictionaries built for Emirati vocabulary and grammar, which the authors say separated authentic-sounding output from text that is grammatical but sounds off to a speaker.

Evaluation used native-speaker review during training (naturalness, tone, cultural appropriateness) and Alyah, a benchmark the authors and the community released for Emirati-dialect capability. Alyah is a fully native multiple-choice benchmark of 1,173 samples collected manually from native Emirati speakers, covering greetings and etiquette, figurative language, heritage knowledge and Emirati poetry. Falcon-Emirati-7B scores 84.83%, which the authors say is ahead of every other Arabic and multilingual model they compared, including several many times its size. The Falcon-H1-Arabic family is excluded from that chart because the new model is built on it. The competing models' individual scores appear only in a chart. The authors conclude that size alone does not buy dialect competence, and that Emirati proficiency has to be trained for on purpose.

A second test used open-ended generation on the same 1,173 questions, scored by an LLM judge (Gemini 3.7 Flash) against ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat and Fanar-2-27B-Instruct, picked as the strongest competitors on the Alyah leaderboard. The judge rated correctness and, independently, whether the answer came back in Emirati rather than MSA. Falcon-Emirati-7B leads on correctness, but the larger gap is dialect fidelity. On partial credit it scores 0.52, against 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat and effectively 0.00 for Fanar-2-27B-Instruct. The authors call that close to two orders of magnitude at the low end. Their reading: the other models often know the answer but reply in MSA by default, even when asked in Emirati, and only Falcon-Emirati-7B reliably answers in the dialect it was asked in.

Fanar-2-27B-Instruct also abstained on 26.2% of questions, versus under 5% for every other model, and had the lowest correctness score of the five at 0.27 (partial credit). The authors read this as a model both less willing and less able to engage with Emirati-specific content. By category, the dialect-fidelity pattern holds across every Alyah category, from greetings to poetry. The one place rivals do relatively better is Greetings & Daily Expressions, where Emirati and MSA overlap most. The post then moves to a pairwise category-by-category comparison, but the text available here cuts off before those results.

Key facts

  • Falcon-Emirati-7B is a dialect-specialized model for Emirati Arabic, built on the 7B variant of Falcon-H1-Arabic (a family with 3B, 7B and 34B sizes).
  • The authors report 84.83% on Alyah, a native multiple-choice benchmark of 1,173 samples, ahead of every other Arabic and multilingual model they compared, in their own comparison.
  • In an LLM-judged open-ended test (Gemini 3.7 Flash), dialect fidelity is 0.52 partial credit versus 0.05 (ALLaM), 0.03 (gemma-3-27b-it), 0.02 (Jais-2-8B-Chat) and effectively 0.00 (Fanar-2-27B-Instruct).
  • Training data came from three sources: native Emirati web data, MSA text about Emirati culture, and synthetic dialect data constrained by glossaries and style rules.
  • Fanar-2-27B-Instruct declined to answer 26.2% of questions, versus under 5% for every other model in the comparison.

Why it matters

Most Arabic models are strong in MSA, but the post argues that is not how people talk. The reported results point to a specific failure: rival models often know the answer but reply in MSA even when asked in Emirati. The authors say size does not fix this, since some of the largest multilingual models score well below smaller, more dialect-aware ones. The work also offers a data recipe for a low-resource, mostly spoken dialect: native web text, cultural MSA text, and glossary-constrained synthetic data.

Who it affects

Teams building chat or assistant products for UAE users, and researchers working on Arabic dialects and low-resource language adaptation. It also bears on other Arabic model builders, since the comparison names ALLaM, gemma-3-27b-it, Jais-2 and Fanar-2 as models that mostly stay in MSA. The post frames the model for understanding and generating Emirati Arabic the way a native speaker would, including tone and cultural context.

How to use it

The visible text gives no download link, licence or release date for Falcon-Emirati-7B, so check the original post before planning any deployment. The authors describe it as a dialect-specialized chat model, with 7B chosen to keep training and inference practical. For evaluation, the post points to Alyah, which the authors and the community released, and to their separate benchmark blog post for construction and category details.

How solid is it

This is a first-party lab post, and the comparisons are the authors' own. The 84.83% Alyah figure is a multiple-choice accuracy; the competing models' scores appear only in a chart. The dialect-fidelity numbers (0.52, 0.05, 0.03, 0.02, 0.00) are partial-credit scores from a single LLM judge, Gemini 3.7 Flash, not accuracy percentages, and the stricter pass/fail values are not given in the text. The five judged models were chosen by the authors as the strongest competitors on the Alyah leaderboard. The available text is cut off in the pairwise comparison section, so those results are not covered here.

Risks and caveats

The authors say building the model involved a lot of trial and error and that there is no established recipe for MSA-to-dialect adaptation. The evaluation rests on one benchmark the authors helped release plus one LLM judge, so the headline gaps are not independently confirmed in the text. No training data sizes, token counts, hyperparameters or compute are given. The post also cautions that automatic metrics alone do not capture naturalness, tone or cultural fit well enough to trust on their own, which is why it also used native-speaker review. Its 'Fanar abstained more' reading is hedged by the authors as what the pattern suggests.

“Size alone doesn't buy you dialect competence.”

— Falcon-Emirati blog post, the authors