Falcon vs. ALLaM: The Fierce Battle Over Arabic AI in the Gulf Skies

Large Language ModelsArabic AI
Illustration generated by AI: Editorial image for فالكون أم علّام؟ معركة شرسة في سماء الخليج لتطويع الذكاء الاصطناعي بالعربية

The Core · TL;DR

  • Falcon-H1 Arabic, launched by the UAE's Technology Innovation Institute on January 5, 2026 with a hybrid Mamba-Transformer architecture, scores 75.36% on the OALL benchmark at 34 billion parameters, surpassing according to the institute much larger global models such as Qwen2.5 72B and Llama-3.3 70B.
  • ALLaM, built from scratch by Saudi Arabia's SDAIA and now operated by HUMAIN through the HUMAIN Chat app, along with the UAE's Jais notably outperforms rivals in understanding Gulf dialects and cultural expressions despite trailing Falcon on general OALL metrics.
  • The Arabic benchmark's shift toward native tests like 3LM, ArabCulture, and AraDice corrected a methodological flaw in the first generation, which relied on Arabic translations of English tests that measured translation quality more than actual Arabic capability, even as major global models like ChatGPT, Gemini, and Claude continue narrowing the gap with specialized Arabic models.

The race for dominance among Arabic large language models (LLMs) has entered a new phase, shifting from mere claims of Arabic "comprehension" to a battle of precise metrics measured on specialized benchmarks. At the heart of this contest stand two sovereign Gulf models representing two distinct philosophies in construction and deployment: the UAE's Falcon and Saudi Arabia's ALLaM, alongside Qatar's Fanar as a third, less prominent player in recent coverage.

Who Is Behind Each Model, and Why It Was Built

Falcon is developed by the Technology Innovation Institute (TII), the applied research arm of Abu Dhabi's Advanced Technology Research Council (ATRC). On January 5, 2026, the institute launched Falcon-H1 Arabic, built on a hybrid architecture combining Mamba and Transformer, an approach that seeks to merge the sequential processing efficiency of State Space Models with the traditional Attention capabilities of Transformers. Methodologically, the more significant point, according to Al Jazeera Net reports, is that Falcon Arabic was built on native Arabic data reflecting the language's full linguistic diversity, rather than through translated data from other languages, a fundamental distinction we will return to when discussing benchmarking.

ALLaM, by contrast, follows a somewhat different institutional path. It was developed by the Saudi Data and Artificial Intelligence Authority (SDAIA) as a model built and trained from scratch to support Modern Standard Arabic alongside Saudi dialects and multiple Arabic vernaculars. However, management of the model later shifted from SDAIA to HUMAIN, meaning ALLaM remains Saudi Arabia's national model built by SDAIA but is now operated by HUMAIN, which currently powers the HUMAIN Chat application. This transition from a government research body to an operational/commercial entity reflects a governance model distinct from the UAE approach, where Falcon remains directly under the umbrella of an applied research institute affiliated with ATRC, without a separate operational intermediary in the same sense.

Strategically, analyses indicate that all three Gulf sovereign models, Falcon, ALLaM, and Fanar, were built to solve the same problem: the weak representation of Arabic in global models. Yet each adopts a different approach reflecting its country's priorities. It is worth noting in this context that the Technology Innovation Institute has been present in the Arabic language model race longer than any other regional entity, with the Falcon family beginning as a multilingual model a year before its specialized Arabic branch emerged, in what is described as a deliberate strategic shift toward linguistic specialization rather than generalization.

What the Benchmark Numbers Actually Show

The decisive test for any claim of linguistic superiority is performance on the Open Arabic LLM Leaderboard (OALL), and here a notable gap emerges in favor of Falcon-H1 Arabic within comparable size categories. At the mid-size model level, the 7-billion-parameter Falcon model achieves an average score of 71.47%, outperforming all models of roughly 10 billion parameters, including Qatar's Fanar-1-9B and Saudi Arabia's HUMAIN ALLaM 7B.

At the larger model level, the biggest Falcon-H1 Arabic release, at 34 billion parameters, scores 75.36%, a result the institute says surpasses far larger global systems in terms of parameter count, such as China's Qwen2.5 72B and Meta's Llama-3.3 70B, despite both models exceeding 70 billion parameters, roughly double Falcon's size. If this comparison holds, it suggests that deep linguistic specialization may offset raw scale disadvantages, a pivotal point in the ongoing debate over the efficiency of smaller models trained on native data versus massive multilingual models.

Falcon-H1 Arabic also records high scores on more specialized and complex Arabic benchmarks than general comprehension tests, including 3LM for scientific and mathematical reasoning in Arabic, ArabCulture for cultural and contextual questions requiring non-linguistic background knowledge, and AraDice for dialect comprehension. This diversification of test types matters methodologically because it separates "surface-level linguistic understanding" from "deep reasoning within Arabic cultural context."

However, this numerical superiority is not the end of the story. According to technical analysis of the leaderboard's own evolution, the shift toward "native" benchmarks such as 3LM, ArabCulture, and AraDice occurred precisely because the first version of OALL relied on Arabic translations of well-known English tests such as MMLU and EXAMS. These translated tests carry phrasing and cultural assumptions originally rooted in English, meaning they were measuring translation quality as much as a model's actual Arabic capability. In other words, the shift to the second generation of benchmarks represents a methodological correction of a flaw that had been masking real differences between models, lending Falcon-H1's current results higher credibility than earlier generations of Arabic models measured against the old translated metrics.

Conversely, specialized sources point to a more complex picture specifically regarding dialects, noting that ALLaM and Jais notably outperform others in understanding Gulf dialects and colloquial cultural expressions. This means Falcon's overall lead on the OALL leaderboard does not necessarily translate into absolute superiority across every subtask, and that selecting the "best" model remains contingent on the target application: is it formal scientific and mathematical reasoning in Modern Standard Arabic, or conversational interaction requiring precise understanding of a specific Gulf dialect and its everyday cultural context?

Where the Gap Against Global Models Remains

Despite the relative progress of sovereign Arabic models on their specialized benchmarks, the broader landscape reveals that competition is no longer confined to Arabic models alone. In recent years, alongside the emergence of specialized models most notably Saudi Arabia's ALLaM and the UAE's Jais, major global models such as ChatGPT, Gemini, and Claude have shown marked improvement in handling Arabic. This means Gulf sovereign models are not only competing among themselves for the top of the OALL leaderboard, but also facing sustained competition from general-purpose models with training and compute resources that dwarf what is available to any single regional project, models that are gradually closing the gap that originally justified the existence of specialized Arabic models.

The consequence of this reality is that the rationale for continued investment in sovereign Arabic models is no longer "closing the linguistic capability gap" on its own, but is gradually shifting toward other considerations: data sovereignty, control over infrastructure and local deployment, and precise specialization in dialects and cultural contexts that may not be a commercial priority for major global players. At its core, the battle between Falcon and ALLaM remains a reflection of a deeper strategic question that transcends leaderboard numbers: is the goal to build an Arabic model that competes globally on raw scale and performance standards, or to build sovereign infrastructure that serves specific national priorities, even if it remains smaller than its international rivals?

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Research

View all in Research