Ask an AI chatbot a question in formal Arabic (Fusha) and it answers beautifully. Ask it the same thing in Gazan, Egyptian, or Moroccan dialect — and it stumbles, mistranslates, or misses the point entirely.
As someone building an Arabic AI newsletter, I run into this every day. Here's what's actually going on under the hood — and why it's a great window into how large language models really work.
1. Models learn from data — and the data is lopsided
An LLM learns language by reading enormous amounts of text. For English, that's a huge chunk of the internet. For formal Arabic, there's a decent amount (news, books, Wikipedia). But for dialects? Very little is written down — dialects live in speech, voice notes, and casual chats, not in the formal text models train on.
Less data for a language variety = weaker performance. This is the well-known "low-resource language" problem, and Arabic dialects are a textbook case.








