Formal Arabic sounds like the news
Most text a voice agent has to say starts life formal: a policy, a template, a reply written by a general model. Read aloud in Modern Standard Arabic, it sounds like a bulletin. A caller in Riyadh hears that the line is a machine before the sentence ends.
Voho Saudi Speak sits between the text and the voice. Give it a formal sentence and a dialect, and it returns what a person from Riyadh, Jeddah or the Eastern Province would say on a call. It is 0.6 billion parameters, small enough to run on a CPU next to the speech model.
The data: real Saudi sentences as targets
The targets are what Saudi speakers actually said: 84,394 transcribed sentences from SADA, spoken by Najdi, Hijazi and Khaleeji speakers, cleaned of noise transcriptions, repeats and duplicates. Each is paired with a formal Modern Standard Arabic version of the same sentence, and pairs where the formal version was identical to the Saudi one were dropped, since they teach nothing.
The test set is SADA’s own held-out split: 1,000 sentences, 462 Najdi, 281 Khaleeji and 257 Hijazi.
Training: a full fine-tune of a 0.6B model
The base is Qwen3-0.6B (Apache 2.0), fully fine-tuned for two epochs on one NVIDIA L4 at a learning rate of 3e-5 with a cosine schedule. The prompt carries the formal sentence and the target dialect, and the loss falls only on the answer, so the model learns to produce Saudi speech rather than to repeat its instructions.
Measuring closeness to speech
chrF++ compares characters and words between the output and what the Saudi speaker said. It suits Arabic dialects, where one word can carry several spellings, better than exact word match.
The floor matters. Dialect and formal Arabic share most of their letters, so leaving the formal sentence unchanged already scores 61.8. Any model worth using has to beat doing nothing, and the untrained base model does not: at 52.2 it scores below the formal text it was given.
Accuracy against download size
At 1.2 GB the fine-tune sits 23.5 points above the untrained model of the same size, and above the dashed line that marks doing nothing at all. Gemini 2.5 Flash is drawn in its own band: Google does not publish its size and it runs only in Google’s cloud, so it is compared on accuracy alone.
Results
Voho Saudi Speak scores 75.7, 14 points above leaving the text formal and 23.5 above the untrained model. Every dialect clears 73: Najdi 77.0, Hijazi 75.8 and Khaleeji 73.2. Gemini 2.5 Flash, a much larger general model given the same instruction, scores 68.7.
Reading the Gemini comparison carefully
The comparison is fair on its own terms and narrow by design. Our model learned the test data’s style of Saudi spelling and phrasing from the training split; Gemini never saw it. So on this test the small model has an advantage that would shrink on Saudi text from elsewhere. The result shows that a 0.6B model, trained for one job, matches the target speech closely; it is not a claim that it knows Saudi Arabic better than Gemini in general.
What it sounds like
Sentences it never saw in training, of the kind a bank or clinic line says every day:
| Dialect | Formal input | Voho Saudi Speak 0.6B |
|---|---|---|
| Najdi | أين أنت الآن؟ أريد أن أحجز موعداً غداً. | وينك الحين أبغى أحجز موعد بكرة |
| Najdi | لا أستطيع رفع الحد إلى أكثر من ألف ريال، يجب أن تزور الفرع. | ما أقدر أرفع الحد إلى أكثر من ألف ريال لازم تزور الفرع. |
| Hijazi | سنرسل لك رمز التحقق الآن، من فضلك أخبرني به. | بنرسلك رمز التحقق دحين من فضلك خبرني فيه |
| Najdi | رقم طلبك هو 48213 وسيصل خلال ثلاثة أيام. | رقم طلبك هو 48213 وبيوصل خلال ثلاث أيام. |
Small enough to run anywhere
The full weights are 1.2 GB. Quantised to GGUF Q4_K_M the model is 397 MB, and the int8 ONNX build 754 MB, small enough for a CPU, an edge box or a browser.
Limitations
It learned from television, so it knows how Saudis talk in dramas and interviews better than how they talk to a bank. As a small model it can drop or change details in long sentences, and occasionally swaps who does what. Anything customer-facing should be checked. It writes without diacritics.
The weights are CC BY-NC-SA 4.0, following SADA. Production Saudi voice runs through the Voho API.
Citation
Please cite the training data when using this model:
Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.
More research