A general model writes more errors than words
OpenAI’s Whisper Small is a good general speech model. On Saudi dialect speech its word error rate is 103.9%: it writes more wrong words than there were words spoken. A rate above 100% means the model is not only mishearing but inventing, filling gaps with Modern Standard Arabic it expects to hear.
That is the gap a Saudi voice agent falls into. Callers in Riyadh, Jeddah and the Eastern Province do not speak the language of the news, and every downstream step (understanding, the action taken, the reply) inherits whatever the first transcript got wrong.
The data: Saudi television, with dialect labels
The training data is SADA, the Saudi Audio Dataset for Arabic published by SDAIA and the National Center for AI: Saudi television speech, each clip labelled with the speaker’s dialect. We kept clips labelled Najdi, Hijazi, Khaleeji, unlabelled Saudi and Modern Standard Arabic, and removed overlapping speakers, non-Saudi dialects and anything over 30 seconds.
That leaves about 187,000 training clips. The test set is SADA’s own held-out split: 4,582 clips the model never saw, 1,704 Najdi, 1,150 Khaleeji, 809 Hijazi, 762 Saudi with no dialect label and 157 in Modern Standard Arabic.
What the model learns to write
The targets are the transcripts with diacritics and tatweel removed. A phone agent never needs vowel marks, and teaching a model to guess them spends capacity on something nobody reads.
Scoring uses the standard normalisation for Arabic speech recognition: alef forms unified, ta marbuta and alef maqsura normalised, punctuation removed. The base model and the fine-tune are scored with exactly the same function, so the comparison measures the model, not the normaliser.
Two epochs on one GPU
Whisper Small (244M parameters, MIT licence) was fine-tuned for two epochs: 11,600 steps at a learning rate of 1e-5 with 500 warm-up steps, in bf16, on a single NVIDIA L4.
Loss falls from 5.2 to under 1 in the first 250 steps as the model stops writing formal Arabic, then settles slowly. It steps down again at the start of the second epoch, around step 5,820, when the model sees each clip for the second time, and ends at 0.44.
Accuracy against download size
Voho Saudi STT Small and OpenAI Whisper Small are the same 967 MB download and were scored by Voho on the same 4,582 clips, so between those two the whole difference is accuracy.
The hollow points are not Voho measurements. They are the word error rates published for the SADA test set in the Open Universal Arabic ASR Leaderboard (Wang, Alhmoud and Alqurishi, 2024, Table 1), placed at each model’s download size on Hugging Face. At 39.2%, the 967 MB Voho model sits above every model in that table, including Whisper Large v3 at three times the size (56.0%) and NVIDIA’s Arabic Conformer with its language model (44.5%).
Two things separate the published points from ours, and both should be read with the chart. The leaderboard scores the whole SADA test set with its own text normalisation, while Voho scores the Saudi-dialect clips only; in the leaderboard Whisper Small reaches 87.3%, against 103.9% in our run, so the published figures are likely the kinder ones. And every published model was tested without having seen SADA, while the Voho model was trained on SADA’s training split. The chart shows what specialising a small model buys on this speech, not that it is the better model everywhere.
Results by dialect
Across all 4,582 test clips, word errors fall from 103.9% to 39.2%, 62% fewer, and character errors from 69.0% to 16.9%. Every dialect improves by a similar margin, so the model has learned Saudi speech rather than one region of it.
Najdi and Hijazi end close together, near 36%. Khaleeji is hardest at 43.2%, and clips with no dialect label are hardest of all at 51.2%: they include more noise and crosstalk. Modern Standard Arabic, which the base model already handled best, still improves from 52.6% to 33.7%.
Characters tell a kinder story than words
Character error is less than half the word error rate, which means many remaining word errors are near-misses: one letter off, two words joined, a prefix split. Read a sample near the median and the pattern is plain: مقصر written as مقسر, تقعدون تضيعون heard as تقدم ضيعون. The meaning is usually recoverable; the spelling is not yet.
| Dialect | What was said | What the model wrote |
|---|---|---|
| Najdi | يا ليتك يا فيصل تعرف وش إللي أبغاه. | ليلتك يا فيصل تعرف وش اللي أبغى |
| Khaleeji | الوالد ما هو مقصر معي الله يطولي بعمره بس | الوالد مهو مقسر معي الله يطولي بعمره بس |
| Khaleeji | طيب يا جماعة لا تقعدون تضيعون الوقت الحين أبغى الشاي حقي. | يا بيتي يا جماعة لا تقدم ضيعون الوقت الحين أبغى الشاي حقي |
Television is not a phone line
The test set is television audio. A phone call arrives at 8 kHz, compressed and noisy, and accuracy on real calls will be lower than these numbers. The training speakers also skew male. Both are the next things to fix, and both need call-centre audio rather than more television.
The weights are CC BY-NC-SA 4.0 because SADA is, so the model is for research and non-commercial use. Production Saudi speech recognition, tuned on telephone audio and deployable in-Kingdom, runs through the Voho API.
What the result establishes
A small open model, trained once on public Saudi data with a standard recipe, closes most of the gap a general model leaves on Saudi dialect speech: from more errors than words to about two in five, and from 69% to 17% of characters. The method is plain enough to repeat on any dialect with labelled speech, which is the point.
Citation
Please cite the training data when using this model:
Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.
More research