Voho is delivering Aramco's AI call center.
Voho Saudi STT Small
Research · 14 September 2026

Hearing Saudi Arabic as it is spoken

A 0.24B speech model trained on about 187,000 Saudi clips cuts word errors from 103.9% to 39.2% against the OpenAI model it started from, across Najdi, Hijazi and Khaleeji.

39.2%
word error rate, from 103.9%
16.9%
character error rate, from 69.0%
4,582
held-out Saudi test clips
11,600
training steps on one GPU

A general model writes more errors than words

OpenAI’s Whisper Small is a good general speech model. On Saudi dialect speech its word error rate is 103.9%: it writes more wrong words than there were words spoken. A rate above 100% means the model is not only mishearing but inventing, filling gaps with Modern Standard Arabic it expects to hear.

That is the gap a Saudi voice agent falls into. Callers in Riyadh, Jeddah and the Eastern Province do not speak the language of the news, and every downstream step (understanding, the action taken, the reply) inherits whatever the first transcript got wrong.

The data: Saudi television, with dialect labels

The training data is SADA, the Saudi Audio Dataset for Arabic published by SDAIA and the National Center for AI: Saudi television speech, each clip labelled with the speaker’s dialect. We kept clips labelled Najdi, Hijazi, Khaleeji, unlabelled Saudi and Modern Standard Arabic, and removed overlapping speakers, non-Saudi dialects and anything over 30 seconds.

That leaves about 187,000 training clips. The test set is SADA’s own held-out split: 4,582 clips the model never saw, 1,704 Najdi, 1,150 Khaleeji, 809 Hijazi, 762 Saudi with no dialect label and 157 in Modern Standard Arabic.

What the model learns to write

The targets are the transcripts with diacritics and tatweel removed. A phone agent never needs vowel marks, and teaching a model to guess them spends capacity on something nobody reads.

Scoring uses the standard normalisation for Arabic speech recognition: alef forms unified, ta marbuta and alef maqsura normalised, punctuation removed. The base model and the fine-tune are scored with exactly the same function, so the comparison measures the model, not the normaliser.

Two epochs on one GPU

Whisper Small (244M parameters, MIT licence) was fine-tuned for two epochs: 11,600 steps at a learning rate of 1e-5 with 500 warm-up steps, in bf16, on a single NVIDIA L4.

Loss falls from 5.2 to under 1 in the first 250 steps as the model stops writing formal Arabic, then settles slowly. It steps down again at the start of the second epoch, around step 5,820, when the model sees each clip for the second time, and ends at 0.44.

00.511.5202,9005,8008,70011,600training steptraining losssecond epoch
Fig. 1Training loss by step. It starts at 5.2, above the top of this axis, and is under 1 within 250 steps as the model abandons formal Arabic; the second step down is the second pass over the data.

Accuracy against download size

Voho Saudi STT Small and OpenAI Whisper Small are the same 967 MB download and were scored by Voho on the same 4,582 clips, so between those two the whole difference is accuracy.

The hollow points are not Voho measurements. They are the word error rates published for the SADA test set in the Open Universal Arabic ASR Leaderboard (Wang, Alhmoud and Alqurishi, 2024, Table 1), placed at each model’s download size on Hugging Face. At 39.2%, the 967 MB Voho model sits above every model in that table, including Whisper Large v3 at three times the size (56.0%) and NVIDIA’s Arabic Conformer with its language model (44.5%).

Two things separate the published points from ours, and both should be read with the chart. The leaderboard scores the whole SADA test set with its own text normalisation, while Voho scores the Saudi-dialect clips only; in the leaderboard Whisper Small reaches 87.3%, against 103.9% in our run, so the published figures are likely the kinder ones. And every published model was tested without having seen SADA, while the Voho model was trained on SADA’s training split. The chart shows what specialising a small model buys on this speech, not that it is the better model everywhere.

20%40%60%80%100%120%200 MB500 MB1 GB2 GB5 GB10 GB20 GBdownload size, published weight file, log scaleword error rate, better is upnot on Hugging Facemeasured by Vohopublished, SADA testVoho Saudi STT Small · 39.2%967 MBOpenAI Whisper Small · 103.9%967 MBWhisper Large v3 · 56.0%Whisper Large v3 Turbo · 60.4%Whisper Medium · 67.7%HuBERT Large Arabic · 67.8%Meta Seamless M4T v2 · 62.5%Meta MMS 1B · 77.5%w2v-BERT 2.0 Arabic · 78.0%NVIDIA Conformer · 44.5%with its language model
Fig. 2Word error rate on Saudi television speech (SADA test) against download size; up is better. Solid points: measured by Voho on 4,582 held-out Saudi clips. Hollow points: published in the Open Universal Arabic ASR Leaderboard for the full SADA test set, a different scoring run, shown for scale.

Results by dialect

Across all 4,582 test clips, word errors fall from 103.9% to 39.2%, 62% fewer, and character errors from 69.0% to 16.9%. Every dialect improves by a similar margin, so the model has learned Saudi speech rather than one region of it.

Najdi and Hijazi end close together, near 36%. Khaleeji is hardest at 43.2%, and clips with no dialect label are hardest of all at 51.2%: they include more noise and crosstalk. Modern Standard Arabic, which the base model already handled best, still improves from 52.6% to 33.7%.

OpenAI Whisper SmallVoho Saudi STT Small
0%25%50%75%100%125%word error rate102.8%35.7%Najdi104.3%36.3%Hijazi106.6%43.2%Khaleeji135.8%51.2%Unlabelled52.6%33.7%Formal (MSA)
Fig. 3Word error rate by dialect, lower is better. Grey is OpenAI Whisper Small, green is the fine-tune, both scored identically on the held-out test set.

Characters tell a kinder story than words

Character error is less than half the word error rate, which means many remaining word errors are near-misses: one letter off, two words joined, a prefix split. Read a sample near the median and the pattern is plain: مقصر written as مقسر, تقعدون تضيعون heard as تقدم ضيعون. The meaning is usually recoverable; the spelling is not yet.

OpenAI Whisper SmallVoho Saudi STT Small
0%25%50%75%100%character error rate68.7%15.3%Najdi67.9%15.4%Hijazi70.5%18.3%Khaleeji97.3%25.1%Unlabelled29.5%12.6%Formal (MSA)
Fig. 4Character error rate by dialect, lower is better.
DialectWhat was saidWhat the model wrote
Najdiيا ليتك يا فيصل تعرف وش إللي أبغاه.ليلتك يا فيصل تعرف وش اللي أبغى
Khaleejiالوالد ما هو مقصر معي الله يطولي بعمره بسالوالد مهو مقسر معي الله يطولي بعمره بس
Khaleejiطيب يا جماعة لا تقعدون تضيعون الوقت الحين أبغى الشاي حقي.يا بيتي يا جماعة لا تقدم ضيعون الوقت الحين أبغى الشاي حقي

Television is not a phone line

The test set is television audio. A phone call arrives at 8 kHz, compressed and noisy, and accuracy on real calls will be lower than these numbers. The training speakers also skew male. Both are the next things to fix, and both need call-centre audio rather than more television.

The weights are CC BY-NC-SA 4.0 because SADA is, so the model is for research and non-commercial use. Production Saudi speech recognition, tuned on telephone audio and deployable in-Kingdom, runs through the Voho API.

What the result establishes

A small open model, trained once on public Saudi data with a standard recipe, closes most of the gap a general model leaves on Saudi dialect speech: from more errors than words to about two in five, and from 69% to 17% of characters. The method is plain enough to repeat on any dialect with labelled speech, which is the point.

Citation

Please cite the training data when using this model:

Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.
Deployment-ready

Start your AI transformation today.

Sign up and build your first agent in the browser, or book a call if you would rather someone walked you through it. Most people do not need the call.

Start

$5 of credit, free

Granted when you sign up, about 70 minutes of live calls. No card to begin.

Then

Plans from SAR 109 a month

Starter puts your agent on your website, Business on a Saudi phone line. Cancel any month.

When you need it

Enterprise terms

Saudi data residency, an uptime SLA and on-premise deployment, on an agreement.