Voho Saudi STT
Small
Speech recognition for Saudi Arabic as it is actually spoken.
Beats OpenAI’s speech model on Saudi Arabic.
Word errors on Saudi speech. Lower wins.
62% fewer errors than OpenAI
vs OpenAI Whisper Small, 4,582 held-out Saudi clips.
- 39.2%
- Word error rate, from 103.9%
- 16.9%
- Character error rate, from 69.0%
- 0.24B
- Parameters
- 967 MB
- Download
What Voho Saudi STT Small is for
Most Arabic speech models are trained on Modern Standard Arabic: the news, not a phone call from Riyadh. On Saudi dialect speech the general model this one starts from writes more wrong words than there were words spoken.
Voho Saudi STT Small is fine-tuned on about 187,000 clips of Saudi speech in Najdi, Hijazi and Khaleeji, and cuts word errors on the Saudi test set by almost two thirds. It is the hearing stage of a Saudi voice agent.
Why teams pick it
Each reason is a number from the evaluation below, against the models named there.
Built for the dialects
Trained on Najdi, Hijazi and Khaleeji speech, and scored on each. Word errors drop from 103.9% to 39.2% on 4,582 held-out Saudi clips.
Close on the characters
Character error falls from 69.0% to 16.9%. Many remaining word errors are near-misses in spelling, such as مقسر for مقصر, rather than a different word.
Small and fast
0.24 billion parameters, under 1 GB. Runs on a CPU or a small GPU with the standard Transformers pipeline.
Honest about where it is weaker
Trained mostly on television speech; phone audio at 8 kHz is harder. For production calls, Voho’s API runs the speech stack end to end.
What it is, in one table
- Parameters
- 0.24B
- Base model
- OpenAI Whisper Small (MIT)
- Input
- Speech audio, up to 30 seconds a clip
- Output
- Arabic text, without diacritics
- Dialects
- Najdi, Hijazi, Khaleeji, plus Modern Standard Arabic
- Download
- 967 MB, safetensors
- Licence
- CC BY-NC-SA 4.0, non-commercial
Measured on held-out data it never trained on
Word and character error rate on the Saudi dialect test set, before and after fine-tuning (lower is better).
| Dialect | Clips | WER before | WER after | CER before | CER after |
|---|---|---|---|---|---|
| All Saudi test clips | 4,582 | 103.9% | 39.2% | 69.0% | 16.9% |
| Najdi (Riyadh, central) | 1,704 | 102.8% | 35.7% | 68.7% | 15.3% |
| Hijazi (Jeddah, Makkah) | 809 | 104.3% | 36.3% | 67.9% | 15.4% |
| Khaleeji (Eastern Province) | 1,150 | 106.6% | 43.2% | 70.5% | 18.3% |
| Saudi, dialect unlabelled | 762 | 135.8% | 51.2% | 97.3% | 25.1% |
| Modern Standard Arabic | 157 | 52.6% | 33.7% | 29.5% | 12.6% |
Scored after the standard Arabic normalisation: diacritics removed, alef forms unified, ta marbuta and alef maqsura normalised, punctuation removed. Both models were scored identically. A word error rate above 100% means more wrong words were written than were spoken.
Real outputs
Typical results, chosen near the median word error rate of the held-out samples, mistakes included.
| Dialect | What was said | What the model wrote |
|---|---|---|
| Najdi | يا ليتك يا فيصل تعرف وش إللي أبغاه. | ليلتك يا فيصل تعرف وش اللي أبغى |
| Khaleeji | الوالد ما هو مقصر معي الله يطولي بعمره بس | الوالد مهو مقسر معي الله يطولي بعمره بس |
| Khaleeji | طيب يا جماعة لا تقعدون تضيعون الوقت الحين أبغى الشاي حقي. | يا بيتي يا جماعة لا تقدم ضيعون الوقت الحين أبغى الشاي حقي |
Every file, its size, and where it runs
| Build | File | Download | Runs on |
|---|---|---|---|
| Full weights | model.safetensors | 967 MB | Transformers on CPU or GPU |
Sizes from Hugging Face's file listing. Speed depends on your hardware; no benchmark is published yet.
Run it in a few lines
Install
pip install transformers torchPython, with Transformers
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="VohoAI/voho-saudi-stt-small")
print(asr("call.wav", generate_kwargs={"language": "arabic", "task": "transcribe"})["text"])How it was trained
- Base model: OpenAI Whisper Small (244M parameters, MIT).
- Data: SADA, the Saudi Audio Dataset for Arabic published by SDAIA and the National Center for AI: Saudi television speech with dialect labels. Clips labelled Najdi, Hijazi, Khaleeji, unlabelled Saudi and Modern Standard Arabic were kept; overlapping speakers, non-Saudi dialects and clips over 30 seconds were removed.
- Targets: transcripts with diacritics removed. Setup: 2 epochs, batch size 32, learning rate 1e-5, bf16, one NVIDIA L4.
Where you can use it
Open weights on Hugging Face
CC BY-NC-SA 4.0, research and non-commercial use
Commercial use
Not under this licence. For production, use the Voho API.
CPU and GPU
Transformers
Offline, on your own servers
Nothing calls out once it is downloaded
Limitations
- Trained mostly on television speech. Telephone audio (8 kHz, compressed, noisy) is harder, and accuracy on real calls will be lower than the table.
- Small model: fast, but less accurate than larger variants.
- Speaker gender in the training data is skewed male.
- Transcripts are written without diacritics.
Citation
Please cite the training data when using this model:
Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.For production Saudi Arabic voice, including in-Kingdom and on-premise deployment, use the Voho API.
Hear it, decide what to say, say it the Saudi way.
The three Voho models are one voice agent: speech to text, the reply, and the reply rewritten the way a Saudi would say it.
- 01 · HearVoho Saudi STT SmallSpeech to text
- 02 · DecideVoho Saudi Chat 4BArabic assistant, spoken Saudi
- 03 · SayVoho Saudi Speak 0.6BFormal Arabic to spoken Saudi
Other models: Voho Saudi Chat 4B · Voho Saudi Speak 0.6B · All models