Voho is delivering Aramco's AI call center.
Beats OpenAI’s speech model on Saudi Arabic.

Voho Saudi STT
Small

Speech recognition for Saudi Arabic as it is actually spoken.

Head to head
held-out test

Beats OpenAI’s speech model on Saudi Arabic.

Word errors on Saudi speech. Lower wins.

Voho39.2%
OpenAI103.9%

62% fewer errors than OpenAI

vs OpenAI Whisper Small, 4,582 held-out Saudi clips.

39.2%
Word error rate, from 103.9%
16.9%
Character error rate, from 69.0%
0.24B
Parameters
967 MB
Download
Overview

What Voho Saudi STT Small is for

Most Arabic speech models are trained on Modern Standard Arabic: the news, not a phone call from Riyadh. On Saudi dialect speech the general model this one starts from writes more wrong words than there were words spoken.

Voho Saudi STT Small is fine-tuned on about 187,000 clips of Saudi speech in Najdi, Hijazi and Khaleeji, and cuts word errors on the Saudi test set by almost two thirds. It is the hearing stage of a Saudi voice agent.

Why it's best

Why teams pick it

Each reason is a number from the evaluation below, against the models named there.

Built for the dialects

Trained on Najdi, Hijazi and Khaleeji speech, and scored on each. Word errors drop from 103.9% to 39.2% on 4,582 held-out Saudi clips.

Close on the characters

Character error falls from 69.0% to 16.9%. Many remaining word errors are near-misses in spelling, such as مقسر for مقصر, rather than a different word.

Small and fast

0.24 billion parameters, under 1 GB. Runs on a CPU or a small GPU with the standard Transformers pipeline.

Honest about where it is weaker

Trained mostly on television speech; phone audio at 8 kHz is harder. For production calls, Voho’s API runs the speech stack end to end.

Specs

What it is, in one table

Parameters
0.24B
Base model
OpenAI Whisper Small (MIT)
Input
Speech audio, up to 30 seconds a clip
Output
Arabic text, without diacritics
Dialects
Najdi, Hijazi, Khaleeji, plus Modern Standard Arabic
Download
967 MB, safetensors
Licence
CC BY-NC-SA 4.0, non-commercial
Evaluation

Measured on held-out data it never trained on

Word and character error rate on the Saudi dialect test set, before and after fine-tuning (lower is better).

DialectClipsWER beforeWER afterCER beforeCER after
All Saudi test clips4,582103.9%39.2%69.0%16.9%
Najdi (Riyadh, central)1,704102.8%35.7%68.7%15.3%
Hijazi (Jeddah, Makkah)809104.3%36.3%67.9%15.4%
Khaleeji (Eastern Province)1,150106.6%43.2%70.5%18.3%
Saudi, dialect unlabelled762135.8%51.2%97.3%25.1%
Modern Standard Arabic15752.6%33.7%29.5%12.6%

Scored after the standard Arabic normalisation: diacritics removed, alef forms unified, ta marbuta and alef maqsura normalised, punctuation removed. Both models were scored identically. A word error rate above 100% means more wrong words were written than were spoken.

Examples

Real outputs

Typical results, chosen near the median word error rate of the held-out samples, mistakes included.

DialectWhat was saidWhat the model wrote
Najdiيا ليتك يا فيصل تعرف وش إللي أبغاه.ليلتك يا فيصل تعرف وش اللي أبغى
Khaleejiالوالد ما هو مقصر معي الله يطولي بعمره بسالوالد مهو مقسر معي الله يطولي بعمره بس
Khaleejiطيب يا جماعة لا تقعدون تضيعون الوقت الحين أبغى الشاي حقي.يا بيتي يا جماعة لا تقدم ضيعون الوقت الحين أبغى الشاي حقي
Formats

Every file, its size, and where it runs

BuildFileDownloadRuns on
Full weightsmodel.safetensors967 MBTransformers on CPU or GPU

Sizes from Hugging Face's file listing. Speed depends on your hardware; no benchmark is published yet.

Install

Run it in a few lines

Install

pip install transformers torch

Python, with Transformers

from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="VohoAI/voho-saudi-stt-small")
print(asr("call.wav", generate_kwargs={"language": "arabic", "task": "transcribe"})["text"])

How it was trained

  • Base model: OpenAI Whisper Small (244M parameters, MIT).
  • Data: SADA, the Saudi Audio Dataset for Arabic published by SDAIA and the National Center for AI: Saudi television speech with dialect labels. Clips labelled Najdi, Hijazi, Khaleeji, unlabelled Saudi and Modern Standard Arabic were kept; overlapping speakers, non-Saudi dialects and clips over 30 seconds were removed.
  • Targets: transcripts with diacritics removed. Setup: 2 epochs, batch size 32, learning rate 1e-5, bf16, one NVIDIA L4.
Availability

Where you can use it

  • Open weights on Hugging Face

    CC BY-NC-SA 4.0, research and non-commercial use

  • Commercial use

    Not under this licence. For production, use the Voho API.

  • CPU and GPU

    Transformers

  • Offline, on your own servers

    Nothing calls out once it is downloaded

Limitations

  • Trained mostly on television speech. Telephone audio (8 kHz, compressed, noisy) is harder, and accuracy on real calls will be lower than the table.
  • Small model: fast, but less accurate than larger variants.
  • Speaker gender in the training data is skewed male.
  • Transcripts are written without diacritics.

Citation

Please cite the training data when using this model:

Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.

For production Saudi Arabic voice, including in-Kingdom and on-premise deployment, use the Voho API.

The voice pipeline

Hear it, decide what to say, say it the Saudi way.

The three Voho models are one voice agent: speech to text, the reply, and the reply rewritten the way a Saudi would say it.

  1. 01 · HearVoho Saudi STT SmallSpeech to text
  2. 02 · DecideVoho Saudi Chat 4BArabic assistant, spoken Saudi
  3. 03 · SayVoho Saudi Speak 0.6BFormal Arabic to spoken Saudi

Other models: Voho Saudi Chat 4B · Voho Saudi Speak 0.6B · All models

Deployment-ready

Start your AI transformation today.

Sign up and build your first agent in the browser, or book a call if you would rather someone walked you through it. Most people do not need the call.

Start

$5 of credit, free

Granted when you sign up, about 70 minutes of live calls. No card to begin.

Then

Plans from SAR 109 a month

Starter puts your agent on your website, Business on a Saudi phone line. Cancel any month.

When you need it

Enterprise terms

Saudi data residency, an uptime SLA and on-premise deployment, on an agreement.