Voho Saudi Speak
0.6B
Turns formal Arabic into Arabic the way Saudis actually say it.
Beats Google’s Gemini at spoken Saudi, at a fraction of the size.
How close it sounds to a real Saudi speaker. Higher wins.
7.0 ahead of Google Gemini
vs Gemini 2.5 Flash, chrF++ on 1,000 held-out Saudi sentences.
- 75.7
- chrF++ to Saudi speech, from 52.2
- 3
- Dialects: Najdi, Hijazi, Khaleeji
- 0.6B
- Parameters
- 397 MB
- Q4 download, runs on a CPU
What Voho Saudi Speak 0.6B is for
A voice agent that reads out written Arabic sounds like a news bulletin. This small model sits between the text and the voice: give it a formal sentence and a dialect, and it returns what a person from Riyadh, Jeddah or the Eastern Province would say on a phone call.
At 0.6 billion parameters it runs on a CPU or next to a speech model on one GPU, and the quantised build is under 400 MB.
Why teams pick it
Each reason is a number from the evaluation below, against the models named there.
Closer to how Saudis talk
On 1,000 held-out sentences its output scores 75.7 chrF++ against what a Saudi speaker actually said, against 61.8 for leaving the formal text as it is and 52.2 for the untrained model.
Three dialects, chosen per call
Najdi for Riyadh and the centre, Hijazi for Jeddah and Makkah, Khaleeji for the Eastern Province. Each scores above 73 chrF++.
Small enough to run anywhere
0.6B parameters. The Q4 GGUF build is 397 MB and the int8 ONNX build 754 MB, for CPUs, edge devices and browsers.
Open, and runs offline
Full weights, GGUF and ONNX builds are all published. Run it inside your own network with no calls out, and inspect exactly what it does.
What it is, in one table
- Parameters
- 0.6B
- Base model
- Qwen3-0.6B (Apache 2.0)
- Input
- A formal Arabic sentence and a dialect
- Output
- The same sentence in spoken Saudi, without diacritics
- Full weights
- 1.2 GB, safetensors
- Smallest build
- 397 MB, GGUF Q4_K_M
- Licence
- CC BY-NC-SA 4.0, non-commercial
Measured on held-out data it never trained on
chrF++ against what a Saudi speaker actually said, on 1,000 held-out sentences (higher is better).
| Dialect | Sentences | Formal text unchanged | Qwen3-0.6B, untrained | Voho Saudi Speak 0.6B | Gemini 2.5 Flash |
|---|---|---|---|---|---|
| All test sentences | 1,000 | 61.8 | 52.2 | 75.7 | 68.7 |
| Najdi (Riyadh, central) | 462 | 62.2 | 51.5 | 77.0 | 70.6 |
| Hijazi (Jeddah, Makkah) | 257 | 65.8 | 57.7 | 75.8 | 69.1 |
| Khaleeji (Eastern Province) | 281 | 56.6 | 47.3 | 73.2 | 64.7 |
"Formal text unchanged" is the floor: dialect and formal Arabic share most of their letters, so leaving a sentence alone already scores well. Gemini 2.5 Flash, a much larger model given the same instruction, is shown for scale. This model learned the test data’s style of Saudi spelling and Gemini did not, so the comparison favours the small model on this test; it does not mean it is better at Saudi Arabic in general.
Real outputs
Sentences it never saw in training, from a bank or clinic line.
| Dialect | Formal input | Voho Saudi Speak 0.6B |
|---|---|---|
| Najdi | أين أنت الآن؟ أريد أن أحجز موعداً غداً. | وينك الحين أبغى أحجز موعد بكرة |
| Najdi | لا أستطيع رفع الحد إلى أكثر من ألف ريال، يجب أن تزور الفرع. | ما أقدر أرفع الحد إلى أكثر من ألف ريال لازم تزور الفرع. |
| Hijazi | سنرسل لك رمز التحقق الآن، من فضلك أخبرني به. | بنرسلك رمز التحقق دحين من فضلك خبرني فيه |
| Najdi | رقم طلبك هو 48213 وسيصل خلال ثلاثة أيام. | رقم طلبك هو 48213 وبيوصل خلال ثلاث أيام. |
Every file, its size, and where it runs
| Build | File | Download | Runs on |
|---|---|---|---|
| Full weights | model.safetensors | 1.2 GB | Transformers on CPU or GPU |
| GGUF F16 | voho-saudi-speak-0.6b-F16.gguf | 1.2 GB | llama.cpp, Ollama; lossless |
| GGUF Q8_0 | voho-saudi-speak-0.6b-Q8_0.gguf | 639 MB | llama.cpp, Ollama; near-lossless |
| GGUF Q4_K_M | voho-saudi-speak-0.6b-Q4_K_M.gguf | 397 MB | llama.cpp, Ollama; any CPU |
| ONNX int8 | onnx/model_int8.onnx | 754 MB | ONNX Runtime; edge and browser |
| ONNX full | onnx/model.onnx | 3.0 GB | ONNX Runtime |
Sizes from Hugging Face's file listing. Speed depends on your hardware; no benchmark is published yet.
Run it in a few lines
Run it locally with Ollama
ollama run hf.co/VohoAI/voho-saudi-speak-0.6b-GGUF:Q4_K_MPython, with Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "VohoAI/voho-saudi-speak-0.6b"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
DIALECT = {"najdi": "النجدية", "hijazi": "الحجازية", "khaleeji": "الخليجية الشرقية"}
def saudi(text, dialect="najdi"):
prompt = f"أعد صياغة هذه الجملة باللهجة السعودية {DIALECT[dialect]} كما يقولها شخص في مكالمة، بدون أي شرح:\n{text}"
ids = tok.apply_chat_template([{"role": "user", "content": prompt}], add_generation_prompt=True,
enable_thinking=False, return_tensors="pt")
out = model.generate(ids, max_new_tokens=96, do_sample=False)
return tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip()
print(saudi("أين أنت الآن؟ أريد أن أحجز موعداً غداً."))How it was trained
- Base model: Qwen3-0.6B (Apache 2.0), full fine-tune, 2 epochs, one NVIDIA L4.
- Targets: 84,394 transcribed sentences of Saudi speech from SADA (Najdi, Hijazi and Khaleeji speakers), cleaned of noise transcriptions, repeats and duplicates, each paired with a formal Modern Standard Arabic version.
- Test set: SADA’s own test split, never seen in training.
Where you can use it
Open weights on Hugging Face
CC BY-NC-SA 4.0, research and non-commercial use
Commercial use
Not under this licence. For production, use the Voho API.
CPU and GPU
Transformers, llama.cpp and Ollama
Offline, on your own servers
Nothing calls out once it is downloaded
Limitations
- Trained on television speech: it knows how Saudis talk in dramas and interviews better than how they talk to a bank.
- Small model: it can drop or change details in long or complicated sentences, and occasionally swaps who does what. Check anything customer-facing.
- Writes without diacritics.
Citation
Please cite the training data when using this model:
Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.For production Saudi Arabic voice, including in-Kingdom and on-premise deployment, use the Voho API.
Hear it, decide what to say, say it the Saudi way.
The three Voho models are one voice agent: speech to text, the reply, and the reply rewritten the way a Saudi would say it.
- 01 · HearVoho Saudi STT SmallSpeech to text
- 02 · DecideVoho Saudi Chat 4BArabic assistant, spoken Saudi
- 03 · SayVoho Saudi Speak 0.6BFormal Arabic to spoken Saudi
Other models: Voho Saudi Chat 4B · Voho Saudi STT Small · All models