Voho is delivering Aramco's AI call center.
Sounds more Saudi than Alibaba’s Qwen. Says it in a fifth of the words.

Voho Saudi Chat
4B

An Arabic assistant that answers in spoken Saudi, not newsreader Arabic.

Head to head
held-out test

Sounds more Saudi than Alibaba’s Qwen. Says it in a fifth of the words.

Replies that read as Saudi. Higher wins.

Voho89.8%
Alibaba Qwen62.3%

27.5 points ahead of Alibaba Qwen

vs Qwen3-4B-Instruct, 400 held-out questions; 6.4 words a reply vs 30.2.

89.8%
Replies read as Gulf, from 62.3%
6.4
Words per reply, from 30.2
2.5 GB
Q4 download, runs on a laptop
Apache 2.0
Commercial use allowed
Overview

What Voho Saudi Chat 4B is for

Ask a general 4B model a question in Saudi and the reply comes back as Gulf Arabic only 62% of the time, with Levantine and Egyptian leaking into the rest, and at 30 words: three times what a person says in one turn on the phone.

Voho Saudi Chat 4B replies the way a person does: consistently Najdi, at phone-call length. It sits between speech-to-text and text-to-speech in a voice agent, deciding what to say.

Why it's best

Why teams pick it

Each reason is a number from the evaluation below, against the models named there.

It sounds Saudi

89.8% of its replies are classified Gulf by an independent dialect classifier, against 62.3% for the model it started from and 94.5% for the reference replies themselves.

It answers like a phone call

6.4 words a reply instead of 30.2. A voice agent that speaks paragraphs loses the caller; this one takes a turn and stops.

You can use it commercially

Apache 2.0 end to end: the base model and every dataset in the mix. Nothing about the licence stops a business adopting it.

It runs where you need it

4 billion parameters. The Q4 build is 2.5 GB and runs locally with Ollama or llama.cpp; the full weights run on one GPU.

Specs

What it is, in one table

Parameters
4.0B
Base model
Qwen3-4B-Instruct-2507 (Apache 2.0)
Input / output
Arabic text in, spoken Najdi Arabic out
Reply length
About 6 words, one spoken turn
Full weights
8.0 GB, safetensors
Smallest build
2.5 GB, GGUF Q4_K_M
Licence
Apache 2.0, commercial use allowed
Evaluation

Measured on held-out data it never trained on

Share of replies an independent dialect classifier (MARBERTv2) assigns to each dialect, on 400 held-out questions. Higher Gulf and lower Modern Standard Arabic is the goal.

Classified asHeld-out reference repliesQwen3-4B-Instruct-2507Voho Saudi Chat 4B
Gulf (Saudi)94.5%62.3%89.8%
Modern Standard Arabic0.2%1.5%0.8%
Egyptian0.2%8.2%0.0%
Levantine4.0%19.8%7.8%
Maghrebi1.0%8.2%1.8%

The reference column is the ceiling: what the held-out replies themselves score. The classifier does not call every Saudi sentence Gulf, so a perfect model would not reach 100%.

ReferenceBase modelVoho Saudi Chat 4B
chrF++ against held-out replies—13.012.0
Mean reply length (words)10.730.26.4

chrF++ is flat against the base: the model says the right kind of thing in the right register, not the same words as the reference, which a dialect classifier measures and a string-overlap metric cannot.

Examples

Real outputs

Questions from the held-out set.

AskedVoho Saudi Chat 4B
هلا والله يا أبو ناصر، تدري أنا أبي أسجل حقوق الملكية؟هلا بك، أبشر. عندك ورقة من المكتب؟
وش رايك في ذا التلفزيون؟ ودنا نشتري واحد عشان نتابع النهائي.والله فكرة، بس أنا ما أحب أشتري شي جديد.
الطالب: يبه، ودي اروح محل الحيوانات اللي في التحلية ونشتري قطه.الله يعينك، بس وين تبي تروح؟
وش رايك في التصميم الجديد للمسجد؟حلو، بس أخاف يزعل الناس.
Formats

Every file, its size, and where it runs

BuildFileDownloadRuns on
Full weightsmodel.safetensors8.0 GBTransformers on one GPU
GGUF Q8_0voho-saudi-chat-4b-Q8_0.gguf4.3 GBllama.cpp, Ollama; near-lossless
GGUF Q6_Kvoho-saudi-chat-4b-Q6_K.gguf3.3 GBllama.cpp, Ollama
GGUF Q5_K_Mvoho-saudi-chat-4b-Q5_K_M.gguf2.9 GBllama.cpp, Ollama
GGUF Q4_K_Mvoho-saudi-chat-4b-Q4_K_M.gguf2.5 GBllama.cpp, Ollama; laptop CPU or GPU

Sizes from Hugging Face's file listing. Speed depends on your hardware; no benchmark is published yet.

Install

Run it in a few lines

Run it locally with Ollama

ollama run hf.co/VohoAI/voho-saudi-chat-4b-GGUF:Q4_K_M

Python, with Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "VohoAI/voho-saudi-chat-4b"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

SYSTEM = "أنت مساعد صوتي سعودي. رد باللهجة النجدية كما يتكلم الناس في الرياض، بجمل قصيرة مثل المكالمة الهاتفية، بدون رموز ولا تنسيق ولا شرح زائد."

messages = [{"role": "system", "content": SYSTEM},
            {"role": "user", "content": "أبي أحجز موعد بكرة الصبح، فيه وقت فاضي؟"}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=128, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

How it was trained

  • Base model: Qwen3-4B-Instruct-2507 (Apache 2.0), LoRA r=32 on all attention and MLP projections, 2 epochs, one NVIDIA L4. Loss on assistant turns only.
  • Data: 7,601 Voho service-call dialogues across eight enterprise verticals (oil and gas, utilities, telecom, banking, government, healthcare, logistics, facilities), 5,551 Voho everyday conversations, 202 rows from 2A2I/Arabic_Aya and 1,181 from arbml/CIDAR, all Apache 2.0.
  • Every dialogue had to pass a Najdi lexicon filter and the MARBERTv2 dialect classifier, which took no part in training. Anything with markdown, tables or a reply longer than a spoken turn was dropped. The dialogue data is published at VohoAI/voho-saudi-dialogues.
  • Keep the system prompt from the example: it is in every training example, and the dialect is noticeably weaker without it.
Availability

Where you can use it

  • Open weights on Hugging Face

    Apache 2.0, commercial use allowed

  • Commercial use

    Allowed under Apache 2.0

  • CPU and GPU

    Transformers, llama.cpp and Ollama

  • Offline, on your own servers

    Nothing calls out once it is downloaded

Limitations

  • Najdi, mostly. Hijazi and Khaleeji replies drift toward Najdi or Modern Standard Arabic.
  • The classifier’s Gulf class covers Saudi, the UAE and Kuwait: a high score means "reads as Gulf", not "reads as Riyadh".
  • Not a knowledge model. It is tuned for register and voice; ground it with retrieval for facts.
  • Writes without diacritics.

For production Saudi Arabic voice, including in-Kingdom and on-premise deployment, use the Voho API.

The voice pipeline

Hear it, decide what to say, say it the Saudi way.

The three Voho models are one voice agent: speech to text, the reply, and the reply rewritten the way a Saudi would say it.

  1. 01 · HearVoho Saudi STT SmallSpeech to text
  2. 02 · DecideVoho Saudi Chat 4BArabic assistant, spoken Saudi
  3. 03 · SayVoho Saudi Speak 0.6BFormal Arabic to spoken Saudi

Other models: Voho Saudi STT Small · Voho Saudi Speak 0.6B · All models

Deployment-ready

Start your AI transformation today.

Sign up and build your first agent in the browser, or book a call if you would rather someone walked you through it. Most people do not need the call.

Start

$5 of credit, free

Granted when you sign up, about 70 minutes of live calls. No card to begin.

Then

Plans from SAR 109 a month

Starter puts your agent on your website, Business on a Saudi phone line. Cancel any month.

When you need it

Enterprise terms

Saudi data residency, an uptime SLA and on-premise deployment, on an agreement.