Voho is delivering Aramco's AI call center.
Voho Saudi Chat 4B
Research · 19 September 2026

Answering in Najdi, at phone-call length

A 4B assistant whose replies read as Gulf 89.8% of the time, against 62.3% for the model it started from, at 6.4 words a turn. Apache 2.0, the dialogue data included.

89.8%
replies read as Gulf, from 62.3%
0.8%
replies in formal Arabic
6.4
words a reply, from 30.2
Apache 2.0
commercial use allowed

The wrong dialect, at the wrong length

Ask a general 4B model a question in Saudi and its reply reads as Gulf Arabic only 62.3% of the time. A fifth of its replies read as Levantine and 8% as Egyptian. They run to 30 words, three times what a person says in one turn on the phone.

For a voice agent both are failures. The wrong dialect tells the caller they are talking to software; the wrong length means the caller waits through a paragraph. Voho Saudi Chat 4B is trained to fix both, and to stay usable by any business: the base model and every dataset are Apache 2.0.

Measuring dialect with a classifier that took no part

Replies are scored by MARBERTv2, an independent Arabic dialect classifier that had no part in training. The metric is the share of replies it reads as Gulf (want high) and as Modern Standard Arabic (want low).

The classifier was checked before any training. Held-out Saudi reference replies read as Gulf 94.5% of the time, and Arabic instruction data read as formal Arabic 68% of the time, so it separates the two cleanly. The ceiling is 94.5%, not 100%: it does not call every genuinely Saudi sentence Gulf, and the reference column on every chart shows that ceiling.

The data: service calls and everyday talk

The training set mixes Voho’s own dialogues with two open datasets: 7,601 service-call dialogues across eight sectors (oil and gas, utilities, telecom, banking, government, healthcare, logistics, facilities), 5,551 everyday conversations, 202 Gulf-tagged rows from Arabic Aya and 1,181 instruction pairs from CIDAR. The test set is 400 held-out questions the model never trained on.

Every dialogue had to pass two checks: a Najdi lexicon filter (Saudi function words such as وش، تبي، الحين present, their formal equivalents absent) and the same MARBERTv2 classifier, with at least 60% of its replies reading as Gulf. CIDAR is capped at about a tenth of the training words, because uncapped it would teach the long formal answers this model exists to remove. Anything with markdown, tables or code was dropped: a speech model reads asterisks aloud.

The run that looked like a win

The first full run trained on the service calls alone. Its headline number rose 17 points, to 79.0% Gulf, and it looked like success. Two other numbers said otherwise: formal Arabic jumped to 11.2% of replies, and the samples showed every question being answered as a support ticket. Asked its opinion on a television for watching the final, it explained that the system supports the television and it could be bought in store.

The fix was coverage, not more dialect. The shipped run adds the everyday conversations (family, food, driving, occasions), 38% of the words in Voho’s dialogues. Gulf rises to 89.8%, formal Arabic falls to 0.8%, and the same question gets: والله فكرة، بس أنا ما أحب أشتري شي جديد.

Gulf (Saudi)Formal ArabicLevantineEgyptianMaghrebi
0%25%50%75%100%Reference replies94.5%Qwen3-4B, untrained62.3%19.8%8.2%8.2%Service calls only79%11.2%Voho Saudi Chat 4B89.8%7.8%
Fig. 1Share of replies the independent classifier assigns to each dialect, 400 held-out questions. Service calls alone raised Gulf but pushed formal Arabic to 11%; adding everyday talk fixed both.

Accuracy against download size

All three points are the same 8.0 GB download, so the comparison is at equal size. Training moved the share of Gulf replies from 62.3% to 89.8%, most of the way to the 94.5% the reference replies themselves score.

50%60%70%80%90%100%200 MB500 MB1 GB2 GB5 GB10 GB20 GBdownload size, published weight file, log scalereplies read as Gulf, better is upReference replies · 94.5%Voho Saudi Chat 4B · 89.8%8.0 GB · 2.5 GB as Q4Earlier run, service calls only · 79.0%not releasedQwen3-4B, untrained · 62.3%8.0 GB
Fig. 2Share of replies read as Gulf against download size, 400 held-out questions; up is better. The dashed line is the ceiling: what the reference replies score. These are the models measured on this test so far.

Shorter is the point

Replies fall from 30.2 words to 6.4. The reference replies average 10.7, so the model errs short, which on a phone line is the side to err on: a caller can always ask for more.

chrF++ against the reference replies is flat, 13.0 for the base and 12.0 for the fine-tune. That is expected and not a regression: the model says the right kind of thing in the right register, not the same words as the reference. A dialect classifier measures that; a string-overlap metric cannot.

Words per reply
0102030words per reply10.7Reference30.2Qwen3-4B, untrained9.4Service calls only6.4Voho Saudi Chat 4B
Fig. 3Mean words per reply. Green is the shipped model.

Training

The base is Qwen3-4B-Instruct-2507 (Apache 2.0), trained with LoRA at rank 32 on every attention and MLP projection, for two epochs at a learning rate of 1e-4 with a cosine schedule, on one NVIDIA L4. The loss falls only on assistant turns, so each multi-turn dialogue teaches every reply in it.

Every training example carries the same short system prompt asking for Najdi at phone-call length. Keep it: without it the dialect is noticeably weaker.

Runs on a laptop

The full weights are 8.0 GB. The Q4_K_M build is 2.5 GB and runs locally with Ollama or llama.cpp.

100 MB200 MB400 MB800 MB1.6 GB3.2 GB6.4 GB8.0 GBsafetensors4.3 GBGGUF Q8_03.3 GBGGUF Q6_K2.9 GBGGUF Q5_K_M2.5 GBGGUF Q4_K_M
Fig. 4Download size of every published build, log scale.

Limitations

It is Najdi first: Hijazi and Khaleeji replies drift toward Najdi or formal Arabic. The classifier’s Gulf class also covers the UAE and Kuwait, so a high score means "reads as Gulf", not "reads as Riyadh". It is tuned for register and voice, not knowledge, so ground it with retrieval for facts. It writes without diacritics.

What the result establishes

A 4B model can be moved from 62% to 90% Gulf replies, and from paragraphs to phone-length turns, with two training runs on one GPU and an evaluation that never touched its own training data. The run that failed is part of the result: an aggregate can rise while the model gets worse at the job, and only reading the samples showed it.

Deployment-ready

Start your AI transformation today.

Sign up and build your first agent in the browser, or book a call if you would rather someone walked you through it. Most people do not need the call.

Start

$5 of credit, free

Granted when you sign up, about 70 minutes of live calls. No card to begin.

Then

Plans from SAR 109 a month

Starter puts your agent on your website, Business on a Saudi phone line. Cancel any month.

When you need it

Enterprise terms

Saudi data residency, an uptime SLA and on-premise deployment, on an agreement.