The wrong dialect, at the wrong length
Ask a general 4B model a question in Saudi and its reply reads as Gulf Arabic only 62.3% of the time. A fifth of its replies read as Levantine and 8% as Egyptian. They run to 30 words, three times what a person says in one turn on the phone.
For a voice agent both are failures. The wrong dialect tells the caller they are talking to software; the wrong length means the caller waits through a paragraph. Voho Saudi Chat 4B is trained to fix both, and to stay usable by any business: the base model and every dataset are Apache 2.0.
Measuring dialect with a classifier that took no part
Replies are scored by MARBERTv2, an independent Arabic dialect classifier that had no part in training. The metric is the share of replies it reads as Gulf (want high) and as Modern Standard Arabic (want low).
The classifier was checked before any training. Held-out Saudi reference replies read as Gulf 94.5% of the time, and Arabic instruction data read as formal Arabic 68% of the time, so it separates the two cleanly. The ceiling is 94.5%, not 100%: it does not call every genuinely Saudi sentence Gulf, and the reference column on every chart shows that ceiling.
The data: service calls and everyday talk
The training set mixes Voho’s own dialogues with two open datasets: 7,601 service-call dialogues across eight sectors (oil and gas, utilities, telecom, banking, government, healthcare, logistics, facilities), 5,551 everyday conversations, 202 Gulf-tagged rows from Arabic Aya and 1,181 instruction pairs from CIDAR. The test set is 400 held-out questions the model never trained on.
Every dialogue had to pass two checks: a Najdi lexicon filter (Saudi function words such as وش، تبي، الحين present, their formal equivalents absent) and the same MARBERTv2 classifier, with at least 60% of its replies reading as Gulf. CIDAR is capped at about a tenth of the training words, because uncapped it would teach the long formal answers this model exists to remove. Anything with markdown, tables or code was dropped: a speech model reads asterisks aloud.
The run that looked like a win
The first full run trained on the service calls alone. Its headline number rose 17 points, to 79.0% Gulf, and it looked like success. Two other numbers said otherwise: formal Arabic jumped to 11.2% of replies, and the samples showed every question being answered as a support ticket. Asked its opinion on a television for watching the final, it explained that the system supports the television and it could be bought in store.
The fix was coverage, not more dialect. The shipped run adds the everyday conversations (family, food, driving, occasions), 38% of the words in Voho’s dialogues. Gulf rises to 89.8%, formal Arabic falls to 0.8%, and the same question gets: والله فكرة، بس أنا ما أحب أشتري شي جديد.
Accuracy against download size
All three points are the same 8.0 GB download, so the comparison is at equal size. Training moved the share of Gulf replies from 62.3% to 89.8%, most of the way to the 94.5% the reference replies themselves score.
Shorter is the point
Replies fall from 30.2 words to 6.4. The reference replies average 10.7, so the model errs short, which on a phone line is the side to err on: a caller can always ask for more.
chrF++ against the reference replies is flat, 13.0 for the base and 12.0 for the fine-tune. That is expected and not a regression: the model says the right kind of thing in the right register, not the same words as the reference. A dialect classifier measures that; a string-overlap metric cannot.
Training
The base is Qwen3-4B-Instruct-2507 (Apache 2.0), trained with LoRA at rank 32 on every attention and MLP projection, for two epochs at a learning rate of 1e-4 with a cosine schedule, on one NVIDIA L4. The loss falls only on assistant turns, so each multi-turn dialogue teaches every reply in it.
Every training example carries the same short system prompt asking for Najdi at phone-call length. Keep it: without it the dialect is noticeably weaker.
Runs on a laptop
The full weights are 8.0 GB. The Q4_K_M build is 2.5 GB and runs locally with Ollama or llama.cpp.
Limitations
It is Najdi first: Hijazi and Khaleeji replies drift toward Najdi or formal Arabic. The classifier’s Gulf class also covers the UAE and Kuwait, so a high score means "reads as Gulf", not "reads as Riyadh". It is tuned for register and voice, not knowledge, so ground it with retrieval for facts. It writes without diacritics.
What the result establishes
A 4B model can be moved from 62% to 90% Gulf replies, and from paragraphs to phone-length turns, with two training runs on one GPU and an evaluation that never touched its own training data. The run that failed is part of the result: an aggregate can rise while the model gets worse at the job, and only reading the samples showed it.
More research