Voho wins a landmark enterprise contract

Questions

The things people ask before they commit.

Dialects, where the data lives, whether it runs inside your own building, what the API does and does not do, and what it costs. Answered the way we would answer them on a call.

01

What it is

Whether this is a different thing from the voice bots you have already been pitched.

01Why not just buy a voice service and connect it ourselves?

You can, but a voice service only speaks. It does not know your customers, it cannot open a ticket in your system, and it will not hand a difficult call to one of your staff with the full story attached. That connecting work is the part that takes months, and it is the part we bring already built.

02Can we see it working before we talk to anyone?

Yes. Every demo runs in your browser at voho.ai/demos with no sign-up: a Saudi Arabic service call, an invoice being read, an engineering archive being searched, building systems being joined up. There is also open example code on GitHub if you would rather read it than watch it.

03Is this a chatbot?

No. A chatbot answers questions. These agents complete the task: they look up the record, raise the request, update the system and read back the reference number. Where your rules say a person must approve something, the agent stops and asks.

04What happens when the agent cannot handle a call?

It hands over to a person, and the transfer carries the caller identity, the summary, the detected intent and every action already taken. The caller never repeats themselves. Deciding when to hand over is a rule you set, not something the model decides on its own.

02

Arabic and dialects

The dialect question, which in this market is the first question.

01Does it really speak Saudi Arabic, or Modern Standard with an accent?

Najdi (the dialect spoken in Riyadh and central Saudi Arabia) is what the voices are built around, not a Modern Standard voice with the edges filed off. Gulf, Egyptian and Modern Standard are available too, and Modern Standard remains the safest choice when you do not know the caller.

02Our customers switch between Arabic and English mid-sentence. Does that break it?

No. Code-switching inside a single sentence is normal in the Gulf and is handled as normal, not as an error. English on its own is fully supported.

03How does it read numbers, dates and reference codes?

Reference numbers are read digit by digit, and there is a text normalisation step that decides how a number, a time or a Hijri date is spoken before it is synthesised. Those are exactly the places a synthesiser guesses wrong, so they are handled deliberately rather than left to chance.

03

Deployment and data

Where it runs, where the data sits, and what your auditors will ask for.

01Where does our data actually live?

Wherever you require. Residency can be pinned to Saudi Arabia or the EU, and everything can run in a private deployment or on your own servers, inside your own building, including with no internet connection at all. It is one of the two questions that gets asked first in this market, so it is a first-class option rather than an exception.

02Can it run with no connection to the outside world?

Yes. In a fully self-hosted deployment the models, the call handling and the records all sit inside your network. Nothing leaves unless you decide a specific piece of work is allowed to.

03How long does it take to go live?

The first workflow is usually live within a month. We build it with you, on your own systems, rather than handing over a platform and a login. Enterprise engagements include a named solutions engineer and a guided 30-day implementation.

04What about PDPL and audit?

Every call produces a record: the turn-by-turn transcript, a summary, and every action the agent took. Audit logs can be streamed to your own SIEM, access is role-based with SSO and SCIM, and PII redaction and retention controls are part of the platform rather than something bolted on afterwards.

04

For engineers

What you get out of the box, what you can take piecemeal, and how it meets the phone system you already have.

01Do we have to bring our own speech-to-text and language model?

No. A Voho agent does the whole loop: it hears the caller in Saudi Arabic, works out what they actually want, takes the action in your systems, stops talking when it is interrupted, hands over to a person when it should, and writes a bilingual transcript and summary at the end. You bring the systems it should act in, not a speech stack assembled from three vendors.

02What if we only want one piece of it?

Then take one piece. There are two ways in: a complete voice agent you configure rather than build, or the Speech API on its own: synthesis over HTTP, a streaming variant, a WebSocket that accepts text while returning audio on the same connection, fourteen voices and Arabic text normalisation. Teams that already have a working call stack and only need a voice that sounds Saudi take the second. Documentation for both is at docs.voho.ai.

03How do actions work?

An action is something the agent does in the world rather than in the conversation, opening a ticket, fetching an order, updating a record. You give it a name, a description of when to use it and the handful of parameters it needs, and the agent decides when the moment has arrived. Actions are what separate an agent from an expensive answering machine, and the return value is written to be read aloud: a short reference number the caller can actually write down.

04Will it work with the phone system we already have?

Yes. Audio is available as 8 kHz mulaw, which is what Cisco, Avaya and SIP trunks already carry, so there is no transcoding step on your side. No new numbers and nothing ripped out.

05What does it connect to on the back end?

SAP, Oracle, ServiceNow, Archibus, core banking, scheduling and helpdesk systems: through their own APIs, with the agent authenticated as a service account whose permissions you control. We sit on top of the systems you have rather than replacing them.

06How fast does it respond on a live call?

Audio starts before the sentence has finished being written. The streaming endpoints exist for exactly that reason: an LLM emits tokens, the first ones are already being spoken while the rest are still being produced. The number worth measuring is time to first audio, and it is the one we report.

05

Commercial

Commercial terms are scoped to call volume, dialect coverage and deployment model. The full breakdown is on the pricing page.

01How does pricing work?

Two models. Pay as you go is billed per minute with no seats, no platform minimum and no annual commitment. Enterprise adds a contractual uptime SLA, deployment control, SSO and SCIM, audit log streaming and a named team. Voice minutes are billed at cost from your chosen speech provider with no markup from us, and telephony is quoted separately by region.

02Do we have to commit before we can try it?

No. The demos need no sign-up, the example code needs only an API key, and pay as you go has no annual commitment. The commitment conversation belongs after you have seen it work on your own traffic.

Still unanswered

Ask us the awkward one.

The technical detail, the compliance question, the thing your last vendor could not do. That is a better first call than a product walkthrough.