Voho wins a landmark enterprise contract
Models and training

An Arabic AI stack, trained on how Saudi Arabia actually speaks.

Speech to text, text to speech and the language layer — selected, tuned and measured for Saudi dialects rather than adapted from English defaults, alongside our own Arabic models in development. Available as a running deployment, or as custom model training on your own calls with the dataset and the evaluation harness left in your hands.

The stack

Four layers, each measured separately.

Most Arabic voice deployments fail in one specific layer while the team argues about another. We instrument all four, so the problem is located rather than guessed at.

Speech to text

Arabic ASR tuned for how the Kingdom actually speaks

Najdi, Hijazi and broader Khaleeji speech over narrowband telephony, with Arabic-English code-switching inside a single utterance. Tuned against the entities that decide a call: Saudi mobile numbers, SAR amounts, national ID and iqama digits, compound Arabic names, districts and your own product vocabulary.

  • Streaming partials with endpointing tuned so callers are not cut off
  • Custom vocabulary and keyword boosting per deployment
  • Constrained parsing for digits, IDs and dates
Text to speech

Saudi voices that do not sound like a newsreader

Voice selection and tuning for register as well as accent, because a Modern Standard delivery reads as distant on an everyday service call. Pronunciation is controlled rather than guessed — diacritics and lexicon entries for names, places and loanwords instead of hoping the model infers the right word.

  • Number, currency, Hijri and Gregorian date handling verified line by line
  • Streaming synthesis measured on time-to-first-byte, not clip duration
  • Custom and cloned voices under written, revocable consent
Language layer

The model as a component with a measured job

Frontier, regional Arabic-first and open-weight models sit behind one thin interface, so the choice is an evidence question rather than a commitment. We measure what actually decides a deployment: dialect comprehension, tool-calling accuracy under ambiguity, instruction adherence in Arabic, and cost per completed call.

  • Model swappable as benchmarks change, without rebuilding the deployment
  • National and regional Arabic models — ALLaM and others — evaluated on your transcripts
  • Robustness tested against the transcription errors your ASR really makes
Voice agents

The assembled system, operated

The three layers above only matter as a working call. Voho runs the assembled agent on your lines — escalation policy, context-preserving transfer to your team, integration with your CRM and scheduling, and a weekly review of real recordings rather than dashboards.

  • Inbound and outbound, with local number provisioning
  • Context carried through to the human on transfer
  • Named owner reviewing live calls every week
Custom model training

Train on your calls. Keep the dataset.

Generic Arabic models are trained on written Modern Standard text and clean speech. Your callers are neither. Training on your own audio is what closes that gap — and the resulting dataset is an asset that stays with you.

Data curation and labelling

Saudi dialect audio collected, transcribed and reviewed by native speakers, built into a dataset that belongs to you. The scarce asset in Arabic AI is not compute — it is correctly labelled dialect data from your own domain.

Dialect and domain fine-tuning

Speech recognition fine-tuned on your recorded calls so the model learns your vocabulary, your accents and your failure cases. Typically the largest single accuracy gain available, and it compounds as your call archive grows.

Custom voice development

A brand voice designed or cloned for your organisation, tuned for Saudi register and validated by blind listening tests with native speakers on real telephony audio — not on studio playback.

Language model adaptation

Adaptation of the language layer to your products, policies, tone and tools — through retrieval, structured prompting or fine-tuning, chosen on measured results rather than on preference.

Evaluation harness

A benchmark built from a hundred of your own calls with reference transcripts, re-run whenever a provider ships a model update. You keep the harness; it is how you hold every vendor, including us, to account.

Deployment and handover

Models served where your compliance team has approved, with documentation and training so your own people can operate, interpret and change the system without a support ticket.

How a training engagement runs

Evidence first, training second.

01Baseline

Benchmark the current stack on your real calls. No training work is proposed until we can show where it fails and by how much.

02Dataset

Curate and label Saudi dialect data from your domain, with a native-speaker review pass and a held-out evaluation set.

03Train

Fine-tune the layer the baseline identified — usually speech recognition first, since its errors propagate into everything downstream.

04Prove

Re-run the harness against the held-out set and report the delta honestly, including where the tuned model is no better.

05Operate

Deploy, then keep measuring. Models drift as your products, prices and providers change.

Deployment and residency

Run it where your compliance team can approve it.

Where audio, transcripts and derived data are processed is answered per data type, with retention periods and a documented deletion path for each. Those answers are prepared for your security review rather than assembled during it.

PDPL-ready data handling

Managed by Voho

We run the stack and the operations. Fastest path to live traffic, and the right default for most first deployments.

Your cloud, in-Kingdom

Deployed into your own tenancy in a region your compliance team has approved, with data residency answered per data type rather than as a blanket statement.

Private, on-premise or air-gapped

In your own data centre, down to fully air-gapped with no outbound connectivity, for regulated workloads that cannot leave your estate. Open-weight models make this a real option rather than a slide — sized to your latency and concurrency rather than to the largest checkpoint available.

National alignment

Built the way the National Strategy for Data & AI asks for.

The strategy's skills pillar expects capability to sit in-Kingdom, its research pillar expects contribution rather than consumption, and its ecosystem pillar backs local suppliers. Training on Saudi data, handing over the dataset and the evaluation harness, and leaving your team able to operate the system is what that looks like in practice.

Read our guide to the National Strategy for Data & AI
Deployment-ready

Start your AI transformation today.

Launch voice agents with the operational rigour your buyers expect — and the deployment speed a startup actually needs.

Onboarding

Live in 30 minutes

Guided setup with a solutions engineer on the call.

Trial

7 days, all features

Full platform access. No feature gates, no sales gate.

Guarantee

30-day outcome

Not earning voice AI revenue in 30 days? We onboard your first client with you.