Inzwi

About

Inzwi means voice.

We build speech technology for Zimbabwean languages — starting with Shona and Ndebele, and starting with the data, because that is the part nobody else was going to build.

Why we exist

Sixteen million speakers,
and nothing that listens.

A Shona speaker can talk to almost no software in their own language. Neither can a Ndebele speaker, and their position is worse — where Shona at least appears in the open speech datasets, Ndebele largely does not appear at all.

The barrier is not model architecture. It is that the speech needed to train on does not exist in a usable form: not enough of it, not consented, not labelled by dialect, and not recorded from the conversational, code-switched speech people actually use.

So we build the corpus and the models together, and the corpus comes first.

What we make

One stack, built in the order that makes it real.

Recognition

Speech to text for Shona and Ndebele, with the language and dialect identified as it transcribes, and code-switched English handled as ordinary speech rather than as an error.

Synthesis

Text to speech built from studio-captured recordings, with an expressive track so a generated voice can laugh, breathe and hesitate instead of only pronouncing.

Understanding

Language models tuned for Zimbabwean languages, so a system can hold a conversation in the language it heard rather than round-tripping through English.

The corpus

The foundation the rest stands on: consented, reviewed, dialect-tagged speech, collected continuously through WhatsApp and the web from contributors who are paid for their work.

How we work

Four things we do not compromise on.

The data decides everything

Model architectures are shared and improving in public. What is not shared is speech that reflects how a language is really used — across its varieties, mixed with English, recorded by people who speak it natively. That is the part worth building, and it is the part we own.

Variety, not just language

Reporting one accuracy number for "Shona" hides which Shona. Every clip we hold carries its dialect, so we can say where a model is strong and where it is thin — and so a speaker from Masvingo or Chipinge is not treated as a rounding error.

Built and governed here

Contributors are Zimbabwean and paid. Consent is explicit and revocable. The corpus is held under the Cyber and Data Protection Act [Chapter 12:07]. Zimbabwean speech is not a resource to be extracted and processed elsewhere.

Honest about what exists

We say what is built, what is in progress, and what is not started. A partner who finds the gap themselves has learned something about us that no benchmark can repair.

Get in touch

Talk to us

Partnerships, API access, research collaboration, or a question about how the corpus is collected.