Privacy Policy
Version v1 — June 2026
The short version
- We collect what we need to operate Inzwi and build the open Shona corpus — your contributions, a small amount of profile metadata, and basic technical info.
- We never sell personal data. We don't run ad networks.
- Contributions are shared in open datasets and used for model training (including commercial models, per our Terms). Personally identifiable fields are stripped or hashed in those public releases.
- Your voice goes into AI training. We don't try to clone individual speakers, but a trained model can produce synthetic voices that share some characteristics with yours.
- You can ask us to delete a contribution before it's been published in a dataset or used to train a released model.
1. Who we are
Inzwi is a Zimbabwean-led open language project. The data controller for the purposes of this Policy is the Inzwi project team, contactable at hello@inzwi.app.
2. What we collect
2.1 Account & onboarding
- Username and (for web accounts) email address
- Password hash (never the plain password) for password-based accounts
- OAuth identifier for Google sign-in, if you use it
- WhatsApp phone number (when you contribute via WhatsApp)
- A separate airtime payout phone number if you provide one
- Display name, optional bio, dialect, age range, gender (gender is optional and only used to balance dataset coverage)
- Your acceptance of these Terms and the version you accepted
2.2 Contributions
- Text you submit — translations, transcriptions, validation comments
- Audio you record — prompted read-aloud and free-form clips
- Validation votes and ratings on others' contributions
- Derived audio metadata (duration, sample rate, SNR, clipping, silence ratio)
2.3 Technical info
- Browser, OS, and device type (for security and quality scoring)
- IP address — used for security audit logs and approximate location only
- WhatsApp message metadata (message id, timestamps) from the Meta Cloud API
2.4 What we don't collect
We don't use third-party advertising or analytics SDKs that track you across sites. We don't collect contacts, photos, location history, or anything from your device beyond what you submit.
3. How we use it
- To run the platform — accounts, leaderboards, validation routing, points, streaks, notifications;
- To build the Inzwi Shona corpus and derived datasets;
- To train, fine-tune, evaluate, and release speech recognition, text-to-speech, and language models for Shona and other Zimbabwean languages;
- To send transactional messages (sign-in security, password resets, nudges you've opted into);
- To distribute airtime / data-bundle rewards to top contributors;
- For research, reporting, and grant applications, with anonymised aggregates where contributor identities aren't necessary.
4. Voice recordings — what to know
Voice is the core artefact Inzwi exists to collect. When you submit a recording:
- The audio file is stored on Inzwi infrastructure and backed up to a private Hugging Face dataset repo;
- It is used as ASR / TTS training data, both internally and in open dataset releases;
- Once a recording has been used in a model that has been released (or in a dataset that has been published), removal is not always technically possible — we will still remove it from internal copies and future releases on request;
- Trained models may produce synthetic voices that resemble general Shona speaker characteristics. We don't build per-speaker voice clones, and we don't publish recordings tied to your personally identifiable information.
5. How we share data
5.1 Open dataset releases
We release the corpus under permissive open licences (e.g. CC BY-SA 4.0) so researchers and developers can use it. In those releases, contributor identifiers are hashed and contact info (email, phone) is stripped. Voice clips are included with linguistic metadata (dialect, declared gender if provided, content type, quality scores) but not with names or contact details.
5.2 Service providers
- Meta WhatsApp Cloud API — to send/receive messages on our WhatsApp bot;
- Hugging Face — to host the dataset (private repo during development, public for releases);
- VPS / cloud hosting for the platform itself;
- Email provider for transactional email;
- Airtime / mobile network APIs to disburse rewards.
These providers process data on our behalf under their own privacy terms. We only share what each provider needs to do its job.
5.3 Commercial use of the corpus
As described in Section 4 of the Terms, the corpus and models trained from it may be used commercially, including by third parties under the open licence. Personally identifiable contributor data is not part of those commercial assets.
5.4 We don't sell personal data
We don't sell email addresses, phone numbers, names, or any contributor-identifying information.
6. Data retention
- Account data is kept while your account is active;
- Contributions are kept for the lifetime of the project — they're the dataset;
- Audit logs (security events) are kept for 12 months;
- Encrypted database backups are rotated daily for 30 days, then monthly for 18 months.
7. Security
Passwords are hashed with bcrypt. The production database is not exposed to the public internet. Sessions are JWT-based with rotation on suspicious activity. Backups are encrypted in transit. We don't pretend any system is perfectly secure — if you spot a vulnerability, please report it to hello@inzwi.app.
8. Your rights
You can:
- See what we hold about you — most of it is visible in Settings;
- Correct profile data via Settings, or by emailing us;
- Request deletion of your account and not-yet-published contributions;
- Opt out of non-essential email (activity nudges, marketing) from Settings;
- Withdraw your consent to future use — by contacting us. Past use of Contributions already incorporated into released datasets or trained models cannot always be reversed.
If you're in a jurisdiction with formal data-protection rights (e.g. GDPR, POPIA), those rights apply to you and you can exercise them by contacting us.
9. International transfers
Inzwi infrastructure is hosted on European-region cloud providers and on Hugging Face. If you contribute from outside those regions, your data is processed there. By contributing you consent to that processing.
10. Children
Inzwi is intended for contributors aged 16+. Children aged 13–15 should only contribute with active parent or guardian consent. We don't knowingly accept contributions from anyone under 13.
11. Changes
We'll update this Policy as the project evolves. Material changes are surfaced in the app — and for changes that affect how your data is used, we'll ask you to acknowledge the new version before your next contribution.
12. Contact
Privacy questions, data-subject requests, security disclosures — hello@inzwi.app.