For voice AI and ASR teams
Brazilian Portuguese, the way it is actually spoken.
Voice recordings across regions, ages and accents, from verified contributors, delivered with transcripts and speaker metadata.
Request a speech sampleShared under NDA once we have your spec
Utterance recordIllustrative
- Region
- Nordeste · Ceará
- Speaker
- 35–44 · female
- Style
- Spontaneous
- Environment
- Quiet indoor
- Transcript
- Included, reviewed
Dataset spec
What is in a delivery.
Speaker mix, speaking style and prompts are set per project and written into the order.
- Language
- Brazilian Portuguese, native speakers
- Coverage
- Regions, ages and accents, balanced to your spec
- Speaking style
- Read and spontaneous. Two-speaker conversation in pilot
- Audio
- 48 kHz · 16-bit PCM · WAV, mono, unprocessed
- Devices
- Contributors’ own smartphones, iOS and Android, recorded in the browser
- Transcripts
- Included with every recording
- Speaker metadata
- Region, age band, gender · self-declared accent, device model, operating system
- Noise labelling
- Quiet, moderate or noisy, with estimated SNR per clip
- Consent
- Signed per contributor
- Monthly capacity
- About 400 h (100 h per week), growing with each pilot cohort
Regional coverage
Five regions, very different Portuguese.
A model trained on São Paulo speech struggles in Recife or Belém. We recruit by region so the accents your users have are in the set.
Norte120registered contributors6 of 7 states
Nordeste480registered contributors9 of 9 states
Centro-Oeste160registered contributors4 of 4 states
Sudeste1,060registered contributors4 of 4 states
Sul180registered contributors3 of 3 states
Verified speakersIdentity checked on Verakin. One person, one profile, declared region and age band.
Audio checksNo clipping above 0.1% of samples, under 30% silence, SNR above 15 dB.
Transcript reviewA native reviewer checks every spontaneous transcript word by word. Target: 98% word accuracy.
Voice is biometric dataConsent signed per contributor. Usage rights and retention defined per project: voice is never used to identify the speaker, and speaker IDs are pseudonymized per delivery.
Hear it before you commit.
Send us your spec: the regions, speaker mix and style you need. After a mutual NDA, we send a sample with transcripts and a project quote.