Chinese to English Voice Translator — Mandarin & Cantonese Audio to Text
Upload Chinese audio (MP3, WAV, M4A) or video — Mandarin (Standard Chinese) works at Tier 1 quality via Whisper Large-v3. Cantonese output usable as first-draft (honest scope — Whisper trained primarily on Mandarin). Set target to English, export TXT/DOCX/SRT/VTT with timestamps and speaker labels. Also handles reverse direction (English → Simplified or Traditional Chinese) and WeChat voice messages. Free 30 minutes on signup.
How to translate Chinese audio to English: upload Chinese audio (MP3, WAV, M4A) or video (MP4, MOV). Auto-detect handles Mandarin (Whisper Tier 1, ~8% WER on FLEURS-ZH per Radford et al. 2022, arXiv:2212.04356). CANTONESE — Whisper trained primarily on Mandarin; Cantonese output treated as first-draft only. Set target to English or 132 others. Export TXT, DOCX, SRT, VTT. Reverse direction outputs Simplified Chinese by default (Traditional available). Free 30 min, no card.
TL;DR — Chinese → English workflow
- 1.Upload Chinese audio or video. WeChat .amr/.m4a supported. Max 5 GB.
- 2.Auto-detect handles Mandarin. Override for Cantonese, Taiwanese, or regional varieties.
- 3.Export. TXT, DOCX, SRT, VTT, JSON with timestamps + speaker labels.
Text vs dubbed audio. This tool outputs English text. For AI-dubbed English audio (voice-cloned Chinese speaker), use ElevenLabs, HeyGen, or Rask AI — different job.
How to translate Chinese audio to English — 4-step walkthrough
Step 1 — upload Chinese audio or video
Audio: MP3, WAV, M4A, OGG, FLAC, AMR (WeChat exports), OPUS. Video: MP4, MOV, WebM, MKV. Up to 5 GB.
Step 2 — auto-detect handles Mandarin (or set Cantonese/regional manually)
Mandarin auto-detects reliably. Override manually for Cantonese (known Whisper weakness), Taiwanese Mandarin, or regional varieties (Shanghainese, Hokkien, Sichuanese).
Step 3 — target = English (or 132 others)
Standard tier uses Google Translate (free); Premium tier uses our internal translation model for professional-grade output.
Step 4 — export
TXT/DOCX for text, SRT/VTT for video subtitles, JSON for downstream processing.
Chinese Voice Translator, Speech Translator, Audio Translator — Same Tool, Three Framings
For Chinese, “voice translator” is the dominant search framing (720/mo primary term). If you searched “chinese to english voice translator”, “translate chinese speech to english”, or “translate chinese voice to english” — you're in the right place.
| Framing | User context | Workflow |
|---|---|---|
| Voice translator | WeChat voice notes, quick spoken messages | Upload .amr/.m4a, get English text in under a minute |
| Speech translator | Chinese business calls, cross-border partnerships, interviews | Upload conversation recording with speaker labels |
| Audio translator | Long files — podcast episodes, lectures, media broadcasts | Upload MP3/WAV up to 5 GB, chunked processing |
Mandarin, Cantonese, and Other Chinese Varieties — Honest Scope
“Chinese” describes multiple spoken languages, not a single language. Whisper Large-v3 was trained primarily on Mandarin. Here's the honest per-variety breakdown so you can set expectations correctly.
Mandarin (Standard Chinese, 普通话 / Pǔtōnghuà) — Tier 1
Whisper training default. FLEURS-ZH benchmark ~8% WER on clean broadcast Mandarin. Best-quality output. Covers Mainland China, Taiwan (Guoyu 国语), Singapore (Huayu 华语), diaspora Mandarin. Taiwan Mandarin (Guoyu) has minor vocabulary differences but transcribes reliably.
Cantonese (广东话 / Gwóngdūng wá) — KNOWN WEAKNESS (Tier 2 first-draft)
Honest scope: Cantonese is a separate spoken variety from Mandarin, with distinct phonology, tonal system (6-9 tones vs Mandarin's 4), and vocabulary. Whisper Large-v3 was not specifically trained on Cantonese — output is usable for content understanding and first-draft translation but expect noticeably lower accuracy than Mandarin.
Practical recommendation: most educated Cantonese speakers also speak Mandarin (Standard Chinese is taught in Hong Kong schools). For business/media/formal contexts, record in Mandarin if possible. For pure Cantonese content (family recordings, Cantonese-only media, colloquial content), treat output as first-draft only or use a Cantonese-specialized transcription service.
Shanghainese (上海话) — Tier 3
Wu Chinese variety. Not specifically trained. Draft-quality only. Most Shanghai speakers also speak Mandarin — record in Mandarin if possible.
Hokkien / Taiwanese Hokkien (闽南语 / 台语) — Tier 3
Min Chinese variety, spoken in Fujian, Taiwan, Singapore, Malaysia. Not specifically trained. Draft-quality only.
Sichuanese (四川话) — Tier 2-3
Mandarin variant with distinct phonology. Modern educated speakers often use Standard Mandarin in media/business — that transcribes reliably. Heavy Sichuanese dialect drops to draft-quality.
Simplified vs Traditional Chinese Characters (For Reverse Direction)
For Chinese → English direction, output is English text — character choice doesn't apply. For reverse direction (English → Chinese), character choice matters:
- Simplified Chinese (简体中文): used in Mainland China, Singapore. Default output. ~1.4 billion readers globally.
- Traditional Chinese (繁體中文 / 繁体中文): used in Taiwan, Hong Kong, Macau, some diaspora communities. Available as post-processing option.
For most business content, Simplified is the safe default. For Taiwan/Hong Kong-specific market localization, Traditional matters. For diaspora communities, check with your audience — older diaspora often prefers Traditional; younger diaspora comfortable with Simplified.
Translate English Audio to Chinese (Reverse Direction)
Same workflow, target = Chinese. Output uses Simplified Chinese by default; Traditional Chinese available.
Common reverse-direction workflows:
- English brand video → Simplified Chinese for Mainland China market
- English content → Traditional Chinese for Taiwan / Hong Kong / Macau distribution
- English e-learning course → Chinese localization
- English YouTube video → Chinese subtitle overlay for Chinese-speaking subscribers
- English business call → Chinese meeting minutes for Chinese-speaking team members
Translate Chinese WeChat Voice Message to English
WeChat is the dominant messaging app in the Chinese-speaking world (1.3+ billion monthly active users) and diaspora communities. Voice notes are common — quick spoken messages that convey more than text would.
3-step WeChat voice workflow
(1) In WeChat, long-press the voice message → Save (varies by WeChat client version and iOS/Android). Exports as .amr or .m4a. (2) Upload to VexaScribe — source auto-detects as Chinese (Mandarin), target = English. (3) Get English text back in under a minute.
Common contexts:
- Chinese-American / Chinese-Canadian diaspora processing voice notes from family in Mainland China
- Business contacts in Shanghai / Beijing / Shenzhen sending voice updates
- Cross-border teams coordinating with Mainland China / Taiwan / Hong Kong offices
- Language learners recording Chinese exchange partners
Quality note: WeChat compresses voice notes with AMR codec (~12 kbps) — lower fidelity than WhatsApp Opus. Expect 5-10% higher WER than clean studio Mandarin, but message content typically comes through clearly on standard speakers.
Tonal Audio Quality Impact on Chinese ASR
Mandarin uses 4 tones (plus neutral tone) — tone contours carry lexical meaning. "mā", "má", "mǎ", "mà" are four different words (mother, hemp, horse, scold). Tonal ASR requires clean audio to distinguish tones reliably.
Quality impact by audio type
- Clean studio recording: Tier 1 quality (~8% WER on FLEURS-ZH)
- Broadcast media (CCTV, TVB, etc.): Tier 1 quality
- Zoom / Teams / phone recording: +5% WER due to compression
- WeChat / WhatsApp voice notes: +5-10% WER due to codec compression + typical mobile-mic noise
- Background music / karaoke / music video: +20-40% WER — music obscures tone contours
- Distant microphone / room echo: +10-20% WER
Translate Chinese Video to English (Video Files)
Video files (MP4, MOV, WebM, MKV) work directly. Common Chinese video → English use cases: Chinese YouTube / Bilibili (哔哩哔哩) / Douyin (抖音) content localization for English audiences, C-drama fan subtitle work, Chinese business video content research, Chinese academic conference recordings. For the full video-flow architecture, see translate video to text.
Supported audio and video formats
Audio
MP3, WAV, M4A, OGG, FLAC, AAC, AMR (WeChat), OPUS. WeChat .amr and .m4a exports supported directly.
Video
MP4, MOV, WebM, MKV, AVI, WMV, FLV. Audio extracted server-side. Max 5 GB.
Free vs Premium Chinese translation
| Feature | Free tier | Premium tier |
|---|---|---|
| Transcription | Whisper Large-v3 (Mandarin Tier 1) | Proprietary higher-accuracy model |
| Translation | Google Translate | Our internal translation model |
| Cost | 30 min free + unlimited translation on all plans | 1 credit per 5,000 chars for translation upgrade |
| Best for | Creator, internal team, first-draft | Client-facing, publication, broadcast |
Common Chinese → English use cases
Cross-border business calls
Sourcing, manufacturing, tech partnerships with Mainland China / Taiwan / Hong Kong / Singapore contacts.
Chinese diaspora family recordings
Adult children processing voice notes/videos from Chinese-speaking relatives.
Chinese-language content research
State media (CCTV, CGTN), interviews, academic sources, policy briefings for English research.
Chinese YouTube / Bilibili / Douyin localization
Content creators translating Mandarin videos for English audiences via SRT overlay.
C-drama / anime fansub drafting
Fan-community translation drafts. Ethical note: fansub for public release requires copyright consideration.
Chinese language learning practice
HSK preparation, immersion learners recording native-speaker exchange partners.
When to use another tool
Cantonese at broadcast quality
Hire a Cantonese-specialized transcription service. Whisper is not reliable for pure Cantonese at broadcast quality.
Dubbed English audio (voice-cloned Chinese speaker)
Use ElevenLabs, HeyGen, or Rask AI. Different output type.
Publication-grade Chinese ↔ English translation
Hire a professional bilingual translator. Machine translation is first-draft only for high-stakes work.
Live Chinese → English conversation
Use Google Translate mobile voice mode or Baidu Translate. This is a batch tool.
Frequently asked questions
How to translate Chinese audio to English?
Upload Chinese audio (MP3, WAV, M4A) or video (MP4, MOV). Auto-detect handles Mandarin (Whisper Tier 1, ~8% WER on FLEURS-ZH). CANTONESE — Whisper trained primarily on Mandarin; Cantonese output treated as first-draft. Set target to English. Export as TXT, DOCX, SRT, VTT. Output defaults to English text; character output for reverse direction (English → Chinese) defaults to Simplified. Free 30 min on signup.
Does it work with Cantonese?
With honest caveats. Whisper Large-v3 was trained primarily on Mandarin (Standard Chinese) — Cantonese (广东话) is a separate spoken variety with distinct phonology and tonal system. Output is usable for content understanding and first-draft translation, but expect noticeably lower accuracy than Mandarin. For broadcast-grade Cantonese work, use a Cantonese-specialized transcription service. Practical recommendation: if your source speaker can switch to Mandarin (which most educated Cantonese speakers can in business/media contexts), record in Mandarin for best quality.
Does it output Simplified or Traditional Chinese?
For reverse direction (English → Chinese), output defaults to Simplified Chinese (used in Mainland China, Singapore). Traditional Chinese (used in Taiwan, Hong Kong, Macau) available as post-processing option. For Chinese → English direction, output is English (Simplified/Traditional character choice doesn't apply since output is English text).
How accurate is Mandarin → English?
Tier 1 on clean Mandarin audio. Whisper Large-v3 on FLEURS-ZH: ~8% WER on clean broadcast Mandarin. Neural translation on Chinese↔English: 20-30 BLEU on standard WMT test sets. Combined pipeline: ~82-88% word-level accuracy on clean Mandarin audio. Drops on: tonal clarity issues (background music, phone call quality, room ambience) — Mandarin's 4-tone system requires clean audio; regional accents (Beijing, Shanghai, Taiwan) with variance; music-heavy audio; multiple simultaneous speakers.
Can I translate WeChat voice messages?
Yes. WeChat is the dominant messaging app in Chinese-speaking world and Chinese diaspora communities. WeChat exports voice notes as .amr or .m4a files. Long-press the voice message → Save (varies by WeChat client version). Upload to VexaScribe — source auto-detects as Chinese, target = English. Common workflow: Chinese diaspora processing voice notes from family in Mainland China / Taiwan / Hong Kong.
Regional Chinese (Shanghainese, Hokkien) supported?
Tier 3 (draft-quality only). Shanghainese (上海话), Hokkien / Taiwanese Hokkien (闽南语/台语), Sichuanese (四川话), and other regional Chinese varieties are separate spoken languages from Mandarin — Whisper Large-v3 was not specifically trained on these. Output is usable for rough understanding but treat as draft only. Most speakers of these varieties also speak Mandarin — if possible, record in Mandarin for reliable transcription.
Can I translate English audio to Chinese?
Yes. Reverse direction: upload English audio, source auto-detects as English, target = Chinese (Simplified default; Traditional available). Output uses standard written Chinese. Common workflow: English brand video → Simplified Chinese for Mainland China market; English content → Traditional Chinese for Taiwan / Hong Kong market.
Is Chinese audio translation free?
Yes on the free tier — 30 minutes of transcription + translation on signup, no credit card. Translation included on all plans. Premium tier (higher-accuracy internal model) available at 1 credit per 5,000 characters for client-facing content.
What audio formats work with Chinese?
MP3, WAV, M4A, OGG, FLAC, AAC, AMR, OPUS all supported. WeChat .amr and .m4a exports work directly. Chinese business audio commonly comes as MP3 or M4A from smartphone recordings; conference calls as WAV or M4A. Max 5 GB per file.
Does tonal audio quality affect accuracy?
Yes significantly. Mandarin uses 4 tones (plus neutral tone) — tonal distinctions carry lexical meaning. Poor audio quality (background music, phone compression, room echo, distant microphone) obscures tone contours and degrades ASR accuracy. Expect 5-15% WER increase on tonally-unclear audio. Best quality: close-mic studio recording or clean broadcast audio.
Can I translate Chinese video to English?
Yes. Upload MP4, MOV, WebM, MKV — VexaScribe extracts audio, transcribes Chinese, translates to English. Export SRT for English subtitles on original Chinese video. Common use case: Chinese YouTube / Bilibili / Douyin content localization for English audiences, C-drama fan subtitle work, Chinese business video content research.
Chinese business call transcription workflow?
Upload the recording (usually MP3 or M4A from phone/Zoom/Teams), source auto-detects as Chinese (Mandarin), target = English. Speaker diarization identifies participants across the call. Export as DOCX with speaker labels for meeting minutes, or SRT if you want timestamped review. Common use case: sourcing/manufacturing/tech calls with Mainland China contacts, cross-border partnership discussions.
Related tools
Translate Audio to Text (Hub)
General audio → translated text in 133 languages.
Translate Japanese Audio to English
Sibling Asian-language page — speech framing.
Translate Video to Text
Video-input flow (Chinese YouTube, Bilibili, Douyin).
Subtitle Translator
Already have Chinese SRT? Translate the file directly.
Translate Spanish Audio to English
Sibling language-pair page — biggest cluster.
How Accurate is Whisper
Chinese WER data + per-language accuracy benchmarks.
Sources
- Radford, A. et al. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356 — Whisper Large-v3 Mandarin WER (~8% on FLEURS-ZH). Cantonese not in FLEURS test set — inferred first-draft-only scope.
- WMT Chinese↔English test sets — neural translation benchmarks (20-30 BLEU on standard test sets).
- Google Cloud Translation supported languages — 133-language coverage (verified August 2026).
Page reviewed and accuracy figures verified August 12, 2026.