Languages / ja / 日本語

Japanese text-to-speech

Japanese is served by five engines: both Qwen3-TTS variants, Chatterbox Multilingual, TADA's multilingual build and Kokoro. Qwen3-TTS CustomVoice includes one natively Japanese preset speaker among its nine.

Read this first: VoiceLabs is pre-launch. The engines described here are what ships at launch, and the studio is in development. The site, accounts, the Pro plan with billing and our MCP server for AI agents are live today.

Japanese carries the same token-density cost as Chinese — roughly two and a half times a Latin character — so the script-aware chunk budget applies here too, and a long Japanese script is split more finely than its character count alone would suggest. This is the difference between a tail that gets spoken and a tail that vanishes with the request still reporting success. Kokoro offers five Japanese preset voices, its second-largest set after English, and is the light option if you are generating a lot of it.

The engines that speak Japanese

  • Qwen3-TTS
  • Qwen3-TTS CustomVoice
  • Chatterbox Multilingual
  • TADA
  • Kokoro

Every engine on that list is on the same plan — there is no per-engine surcharge and no tier that unlocks one. You choose per generation.

One caveat on TADA: it ships in two sizes and only the larger, multilingual build speaks Japanese. Its smaller and faster model is English-only, so a Japanese job has to be on the right one.

Voices you can start from

PRESET VOICES

Kokoro contributes 5 preset Japanese voices. Qwen3-TTS CustomVoice adds 1 natively Japanese preset speaker.

YOUR OWN VOICE

Chatterbox Multilingual clones from a short reference clip in all 23 languages, Japaneseincluded — which is the answer whenever no preset is close enough. Cloning requires that the voice is your own or that you have the speaker’s explicit permission, on every plan.

Long Japanese scripts

Japanese is written in Japanese script, which costs more tokens per character than Latin script does. That matters because engines have a token ceiling they do not expose, and a splitter that counts characters overshoots it — the generation stops mid-sentence, the rest is discarded, and the request still reports success. Our splitter prices each chunk in estimated tokens with a per-script cost, so a long Japanese script is split more finely than its character count alone would suggest, and it completes.

One plan, every engine and every language: $8 a month for 2 hours of generated audio, with a 7-day free trial.

Engine coverage on this page is pinned to the studio’s own language map by a cross-package test, so it cannot drift from what you are actually offered. See the engines page for what maintaining them involves, and audio rights for what you may do with the result.