Pocket TTS Turkish

Pocket TTS Turkish is a 6-layer Turkish text-to-speech model built on Kyutai's Pocket TTS. It runs faster than real time on a single CPU thread, clones a voice from a few seconds of audio and supports five emotion tags.

  • About 110 million parameters (codec included), 440 MB download, 24 kHz mono output
  • About 5 times faster than real time on one CPU thread, about 20 times on a GPU
  • Six built-in voices, or clone a voice from a short recording
  • Emotion tags: mutlu (happy), üzgün (sad), kızgın (angry), şaşkın (surprised), sakin (calm)
  • Released under CC-BY-4.0

Try it in the browser: Pocket TTS Turkish demo. Python package and command line tool: pocket-tts-turkish on GitHub.

Voices

Each voice reads: "Merhaba, ben bir Türkçe konuşma sentezleme modeliyim. Metinleri bilgisayarınızda, internet bağlantısı olmadan seslendirebilirim."

voice character audio
female_1 bright and energetic (default)
female_2 warm, clear and friendly
female_3 confident, clear articulation
male_1 calm and deep, narrator style
male_2 energetic and bright
male_3 confident, crisp articulation

The built-in voices are synthetic. They were designed from text descriptions and are not recordings of real people.

Usage

With the pocket-tts-turkish package (recommended)

pip install pocket-tts-turkish
from pocket_tts_turkish import TurkishTTS, save_wav

tts = TurkishTTS.from_pretrained()  # downloads this model once
audio = tts.generate("Randevunuz 15.10.2026 saat 14:30'da. Ücret 1.250 TL.", voice="female_1", emotion="sakin")
save_wav("randevu.wav", audio, tts.sample_rate)

voice = tts.voice_from_file("kayit.wav")  # clone a voice from your own recording
audio = tts.generate("Bu cümleyi benim sesimle okuyun.", voice=voice)

For real-time playback, stream yields the audio in short pieces while it is being generated; the first piece is ready in under 0.1 seconds on one CPU thread, after a first call of about 0.4 seconds that prepares the voice:

for chunk in tts.stream("Randevunuz 15.10.2026 saat 14:30'da.", voice="female_1"):
    ...  # a NumPy float32 array at 24 kHz; send it to your audio output

The same is available from the command line:

pocket-tts-turkish generate --text "Merhaba, size nasıl yardımcı olabilirim?" --voice male_1 -o merhaba.wav

See the GitHub README for all options.

With pocket-tts directly

The model also runs with Kyutai's pocket-tts library:

uvx pocket-tts generate --config hf://wite-tech/pocket-tts-turkish/config.yaml \
    --voice hf://wite-tech/pocket-tts-turkish/voices/female_1.wav \
    --text "[sakin] lütfen endişelenmeyin, her şey yolunda."
from pocket_tts import TTSModel

model = TTSModel.load_model(config="hf://wite-tech/pocket-tts-turkish/config.yaml")
state = model.get_state_for_audio_prompt("hf://wite-tech/pocket-tts-turkish/voices/female_1.wav")
audio = model.generate_audio(state, "[sakin] lütfen endişelenmeyin, her şey yolunda.")

Without the package, write numbers as words, generate one sentence at a time, start each sentence with a lowercase letter, and use a voice reference that stops between two words. The package does all of this for you.

Evaluation

The model was compared with eight public Turkish-capable systems: FreyaTTS-small, Piper (tr_TR-dfki-medium), MMS-TTS, VoxCPM2, Trendyol-TTS, Chatterbox Multilingual, Qwen3-TTS 0.6B Turkish and XTTS-v2. Two public test sets were used: Freya-TR-Eval (T1, 495 everyday conversational sentences) and the FLEURS Turkish test set (T2, 200 long read sentences). Every system received the same text, with numbers already written as words. The word error rate (WER) is the share of words that Whisper large-v3 transcribes differently from the input text, measured on audio band-limited to 8 kHz as in the Freya-TR-Eval recipe; lower is better. Speed was measured on one RTX 5090.

Speed and accuracy

Speed against intelligibility

Each bubble is one system. Its horizontal position is the real-time factor, the synthesis time divided by the length of the audio, on a log scale: further left is faster, and 0.1 means ten times faster than real time. Its height is the WER on T1, and its area follows the number of parameters. The best place to be is the lower left.

Pocket TTS Turkish reaches 1.9% WER at a real-time factor of 0.048, about 20 times faster than real time, with 110 million parameters. The two systems with a lower WER, Trendyol-TTS (1.1%) and Qwen3-TTS (1.7%), are 8 to 22 times larger and 7 to 9 times slower. The two systems that are faster, Piper and MMS-TTS, make more errors (3.2% and 6.1%) and cannot clone a voice.

What a real-time voice agent needs

What a real-time voice agent needs, and which systems deliver it

The columns are nine requirements for a voice agent that answers in real time: it streams audio, clones a reference voice, has an emotion control, runs on a CPU faster than real time, keeps the WER at or below 2.5% on T1 and 5% on T2, synthesizes a 5-second reply in under half a second, uses less than 2 GB of GPU memory, and has a licence that allows commercial use. A filled circle means the system meets the requirement and a cross means it does not; the number on the right is the total.

Pocket TTS Turkish is the only system that meets all nine; the next best, VoxCPM2, meets six. A requirement counts as met only where it was measured, so systems that were not run on a CPU are marked as not running on one.

Training

  • Starting point: Kyutai's Pocket TTS, the English 24-layer model.
  • Tokenizer: a new Turkish SentencePiece tokenizer with 4,000 pieces, including the emotion tags; the text embeddings were trained from scratch.
  • Teacher: the 24-layer model fine-tuned on Turkish for 250,000 steps.
  • Student: this 6-layer model, initialized from teacher layers and distilled for 200,000 steps, with classifier-free guidance distilled into the weights.
  • Data: 306 hours of synthetic Turkish speech in seven voices, six of which are included here, neutral and with the five emotions, generated locally with one of our own Turkish text-to-speech models. The texts are about 233,000 sentences from Tatoeba (CC BY 2.0 FR), FineWeb-2 Turkish (ODC-By 1.0) and Turkish Wikipedia (CC BY-SA 4.0), with numbers written as words. Audio and text were aligned with mpoyraz/wav2vec2-xls-r-300m-cv7-turkish.

Limitations

  • Turkish only.

Responsible use

The model can imitate a voice from a few seconds of audio. Only clone a voice with the explicit consent of the person it belongs to. Do not use the model for impersonation, fraud, fraudulent calls or misinformation, or to present generated speech as a real recording of a person. Use must comply with applicable laws. The same terms apply to the upstream Pocket TTS model.

License and attribution

This model is released under CC-BY-4.0. It is derived from Pocket TTS by Kyutai, also released under CC-BY-4.0. Changes from the original: a Turkish tokenizer, fine-tuning on Turkish speech and distillation to six layers. The training texts come from Tatoeba, FineWeb-2 and Wikipedia under the licences listed above.

Citation

@misc{pocket_tts_turkish_2026,
  title  = {Pocket TTS Turkish},
  author = {Khaled Moawad and Ammar Rashed and Ekrem Çetinkaya},
  year   = {2026},
  url    = {https://huggingface.co/wite-tech/pocket-tts-turkish}
}

Please also cite the work this model is built on:

@article{rouard2025continuous,
  title         = {Continuous Audio Language Models},
  author        = {Rouard, Simon and Orsini, Manu and Roebel, Axel and Zeghidour, Neil and D\'efossez, Alexandre},
  year          = {2025},
  eprint        = {2509.06926},
  archivePrefix = {arXiv},
  primaryClass  = {cs.SD}
}
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wite-tech/pocket-tts-turkish

Finetuned
(27)
this model

Space using wite-tech/pocket-tts-turkish 1

Paper for wite-tech/pocket-tts-turkish