Instructions to use wite-tech/pocket-tts-turkish with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use wite-tech/pocket-tts-turkish with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("wite-tech/pocket-tts-turkish") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Pocket TTS Turkish
Pocket TTS Turkish is a 6-layer Turkish text-to-speech model built on Kyutai's Pocket TTS. It runs faster than real time on a single CPU thread, clones a voice from a few seconds of audio and supports five emotion tags.
- About 110 million parameters (codec included), 440 MB download, 24 kHz mono output
- About 5 times faster than real time on one CPU thread, about 20 times on a GPU
- Six built-in voices, or clone a voice from a short recording
- Emotion tags:
mutlu(happy),üzgün(sad),kızgın(angry),şaşkın(surprised),sakin(calm) - Released under CC-BY-4.0
Try it in the browser: Pocket TTS Turkish demo. Python package and command line tool: pocket-tts-turkish on GitHub.
Voices
Each voice reads: "Merhaba, ben bir Türkçe konuşma sentezleme modeliyim. Metinleri bilgisayarınızda, internet bağlantısı olmadan seslendirebilirim."
| voice | character | audio |
|---|---|---|
female_1 |
bright and energetic (default) | |
female_2 |
warm, clear and friendly | |
female_3 |
confident, clear articulation | |
male_1 |
calm and deep, narrator style | |
male_2 |
energetic and bright | |
male_3 |
confident, crisp articulation |
The built-in voices are synthetic. They were designed from text descriptions and are not recordings of real people.
Usage
With the pocket-tts-turkish package (recommended)
pip install pocket-tts-turkish
from pocket_tts_turkish import TurkishTTS, save_wav
tts = TurkishTTS.from_pretrained() # downloads this model once
audio = tts.generate("Randevunuz 15.10.2026 saat 14:30'da. Ücret 1.250 TL.", voice="female_1", emotion="sakin")
save_wav("randevu.wav", audio, tts.sample_rate)
voice = tts.voice_from_file("kayit.wav") # clone a voice from your own recording
audio = tts.generate("Bu cümleyi benim sesimle okuyun.", voice=voice)
For real-time playback, stream yields the audio in short pieces while it is being generated; the
first piece is ready in under 0.1 seconds on one CPU thread, after a first call of about 0.4 seconds
that prepares the voice:
for chunk in tts.stream("Randevunuz 15.10.2026 saat 14:30'da.", voice="female_1"):
... # a NumPy float32 array at 24 kHz; send it to your audio output
The same is available from the command line:
pocket-tts-turkish generate --text "Merhaba, size nasıl yardımcı olabilirim?" --voice male_1 -o merhaba.wav
See the GitHub README for all options.
With pocket-tts directly
The model also runs with Kyutai's pocket-tts library:
uvx pocket-tts generate --config hf://wite-tech/pocket-tts-turkish/config.yaml \
--voice hf://wite-tech/pocket-tts-turkish/voices/female_1.wav \
--text "[sakin] lütfen endişelenmeyin, her şey yolunda."
from pocket_tts import TTSModel
model = TTSModel.load_model(config="hf://wite-tech/pocket-tts-turkish/config.yaml")
state = model.get_state_for_audio_prompt("hf://wite-tech/pocket-tts-turkish/voices/female_1.wav")
audio = model.generate_audio(state, "[sakin] lütfen endişelenmeyin, her şey yolunda.")
Without the package, write numbers as words, generate one sentence at a time, start each sentence with a lowercase letter, and use a voice reference that stops between two words. The package does all of this for you.
Evaluation
The model was compared with eight public Turkish-capable systems: FreyaTTS-small, Piper
(tr_TR-dfki-medium), MMS-TTS, VoxCPM2, Trendyol-TTS, Chatterbox Multilingual, Qwen3-TTS 0.6B Turkish
and XTTS-v2. Two public test sets were used: Freya-TR-Eval (T1, 495 everyday conversational
sentences) and the FLEURS Turkish test set (T2, 200 long read sentences). Every system received the
same text, with numbers already written as words. The word error rate (WER) is the share of words
that Whisper large-v3 transcribes differently from the input text, measured on audio band-limited to
8 kHz as in the Freya-TR-Eval recipe; lower is better. Speed was measured on one RTX 5090.
Speed and accuracy
Each bubble is one system. Its horizontal position is the real-time factor, the synthesis time divided by the length of the audio, on a log scale: further left is faster, and 0.1 means ten times faster than real time. Its height is the WER on T1, and its area follows the number of parameters. The best place to be is the lower left.
Pocket TTS Turkish reaches 1.9% WER at a real-time factor of 0.048, about 20 times faster than real time, with 110 million parameters. The two systems with a lower WER, Trendyol-TTS (1.1%) and Qwen3-TTS (1.7%), are 8 to 22 times larger and 7 to 9 times slower. The two systems that are faster, Piper and MMS-TTS, make more errors (3.2% and 6.1%) and cannot clone a voice.
What a real-time voice agent needs
The columns are nine requirements for a voice agent that answers in real time: it streams audio, clones a reference voice, has an emotion control, runs on a CPU faster than real time, keeps the WER at or below 2.5% on T1 and 5% on T2, synthesizes a 5-second reply in under half a second, uses less than 2 GB of GPU memory, and has a licence that allows commercial use. A filled circle means the system meets the requirement and a cross means it does not; the number on the right is the total.
Pocket TTS Turkish is the only system that meets all nine; the next best, VoxCPM2, meets six. A requirement counts as met only where it was measured, so systems that were not run on a CPU are marked as not running on one.
Training
- Starting point: Kyutai's Pocket TTS, the English 24-layer model.
- Tokenizer: a new Turkish SentencePiece tokenizer with 4,000 pieces, including the emotion tags; the text embeddings were trained from scratch.
- Teacher: the 24-layer model fine-tuned on Turkish for 250,000 steps.
- Student: this 6-layer model, initialized from teacher layers and distilled for 200,000 steps, with classifier-free guidance distilled into the weights.
- Data: 306 hours of synthetic Turkish speech in seven voices, six of which are included here,
neutral and with the five emotions, generated locally with one of our own Turkish text-to-speech
models. The texts are about 233,000 sentences from Tatoeba (CC BY 2.0 FR), FineWeb-2 Turkish
(ODC-By 1.0) and Turkish Wikipedia (CC BY-SA 4.0), with numbers written as words. Audio and text
were aligned with
mpoyraz/wav2vec2-xls-r-300m-cv7-turkish.
Limitations
- Turkish only.
Responsible use
The model can imitate a voice from a few seconds of audio. Only clone a voice with the explicit consent of the person it belongs to. Do not use the model for impersonation, fraud, fraudulent calls or misinformation, or to present generated speech as a real recording of a person. Use must comply with applicable laws. The same terms apply to the upstream Pocket TTS model.
License and attribution
This model is released under CC-BY-4.0. It is derived from Pocket TTS by Kyutai, also released under CC-BY-4.0. Changes from the original: a Turkish tokenizer, fine-tuning on Turkish speech and distillation to six layers. The training texts come from Tatoeba, FineWeb-2 and Wikipedia under the licences listed above.
Citation
@misc{pocket_tts_turkish_2026,
title = {Pocket TTS Turkish},
author = {Khaled Moawad and Ammar Rashed and Ekrem Çetinkaya},
year = {2026},
url = {https://huggingface.co/wite-tech/pocket-tts-turkish}
}
Please also cite the work this model is built on:
@article{rouard2025continuous,
title = {Continuous Audio Language Models},
author = {Rouard, Simon and Orsini, Manu and Roebel, Axel and Zeghidour, Neil and D\'efossez, Alexandre},
year = {2025},
eprint = {2509.06926},
archivePrefix = {arXiv},
primaryClass = {cs.SD}
}
- Downloads last month
- -
Model tree for wite-tech/pocket-tts-turkish
Base model
kyutai/pocket-tts
