Changelog
- Plans
- Updates
- 2022-06-06 Silero TTS in 20 Languages With 174 Speakers
- 2022-04-12 Silero TTS in High Resolution, 10x Faster and More Stable
- 2022-02-28 Experimental Pip Package
- 2022-02-24 English V6 Release
- 2021-12-09 Improved Text Recapitalization and Repunctuation Model for 4 Languages
- 2021-10-06 Text Recapitalization and Repunctuation Model for 4 Languages
- 2021-09-03 German V4 and English V5 Models
- 2021-08-09 German V3 Large Model
- 2021-06-18 Large V2 TTS release, v4_0 Large English STT Model
- 2021-04-21 Large V2 TTS release, v4_0 Large English STT Model
- 2021-04-20 Add v3 STT English Models
- 2021-03-29 Add v1 TTS Models
- 2021-03-03 Add xxsmall Speed Metrics
- 2021-03-03 Ukrainian Model V3 Released
- 2021-02-15 Some Organizational Issues
- 2020-12-04 Add EE Distro Sizing and New Speed Metrics
- 2020-11-26 Fix TensorFlow Examples
- [2020-11-03 [Experimental] Ukrainian Model V1 Released](#2020-11-03-experimental-ukrainian-model-v1-released)
- 2020-11-03 English Model V2 Released
- 2020-10-28 Minor PyTorch 1.7 fix
- 2020-10-19 Update wiki
- 2020-10-03 Batched ONNX and TF Models
- 2020-09-29 Update English benchmarks
- 2020-09-29 Published on TF Hub
- 2020-09-28 Added Timestamps To Decoder
- 2020-09-27 Examples, Usability, TF Example
- 2020-09-23 Fixed broken TF Model Archives
- 2020-09-23 Silero Models now on Torch Hub
- 2020-09-22 Tensorflow SavedModels, Colab VAD, tf.js
- 2020-09-19 Fx Minor Bugs
- 2020-09-17 TorchHub
- 2020-09-16 V1 Release
- 2020-09-15 Quality Benchmarks for German
- 2020-09-12 Quality Benchmarks for English
- 2020-09-11 Initial upload
Plans
Updates
2022-06-06 Silero TTS in 20 Languages With 174 Speakers
- Huge release - 20 languages, 173 voices;
- 1 new high quality Russian voice (eugeny);
- The CIS languages: Kalmyk, Russian, Tatar, Uzbek и Ukrainian;
- Romance and Germanic languages: English, Indic English, Spanish, German, French;
- 10 Indic languages;
- Russian automated stress model vastly improved (please see this link for details);
- All models inherit all of the previous SSML perks;
2022-04-12 Silero TTS in High Resolution, 10x Faster and More Stable
- Huge release - Russian only for now;
- Model size reduced 2x;
- New models are 10x faster;
- We added flags to control stress;
- Now the models can make proper pauses;
- High quality voice added (and unlimited "random" voices);
- All speakers squeezed into the same model;
- Input length limitations lifted, now models can work with paragraphs of text;
- Pauses, speed and pitch can be controlled via SSML;
- Sampling rates of 8, 24 or 48 kHz are supported;
- Models are much more stable — they do not omit words anymore;
2022-02-28 Experimental Pip Package
- Models are downloaded on demand both by pip and PyTorch Hub;
- If you need caching, do it manually or via invoking a necessary model once (it will be downloaded to a cache folder);
- Please see these docs for more information;
- PyTorch Hub and pip package are based on the same code. Hence all examples, historically based on torch.hub.load can be used with a pip-package;
2022-02-24 English V6 Release
- New en_v6 models;
- Quality improvements for English models;
2021-12-09 Improved Text Recapitalization and Repunctuation Model for 4 Languages
- The model now can work with long inputs, 512 tokens or ca. 150 words;
- Inputs longer than 150 words are automatically processed in chunks;
- The bugs with newer PyTorch versions have been fixed;
- Model was trained longer with larger batches;
- Model size slightly reduced to 85 MB;
- The rest of model optimizations were deemed too high maintenance;
2021-10-06 Text Recapitalization and Repunctuation Model for 4 Languages
- Inserts capital letters and basic punctuation marks (dot, comma, hyphen, question mark, exclamation mark, dash for Russian);
- Works for 4 languages (Russian, English, German, Spanish) and can be extended;
- By design is domain agnostic and is not based on any hard-coded rules;
- Has non-trivial metrics and succeeds in the task of improving text readability;
2021-09-03 German V4 and English V5 Models
- German V4 large jit and onnx models;
- English V5 small (jit and onnx), small_q (only jit) and xlarge (jit and onnx) models;
- Vast quality improvements (metrics to be added shortly) on the majority of domains;
- English xsmall models coming soon (jit and onnx);
2021-08-09 German V3 Large Model
- German V3 Large jit model trained on more data - large quality improvement;
- Metrics coming soon;
2021-06-18 Large V2 TTS release, v4_0 Large English STT Model
- Added v4_0 large English model with metrics;
- V2 TTS models with x4 faster vocoder;
- Russian models now feature automatic stress and ё, homonyms are not handled yet;
- A multi-language multi-speaker model;
2021-04-21 Large V2 TTS release, v4_0 Large English STT Model
Huge update for English!
- Polish docs;
- Add xsmall and xsmall_q model flavours for en_v3;
- Polish performance benchmarks page a bit;
2021-04-20 Add v3 STT English Models
Huge update for English!
- Default model (jit or onnx) size is reduced almost by 50% without sacrificing quality (!);
- New model flavours: jit_q (smaller quantized model), jit_skip (with exposed skip connections), jit_large (higher quality model), onnx_large (!);
- New smallest model jit_q is only 40M in size (!);
- Tensorflow checkpoints discontinued;
- New performance benchmarks - default models are on par with previous models and Google, large models mostly outperform Google (!);
- Even more quality improvements coming soon (!);
- CE benchmarks coming soon;
- xsmall model was created (2x smaller than the default), but I could not quantize it. I am looking into creating a xxsmall model;
- Still working on making EE models fully JIT-traceable;
2021-03-29 Add v1 TTS Models
- Added v1 TTS models;
- Add TTS performance benchmarks;
- Polish existing wiki;
- Progress on an additional method of model compression;
2021-03-03 Add xxsmall Speed Metrics
- See metrics here
- Note that this is only an acoustic model, full end-to-end system metrics differ, though xxsmall metrics trickle down for CPU systems
2021-03-03 Ukrainian Model V3 Released
- Fine tuned from a commercial production Russian model
- Trained on a larger corpus (around 1,000 hours)
- Model flavors: jit (CPU or GPU), jit_q (quantized and CPU only) and onnx (ONNX)
- Huge model speed improvements for CPU inference (!roughly 3x faster!) compared to the previous one, comparable with xxsmall from here
- Will be dropping TF support altogether
- No proper quality benchmarks for an experimental model though
2021-02-15 Some Organizational Issues
- Migrate to our own model hosting
- Solve large parasite traffic / DDOS issue, source still unknown
- Remove the CDN
- Some community / answers tidying up
- Major progress on silero-vad
- Major progress on TTS, preparing for a release
2020-12-04 Add EE Distro Sizing and New Speed Metrics
- https://github.com/snakers4/silero-models/wiki/Performance-Benchmarks
2020-11-26 Fix TensorFlow Examples
2020-11-03 [Experimental] Ukrainian Model V1 Released
- An experimental model
- Trained from a small community contributed corpus
- New Full model size reduced to 85 MB
- New - quantized model is ony 25 MB
- No TF or ONNX models
- Will be re-released a fine-tuned model from a larger Russian corpus upon V3 release
2020-11-03 English Model V2 Released
- A minor release, i.e. other models not affected
- English model was made much more robust to certain dialects
- Performance metrics coming soon
2020-10-28 Minor PyTorch 1.7 fix
- torch.hub.load signature was changed
2020-10-19 Update wiki
- Add article on Methodology, update wiki
2020-10-03 Batched ONNX and TF Models
- Extensively clean up and simplify ONNX and TF model code
- Add batch support to TF and ONNX models
- Update examples
- (pending) Submit new models to TF Hub and update examples there
2020-09-29 Update English benchmarks
new pruned double-lm quality benchmarks
2020-09-29 Published on TF Hub
- https://tfhub.dev/silero
- https://tfhub.dev/silero/silero-stt/en/1
- https://tfhub.dev/silero/collections/silero-stt/1
2020-09-28 Added Timestamps To Decoder
- Now standard decoder has an option to return word timestamps
- Please see Colab examples for PyTorch
2020-09-27 Examples, Usability, TF Example
- Polish and simplify the main readme
- Remove folders from inside of TF archives
- Polish model url naming, purge CDN cache
- Add TF Colab example
- Remove old ONNX example
- Submit to TF and ONNX hub
2020-09-23 Fixed broken TF Model Archives
- Fixed weird archiving issue for Windows, purged CDN cache
2020-09-23 Silero Models now on Torch Hub
- https://pytorch.org/hub/snakers4_silero-models_stt/
2020-09-22 Tensorflow SavedModels, Colab VAD, tf.js
- Add VAD to colab example
- Add proper SavedModel Tensorflow checkpoints for all of the languages
- Add experimental tf.js checkpoints for English
2020-09-19 Fx Minor Bugs
Fix minor colab bugs
2020-09-17 TorchHub
- Added loading via TorchHub
- Added issue templates
- Fix typos
2020-09-16 V1 Release
- Examples and docs further polished
- Added a more thorough colab example with file upload / speech recording
- Added performance benchmarks
- Added quality benchmarks for Spanish in the wiki
2020-09-15 Quality Benchmarks for German
Added quality benchmarks for German in the wiki
2020-09-12 Quality Benchmarks for English
Added quality benchmarks for English in the wiki
2020-09-11 Initial upload
First commit and first models uploaded:
- English
- German
- Spanish
---
README
[](mailto:[email protected]) [](https://t.me/silero_speech) [](https://github.com/snakers4/silero-models/blob/master/LICENSE)
[](https://badge.fury.io/py/silero) [](https://pepy.tech/projects/silero)
- Silero Models
- Installation and Basics
- Text-To-Speech
- Models and Speakers
- V5 Turkic
- V5 Caucasian
- V5
- V5 CIS Base Models
- V5 CIS Ext Models
- V4
- V3
- Dependencies
- PyTorch
- Standalone Use
- SSML
- Cyrillic languages v4
- Indic languages v4
- Example
- Supported languages
- Contact
- Licence
- Citations
- Further reading
- English
- Chinese
- Russian
Silero Models
Our TTS models satisfy the following criteria:
- Fully end-to-end;
- Large library of voices;
- Natural-sounding speech;
- One-line usage, minimal, portable;
- Impressively fast on CPU and GPU;
- For the Russian language - automated stress and homographs;
Installation and Basics
You can basically use our models in 3 flavours:
- Via PyTorch Hub: torch.hub.load();
- Via pip: pip install silero and then from silero import silero_tts;
- Via caching the required models and utils manually and modifying if necessary;
Models are downloaded on demand both by pip and PyTorch Hub. If you need caching, do it manually or via invoking a necessary model once (it will be downloaded to a cache folder). Please see these docs for more information.
PyTorch Hub and pip package are based on the same code. All of the torch.hub.load examples can be used with the pip package via this basic change:
from silero import silero_tts
model, example_text = silero_tts(language='ru',
speaker='v5_ru')
audio = model.apply_tts(text=example_text)Text-To-Speech
Models and Speakers
All of the provided models are listed in the models.yml file. Any metadata and newer versions will be added there.
#### V5 Turkic
- This model does not support stress at all;
- This model supports SSML;
- This model supports 8000, 24000, 48000 sampling rates;
| ID | Speakers | Language | Colab |
|-----------|------------------------------|-------------------------|-------|
| v5_turkic | chv_0,chv_1 | chv (Chuvash) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | crh_0, crh_1, crh_2 | crh (Crimean Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | gag_0, gag_1, gag_2 | gag (Gagauz) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | kaa_0, kaa_1 | kaa (Karakalpak) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | kir_0 | kir (Kyrgyz) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | kjh_0 | kjh (Khakas) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | sah_0, sah_1, sah_2 | sah (Yakut ) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | sty_0 | sty (Siberian Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | tat_0 | tat (Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | tgk_0 | tgk (Tajik) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | tuk_0, tuk_1, tuk_2 | tuk (Turkmen) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | tyv_0 | tyv (Tuvan) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | uzb_0 | uzb (Uzbek) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_turkic | xal_0, xal_1, xal_2 | xal (Kalmyk) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
#### V5 Caucasian
- This model does not support stress at all;
- This model supports SSML;
- This model supports 8000, 24000, 48000 sampling rates;
| ID | Speakers | Language | Colab |
|--------------|---------------------------------------------|-----------------------|-------|
| v5_caucasian | abq_0, abq_1, abq_2 | abq (Abaza) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | ady_0, ady_1, ady_2, ady_3 | ady (Adygean) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | agx_0 | agx (Aghul) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | ava_0 | ava (Avar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | che_0, che_1, che_2, che_3, che_4 | che (Chechen) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | darg_0, darg_1 | darg (Dargin) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | inh_0, inh_1 | inh (Ingush) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | kbd_0, kbd_1, kbd_2 | kbd (K.-Ciscassian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | krc_0, krc_1, krc_2 | krc (K.-Balkar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | kum_0, kum_1, kum_2 | kum (Kumyk) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | lbe_0 | lbe (Lak) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | lez_0 | lez (Lezgin) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | oss_0 | oss (Ossetian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | tab_0, tab_1 | tab (Tabasaran) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
| v5_caucasian | tkr_0 | tkr (Tsakhur) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_turkic_caucasian.ipynb) |
examples_tts_turkic_caucasian.ipynb
#### V5
V5 models support SSML. Also see Colab examples for main SSML tag usage.
Russian-only models support automated stress and homographs. v5_2_ru cointains minor fixes and removes numpy and scipy dependencies.
v5_3_ru cointains minor fixes. v5_4_ru also supports questions.
| ID | Speakers | Auto-stress / Homographs / Questions | Language | SR | Colab |
| ------- | --------------------------------------------- | ----------- | -------------- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| v5_5_ru | aidar, baya, kseniya, xenia, eugene | ✅ / ✅ / ✅ | ru (Russian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v5_4_ru | aidar, baya, kseniya, xenia | ✅ / ✅ / ✅ | ru (Russian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v5_3_ru | aidar, baya, kseniya, xenia, eugene | ✅ / ✅ / ❌ | ru (Russian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v5_2_ru | aidar, baya, kseniya, xenia, eugene | ✅ / ✅ / ❌ | ru (Russian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v5_ru | aidar, baya, kseniya, xenia, eugene | ✅ / ✅ / ❌ | ru (Russian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
#### V5 CIS Base Models
- All of the below models support 8000, 24000, 48000 sampling rates and contain no auto-stress or homographs;
- v5_cis_base models assume that proper stress should be added for each word for all languages, i.e. к+ошка;
- v5_cis_base_nostress models assume that proper stress should be added for each word ONLY for slavic languages (i.e. ru, bel, ukr);
- All of the below models are published under MIT licence;
- V5 UTMOS and throughput metrics;
- V5 models support SSML. Also see Colab examples for main SSML tag usage;
- Use cases for the model;
- Minimal system requirements: a PyTorch-compatible system, a modern processor with AVX2 instruction set for x86/64 platform.
| ID | Speakers | Language | Colab |
| ------------------------------------- | -------------------------------------------- | -------------------- | -------------------- |
| v5_cis_base, v5_cis_base_nostress | aze_gamat | aze (Azerbaijani) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | hye_zara | hye (Armenian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | bak_aigul, bak_alfia, bak_alfia2 | bak (Bashkir) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | bak_miyau, bak_ramilia | bak (Bashkir) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | bel_anatoliy, bel_dmitriy, bel_larisa | bel (Belarus) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | kat_vika | kat (Georgian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | kbd_eduard | kbd (Kab.-Cherkes) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | kaz_zhadyra, kaz_zhazira | kaz (Kazakh) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | xal_kejilgan, xal_kermen | xal (Kalmyk) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | kir_nurgul | kir (Kyrgyz) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | mdf_oksana | mdf (Moksha) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | all of these speakers, but with ru_ prefix | ru (Russian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | tgk_onaoy, tgk_safarhuja | tgk (Tajik) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | tat_albina, tat_marat | tat (Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | udm_bogdan | udm (Udmurt) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | uzb_saida | uzb (Uzbek) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | ukr_igor, ukr_roman | ukr (Ukrainian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | kjh_karina, kjh_sibday | kjh (Khakas) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | chv_ekaterina | chv (Chuvash) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | erz_alexandr | erz (Erzya) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_base, v5_cis_base_nostress | sah_zinaida | sah (Yakut) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb)
<details>
<summary>Supported alphabets</summary>
Please note that Georgian and Armenian are in fact internally supported via direct translation into cyrillic script inside of the package. Azerbaijani and Uzbek support both alphabets (Cyrillic and Latin).
| ID | Название | Алфавит(ы) |
|-----|---------------|--------------------------------------------|
| aze | aze (Azerbaijani) | abcçdeәfgğhxıijkqlmnoöprsştuüvyz |
| aze | aze (Azerbaijani) | абвгғдеәжзиыјкҝлмноөпрстуүфхһчҹш |
| hye | hye (Armenian) | աբգդեզէըթժիլխծկհձղճմյնշոչպջռսվտրցւփքօֆև |
| bak | bak (Bashkir) | абвгдежзийклмнопрстуфхцчшщъыьэюяёғҙҡңҫүһәө |
| bel | bel (Belarus) | абвгдежзйклмнопрстуфхцчшыьэюяёіў |
| kat | kat (Georgian) | აბგდევზთიკლმნოპჟრსტუფქღყშჩცძწჭხჯჰ |
| kbd | kbd (Kab.-Cherkes) | абвгдежзийклмнопрстуфхцчшщъыьэюяёӏ |
| kaz | kaz (Kazakh) | абвгдежзийклмнопрстуфхцчшщыьэюяіғқңүұһәө |
| xal | xal (Kalmyk) | абвгдежзийклмнопрстуфхцчшщъыьэюяҗңүһәө |
| kir | kir (Kyrgyz) | абвгдежзийклмнопрстуфхцчшыьэюяёңүө |
| mdf | mdf (Moksha) | абвгдежзийклмнопрстуфхцчшщъыьэюяё |
| ru | ru (Russian) | абвгдеёжзийклмнопрстуфхцчшщъыьэюя |
| tgk | tgk (Tajik) | абвгдежзийклмнопрстуфхчшъэюяёғқҳҷӣӯ |
| tat | tat (Tatar) | абвгдежзийклмнопрстуфхцчшъыьэюяҗңүһәө |
| udm | udm (Udmurt) | абвгдежзийклмнопрстуфхцчшщъыьэюяёӝӟӥӧӵ |
| uzb | uzb (Uzbek) | абвгдежзийклмнопрстуфхцчшъьэюяёўғқҳ |
| uzb | uzb (Uzbek) | abcdefghijklmnopqrstuvxyz |
| ukr | ukr (Ukrainian) | абвгґдеєжзиіїйклмнопрстуфхцчшщьюя |
| kjh | kjh (Khakas) | абвгдежзийклмнопрстуфхцчшщъыьэюяёіғңҷӧӱ |
| chv | chv (Chuvash) | абвгдежзийклмнопрстуфхцчшщъыьэюяёҫӑӗӳ |
| erz | erz (Erzya) | абвгдежзийклмнопрстуфхцчшщъыьэюяё |
| sah | sah (Yakut) | абвгдежзийклмнопрстуфхцчшщъыьэюяёҕҥүһө |
</details>
#### V5 CIS Ext Models
- All of the below models support 8000, 24000, 48000 sampling rates and contain no auto-stress or homographs;
- v5_cis_ext models assume that proper stress should be added for each word for all languages, i.e. к+ошка;
- v5_cis_ext_nostress are coming soon;
- All of the below models are published under CC-NC-BY licence;
- V5 models support SSML. Also see Colab examples for main SSML tag usage.
| ID | Speakers | Language | Colab |
| ------------ | --------------------------------------------------------------------- | ----------------- | -------------------- |
| v5_cis_ext | kaz_abai, kaz_aidana, kaz_aisha, kaz_bakir, kaz_danara | kaz (Kazakh) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | xal_delghir, xal_erdni | xal (Kalmyk) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | tat_adiba, tat_alsou, tat_amir, tat_azat, tat_batir | tat (Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | tat_bulat, tat_damir, tat_guzel, tat_ildar, tat_ilgiz | tat (Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | tat_karim, tat_mansur, tat_murat, tat_rasima, tat_rustem | tat (Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | tat_timur, tat_zifa, tat_zufar, tat_zulfiya | tat (Tatar) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | uzb_anora, uzb_dilnavoz | uzb (Uzbek) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | ukr_kateryna, ukr_lada, ukr_mykyta, ukr_oleksa, ukr_tetiana | ukr (Ukrainian) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
| v5_cis_ext | chv_aihwa, chv_alima | chv (Chuvash) | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts_cis.ipynb) |
#### V4
V4 models support SSML. Also see Colab examples for main SSML tag usage.
<details>
<summary>V4 models: v4_ru, v4_cyrillic, v4_ua, v4_uz, v4_indic </summary>
| ID | Speakers |Auto-stress | Language | SR | Colab |
| ------------- | ----------- | ----------- |---------------------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| v4_ru | aidar, baya, kseniya, xenia, eugene, random | yes | ru (Russian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v4_cyrillic | b_ava, marat_tt, kalmyk_erdni... | no | cyrillic (Avar, Tatar, Kalmyk, ...) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v4_ua | mykyta, random | no | ua (Ukrainian) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v4_uz | dilnavoz | no | uz (Uzbek) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v4_indic | hindi_male, hindi_female, ..., random | no | indic (Hindi, Telugu, ...) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
</details>
#### V3
V3 models support SSML. Also see Colab examples for main SSML tag usage.
<details>
<summary>V3 models: v3_en, v3_en_indic, v3_de, v3_es, v3_fr, v3_indic </summary>
| ID | Speakers |Auto-stress | Language | SR | Colab |
| ------------- | ----------- | ----------- |---------------------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| v3_en | en_0, en_1, ..., en_117, random | no | en (English) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v3_en_indic | tamil_female, ..., assamese_male, random | no | en (English) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v3_de | eva_k, ..., karlsson, random | no | de (German) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v3_es | es_0, es_1, es_2, random | no | es (Spanish) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v3_fr | fr_0, ..., fr_5, random | no | fr (French) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
| v3_indic | hindi_male, hindi_female, ..., random | no | indic (Hindi, Telugu, ...) | 8000, 24000, 48000 | [](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb) |
</details>
Dependencies
Basic dependencies for Colab examples:
- torch, 1.10+ for v3 models/ 2.0+ for v4 and v5 models;
- torchaudio, latest version bound to PyTorch should work (required only because models are hosted together with STT, not required for work);
- omegaconf, latest (can be removed as well, if you do not load all of the configs);
PyTorch
[](https://colab.research.google.com/github/snakers4/silero-models/blob/master/examples_tts.ipynb)
[](https://pytorch.org/hub/snakers4_silero-models_tts/)
V5
import torchlanguage = 'ru'
model_id = 'v5_ru'
sample_rate = 48000
speaker = 'xenia'
device = torch.device('cpu')
model, example_text = torch.hub.load(repo_or_dir='snakers4/silero-models',
model='silero_tts',
language=language,
speaker=model_id)
model.to(device) # gpu or cpu
audio = model.apply_tts(text=example_text,
speaker=speaker,
sample_rate=sample_rate)
Standalone Use
- Standalone usage only requires PyTorch 1.12+ and the Python Standard Library;
- Please see the detailed examples in Colab;
V5
import os
import torchdevice = torch.device('cpu')
torch.set_num_threads(4)
local_file = 'model.pt'
if not os.path.isfile(local_file):
torch.hub.download_url_to_file('https://models.silero.ai/models/tts/ru/v5_ru.pt',
local_file)
model = torch.package.PackageImporter(local_file).load_pickle("tts_models", "model")
model.to(device)
example_text = 'Меня зовут Лева Королев. Я из готов. И я уже готов открыть все ваши замки любой сложности!'
sample_rate = 48000
speaker='baya'
audio_paths = model.save_wav(text=example_text,
speaker=speaker,
sample_rate=sample_rate)
SSML
Check out our TTS Wiki page.
Cyrillic languages v4
To be superseded with v5 model(s) soon.
Supported tokenset:!,-.:?iµöабвгдежзийклмнопрстуфхцчшщъыьэюяёђѓєіјњћќўѳғҕҗҙқҡңҥҫүұҳҷһӏӑӓӕӗәӝӟӥӧөӱӳӵӹ
| Speaker_ID | Language | Gender |
| ------------ | --------------- | ------ |
| b_ava | Avar | F |
| b_bashkir | Bashkir | M |
| b_bulb | Bulgarian | M |
| b_bulc | Bulgarian | M |
| b_che | Chechen | M |
| b_cv | Chuvash | M |
| cv_ekaterina | Chuvash | F |
| b_myv | Erzya | M |
| b_kalmyk | Kalmyk | M |
| b_krc | Karachay-Balkar | M |
| kz_M1 | Kazakh | M |
| kz_M2 | Kazakh | M |
| kz_F3 | Kazakh | F |
| kz_F1 | Kazakh | F |
| kz_F2 | Kazakh | F |
| b_kjh | Khakas | F |
| b_kpv | Komi-Ziryan | M |
| b_lez | Lezghian | M |
| b_mhr | Mari | F |
| b_mrj | Mari High | M |
| b_nog | Nogai | F |
| b_oss | Ossetic | M |
| b_ru | Russian | M |
| b_tat | Tatar | M |
| marat_tt | Tatar | M |
| b_tyv | Tuvinian | M |
| b_udm | Udmurt | M |
| b_uzb | Uzbek | M |
| b_sah | Yakut | M |
| kalmyk_erdni | Kalmyk | M |
| kalmyk_delghir | Kalmyk | F |
Indic languages v4
#### Example
(!!!) All input sentences should be romanized to ISO format using aksharamukha. An example for hindi:
V3
import torch
from aksharamukha import transliterateLoading model
model, example_text = torch.hub.load(repo_or_dir='snakers4/silero-models',
model='silero_tts',
language='indic',
speaker='v4_indic')orig_text = "प्रसिद्द कबीर अध्येता, पुरुषोत्तम अग्रवाल का यह शोध आलेख, उस रामानंद की खोज करता है"
roman_text = transliterate.process('Devanagari', 'ISO', orig_text)
print(roman_text)
audio = model.apply_tts(roman_text,
speaker='hindi_male')
#### Supported languages
| Language | Speakers | Romanization function
-- | -- | --
hindi | hindi_female, hindi_male | transliterate.process('Devanagari', 'ISO', orig_text)
malayalam | malayalam_female, malayalam_male |transliterate.process('Malayalam', 'ISO', orig_text)
manipuri | manipuri_female |transliterate.process('Bengali', 'ISO', orig_text)
bengali | bengali_female, bengali_male | transliterate.process('Bengali', 'ISO', orig_text)
rajasthani | rajasthani_female, rajasthani_female | transliterate.process('Devanagari', 'ISO', orig_text)
tamil | tamil_female, tamil_male |transliterate.process('Tamil', 'ISO', orig_text, pre_options=['TamilTranscribe'])
telugu | telugu_female, telugu_male | transliterate.process('Telugu', 'ISO', orig_text)
gujarati | gujarati_female, gujarati_male | transliterate.process('Gujarati', 'ISO', orig_text)
kannada | kannada_female, kannada_male |transliterate.process('Kannada', 'ISO', orig_text)
Contact
Try our models, create an issue, join our chat, email us, and read the latest news.
Licence
All of the models are published under the main repo license (i.e. CC-NC-BY) except for the base cis-tts models, which are under MIT.
Citations
@misc{Silero Models,
author = {Silero Team},
title = {Silero Models: pre-trained text-to-speech models made embarrassingly simple},
year = {2025},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/snakers4/silero-models}},
commit = {insert_some_commit_here},
email = {[email protected]}
}Further reading
English
- STT:
- Towards an Imagenet Moment For Speech-To-Text - link
- A Speech-To-Text Practitioners Criticisms of Industry and Academia - link
- Modern Google-level STT Models Released - link
- TTS:
- Multilingual Text-to-Speech Models for Indic Languages - link
- Our new public speech synthesis in super-high quality, 10x faster and more stable - link
- High-Quality Text-to-Speech Made Accessible, Simple and Fast - link
- VAD:
- One Voice Detector to Rule Them All - link
- Modern Portable Voice Activity Detector Released - link
- Text Enhancement:
- We have published a model for text repunctuation and recapitalization for four languages - link
Chinese
- STT:
- 迈向语音识别领域的 ImageNet 时刻 - link
- 语音领域学术界和工业界的七宗罪 - link
Russian
- STT
- OpenAI решили распознавание речи! Разбираемся так ли это … - link
- Наши сервисы для бесплатного распознавания речи стали лучше и удобнее - link
- Telegram-бот Silero бесплатно переводит речь в текст - link
- Бесплатное распознавание речи для всех желающих - link
- Последние обновления моделей распознавания речи из Silero Models - link
- Сжимаем трансформеры: простые, универсальные и прикладные способы cделать их компактными и быстрыми - link
- Ультимативное сравнение систем распознавания речи: Ashmanov, Google, Sber, Silero, Tinkoff, Yandex - link
- Мы опубликовали современные STT модели сравнимые по качеству с Google - link
- Понижаем барьеры на вход в распознавание речи - link
- Огромный открытый датасет русской речи версия 1.0 - link
- Насколько Быстрой Можно Сделать Систему STT? - link
- Наша система Speech-To-Text - link
- Speech-To-Text - link
- TTS:
- Теперь silero-tts v5 на русском языке умеет задавать вопросы - link
- Наш синтез для 20 языков теперь работает локально под Windows как экранная читалка (SAPI5) и в Балаболке - link
- Мы добавили поддержку ещё 19 языков России и СНГ в проект silero-stress - link
- Мы опубликовали стабильный, быстрый, качественный и доступный синтез для 20 языков России - link
- Мы опубликовали silero-tts v5 на русском языке - link
- Мы решили задачу омографов и ударений в русском языке - link
- Делаем быстрый, качественный и доступный синтез на языках России — нужно ваше участие - link
- Теперь наш синтез также доступен в виде бота в Телеграме - link
- Может ли синтез речи обмануть систему биометрической идентификации? - link
- Теперь наш синтез на 20 языках - link
- Теперь наш публичный синтез в супер-высоком качестве, в 10 раз быстрее и без детских болячек - link
- Синтезируем голос бабушки, дедушки и Ленина + новости нашего публичного синтеза - link
- Мы сделали наш публичный синтез речи еще лучше - link
- Мы Опубликовали Качественный, Простой, Доступный и Быстрый Синтез Речи - link
- VAD:
- Новый релиз публичного детектора голоса Silero VAD v6 - link
- Наш публичный детектор голоса стал лучше - link
- А ты используешь VAD? Что это такое и зачем он нужен - link
- Модели для Детекции Речи, Чисел и Распознавания Языков - link
- Мы опубликовали современный Voice Activity Detector и не только -link
- Text Enhancement:
- Восстановление знаков пунктуации и заглавных букв — теперь и на длинных текстах - link
- Мы опубликовали модель, расставляющую знаки препинания и заглавные буквы в тексте на четырех языках - link
---