Update TTS readme (#224)

open-mmlab · Jun 25, 2024 · 5dfe9fd · 5dfe9fd
1 parent f96a153
commit 5dfe9fd
Show file tree

Hide file tree

Showing 8 changed files with 4 additions and 4 deletions.
diff --git a/egs/tts/README.md b/egs/tts/README.md
@@ -3,14 +3,14 @@
 
 ## Quick Start
 
-We provide a **[beginner recipe](VALLE/)** to demonstrate how to train a cutting edge TTS model. Specifically, it is Amphion's re-implementation for [Vall-E](https://arxiv.org/abs/2301.02111), which is a zero-shot TTS architecture that uses a neural codec language model with discrete codes.
+We provide a **[beginner recipe](VALLE_V2/)** to demonstrate how to train a cutting edge TTS model. Specifically, it is Amphion's re-implementation for [VALL-E](https://arxiv.org/abs/2301.02111), which is a zero-shot TTS architecture that uses a neural codec language model with discrete codes.
 
 ## Supported Model Architectures
 
 Until now, Amphion TTS supports the following models or architectures,
 - **[FastSpeech2](FastSpeech2)**: A non-autoregressive TTS architecture that utilizes feed-forward Transformer blocks.
 - **[VITS](VITS)**: An end-to-end TTS architecture that utilizes conditional variational autoencoder with adversarial learning
-- **[Vall-E](VALLE)**: A zero-shot TTS architecture that uses a neural codec language model with discrete codes.
+- **[VALL-E](VALLE_V2)**: A zero-shot TTS architecture that uses a neural codec language model with discrete codes. This model is our updated VALL-E implementation as of June 2024 which uses Llama as its underlying architecture. The previous version of VALL-E release can be found [here](VALLE)
 - **[NaturalSpeech2](NaturalSpeech2)** (👨‍💻 developing): An architecture for TTS that utilizes a latent diffusion model to generate natural-sounding voices.
 
 ## Amphion TTS Demo

diff --git a/egs/tts/valle_v2/README.md → egs/tts/VALLE_V2/README.md b/egs/tts/valle_v2/README.md → egs/tts/VALLE_V2/README.md
diff --git a/egs/tts/valle_v2/demo.ipynb → egs/tts/VALLE_V2/demo.ipynb b/egs/tts/valle_v2/demo.ipynb → egs/tts/VALLE_V2/demo.ipynb
@@ -22,9 +22,9 @@
    "source": [
     "# put your cheackpoint file (.bin) in the root path of AmphionVALLEv2\n",
     "# or use your own pretrained weights\n",
-    "ar_model_path = 'ckpts/valle_ar_mls_196000.bin'  #huggingface-cli download jiaqili3/vallex valle_ar_mls_196000.bin valle_nar_mls_164000.bin --local-dir ckpts\n",
+    "ar_model_path = 'ckpts/valle_ar_mls_196000.bin'  # huggingface-cli download amphion/valle valle_ar_mls_196000.bin valle_nar_mls_164000.bin --local-dir ckpts\n",
     "nar_model_path = 'ckpts/valle_nar_mls_164000.bin'\n",
-    "speechtokenizer_path = 'ckpts/speechtokenizer_hubert_avg' # huggingface-cli download fnlp/SpeechTokenizer speechtokenizer_hubert_avg/SpeechTokenizer.pt speechtokenizer_hubert_avg/config.json --local-dir ckpts"
+    "speechtokenizer_path = 'ckpts/speechtokenizer_hubert_avg' # huggingface-cli download amphion/valle speechtokenizer_hubert_avg/SpeechTokenizer.pt speechtokenizer_hubert_avg/config.json --local-dir ckpts"
    ]
   },
   {

diff --git a/egs/tts/valle_v2/example.wav → egs/tts/VALLE_V2/example.wav b/egs/tts/valle_v2/example.wav → egs/tts/VALLE_V2/example.wav
diff --git a/egs/tts/valle_v2/exp_ar_libritts.json → egs/tts/VALLE_V2/exp_ar_libritts.json b/egs/tts/valle_v2/exp_ar_libritts.json → egs/tts/VALLE_V2/exp_ar_libritts.json
diff --git a/egs/tts/valle_v2/exp_nar_libritts.json → egs/tts/VALLE_V2/exp_nar_libritts.json b/egs/tts/valle_v2/exp_nar_libritts.json → egs/tts/VALLE_V2/exp_nar_libritts.json
diff --git a/egs/tts/valle_v2/train_ar_libritts.sh → egs/tts/VALLE_V2/train_ar_libritts.sh b/egs/tts/valle_v2/train_ar_libritts.sh → egs/tts/VALLE_V2/train_ar_libritts.sh
diff --git a/egs/tts/valle_v2/train_nar_libritts.sh → egs/tts/VALLE_V2/train_nar_libritts.sh b/egs/tts/valle_v2/train_nar_libritts.sh → egs/tts/VALLE_V2/train_nar_libritts.sh