FastSpeech-Pytorch

The Implementation of FastSpeech Based on Pytorch.

Model

My Blog

Train

Download and extract LJSpeech dataset.
Put LJSpeech dataset in data.
Run preprocess.py.
If you want to get the target of alignment before training(It will speed up the training process greatly), you need download the pre-trained Tacotron2 model published by NVIDIA here. The outputs of mel spectrogram and alignment are shown as follow:

5. Put the pre-trained Tacotron2 model in `Tacotron2/pre_trained_model` 6. Run `alignment.py`, it will take long time. 7. Change ```python pre_target = True``` in `hparam.py`. 8. Run `train.py`.

Dependencies

python 3.6
pytorch 1.1.0
numpy 1.16.2
scipy 1.2.1
librosa 0.6.3
inflect 2.1.0
matplotlib 2.2.2

Notes

If you don't prepare the target of alignment before training, the process of training would be very long.
In the paper of Transformer-TTS, authors use pre-trained Transformer-TTS to provide the target of alignment. I didn't have a well-trained Transformer-TTS model so I use Tacotron2 instead.
If you want to use another model to get targets of alignment, you need rewrite alignment.py.
The returned value of alignment.py is a tensor whose value is the multiple that encoder's outputs are supposed to be expanded by.
For example:

test_target = torch.stack([torch.Tensor([0, 2, 3, 0, 3, 2, 1, 0, 0, 0]),
                           torch.Tensor([1, 2, 3, 2, 2, 0, 3, 6, 3, 5])])

Name		Name	Last commit message	Last commit date
Latest commit History 46 Commits
Tacotron2		Tacotron2
alignment_targets		alignment_targets
data		data
dataset		dataset
img		img
logger		logger
paper		paper
text		text
transformer		transformer
visualize_loss		visualize_loss
FastSpeech.py		FastSpeech.py
Networks.py		Networks.py
README.md		README.md
alignment.py		alignment.py
audio.py		audio.py
data_utils.py		data_utils.py
hparams.py		hparams.py
loss.py		loss.py
optimizer.py		optimizer.py
preprocess.py		preprocess.py
train.py		train.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

FastSpeech-Pytorch

Model

My Blog

Train

Dependencies

Notes

Reference

About

Releases

Packages

Languages

dachengai/FastSpeech

Folders and files

Latest commit

History

Repository files navigation

FastSpeech-Pytorch

Model

My Blog

Train

Dependencies

Notes

Reference

About

Resources

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages