Neural VTLN for Speaker Adaptation in TTS

Schnell, BastianGarner, Philip N.2020-02-182020-02-182020-02-18201910.21437/ssw.2019-6https://infoscience.epfl.ch/handle/20.500.14299/166352Vocal tract length normalisation (VTLN) is well established as a speaker adaptation technique that can work with very little adaptation data. It is also well known that VTLN can be cast as a linear transform in the cepstral domain. Building on this latter property, we show that it can be cast as a (linear) layer in a deep neural network (DNN) for speech synthesis. We show that VTLN parameters can then be trained in the same framework as the rest of the DNN using automatic gradients. Experimental results show that the DNN is capable of predicting phone-dependent warpings on artificial data, and that such warpings improve the quality of an acoustic model on real data in subjective listening tests.Neural VTLN for Speaker Adaptation in TTStext::conference output::conference proceedings::conference paper