We release Lip2Wav, a 120-hour single-speaker dataset, and a new architecture that learns individual speaking styles to generate natural speech from lip movements, four times more intelligible than prior work.
Need help?
Contact usWe release Lip2Wav, a 120-hour single-speaker dataset, and a new architecture that learns individual speaking styles to generate natural speech from lip movements, four times more intelligible than prior work.