Generative-based Target Speech Extraction with Speech Discretization and Vocoder
Abstract
Target speech extraction (TSE) is a task aimed at isolating the speech of a specific target speaker from an audio mixture, using an auxiliary recording of that target speaker. Most existing TSE methods employ discriminative-based models to estimate the target speaker’s proportion in the mixture, but they are plagued by the inability to effectively eliminate residual interfering speech. In this paper, we present a novel generation-based TSE approach by combining speech discretization and vocoder techniques. By predicting a sequence of discrete tokens with the auxiliary audio and employing a vocoder that takes discrete tokens as input, the target speech can be effectively re-synthesized. Our experiments conducted on the WSJ0-2mix and Libri2mix datasets demonstrate that our proposed method yields high-quality target speech without interference.
Test samples of Target Speech Extraction
The sound files blow are some raw noisy wavs and the extracted speech from different methods. All the samples are driven from the test set of Libri2Mix (16k min). We use UniCATs-HuBERT-4096 as the discrete vocoder.
(1) Sample1:
Noisy wav:
Reference wav:
DPCCN:
Our discrete extraction:
(UniCATs-HuBERT-4096)
(2) Sample2:
Noisy wav:
Reference wav:
DPCCN:
Our discrete extraction:
(UniCATs-HuBERT-4096)
(3) Sample3:
Noisy wav:
Reference wav:
DPCCN:
Our discrete extraction:
(UniCATs-HuBERT-4096)
(4) Sample4:
Noisy wav:
Reference wav:
DPCCN:
Our discrete extraction:
(UniCATs-HuBERT-4096)