September 1, 2014

Application of Convolutional Neural Networks to Speaker Recognition in Noisy Conditions

Citation

McLaren, M., Lei, Y., Scheffer, N., & Ferrer, L. (2014). Application of convolutional neural networks to speaker recognition in noisy conditions. INTERSPEECH.

Abstract

This paper applies a convolutional neural network (CNN) trained for automatic speech recognition (ASR) to the task of speaker identification (SID). In the CNN/i-vector front end, the sufficient statistics are collected based on the outputs of the CNN as opposed to the traditional universal background model (UBM). Evaluated on heavily degraded speech data, the CNN/i-vector front end provides performance comparable to the UBM/i-vector baseline. The combination of these approaches, however, is shown to provide improvements of 26% in miss rate to considerably outperform the fusion of two different features in the traditional UBM/i-vectors approach. An analysis of the language- and channel-dependency of the CNN/i-vector approach is also provided to highlight future research directions.

Index Terms: Deep neural networks, Convolutional neural networks, Speaker recognition, i-vectors, noisy speech

↓ Download PDF

Application of Convolutional Neural Networks to Speaker Recognition in Noisy Conditions

Abstract

Read more from SRI

Researchers develop materials that can take on the toughest conditions

Podcast: Re-imagining instructional quality and coaching

SRI’s Genome Explorer: Enhanced genome browser delivers better user experience