Author: Mitchell McLaren

November 18, 2022

Toward Fail-Safe Speaker Recognition: Trial-Based Calibration with a Reject Option

In this work, we extend the TBC method, proposing a new similarity metric for selecting training data that results in significant gains over the one proposed in the original work.
September 1, 2018

Analysis of Complementary Information Sources in the Speaker Embeddings Framework

In this study, our aim is analyzing the behavior of the speaker recognition systems based on speaker embeddings toward different front-end features, including the standard MFCC, as well as PNCC, and PLP.
September 1, 2018

Robust Speaker Recognition from Distant Speech under Real Reverberant Environments Using Speaker Embeddings

This article focuses on speaker recognition using speech acquired using a single distant or far-field microphone in an indoors environment.
June 1, 2018

Approaches to multi-domain language recognition

Approaches found to provide robustness in multi-domain LID include a domain-and-language-weighted Gaussian backend classifier, duration-aware calibration, and a source normalized multi-resolution neural network backend.
June 1, 2018

How to train your speaker embedding extractor

In this study, we aim to explore some of the fundamental requirements for building a good speaker embeddings extractor. We analyze the impact of voice activity detection, types of degradation, the amount of degraded data, and number of speakers required for a good network.
December 1, 2017

Language Diarization for Semi-supervised Bilingual Acoustic Model Training

In this paper, we investigate several automatic transcription schemes for using raw bilingual broadcast news data in semi-supervised bilingual acoustic model training.
August 1, 2017

Calibration Approaches for Language Detection

In this paper, we focus on situations in which either (1) the system-modeled languages are not observed during use or (2) the test data contains OOS languages that are unseen during modeling or calibration.
August 1, 2017

Improving Robustness of Speaker Recognition to New Conditions Using Unlabeled Data

We benchmark these approaches on several distinctly different databases, after we describe our SRICON-UAM team system submission for the NIST 2016 SRE.
September 1, 2016

On the Issue of Calibration in DNN-Based Speaker Recognition Systems

This article is concerned with the issue of calibration in the context of Deep Neural Network (DNN) based approaches to speaker recognition. We propose a hybrid alignment framework, which stems from our previous work in DNN senone alignment, that uses the bottleneck features only for the alignment of features during statistics calculation.
September 1, 2016

The 2016 Speakers in the Wild Speaker Recognition Evaluation

This article provides details of the SITW speaker recognition challenge and analysis of evaluation results. We provide an analysis of some of the top performing systems submitted during the evaluation and provide future research directions.
September 1, 2016

The Speakers in the Wild (SITW) Speaker Recognition Database

The Speakers in the Wild (SITW) speaker recognition database contains hand-annotated speech samples from open-source media for the purpose of benchmarking text-independent speaker recognition technology.
June 1, 2016

Exploring the role of phonetic bottleneck features for speaker and language recognition

Using bottleneck features extracted from a deep neural network (DNN) trained to predict senone posteriors has resulted in new, state-of-the-art technology for language and speaker identification.