Robust speech representation of voiced sounds based on synchrony determiniation with PLLS

Citation

P. Pelle, C. Estienne and H. Franco, “Robust speech representation of voiced sounds based on synchrony determination with PLLs,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2011), pp. 5424–5427.

Abstract

We propose to include synchrony effects, known to exist in the auditory system, to represent voiced parts of the speech signal in a robust way.  The system decomposes the input signal by means of a bandpass filter bank, and utilizes a bank of phase locked loops (PLLs) to obtain information on the frequencies present at a specific time.  This information about the frequency distribution is transformed into a spectral-like representation based on synchrony effects.  Noisy speech recognition experiments are performed using this synchrony-based spectrum, which is transformed into a small set of coefficients by using a ransformation similar to that utilized for mel cepstrum features.  We show that recognition performance compared to mel cepstrum features is advantageous, when measured over a range of SNR conditions, especially in the high noise level case.

Index Terms: speech features, robustness, PLL, noise, auditory system.


Read more from SRI

  • Banner and attendees at the IEEE Hard Tech Venture Summit

    Cultivating hard tech startups that scale

    IEEE’s Hard Tech Venture Summit convened innovators at SRI to refine strategies and build new networks.

  • Patient going into a MRI

    Bringing surgical tools inside the MRI

    Drawing on SRI’s unique innovation ecosystem, the startup Medical Devices Corner is seeking to improve cancer surgery by advancing MRI-safe teleoperation.

  • Christopher Mims and Susan Patrick

    PARC Forum: How to AI

    The Wall Street Journal tech columnist Christopher Mims and SRI Education’s Susan Patrick discuss how AI can strengthen human agency.