Publications
-
Meeting Structure Annotation: Data and Tools
We present a set of annotations of hierarchical topic segmentations and action item sub-dialogues collected over 65 meetings from the ICSI and ISL meeting corpora, designed to support automatic meeting…
-
Spoken Language Understanding
SLU systems contain an automatic speech recognition (ASR) component and must be robust to noise due to the spontaneous nature of spoken language and the errors introduced by ASR. SLU…
-
Does Active Learning Help Automatic Dialog Act Tagging in Meeting Data?
We ask if active learning with lexical cues can help for this task and this domain. To better address this question, we explore active learning for two different types of…
-
Pushing the Envelope — Aside
Despite successes, there are still significant limitations to speech recognition performance. For this reason, authors have proposed methods that incorporate different (and larger) analysis windows, which are described in this…
-
Comparing HMM, Maximum Entropy, and Conditional Random Fields for Disfluency Detection
We compare a generative hidden Markov model (HMM)-based approach and two conditional models — a maximum entropy (Maxent) model and a conditional random field (CRF) — for detecting disfluencies in…
-
Distinguishing Deceptive from Non-Deceptive Speech
We present results from a study seeking to distinguish deceptive from non-deceptive speech using machine learning techniques on features extracted from a large corpus of deceptive and non-deceptive speech. We…
-
Improved Discriminative Training Using Phone Lattices
We present an efficient discriminative training procedure utilizing phone lattices. Different approaches to expediting lattice generation, statistics collection, and convergence were studied.
-
Generation of fast interpreters for Huffman compressed bytecode
Our approach uses canonical Huffman codes to generate compact opcodes with custom-sized operand fields and with a virtual machine that directly executes this compact code. In effect, this automatically creates…
-
Speech Translation for Low-Resource Languages: The Case of Pashto
We present a number of challenges and solutions that have arisen in the development of a speech translation system for American English and Pashto, highlighting those specific to a very…
-
Leveraging Speaker-dependent Variation of Adaptation
This work introduces an automatic procedure for determining the size of regression class trees for individual speakers using an ensemble of speaker-level features to control the number of transformations, if…
-
A Personalized Time Management Assistant: Research Directions
This paper presents ongoing work to build the Personalized Time Manager (PTIME) system, a persistent assistant that builds on our previous work on a personalized calendar agent (PCalM) (Berry et…
-
A Robust Method for Tracking Scene Text in Video Imagery
We describe an approach that tracks planar regions of scene text that can undergo arbitrary 3-D rigid motion and scale changes. Our approach computes homographies on blocks of contiguous frames…