K. Kamangar, D. Hakkani-Tür, G. Tur and M. Levit, “An iterative unsupervised learning method for information distillation,” in Proc. 2008 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), pp. 4949–4952.
Information distillation techniques are used to analyze and interpret large volumes of speech and text archives in multiple languages and produce structured information of interest to the user. In this work, we propose an iterative unsupervised sentence extraction method to answer open-ended natural language queries about an event. The approach consists of finding the subset of sentences that are very likely to be relevant or irrelevant for the query from candidate documents, and iteratively training a classification model using these examples. Our results indicate that performance of the system may be improved by around 30 pct. relative in terms of F-measure, by using the proposed method.