Machine learning for text using latent information
Kechen Qin · 2022
Artificial intelligence (AI) is the broad science of mimicking human abilities, and machine learning is a specific subset of AI that trains a machine how to learn. The most common approach in machine learning is supervised learning, where one asks a model to learn a mapping from an input to an output variable, e.g., a document x to a category y it belongs to. However, (x,y) pair is not always enough for describing the input-output relationship. Some information is unobserved, but plays an important role in modeling the relationship. For example, in question answering one is typically given the answer y for a training query x, but not the unobserved reasoning path leading to the answer. Similarly, in dialogue summarization, one may be given the dialogue transcript x and the labeled summary-worthy sentences y, but not the latent topic discussed in the conversation, which is helpful but missing from the given information. This thesis aims to develop and study novel paradigms that leverage latent information to bridge the gap between observed input data and the learning target. We apply latent variable modeling approaches on various real-world tasks. Our experimental results show that state-of-the-art results can be improved without adding much complexity to the model by just considering the latent information. In addition, studying latent information can help humans better understand the prediction process and thus improve the interpretability of the model. Finally, evaluation on downstream tasks provides evidences of the inferred latent information being successfully utilized to solve related tasks.--Author's abstract