Machine learning in interface design for an audio browser

Amanda J. Stent, I. V. Ramakrishnan, Zan Sun · 2006

Machine learning techniques are helpful for problems with unclear solutions, or where improvements come from experience. Machine learning has been widely used in the fields of data mining, natural language processing, pattern recognition and so on. The focus of this thesis is to apply machine learning techniques in our new speech-driven Web browser system—HearSay2. The task of HearSay2 is to provide an efficient and clear representation of Web pages to users, using speech plus text input and speech output. HearSay2 analyzes a Web page and organizes its contents into a hierarchical structure of related concepts. Compared to the common screen reader systems, HearSay2 can facilitate users' experience by presenting a structured view of a page, and allowing users to jump to the desired section without being overwhelmed by the preceding content. HearSay2 uses special algorithms that were designed for analyzing Web pages automatically. They were based on the observation that semantically related content are usually presented in a similar style. Once the structure of the Web page is obtained, the content is populated to a set of pre-designed VoiceXML templates that make up the voice interface of HearSay2. In this thesis I focus on the tasks HearSay2 must perform after Web page analysis finishes, which are: (a) to summarize and label the textual content, (b) to extract the instances of related concepts for domain-specific navigation, and (c) to choose an appropriate dialog template for each piece of content. These tasks are mostly related to the ease of use of the interface. Instead of using heuristic rules for these tasks, machine learning can solve them efficiently without losing generality. In this thesis, I applied state-to-art machine learning techniques including SVM, Bayesian methods, and decision trees to perform these tasks.

Read the paper · More papers on PaperTik