PANDA - Discovering Part Name in Noisy Text Data

Anne Kao, Nobal Bikram Niraula, Daniel Whyatt · 2018

Part identification plays a key role in vehicle prognostics and health management. Part identifiers are often expressed as nomenclature and buried in noisy free text data found in maintenance reports, supply chain management records, service and support communication logs, and manufacturing quality data. There is little consistency in how part names are actually described in noisy free text, with variations spawned by typos, ad hoc abbreviations, acronyms, and incomplete names. This makes search and analysis of parts involved in this data extremely challenging. In this paper, we will discuss our method and tool PANDA (PArt Name Discovery Analytics), based on a unique method that exploits statistical, linguistic and machine learning techniques in a unique way to discover part names in noisy free text. The algorithm is very scalable and efficient, and provides actionable results for analysis of vehicle health management.

Read the paper · More papers on PaperTik