Unsupervised Context-Sensitive Spelling Correction of Clinical Free-Text with Word and Character N-Gram Embeddings

Pieter Fivez, Simon Šuster, Walter M. P. Daelemans · 2017

We present an unsupervised contextsensitive spelling correction method for clinical free-text that uses word and character n-gram embeddings.Our method generates misspelling replacement candidates and ranks them according to their semantic fit, by calculating a weighted cosine similarity between the vectorized representation of a candidate and the misspelling context.We greatly outperform two baseline off-the-shelf spelling correction tools on a manually annotated MIMIC-III test set, and counter the frequency bias of an optimized noisy channel model, showing that neural embeddings can be successfully exploited to include context-awareness in a spelling correction model.Our source code, including a script to extract the annotated test data, can be found at https://github.com/pieterfivez/bionlp2017.

Read the paper · More papers on PaperTik