Learning Representations of Orthographic Word Forms

Christopher T. Kello, Daragh E. Sibley · eScholarship (California Digital Library) · 2006

Presented is an extension of the simple recurrent network (SRN), termed the sequence encoder, which learns fixedwidth representations of variable-length sequences.This architecture was used to learn orthographic representations for nearly 75,000 English words, of which nearly 69,000 were multisyllabic.Analyses showed that sequence encoder representations are shaped by the dependencies among letters in English word forms that reflect orthographic structure.The model was used to predict participant ratings of the orthographic legality of pseudowords, and results showed that the model accounted for a substantial amount of variance in the ratings.

Read the paper · More papers on PaperTik