PresiUniv at TSAR-2022 Shared Task: Generation and Ranking of Simplification Substitutes of Complex Words in Multiple Languages

Peniel Whistely, Sandeep Mathias, Galiveeti Poornima · 2022

In this paper, we describe our system, Pre-siUniv, to generate and rank candidate simplifications using publicly available pre-trained language models (BERT, BETO, and BERTimbeau), word embeddings (Eg.FastText, NILC), and part-of-speech taggers (NLTK PoS Tagger, Stanford PoS Tagger and Mac-Morpho), to generate and rank candidate contextual simplifications for a given complex word.In this shared task, our system was placed first in the Spanish track, 5th in the Brazilian-Portuguese track, and 10th in the English track.We upload our codes and data for this project to aid in replication of our results.We also analyze some of the errors and describe design decisions which we took while writing the paper.

Read the paper · More papers on PaperTik