Tamil Text Error Correction with Multi-lingual T5 Model
Vaan Amuthu Elango, Peeta Basa Pati · 2023
Due to a variety of reasons, such as misspelling, typographical errors, and illiteracy, errors to arise in writing. Additionally, errors result from the conversion image or voice to text. MT5 is a well-known pre-trained multilingual transformer model and is employed correct errors in Tamil text. To improve model's capability, transfer learning is performed with Tamil dataset. The model's ability to handle various kinds and levels of errors for Tamil text is investigated. Character Error Rate (CER) & Word Error Rate (WER) are employed as the metrics to evaluate the model's performance. The model has shown promising results with improvement of 97.7% for CER and 89.3% for WER on the employed dataset.