Extracting Named Entities from Russian-Language Documents with Varying Degrees of Structural Clarity

Maria D. Averina, O. A. Levanova · Automatic Control and Computer Sciences · 2024

Abstract— This study addresses the task of recognizing named entities in Russian texts using the CRF model. We analyze two datasets: well-structured refinancing documents and loosely structured court transcripts. We test the model with various text features and CRF parameters (optimization algorithms). On average, the best F-measure for well-structured documents is 0.99, while for loosely structured ones, it is 0.86.

Read the paper · More papers on PaperTik