Document-level event extraction from Italian crime news using minimal data

Giovanni Bonisoli, David Vilares, Federica Rollo, Laura Po · Knowledge-Based Systems · 2025

Event extraction from unstructured text is a critical task in natural language processing, often requiring substantial annotated data. This study presents an approach to document-level event extraction applied to Italian crime news, utilizing large language models (LLMs) with minimal labeled data. Our method leverages zero-shot prompting and in-context learning to effectively extract relevant event information. We address three key challenges: (1) identifying text spans corresponding to event entities, (2) associating related spans dispersed throughout the text with the same entity, and (3) formatting the extracted data into a structured JSON. The findings are promising: LLMs achieve an F1-score of approximately 60% for detecting event-related text spans, demonstrating their potential even in resource-constrained settings. This work represents a significant advancement in utilizing LLMs for tasks traditionally dependent on extensive data, showing that meaningful results are achievable with minimal data annotation. Additionally, the proposed approach outperforms several baselines, confirming its robustness and adaptability to various event extraction scenarios. • Novel use of LLMs for event extraction from Italian news with minimal annotated data. • In-context learning outperforms zero-shot prompting for identifying event spans. • Mixtral achieves top performance in event extraction from Italian crime news. • LLMs outperform QA models, proving robust in data-scarce environments. • In-context learning performance depends on the quality of selected examples.

Read the paper · More papers on PaperTik