Large Language Models for translating contract-related texts to logical predicates: Prompting, fine-tuning or dedicated library?

María Navas-Loro, Hideaki Takeda, Ken Satoh · Expert Systems with Applications · 2026

• Two-step approach (Named Entity Recognition + Rules) performed better than a direct code generation approach for natural language to PROLEG translation. • Regarding Large Language Models (LLMs), code LLMs (designed for code generation) did not show better results than general-purpose LLMs. • Llama 3 70B obtained better results for direct code generation than other models, including Llama 3.1 70B. • A NER library (GliNER) showed better performance than state-of-the-art Large Language Models, among which Llama 3.1 70B achieves the best results. This study presents a comparative analysis of methodologies for automating the extraction of legal information from contract texts and expressing it as PROLEG logical clauses, focusing on two approaches: Direct Code Generation (DCG) via Large Language Model (LLM) prompting, and Named Entity Recognition (NER) with rule-based transformation. For the NER-based approach, we evaluated three implementations: (1) LLM prompting for entity recognition, (2) a fine-tuned LLM for NER, and (3) a specialized NER library. Our findings reveal that while LLMs demonstrate versatility in general NLP tasks, the highest precision and contextual adaptation were achieved using a NER library specifically tailored to handle named entities. For finding named entities relevant to the contract (including, for instance, the buyer or the rescission date of a contract), this approach outperformed both DCG and NER using LLMs. The results underscore the importance of tailored lexical resources and rule-based post-processing in legal NLP applications, suggesting a paradigm for optimizing automated contract analysis systems, as opposed to the current trend of using general-purpose LLMs for most NLP tasks.

Read the paper · More papers on PaperTik