Automatic translation of natural language requirements into CTL specifications using Large Language Models: A multi-approach evaluation
Rim Zrelli, Henrique Amaral Misson, Marwa Ben Attia, Felipe Göhring de Magalhães, Abdo Shabah, Gabriela Nicolescu · Journal of Systems and Software · 2026
Translating natural language (NL) requirements into formal specifications such as Computation Tree Logic (CTL) is essential for improving the efficiency and scalability of formal verification, especially in safety-critical systems. This study evaluates the ability of Large Language Models (LLMs) to automate this process. We compare three approaches: fine-tuning the Mistral model, using GPT-4 in a few-shot learning setup, and a hybrid that feeds a BERT pattern classifier’s prediction to GPT-4. Using the Natural2CTL dataset, we assess strict logical accuracy and an ambiguity-tolerant accuracy, complemented by auxiliary semantic and structural operator similarity measures. Fine-tuning yields the strongest strict correctness and operator-structure fidelity, while the hybrid narrows the gap to fine-tuning and substantially improves over few-shot prompting alone. Residual errors across methods concentrate in path-quantifier selection, temporal granularity, and scoping in multi-clause requirements. Overall, LLMs can draft CTL candidates that are usable after lightweight normalisation, but they should be integrated into human-in-the-loop workflows with basic automated checks before use in high-assurance settings. • Benchmarks three LLM-based methods for NL-to-CTL translation. • Fine-tuned Mistral achieves 47.6% strict logical accuracy and 71.4% ambiguity-tolerant accuracy for CTL specification generation. • GPT-4 few-shot learning offers rapid prototyping but lower syntactic precision. • BERT-GPT hybrid balances pattern recognition and generative translation. • LLM automation reduces expert effort, but expert review remains critical for safety.