RLSF: Fine-tuning LLMs via Symbolic Feedback

Piyush Jha, Prithwish Kumar Jana, Pranavkrishna Suresh, Arnav Arora, Vijay S Ganesh · Frontiers in artificial intelligence and applications · 2025

Large Language Models (LLMs) have transformed AI but often struggle with tasks that require domain-specific reasoning and logical alignment. Traditional fine-tuning methods do not leverage the vast amount of symbolic domain-knowledge available to us via symbolic reasoning tools (e.g., provers), and are further limited by sparse rewards and unreliable reward models. We introduce Reinforcement Learning via Symbolic Feedback (RLSF), a novel fine-tuning paradigm where symbolic reasoning tools (e.g., solvers, provers, and algebra systems) provide fine-grained feedback to LLMs. RLSF uses poly-sized certificates (e.g., proofs) generated by symbolic tools to identify and correct errors in model outputs, offering token-level guidance without requiring differentiable reasoning systems. This paradigm bridges the gap between symbolic reasoning and LLM fine-tuning, enabling precise alignment with domain-specific constraints while addressing key limitations of traditional reward signals. Via extensive evaluations, we show that our RLSF-based fine-tuning of LLMs outperforms traditional approaches on five different applications (that have some associated logical or domain constraints), namely, program synthesis from natural language pseudo-code to programming language (+31.43% in functional correctness for Google’s CodeGemma-2b compared to supervised fine-tuning, +17.01% in functional correctness compared to GPT-3.5 – 100× larger), three chemistry tasks (+5.5% exact match for molecule generation, +19.4% exact match for forward synthesis, +33.7% exact match for retrosynthesis, using Meta’s Galactica-1.3b, compared to GPT-4 – 1000× larger), and solving the Game of 24 (+25% success rate using Meta’s Llama2-7b compared to traditional methods, and +7% success rate compared to GPT-3.5 – 25× larger). A key takeaway is that fine-tuning via RLSF enables relatively smaller LLMs to significantly outperform closed-source models that are orders of magnitude larger.

Read the paper · More papers on PaperTik