Enhancing Mathematical Problem Solving in Large Language Models through Tool-Integrated Reasoning and Python Code Execution

Siyue Li · 2024

Mathematical problem solving remains a significant challenge for large language models (LLMs) due to the inherent complexity of mathematical reasoning and the precision required in calculations. This paper introduces a novel approach that leverages Tool-Integrated Reasoning (TIR) and Python code execution to address these challenges and enhance the mathematical reasoning capabilities of LLMs. We fine-tune the DeepSeekMath-Base 7B model using a two-stage process. The first stage involves training on a diverse dataset of natural language mathematical problems using a Chain of Thought (CoT) template, improving the model’s understanding of reasoning paths. In the second stage, we introduce a synthetic dataset that includes examples of tool-integrated reasoning, enabling the model to generate and execute Python code as part of its problem-solving process. Our results demonstrate significant performance improvements, with the model achieving an accuracy of 0.7782 and a maj@N score of 0.7344, outperforming existing models in the domain. This research not only advances the state of mathematical reasoning in LLMs but also provides a framework for integrating external tools to improve reasoning and decision-making processes in AI systems.

Read the paper · More papers on PaperTik