From Self-Learning to Self-Evolving Architectures in Large Language Models: A Short Survey

Ranjan Sapkota, Konstantinos I. Roumeliotis, Sushil Pokhrel, Manoj Karkee · 2025

This short survey provides a focused overview of the evolution of self-learning in large language models (LLMs), with an emphasis on emerging self-evolving LLM architectures. We trace the progression from traditional approaches such as supervised fine-tuning and reinforcement learning from human feedback (RLHF) to more advanced methods including self-instruct, self-play, and reinforcement learning with verifiable rewards (RLVR). Approaches are categorized by their data dependency, feedback mechanisms, curriculum adaptability, and scalability. A central focus is the Absolute Zero Reasoner (AZR), a recent framework that exemplifies absolute zero reasoning a fully autonomous, closed-loop system that requires no external training data. AZR introduces a dual-policy architecture where the model generates tasks, solves them, and verifies outputs using code execution. It achieves state-of-the-art performance in coding and math, while exhibiting emergent reasoning abilities such as deduction, induction, and abduction. We highlight AZR’s unique strengths in curriculum self-optimization, domain generalization, and scalable lifelong learning. The survey concludes with open challenges in alignment, interpretability, and policy, and outlines future directions in verifiable AI, meta-learning, and responsible AGI development.

Read the paper · More papers on PaperTik