Reinforcement Learning in de novo Molecular Design: Comparative Review of REINVENT 4, MolDQN, and ORGANIC
Arjun Subramanian, Connor Taylor · Ambient intelligence and smart environments · 2025
The capacity of computer-aided drug development to effectively explore vast chemical regions has led to an explosion in recent years. Reinforcement learning (RL), which learns through trial-and-error interactions, is uniquely suited to drug discovery and adaptive decision-making in Intelligent Environments (IEs). Three prominent reinforcement learning-based models for de novo molecular generation are examined in this review: REINVENT, a policy-based method that optimizes molecular design using an RNN and a prior network; MolDQN, a DQN-based method that presents molecular modification as a sequential decision process; and ORGANIC, a GAN-based model that learns to generate molecules by differentiating between high- and low-quality candidates. Comparing model architectures and key findings provide valuable information about the real-world uses and future direction of reinforcement learning in drug design. The adaptive and multi-modal nature RL lends itself to applications IE including personalized medicine and environmental toxicology.