Multi-Objective Molecular Property Optimization and Docking-Guided EGFR Inhibitor Discovery Using Deep Reinforcement Learning and Graph-Based Representations
Kanderi Johith Kumar, Kandlapalli Aravind Sai, Katragadda Megha Shyam, Malireddy Charan Kumar Reddy, Andrew Tom, Shinu M. Rajagopal, Thejus Varghese Thomas · Procedia Computer Science · 2026
Designing potent and selective Epidermal Growth Factor Receptor (EGFR) inhibitors requires simultaneous optimization of conflicting objectives including binding affinity, synthetic accessibility, and drug-likeness. This study introduces a unified computational framework coupling an Enhanced Dueling Deep Q-Network (DQN) with bidirectional LSTM and multi-head self-attention graph encodings for automated EGFR inhibitor generation. Trained on 12,000 curated EGFR-active and inactive compounds, the model jointly optimizes synthetic accessibility scores (SAS), quantitative estimates of drug-likeness (QED), and binding affinity predictions, achieving validation R² of 0.99±0.01 for SAS and 0.95±0.02 for QED with 8-12% improvements over baseline reinforcement learning and graph neural approaches. Reward-guided molecular exploration generated 50 top candidates with mean docking affinities of -13.8±0.5 kcal/mol against EGFR (PDB:2GS7), with 84% exceeding known inhibitor benchmarks. The framework addresses sparse reward challenges through dynamic parameter adjustment and experience replay mechanisms, significantly improving exploration-exploitation balance during training. Integration within the ChemViz platform provides real-time predictive analytics and interactive 2D/3D molecular visualizations, accelerating lead compound discovery while maintaining chemical validity and synthetic feasibility. This generalizable multi-objective optimization strategy offers substantial improvements in computational drug discovery efficiency, providing a robust platform for automated molecular design across diverse therapeutic targets.