Reviewer #1 (Public review): Online reinforcement learning of state representation in recurrent network supported by the power of random feedback and biological constraints
2025
Recurrent neural network and its readout (cortex–striatum) can learn state representation and value using online random-weight feedback of temporal-difference reward-prediction-error (dopamine) through feedback alignment or biological non-negative-weight constraint-induced loose alignment.