Render, Encode, Plan: A simple pipeline for hybrid RL-DL learning inside Unreal Engine
Daniele Della Pietra, Nicola Garau · Computers & Graphics · 2025
Learning is an iterative process that requires multiple forms of interaction with the environment. During learning, we experience the world through the repetition of observations and actions, gaining an insight into which combination of these leads to the best results, according to our goals. The same paradigm has been applied to traditional reinforcement learning (RL) over the years, with impressive results in 3D navigation and planning. On the other hand, the computer vision community has been focusing mostly on vision-related tasks (e.g. classification, segmentation, depth estimation) using deep learning (DL). We present REP: Render, Encode, Plan , a unified framework to train embodied agents of different kinds (humanoids, vehicles, and drones) inside Unreal Engine, showing how a combination of RL and DL can help to shape intelligent agents that can better sense the surrounding environment. The main advantage of our method is the combination of different sensory modalities, including game state observations and vision features, that allow the agents to share a similar structure in their observations and rewards, while defining separate rewards based on their goals. We demonstrate impressive generalization capabilities on large-scale realistic 3D environments and on multiple dynamically changing scenarios, with different goals and rewards. All code, complete experiments, and environments will be available at https://mmlab-cv.github.io/REP/ . • Hybrid RL-DL : REP trains agents in UE5 using visual and physical observations. • Multi-agent, multi-modal : Supports drones, cars, humans, and multiple tasks. • UE5-integrated : Uses NNE and shared memory for real-time training and inference. • Robust generalization : Solves diverse tasks in large, dynamic 3D environments.