Reinforcement Learning for Autonomous Agents via Deep Q Networks (DQN) within ViZDoom environment

Enoch Solomon · 2024

In the realm of game playing, deep reinforcement learning predominantly relies on visual input to map states to actions. The visual data extracted from the game environment serves as the primary foundation for state representation in reinforcement learning agents. However, humans leverage additional sensory inputs, such as audio cues, which play a pivotal role in perception and decision-making. Hence, the inclusion of raw audio alongside visual information holds promise in providing valuable insights to reinforcement learning agents. This study advocates to the enhancement of visual data in state representation by integrating with raw audio samples as complementary information. By melding raw audio with visual cues, our objective is to enrich the decision-making process of the agent at each stage. Experimental assessments were conducted employing Deep Q Networks (DQN) within ViZDoom environment. The results of our experiments reveal that augmenting visual information with raw audio samples yields superior rewards and expedites the learning rate compared to relying solely on visual data. This study underscores the potential advantages of incorporating multisensory information, particularly raw audio, into the state representation of reinforcement learning agents.

Read the paper · More papers on PaperTik