Parallel Transfer Learning: Accelerating Reinforcement Learning in Multi-Agent Systems

Adam Taylor · Trinity's Access to Research Output (TARA) (Trinity College Dublin) · 2015

Learning-based approaches to autonomic systems allow systems to adapt their behaviour to best suit their operating environment.One of the more widely used learning methods is Reinforcement Learning (RL).RL agents learn by repeatedly executing actions and observing their results, and over time a representation of how to behave well is developed.A significant issue with this approach is that it takes a long time to reach its best performance.Every action has to be experienced several times in each particular circumstance for its value to be representative of its converged value.The presence of multiple agents increases the number of experiences required.Multi-Agent Systems (MAS) are systems in which multiple agents affect a shared environment.MAS are inherently large-scale as they are composed of many interacting agents.In MAS, the cumulative effects of agents' actions make the environment more variable.Greater variability in the outcomes of actions requires more learning as a single sample becomes less representative of the converged value.An RL system's performance is necessarily sub-optimal while it is learning.Each learning experience is expensive, taking time and affecting the system that is being controlled.Experiences should be used as efficiently as possible to reduce the time spent learning, improving overall performance.The less time needed to learn, the better the performance of a system over its lifetime.Transfer Learning (TL) is a method of using additional knowledge to accelerate learning.It operates by taking knowledge from a source task (a process that supplies information) and reusing it in a target problem, with the aim of reducing the amount of vi learning to be done.Inter-Task Mappings (ITM) are used to allow more diverse tasks to share knowledge, effectively translating knowledge so that it is mutually intelligible.TL has shown promising results, but it requires a learnt source of information prior to the execution of the target system.This means it can not operate in real-time.This thesis addresses TL's limitations by allowing the source of learnt information and the task to be accelerated to run concurrently.The main contribution, Parallel Transfer Learning (PTL), enables different agents to support each other's learning through mutual knowledge exchange.This is particularly beneficial in MAS, as agents are naturally concurrent and typically learn broadly similar things when in the same environment, so there is likely useful information to share.PTL can reuse useful knowledge in several different ways, each designed for particular types of environment.Detecting the type of environment, and if it is changing, allows PTL to self-configure for a particular system at a given time.PTL accomplishes this by modelling its performance over time and reacting to divergences in performance levels.PTL can transfer information between more diverse agents using ITM, which can be learnt in real-time by having the source and target share information.PTL's evaluation is twofold: fundamental aspects are examined in simple environments, while overall effectiveness is evaluated in an example MAS, the Smart Grid.PTL is evaluated against standard RL, to quantify any improvement in learning time.The results show that transferred information can accelerate learning and it is of particular benefit when agents learn in different situations.When agents are in similar situations, comparable performance can be achieved in 11.11% of the time in one application and 40% in another.The learnt ITM allow improvement in homogeneous tasks and can find an effective mapping in the heterogeneous case.PTL can perform well in changing environments as long as the change stops and knowledge can by learnt.

Read the paper · More papers on PaperTik