Ready or Not, Here We Go

Timothy G. Dyster, Sameer Anil Sheth, Guy M. McKhann · Neurosurgery · 2016

For the last few decades, computational neuroscientists have devoted substantial resources to improving the performance of artificial intelligence (AI) on classic games, including chess, checkers, backgammon, and Scrabble. Expert-level play by AI has been achieved largely via algorithms that test all possible combinations of moves and outcomes and choose the move combination that optimizes the score. This brute force strategy is computationally expensive, however, and whereas it may be feasible for games with relatively limited numbers of moves, it is difficult to extend to increasingly complex games. In the ancient Chinese game Go, the number of possible board configurations increases rapidly as the game progresses, making it exceedingly more complicated than chess, in which the number of possible board configurations decreases with time. Thus, although the IBM supercomputer Deep Blue defeated reigning chess world champion Garry Kasparov 10 years ago, human Go experts have consistently outperformed AI—until now. In a recently published article in Nature, Silver et al1 described a novel decision-making algorithm that allowed AI to learn the game of Go and defeat the human European Go champion. This computing achievement, one of the “grand challenges” of AI,2-4 had not been projected to occur for another several years. Tackling this challenge was a fitting endeavor for the team from Google DeepMind, given the vastly larger parameter space of Go compared with chess, a difference measurable in factors of a googol (10100). The authors' approach uses deep neural networks (DNNs), computational techniques originally developed for processing complex visual information such as image classification or facial recognition. DNNs consist of layers of “neurons” (computational units) tiled on top of each other. Increasingly deep layers contain increasingly abstract representations of the actual image. In this rendition, the researchers passed the Go board configuration (ie, the pattern of tokens on the board) into the DNN as a 19 × 19 image with several layers. Each point in the 19 × 19 image represented a location on the game board, and each of the 48 layers of the image contained a different sphere of information relevant to game play (eg, which player occupied a location, how many turns had passed since a location had been occupied, and whether placing a piece at a location was legal by game rules). The DNN was then trained by a combination of machine learning strategies (Figure, A). These strategies allow an AI to “learn” a task by making future predictions based on previously observed patterns without being explicitly instructed how to respond. This step is crucial to imbue the AI with flexibility and efficiency. The first training step consisted of supervised learning, in which the DNN was instructed by moves chosen by an expert. The authors exposed the algorithm to 30 million positions, from which it could start building a decision policy. This allowed the AI to create a policy network that could somewhat accurately predict and mimic human expert game play. The second training step consisted of reinforcement learning, a type of machine learning in which actions that increase the likelihood of attaining a desired goal (eg, winning the game) are reinforced and those that stray from the goal are punished. Having completed the supervised learning step, the AI then played iteratively against itself over a million rounds, with the objective of maximizing the chance of winning against previous versions of its own policy. In the final step, regression was applied to the self-play data to create a value network that could predict the final outcome of a game based on any board position that the AI had encountered during self-play. After completion of this training pipeline, the AI, AlphaGo, was able to evaluate potential moves by combining its policy and value networks and then select its next move according to the results of an optimized search algorithm (Figure, B).Figure: Neural network training pipeline and architecture. A, during training, the deep neural network undergoes supervised learning (SL) and reinforcement learning (RL) to create a policy for selecting moves, which is then optimized by many iterations of self-play. Afterward, a value network is created by regression to predict game outcome on the basis of the board configuration. B, schematic representation of the neural network architecture used in AlphaGo. The board configuration is represented as a multilayer image. Next moves are selected by applying policy and value networks to generate a probability distribution from the image. Reprinted by permission from Macmillan Publishers Ltd: Nature (Silver D, Huang A, Maddison CJ, et al. Mastering the game of go with deep neural networks and tree search. Nature. 2016;529(7587):484-489), copyright 2016.The researchers evaluated AlphaGo by an internal tournament against other previously developed Go programs and against the human European Go champion, Fan Hui. Against other Go programs, AlphaGo won 494 of 495 games (99.8%), suggesting that AlphaGo achieved a significantly more expert level of play than prior programs. In the match against Fan Hui, AlphaGo won with a perfect 5-0 record. This was the first time AI defeated a human expert in the full game of Go without a handicap. One tantalizing question this study brings up is whether the decision-making algorithm used by AlphaGo or other similar DNN AIs actually resembles that of our own brain. The fact that AlphaGo can compete successfully with human experts suggests that its computational approach may indeed be similar to our own. Providing further support for this notion, there is strong evidence that our brains use reinforcement learning algorithms to help us make optimal decisions. As neurosurgeons, we have the ability to address fundamental questions about the computational properties of the human brain. Whether with acute intraoperative recordings during deep brain stimulation procedures, chronic inpatient recordings in patients with epilepsy, or long-term recordings in patients with telemetry-enabled brain stimulators, our opportunities for studying human neurophysiology continue to increase. In collaboration with computational neuroscientists, engineers, and modeling theorists, we can test predictions using direct recordings of the human brain spanning the scale from single neurons to widely distributed neural populations. We have certainly made progress in understanding the neural basis of decision making from rodent and nonhuman primate studies.5-7 However, it is not unreasonable to consider that our species' highest cognitive functions may be fundamentally different than those of even our phylogenetically closest neighbors. One of the challenges and great opportunities for our field in the next decade will be refining our understanding of the complex computations that cascade through our brain and make us who we are.

Read the paper · More papers on PaperTik