Low energy implementation of feedforward neural network with backpropagation algorithm using a spin orbit torque driven skyrmionic device

U. Saxena, D. Kaushik, M. Bansal, H. Chandel, U. Sahu, D. Bhowmik · 2018 IEEE International Magnetics Conference (INTERMAG) · 2018

Non-volatility of spintronic devices opens up possibility of implementation of Artificial Neural Networks (ANN) in hardware and thereby take advantage of the parallel architecture of the system similar to human brain [1, 2, 3]. Spin orbit torque driven domain wall based devices that act as synapse have been shown to be much more energy efficient than CMOS devices that implement neural networks [2], [4]. Here, we have considered the role of defects in ferromagnetic layer and thereby proposed a spin orbit torque driven skyrmionic device, which consumes even lower energy compared to domain wall based device for synaptic behavior. We perform micromagnetic simulations on mumax3 [5]to demonstrate motion of Neel skyrmion and transverse Neel domain wall in a ferromagnet layer of thickness 1 nm, driven by spin orbit torque from current flowing through layer of heavy metal (Pt, spin Hall angle = 0.07) underneath [6, 8, 9]. We observe (as shown in Fig. 1a)that in the presence of triangular notch defects of 6 nm depth and modified anisotropy 1.2 MJ$/ \mathrm {m}^{3}($the magnet has uniform perpendicular anisotropy of 0.8 MJ$/ \mathrm {m}^{3}$elsewhere) on the two edges of the magnet at a separation of 60 nm each, domain wall is pinned up to a current density of 24 MA/cm2through the heavy metal. But a skyrmion is depinned by current density as low as 1 MA/ cm2though its velocity at such small current density is also very small [6], [7]. Based on this result, we propose a skyrmionic device as shown in Fig. 1b, which can act as a synapse. A Magnetic Tunnel Junction (MTJ) structure is present at the right end of the device. Based on magnitude and duration of current pulse through heavy metal layer (applied between terminal T3 and T1) certain number of skyrmions moves to the region of the ferromagnet below the MTJ. Hence Tunneling Magneto-Resistance of MTJ measured between terminal T2 and T1 is a function of the current between T3 and T1 and corresponds to the weight of the synapse, which can be controlled by the current and stored subsequently since skyrmions won't move once current pulse is removed just like domain walls (Fig. 1b). The resistance vs current behavior of the skyrmionic synapse (Fig. 1c), obtained through micromagnetic simulations, shows that in the presence of defects a very small range of current $( 3.5 \mu \mathrm {A}$to $6.5 \mu \mathrm {A})$can modulate the resistance from its lowest to highest value, corresponding to weight of synapse varying from -1 to 1. On the other hand a much larger range of current ($- 300 \mu \mathrm {A}$to $380 \mu \mathrm {A})$is needed to do the same for domain wall based synapse proposed in Ref. 4 if defects are present. Next we solve the standard digit recognition problem (Fig. 2b)by simulating a feedforward neural network with backpropagation algorithm [10](Fig. 2a)using the domain wall device (Ref. 4) or skyrmion based device (Fig. 1b)as synapses. A three layer feedforward network is trained to identify digits 0–4 across 10 computer generated variations for each digit, that mimic variations due to different handwriting. 35,000 iterations are needed to train the network and total number of synapses is 425. Error generated at the output layer after every iteration needs to be used to change the weights of all the synapses for subsequent iteration through the back-propagation algorithm. This is implemented by passing currents proportional to the generated error through the heavy metal layer of the synaptic devices (between terminal T3 and T1) and move the domain wall/ skyrmion to change the resistance of the MTJ (between T2 and T1) and thus change the weight value of the synapse.The energy dissipation due to this current flow in the heavy metal layer is called the “write” energy consumption, which is the most dominant energy contribution in the system. Fig. 2cis our key result which compares the total “write” energy consumption to train the network with synaptic devices being of the following four types: i. domain wall without defect (write current pulse width: 1 ns) ii. Skyrmion without defect (write pulse width: 15 ns) iii. Domain wall with defect (write pulse width: 0.5 ns) iv. Skyrmion with defect (write pulse width: 550 ns). We see that in the absence of defect, domain wall synapses consume less energy than skyrmion based synapses because domain wall moves faster than skyrmion for a given current density (Fig. 1a). But in presence of defects, since domain wall is pinned till current density of $\sim 24$MA/ cm2while skyrmions do not get pinned at current density even much smaller than that, if we “write” the skyrmion synapses with current pulses of very long duration $( \ge 500$ns) and very small magnitude $(3- 6 \mu \mathrm {A})$the skyrmion synapse based network can be trained at 2 orders of magnitude lower energy consumption than domain wall synapse based network. Thus in this paper we propose and simulate a skyrmion based synaptic device and show it to be much more energy efficient than domain wall based device. It will be particularly suitable for solving problems where time is not a major constraint.

Read the paper · More papers on PaperTik