vix.ing · top · new · best · stats · spec

Reward-based learning under hardware constraints - Using a RISC\n processor embedded in a neuromorphic substrate

2013/03/26 by Simon Friedmann, Nicolas Frémaux, Friedmann, Simon +7
Engineering · Neuroscience · #Advanced Memory and Neural Computing #Neural dynamics and brain function #CCD and CMOS Imaging Sensors

paper · pdf · doi:10.48550/arxiv.1303.6708

Abstract

In this study, we propose and analyze in simulations a new, highly flexible\nmethod of implementing synaptic plasticity in a wafer-scale, accelerated\nneuromorphic hardware system. The study focuses on globally modulated STDP, as\na special use-case of this method. Flexibility is achieved by embedding a\ngeneral-purpose processor dedicated to plasticity into the wafer. To evaluate\nthe suitability of the proposed system, we use a reward modulated STDP rule in\na spike train learning task. A single layer of neurons is trained to fire at\nspecific points in time with only the reward as feedback. This model is\nsimulated to measure its performance, i.e. the increase in received reward\nafter learning. Using this performance as baseline, we then simulate the model\nwith various constraints imposed by the proposed implementation and compare the\nperformance. The simulated constraints include discretized synaptic weights, a\nrestricted interface between analog synapses and embedded processor, and\nmismatch of analog circuits. We find that probabilistic updates can increase\nthe performance of low-resolution weights, a simple interface between analog\nsynapses and processor is sufficient for learning, and performance is\ninsensitive to mismatch. Further, we consider communication latency between\nwafer and the conventional control computer system that is simulating the\nenvironment. This latency increases the delay, with which the reward is sent to\nthe embedded processor. Because of the time continuous operation of the analog\nsynapses, delay can cause a deviation of the updates as compared to the not\ndelayed situation. We find that for highly accelerated systems latency has to\nbe kept to a minimum. This study demonstrates the suitability of the proposed\nimplementation to emulate the selected reward modulated STDP learning rule.\n

Citations

Related