Learning by minimizing surprise — predictable pulses vs chaotic punishment
Transcript
This chapter, from the episode video's captions · 518 words
39:27get a cookie literally that is going to fuel dopamineergic neurons in our brain that are going to send signals of dopamine, pleasure, and all this other kind of stuff. So, we've got like emotional signals for reward. We've got other crazy signals for reward. This is a brain in a petri dish. It does not have dopamineergic neurons. It does not have these centers and this complex brain anatomy to have reward. >> Yes. >> So how are we going to do that? The re and this I think is actually the coolest part of the paper to me. >> Okay. >> Okay. It's this idea of a closed loop feedback architecture and the learning mechanism is literally minimizing surprise.
40:08>> Here's what they're doing. If there is a successful hit, meaning the the pong paddle thing hit the hit the ball. Yeah. >> Okay. And we're good to go. What it's going to do is >> all of the electrodes in my multi-elerode array is going to send out a predictable electrical pulse. It's about 10 volts on every single electrode. >> Okay. >> At 10 hertz. So 10 every second. Yeah. >> For about 50 milliseconds. Very short, very highly structured. >> Okay. >> Okay. >> If on the other hand, it misses, >> it's going to get 4 seconds at 20 volts
40:50of highly chaotic electrical activity. >> Mhm. >> And the idea is brains hate that. >> Yes. >> They want to be able to predict the future. Yes. >> Okay. And when you're giving them highly unpredictable stuff, even this naive network of biological neurons >> wants to decrease the amount of times it gets that. >> So if we if we look at the GIF, I just wanted to show one thing. So here we've got the brain coming in. You see that bump in electrical activity everywhere. And now it's about to miss. >> Y >> it's going to miss for 4 seconds. It's just like there's noise everywhere. >> Yes. >> Okay. We're feeding a bunch of electrical noise. The neurons hate that.
41:30Yes. >> Now the it's going to go it's going to actually Oh, I guess it's going to recycle. Let's recycle. >> So, it'll recycle. When it gets there, >> bunch of predictable electrical signal. That's all we're doing. >> Yes. >> Our reward mechanism is simply if you got it right, we're going to reward you with a bunch of predictable activity. >> Yep. >> 10 hertz, 50 milliseconds. And if you get it wrong, four seconds of random chaotic activity. >> And I I think this dovetales with some sort of historical uh science research that's been done around >> the you know when we look at how do we
42:11get ordered systems with entropy and this idea of everything goes to soup and the way in which that works is by minimizing surprise. Yes. this concept it's like a biological it is how life fundamentally is able to instantiate itself >> and that we're basically hijacking that natural kind of concept by artificially
From Can Human Neurons Really Play Doom? The Science Behind Wetware
Did a dish of human neurons really learn to play Doom—or is the wetware story more hype than breakthrough?