Results — the neurons really did get better at Pong
Transcript
This chapter, from the episode video's captions · 849 words
42:37>> punishing the wrong decision >> with chaos which makes the system work harder >> and it is naturally going to want to not work harder and so that's why it chooses the right decision. It's not, it doesn't really know that it's the right decision. It is optimizing for the less >> en the more energyefficient. >> Yes. >> Pathway. >> Yes. The less surprising. >> The less surprising. >> Yeah. It's really the less surprising outcome is what it wants. Right. And now let's look at the results. >> Okay. >> Okay. With this training paradigm, let's look at the results. The apparent learning is demonstrated in 5 minutes. So here what we're looking at is two sets of box plots. Mhm.
43:17>> The green box plot is what happens from 0 to 5 minutes. >> Okay. >> Um we're looking at the change in rally length, which is like the, you know, the how many times the rally happens for for the pong. It's just like in tennis, like how many times do I keep the ball alive? Um and you can see from 6 to 20 minutes, there's a giant change in the size of the rallies from 0 to 5. So within 5 minutes, this thing is learning quite well. there's a up change in rally and that's for the stimulus for the controls which is silent and no feedback there's no change. Okay, so we're getting a significant change only when they're actually playing the game and they have this feedback. Now, if we go to the next
43:57>> photo, what we're also seeing is that human cortical neurons do a lot better >> than mouse cortical neurons. >> Okay. And that I don't know that makes me feel good that my neurons are better than a mouse's neurons. >> They are better, >> right? So, and that kind of makes sense. We're We've had evolution and our brains arguably are probably better than mouse. >> And the thing is there was an improvement in both cases. It's just our improvement had a >> had a had a big that's that's a good point. Very good. Yeah. It's that ours had a much bigger improvement than the >> compared to the mice, >> right? So now that we have that, let's actually look at this free energy principle that you were alluding to a bit earlier. >> Oh, interesting. Okay.
44:38>> That's what it's called. The free energy principle. This is was it was first pioneered by Carl Fristen and the idea is what you were saying all self-organizing biological systems want to minimize something called the free energy the variational free energy in other sense what we want to do is maximize the information entropy and we want to minimize the surprise and the prediction error okay >> what we can look at is how much surprise there is for what is the probability of a certain state. >> Mhm. >> If the probability of a certain state is very low, there's a lot of surprise, right? And you can quantify that in the same way that we actually quantify
45:18entropy in physics, which is the log of the probability of stuff. >> And effectively what the free energy principle is saying is that biological systems want to maintain homeostasis, meaning like constancy >> over all the random stuff. And we want to minimize that surprise, right? >> And what do we what do we do if we want to do that? Well, we've got some internal model of what the world is. And we are going to tune the parameters of that internal model such that that internal model mimics more closely the world that we live in. Right? That is the idea. It's a stark departure, I should say,
46:00and highlight from traditional reinforcement learning. >> Okay. >> Okay. In tra in traditional reinforcement learning, you've got some algorithm and that algorithm explicitly gives negative feedback through back propagation for wrong trials. Here we are letting chaos itself >> be the negative feedback. >> Yeah. Yeah. Yeah. >> Now, one of the things that's kind of crazy is that the algorithms for reinforcement learning require thousands of epochs of trial and error >> to update these weights through back propagation. Right? Biological systems will inherently self-organize to minimize this free energy and they'll do it very very quickly. They've just got
46:41really nice cellular mechanisms to find that sweet spot. >> This goes back >> in their internal model, >> the energy efficiency piece again. >> Yes. Exactly. Right. And so maybe that's why we want to use the wet wear >> in the first place. Right. So let's go back to this FEP in a dish. Our free energy principle in a dish. Effectively what's happening is you've got predictions and then you've got prediction errors. The predictions are where I want to move my >> paddle. And the prediction error is the chaos that that I get when I get it wrong. Right? I'm trying to change the
From Can Human Neurons Really Play Doom? The Science Behind Wetware
Did a dish of human neurons really learn to play Doom—or is the wetware story more hype than breakthrough?