Scaling laws + the stability bottleneck in deep nets

Transcript
This chapter, from the episode video's captions · 180 words
1:09:43>> Um scaling laws are very effectively more compute, more parameters, better performance. >> If you have bigger, better stuff, you spend a lot of money. All more chips, bigger, better, faster chips means you can get to the next stage of >> Exactly. of >> whatever you're trying to scale up. >> Yeah. Yeah. Exactly. But there is a caveat, right? Which is at massive scale, you've got like hundreds of billions of parameters. I think now we're getting to trillions of parameters. You're going to hit something called a stability limit. Okay? And in deep neural networks, it's all about signal propagation, right? Which is like there's a signal of did I get something right or did I get something wrong? And as that propagates
1:10:24through my network, it's very chaotic, right? And it can lead to amplification, vanishing gradient problem, exploding gradient problem. These are stuff that we're going to talk about. So there's some kind of limit, right? There might be some kind of like thermodynamic limit to how much we can do. And the solution
From Roman Concrete, Brain "Cognitive Legos," DeepSeek, and Econophysics
Roman concrete, compositional brains, DeepSeek scaling, and market impact physics.