Roman Concrete, Brain "Cognitive Legos," DeepSeek, and Econophysics
EP 21
·1:25:04

The memory wall: kernels, throughput, and systems tricks

Watch Roman Concrete, Brain "Cognitive Legos," DeepSeek, and Econophysics

Transcript

This chapter, from the episode video's captions · 824 words

1:25:04tech companies within China get to that point. They're getting they're trying to get there. >> They're trying to get there. >> They're trying to get there, but they have not yet. And that's still the bottleneck. >> Yeah. And I think I think very recently the Trump administration said that they're going to allow Nvidia to sell to China but they'd have to pay like a 25% tariff that would go to the government. So the government would like share in the profit. Something like that. I'm I'm not quite clear. But in any case, what we're trying to do is like choke out, >> right, >> the Chinese AI industry, right, by starving them of GPUs. So the Deep Sea guys have to be very clever. >> Yes. >> In how they implement this because

1:25:45here's the idea, right? Modern GPUs um like the Nvidia H100s, they're characterized by a massive disparity between how much they can compute. So the compute capability which is in number of flops, floatingoint operations and their memory bandwidth in like terabytes per second. Okay, this is called the memory wall which it's refers to the latency and the energy costs associated with like moving data from memory which is this high bandwidth RAM to like your spam registers which is the thing that >> h the computing part of the GPU has access to right and so it's really slow to like transfer stuff

1:26:26>> and if you're starved of GPUs you need to come up with clever ways to if you want to do something very cool like this manifold old hyperconnection thingy thingy thingy, but you're not you don't have a lot of GPUs. You got to get clever, right? And that part of the paper I actually thought was really cool, right? So, you've got this bottleneck, right? And >> if you want to naively do something like this hyperconnection stuff, >> you know, you're trying to before you just had a single lane, now you've got let's say four lanes. That's four times the compute that you have to do because you have to figure out you have to you have to read your your input then you have to churn the numbers four times

1:27:07>> right >> and then that's expanding a lot of a lot of >> comput four to four to five times higher than the original ResNet >> right >> okay and when you're starved with GPUs four to five times higher input output is not >> something that >> it's too inefficient >> yeah you can't just by >> right >> like 500 more right so what deepseat did was reduce that computational overhead to something that was negligible six to 6.7%. really >> instead of four to five times. And here's how they did it. They use this language called tile lang. >> Okay. >> Okay. And what they're doing is literally overcoming the memory wall by

1:27:49something called kernel fusion. So whenever you're doing AI, let's say you do like some kind of operation, you have to like start up a kernel. >> Yes. >> It's it's just like it's kind of like a processing unit. >> Yes. >> Okay. And individual kernels are started up for individual tasks. like there's matrix multiplication, there's the activation and stuff like that. But what these guys did was customize it so that all of the math is done in a single kernel faster arithmetic and the intermediate results don't go to memory back and forth. So you don't like you don't like do you don't like go to memory, get some data, do something, put it back in memory, go somewhere else, put it back. It's kind of like if you're

1:28:29like a chef, like an industrial chef, you don't go to the fridge all the time, right? It's like you get all your ingredients, you put it on your stove top. The fridge is like your high bandwidth memory. The stove top is like your SRAMM, which is the countertop. So that's it's it's limited, but you want to do everything you can in this in this little time. >> And the kernel is your instructions. But if you can make your kernel very complicated and clever, yes, >> then I don't have to go visit the fridge all the time, right? And that's exactly what they did. reduce the global memory reads by a huge amount. >> That's fascinating. Which then which then opens up the memory like the

1:29:09bandwidth on memory because you're not constantly going back to store and retrieve in time as you're doing any number of different operations or processes. >> Yeah. And the proof the proof is in the pudding. They've got a smoking gun. They use the 27 billion parameter model. The loss is lower. The gradients aren't like blowing up. Um, so it it is it is very

From Roman Concrete, Brain "Cognitive Legos," DeepSeek, and Econophysics

Roman concrete, compositional brains, DeepSeek scaling, and market impact physics.