EP 51 · 1:11:25

The future of high-temperature computing

From The Tech Elon Has Been Waiting For

Episode
16/26
Watch The Tech Elon Has Been Waiting For
Transcript

491 words · auto-generated from the episode video

1:11:25>> right? >> So now let's talk about the future. Okay, what would this current research have to do with how we move forward in computation as a society? We've talked about the vonoman bottleneck for a long time. This is the tyranny of the processor having to be separate from our memory. Modern architectures waste a lot of power and time trying to move data between our CPU which is where all of the computation is happening and the RAM which is where the data and the memory is stored. This was vonoman came up with this and it's the bedrock of almost all of the computation that we have in our society today. Now, with

1:12:06neural networks, this is especially bad because you've got to load the weights and then you've got to do the computation, then you've got to load a new set of weights. So, you got to dump the old weights, right? And all of this takes a lot of time. It takes a lot of energy, dissipates heat, and what exactly is the computation that happens in these AI systems? When you've got a neural network, what we're really doing fundamentally is we're relying on matrix times vector multiplication. The vector is your state and the matrix is the weights of your network. You multiply the matrix. You multiply the matrix to the vector and then you get a new vector

1:12:46that tells you how the state is changing. At the end of the whole thing, you get a giant vector. In in in the example of of LLMs, you have a giant vector with all of the next possible tokens and you pick usually the one with the highest probability. Okay, how does matrix multiplication work? How do you multiply a matrix by a vector? For those from linear algebra, you already know this, but let's just do a little bit of a review. Okay, let's say I've got a 3x3 matrix on the left there. I've got 112 213142 and I want to multiply it by a vector 312. 312 is the vector that's my state of like next tokens, let's say. And the matrix tells us how to change that

1:13:29vector to create a new probability vector with the new tokens. Okay, what you do is you go row by row and you take each row, you multiply it by the vector. So the one gets multiplied by the three, the one gets multiplied by the one, two gets multiplied by the two, and you add it all up. That's your first row. Then you go to the second row, add it all up. Third row, multiply, add it all up, and you get your new vector. Okay, that's fundamentally what most of the AI computation that we all know and love. I guess that's what that is. Okay, that's what's happening. >> Keyword, I guess. >> Yeah, that's what's happening. Now,

From the episode
  1. EP 51

    The Tech Elon Has Been Waiting For

    A graphene-based memory device works at 1,300°F, opening new possibilities for extreme-environment electronics, in-memory AI, planetary exploration, and data centers in space.

    The Tech Elon Has Been Waiting For

Materials ScienceSpaceComputing