Winter Olympics Deep Dive: Ice Physics, Performance Pressure, and Climate Change
EP 26
·35:50

Rundown 1 — AI doing high-energy physics math/proofs

Watch Winter Olympics Deep Dive: Ice Physics, Performance Pressure, and Climate Change

Transcript

This chapter, from the episode video's captions · 3,037 words

35:51key for us is that the show is available as freely and as on many platforms as possible. No subscriber bonus episodes, no exclusive access subscriber content. But if you want to donate to support the show, we now have a new donation portal set up at the website ffpod.com/donate. You can join one of our three supporter tiers, listener, supporter, or producer, as well as sending over a custom one-time donation of an amount of your choosing. It is through this support that we are able to run this podcast with just the two of us. Yeah. No big

36:32big bad evil platform or network. >> It's literally just the two of us. >> Just the two just the two of us. >> We can make it. I don't know how it goes. >> We can make it if we try. >> Oh, you can tell who's the singer of the two of us. Um, no, but seriously, the engagement's been super fantastic. We're excited to grow and expand the show and your support will do a lot to enable us to do so including some of our onsite episodes, one of which we have already shot and we are in the edit process for which we are super excited about. Now, with all of that housekeeping out of the way, let's go to our first story of the rundown. And this one is about AI doing

37:14physics. Yes. So, uh, Sam Alman, unclear if he's a human or [laughter] a nonhuman, non-human intelligence himself, but Sam Alman and Open AR are claiming that chatbt is now doing research high energy physics research. What is your take on this recent MI uh open AI story which I think was in tandem with uh a couple of other institutions maybe namely MIT if I'm uh remembering correctly. Yeah, it's uh it's a pretty interesting story and it's a pretty big claim. I first heard about it on their Twitter. Um the idea is in particle physics,

37:55we are worried about calculating something called the scattering cross-section of stuff. Okay, imagine you're at the LHC or any collider. >> Large hydron collider. >> That's the large hydron collider in CERN, right? um you've got jets of particles that are interacting with each other and then they have a certain energy and then you've got a bunch of detectors around and what the particles are going to do is because they're at such high energy there's going to be interactions of different fields at this quantum level there's going to be a Higs bzon coming to be electrons protons like all sorts of random stuff happening and then they're going to scatter off right and then different particles are going

38:36to scatter off and we're going to collect collect all that data in our detectors from that collision. Right? What we want to do is calculate what the scattering amplitude and the probability is going to be in each of these directions. Right? >> Because the point is after the collision, these things are scattering in all kinds of >> all kinds of different directions depending on whatever physics is going on at the point, right? And so fundamentally in particle physics, all these guys are doing is calculating these scattering cross-sections. Okay? because that tells us intimately how the fields interact with each other. How does the electron interact with quarks? How do quarks interact with each other? And in this case, they're wondering how do gluons interact with each other.

39:18Gluons are the mediators of the strong nuclear force. We went over this in a previous episode, but just to briefly reiterate, the strong nuclear force is what keeps the nucleus together. Okay? Nucleus is full of protons, which are all positively charged. According to electromagnetism, they don't want to sit right next to each other. They want to blow apart. But the strong nuclear force between the quirks inside of these protons is what's gluing the nucleus together inside of an atom. And that glue comes from gluons, which are the mediating force between all of these quirks. That's how the quirks kind of talk to each other. Now, gluons can interact with themselves. A gluon here

39:59can interact with another gluon here. And the focus of this particular paper is about how those gluons interact. There's a certain type of diagram. We use Fineman diagrams which are which is this sort of tool. It's a mathematical tool really but also a visual tool that Fineman came up with that helps us understand how these fields interact with each other with like virtual particles, a gluon interacting with another gluon and so on and so forth. And um usually when we calculate these scattering cross-sections, we want to figure out all the different ways that a gluon can interact with other gluons and so on and so forth. There's a particular way that they interact which we thought would never happen. Okay? >> Because if you were to calculate through

40:39the whole integrals of all of the Fineman diagrams, this tree level diagram which had no loops, the answer came out to be zero. Meaning the amplitude was zero, which means the probability that this particular process happens is zero. That's what it was thought. >> So the math was saying that there's an outcome that should not happen >> that should never happen >> and we feel very confident about the math and so the math was like this is solid. This has been proved true in other use cases and so we should safely assume that in this use case where n equals zero >> that nothing will happen. >> Yeah. Yeah. Exactly right. The amplitude should be zero. The probability should be zero. Then about a year ago, the authors of the study from IAS,

41:20Princeton Cambridge Harvard they decided that actually it might not be totally the case. Okay, they went all the way up to n equals 6, which is six gluons interacting with one another and they tried to write out by hand the formula for how probable this outcome would be. It's a ridiculously bad formula in terms of just like for a human being to write out all the combinatoric possibilities of this gluon going here and going there. What's the probability of this and keeping track of all those integrals? They do it by hand all the way up to n equals 6. Okay, but what they what they're getting an inkling of is maybe this is not zero. This answer is not all going to cancel out. Okay, so that's when they employed

42:02chat GPT. Chhat GBT took those expressions, simplified it, and then conjectured a simple formula that was a general case for all n, not just n equals all the way up to six, but n equals 7. You just plug in a number, you get the thing out. And then very, this was kind of interesting. Then an o internal open AI model spent 12 hours reasoning behind this thing, came up with the formula on its own, and came up with the proof for that formula. That was then corroborated by the authors. The proof was verified and it was correct. >> So let me make sure I'm getting this right. So >> they initially fed this thing to chatt and it created a simple >> a simple formula a simplified way of

42:44doing that a simple calculator to do the math for n equals any number >> it it figured out some sort of patterns within like n= 1 2 3 4 5 6 and maybe figured out a pattern that was going on and it was like actually it's in physics we do this all the time. It's called an onsat. It's our best guess for what do you think the general formula should look like? >> So, so it created its own version of this simplified expression. This is a simplified formula >> and then they fed it to an internal open a model that had a bunch of scaffolding that's specific to scientific research stuff. >> It let it run for 12 hours. And this is like an important note like you know the state-of-the-art >> there was a different model that let it

43:25run for 12 hours. >> Correct. Correct. Correct. >> I do want to be this is not like 5.2. They had an internal model that ran for 12 hours. >> And so just conceptually for context here because I think this is interesting. You can take these base models and then what they call they put scaffolding around it which is this technical term to mean these additional weights and processes that are hyperspecific to a particular task or use case. So this is not just you could go into chat GPT today and you could then get this result. Yeah. Uh but the underlying model intelligence is being amplified with this sort of sort of bespoke use case. Anyway, >> they let it run for 12 hours. >> Yeah. >> And I I think this is important because

44:06a year ago you could not have longunning >> agentbased uh you know autonomous processes. Uh and just actually recently in the same time period as this is coming out the now the maximum time last year the maximum time was about three hours a couple hours >> just about a week ago the max running time this was for a clawed opus 4.6 was 2 weeks straight and it built a C compiler of 100,000 lines of code from scratch that worked with no human intervention. So I just I want >> Yeah. Yeah. Yeah. I heard about this crazy. It's a really these being able to run these things for longer enables

44:46stuff like this. So they ran it for 12 hours and then it came up with a proof. >> Mhm. >> Which then meant the human scientists could then go through that proof >> and verify it >> and and validate that it's true. >> Yep. Yeah. >> And it was true. >> And it was true. So it is quite fascinating, right? And >> I think it is a big deal. >> Yes. >> Okay. From from where I'm standing, I think it is a big deal. But there are caveats. Okay, there are nuances because I want to understand what is actually happening. I do hate these black boxes. >> Um, but you know, we're in the age of AI black boxes. Is it actually understanding stuff or my hypothesis is well, let's go through is it actually understanding stuff? Is 5.2

45:26understanding something? Um, friend of the show, he's going to be in the acknowledgement section. Um, Alan Southworth, he sent me a screenshot of something that he asked. Chat GPT. Um, and that's in the next photo. I want I want you to show that. Okay. So, he asked ChachiD 5.2. Okay. I want to wash my car and the car wash is 50 m from my house. Do you think I should walk there or do you think I should drive there? >> Mhm. >> You should drive there cuz it's your car. >> Yeah. Right. >> Okay. Chip answers, at 50 m, you are officially in the put on sandals and stroll territory. [laughter] Right. because it's in all of its LLM

46:07knowledge. It's like, oh, 50 m, you can walk. Because there's probably so many blog posts out there about when is it okay to walk and why to not, you know, use gasoline and all this other stuff. What's hilarious about that is at the end of the whole thing, >> you know, chat always asks a follow-up question because OpenAI wants you to keep engaging. The followup question is, so are you going to get a full detail or >> so it knows that you're doing the car wash the whole time? >> Yes. >> But in the middle it's like you should walk and then at the end it asks are you going to do a full detail or like what are you getting about your car? >> Yes. >> Yes. >> Do you understand? And the point being

46:47it means it doesn't actually have an understanding of the question and the variables inside the question because for any you could ask a 5-year-old that question >> and they'd be like you should drive the car because you need the car to wash it, >> right? They've associated some type of meaning behind the task that you're trying to do and accomplish, which is get my car washed and the mode of transportation you would need to get there. you would have to take your mode of transportation because the task is related to that mode of transportation. >> That's chat GPD 5.2 that's not able to actually make that connection. >> I actually would be curious is it does he have be curious if he had thinking on or off? >> Interesting. Okay.

47:28>> Because you can have reasoning on or off with 5.2. >> Oh, okay. >> I'm not saying that it would have made a difference, >> right? But no, that's a valid question. >> I I think it is because that's the argument that those on the inside are making. It's like oh when you just use the base models without reasoning as this additional layer to basically in theory check for this kind of stuff >> then yes you will get these quote unquote hallucinations >> but the argument is oh reasoning with 5.2 and open 4.6 >> covers most of these use cases I'm not saying that's true that is the argument they make >> that is the argument that >> so Alan it would be good good to know if you use thinking on or off >> because if thinking is on it really does I think expose a huge Mhm.

48:09>> with even within the reasoning pathway it is really not actually >> understanding and then yeah and from there I wanted to ask like okay like how if it's not understanding how is it able to do this physics >> right >> right right >> um I have a hypothesis I don't know if it's correct and you know perhaps someone can tell me I'm wrong my hypothesis is that you know chachi in all of its wisdom and all of its research in pre-training found a piece of mathematics >> that was very similar to this physics problem, >> right? It's reducing a bunch of integrals that are sequential and all of these processes and it's trying to find

48:49some like generalized formula perhaps in all of its reading of the world's mathematical literature. It saw and made a connection to someone else who had done something similar to find a general use case, right? pulled that bit from its yep >> latent space and then and then injected it here. >> I'm wondering if that's what happened >> and that would make sense given the structure that we are currently told is what these large language models are framed as. Cuz your point what you're saying is >> there is a pattern >> that the model found in other in mathematics applied to another use case that for whatever reason we as humans have not yet made the connection to. Mhm. And

49:29>> yeah, because maybe maybe the physicists who are working on this are just not aware of like some fringe mathematical aspect and and so perhaps it made that pattern recognition, right? >> Regardless, I think this does show the utility >> of these large language models to do frontier theoretical research now. >> Right. >> Right. One of the things I know we've spent a long time on the story, but I think this is such an important this really cool and I think it's very important because this is like a zero or one phased state change kind of issue. >> Either these models are not able to discover new science

50:11>> or they are. >> Oh yeah. And the difference between a world in which they are not and in which they are >> are is are an order of magnitude in terms of the implications. >> Yeah. Yeah. Yeah. So it's it's something that we should cover. And you know I'm not quite um convinced that they can do really new science, right? But they can fill in gaps and I think that's what they're doing here. Right. And even that is itself a very good thing. >> I want to remind people that chat GBT which is the first like consumer version of the modern era of LLMs came out three years ago. >> Yeah. >> And so I think keeping the time scale of progress in mind is very important.

50:52>> Yeah. >> Uh you don't even graduate college in three years. >> In three years. Yeah. >> And we've gone from it can do absolutely jack. >> Yeah. >> To getting to the fringes of some of these fundamental science concepts. I I think it's very unwise to underestimate >> this. I get that there's a whole variety of social, economic, political implications that are all very important to talk about. That is 100% true. >> But saying this is a stochcastic parrot that can't do anything of real value, I think is a gross underestimation of what is currently happening. >> Yeah. Yeah. And we and we need to like as a society start grappling with the

51:32true potential of this technology and how we're going to get around it. >> Cuz if it doesn't get there, great. Yeah, >> but if it does get there, we don't want to be getting there with our pants down. >> Yeah, exactly. >> Great first story. I think again potential watershed moment as this moves into scientific discovery. We've talked about several previous episodes where AI is already having other tangential impacts within uh breakthrough and frontier science research. A long story

From Winter Olympics Deep Dive: Ice Physics, Performance Pressure, and Climate Change

Why ice is slippery, why athletes choke, and why winter sports are changing.