What Claude Actually Did to the Riemann Hypothesis
EP 53
·1:07:20

Claude creates hostile referees

Watch What Claude Actually Did to the Riemann Hypothesis

To stress-test a promising proof attempt that reached 50% of the way toward a bound on the Riemann hypothesis, Claude spun up three sub-agents acting as hostile peer referees. Referee A caught a real technical flaw involving ill-conditioning of a matrix, which the proof then corrected. A second referee independently confirmed the argument contained no circular logic, meaning the proof was not secretly assuming the Riemann hypothesis in order to prove it.

  • The 50% figure matters because the hosts note it would shatter an existing record if confirmed.
  • The referee sub-agents were labeled A, B, C rather than the traditional Referee 1, Referee 2, prompting a brief aside about AI having no sense of academic ranking.
  • Krishna compares the AI referees to real peer review, where a constructive referee flags a potential confound and suggests how to test for it rather than simply rejecting the work.

Transcript

This chapter, from the episode video's captions · 359 words

1:07:21>> You just got here. >> Yeah. Yeah. So, um, it notes that like 50% would shatter the record. And it goes back to the human to Jared. And it's like, u, my prior is that this is wrong. >> Okay. So, here's what I'm going to do. I'm going to rig up hostile referees, three hostile referees, more sub agents that are going to critique this proof. >> Um, those referees, go ahead and critique it. >> Referee A, usually it's referee 2. In this case, I guess the AI just [laughter] doesn't have a doesn't have a a ranking, right? So, here, referee A discovers the technical flaw. And um I didn't I didn't look in they didn't they didn't actually show what the referee was saying, but I wonder if it was like

1:08:01as scathing as the stuff that we get in the emails. Like, you should just quit the job [laughter] and go work in a McDonald's, right? Like that's the stuff that we get in the emails. But I I wonder what referee A said to this sub agent E2. But it discovered a technical flaw regarding some in ill conditioning of the matrix and it corrected it. >> Mhm. [clears throat] >> And then it ran up these inequalities and now that thing was fixed. So the referee was doing what normal referees do, which is like, hey, this might be wrong. You know, a nice referee at least would like be like, this might be wrong. This is how you could change your experiment to test for this um >> confound and things like that. another referee independently verified that the calculations don't secretly import the

1:08:43Remon hypothesis. Like, are you just trying to prove it by like assuming the thing that you're trying to prove? That doesn't make any sense. So, there's no circular circular logic. Okay, so now we're up to 50%. >> Mhm. >> It goes back to the human prompter to Jared >> and it's like, we got we got to 50%. [snorts] >> Jared then asks, "What would you do next?" And Claude responds, "Push it to

From What Claude Actually Did to the Riemann Hypothesis

Claude takes a real run at the Riemann Hypothesis, forcing us to ask what agentic AI can now do in mathematics, before we open the summer transfer window for America’s scientists.