Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And, uh, to that exact point, like how, us contrast and compare, so RLVR, uh, versus, uh, which is reinforcement running with verifiable rewards against what it is that you do. Is that, are those just completely different approaches? Because they both seem at the same thing, which is to basically get to, uh, perfection.
A Yeah, I think we, um, we, the world is realizing that we need verification. Verification means very different things in math. I think in early 25 or late 24, Ah, it means the numerical answer associated with each problem. Now, the thing is, like, one is, you know, reward hacking. Like, you know, we have seen from, say, Frontier Math and other benchmark, which only compels a numerical answer that it doesn't actually necessarily reflect the model's capability in the logical reasoning. So it's able to get to the answer without reasoning through it, which is quite fascinating. I mean, there's always these, I did Math Olympia before, and there's always this, like, classmate who's really good at guessing the answer. I don't know, like, AIME, which is this exam, um, that all the answers are between zero, zero, zero to nine, nine, nine. Like, I remember there's one year where my friend told me that, like, you know, he just basically guessed three questions correctly. Whereas the rest of us need to, like, read something through, and like, Jesus Christ, he just put, like, a zero in there, and then somehow that answer is indeed zero. It sounds very unfair, and you know, in a way, in high school, the teacher will, like, ask you to show your work. So, for a while, I think, like, verifiable reward means that final numerical, like, output. I think that like people are now realizing it doesn't…
AI assessment note: “Verification means very different things... verifiable reward means that final numerical, like, output”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Great. Uh, and what kind of math are we talking about? Is that high school math? Is that competitive math? Or is that deep research math?
A Yeah, I think people generally start with competitive math because it's kind of, you know that you have like a non-solution, and then you, um, kind of start hill climbing the infinite, sort of like infinitely high mountain of math. Um, there are actually two axis of difficulty. One is how creative the solution is, and The other one, roughly speaking, is how abstract the mathematical object is. So say a qualifying exam can be incredibly abstract, but the sort of creativity required to solve each problem might not be that high. It might be very standard. On the other hand, IMO problem, while it's very sort of easy to understand, even by high school students, not very abstract, but it's incredibly creative.
AI assessment note: “I think people generally start with competitive math because it's kind of”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Fascinating. And is the system solving problems in a predictable way? Where I'm going with this is the whole move 37 discussion where you find AI, um, solving problems in a, in a almost alien kind of ways. Is that part of what you're doing? Is that what you're seeing? Or is that something that's coming up in the future?
A It's interesting because there's, I think, like a hindsight problem. So it's like we Didn't know how to do it. I think, like, I definitely have no hope in solving those conjectures, and our founding mathematician professor Ohno didn't know how to do it either. Okay, now we saw the Lean code, right, like thousands of lines of Lean code, and sort of read through it, understand it, and, uh, maybe we see some of the techniques as sort of standard, but I think the application of them and the combination of them is also Not entirely. I think it's somewhat at the level of a junior math professor, say, like, you know, a postdoc or a junior researcher. Obviously, a lot of junior researchers do amazing work, and they have their move-thirty-seven moments in some very long-standing open questions, but there are a lot of sort of day-to-day, um, of research that it feels like it's at that level. There's also this question of, like, because it's solving the problem in Lean, which is a Kind of machine language, not a human natural language. Uh, the proofs actually look quite different. So we actually analyzed all 12 problem solutions of the Putnam exam, and we found that a lot of the solutions actually differ from the human solution. So because it is a, you know, lean based system, it is really good at sort of routine bookkeeping, and it will actually choose a lot of the more mechanistic, you …
AI assessment note: “we found that a lot of the solutions actually differ from the human solution”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q And do you need a Lean equivalent for each one of those domains as you expand?
A It's a very interesting question. I think so, uh, as you can see even within math, right, sometimes creation of domain-specific language, like the vector language for Euclidean geometry, have its gains. Um, there could be the case where in other domains, um, something that is not exactly the abstraction of Lean are the right sort of medium. But in that case, you know, you can sort of do code translation, and you can kind of build out You know, the sort of, uh, stack that's required to use your, like, Lean-based Serum Proving Engine. I think the sort of gap between, say, for example, Lean and another, like, strongly typed language like Rust is a lot closer than the gap between Lean and English. And I think that's a lot of the commercial value. I mean, if you can sort of reason in between informal and formal space, That, I think, is going to unlock a lot of the things beyond just the power of a formal theorem prover.
AI assessment note: “I think so, uh, as you can see even within math, right, sometimes creation”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q How generalizable do you think the approach that you guys are using is?
A I think the, the one thing, if you talk about perfect AI, I think people's first reaction is, wow, that's really valuable. I think a lot of the, you know, different labs are trying to, uh, reduce a hallucination, uh, or increase sort of the accuracy, um, so many, many different ways. If you have a lot of the sort of industries where mistakes are extremely costly, um, that's a block to AI deployment if you don't have that sort of provable guarantee. And now that's the value of like, say, catching the edge case. And there's this additional value of trusting that your edge case can be covered. So two, two additional layers of value, um, to sort of, you know, reliable, consistently correct AI. In terms of, like, how general this is, I think we start with math. Our worldview is math reasoning is a true reasoning layer of AGI. And I think a lot of the labs share that view, um, labs across US, China, Europe. And, you know, from math, you kind of get to code. Uh, math gives you Proof of property, and code gives you output. Output and property affiliated to it are two quite important parts of the digital world, and so from math you go to code, and from code you can run a lot of real-world experiments in the software stack. Then you can have a lot of other things. We don't claim to be doing things that are in the physical world at all. We are obviously not doing things that are non-verif…
AI assessment note: “In terms of, like, how general this is, I think we start with math.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Amazing. And then, um, on, um, some other incredible things that you've done. So you did, um, MIT in three years, I believe. Then you went to the UK on a Rhodes scholarship to study neuroscience. Why neuroscience? Was it all part of a grand plan towards AI, or was it just your interest naturally carried you?
A Yeah. I think the Rhodes program, like, did a really good job and Probably too good a job to encourage us to just like shift direction. Like it has this sort of, I mean, broad sort of belief that you need a lot of disciplines and studies to help you become a global leader. Um, that's what the Rhodes Scholarship is trying to sort of nurture. People who have a background in STEM, they will encourage you to go into liberal arts. I wasn't fully encouraged to go into liberal arts, so I picked something that's kind of in the middle, like neuroscience. Obviously, I think at Oxford, I have a lot of math friends, and so, kind of, still math was part of the equation. Um, I was also trying to apply math in my neuroscience study, specifically, maybe because I just am afraid of animal experiments. Like, I'm not gonna, I, I'm, I'm probably just gonna stay in, you know, data analysis, computational neuroscience. I was quite interested in topological data analysis, and persistent homology, and Uh, but later I realized, I think it was like two, three months after the school year started, I realized there was something called UCL Gatsby. UCL Gatsby is this, like, premium, like, you know, AI hub in London, and, um, Oxford to London is a short train, and there's so many world-class faculties there, like, um, doing really cool research in, say, like, the radical, uh, machine learning, um, like, you…
AI assessment note: “I picked something that's kind of in the middle, like neuroscience.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q And then, uh, still bearing in mind that you're a very young company, what is the current state of the product? Are you mostly focused on MVP kind of Product that can solve this amazing problem, but that's not industrialized yet. What part is research versus what part is engineering and product so far?
A We focus a lot more on, I say, so there's this sort of, like, you know, team of really strong machine learning researchers and engineers, and they're all, like, both researcher and engineer in one. They're, like, really amazing. Um, we have a lot of really good people from Meta, from Google Brain, from Anthropic, et cetera, and we just, um, we keep hiring, you know, more and more, uh, sort of Frontier Lab researchers. Um, this part, I think, is like, you know, focused on developing the core capability of the system. So we want to basically push the, the goalposts, like, forward, right? So from Putnam Perfect Score, that was four months in, then two months later was the four research conjectures, and then, you know, during this middle, we also have tested something that is transfer learning from math to code verification, so another evaluation on a community-recognized benchmark. We want to kind of pry Where we can get, because we have really great mathematicians telling us how we should think about certain research problem targets, and we currently have really hard research math problems in-house that we are tackling. It's showing some problems. It's also obviously getting stuck. Um, so there's this part, and this part is like, you know, the current focus of the company. Once we know where the frontier is, then we can try to say, okay, let's make it robust. Let's make it sort o…
AI assessment note: “Once we know where the frontier is, then we can try to say, okay, let's make it robust.”
Answered raw tape
D 5 · C 3 · P 3 · Cm 3 3.60
Q Fascinating. Thank you for that. So let's, um, actually go to the, into the product now. So we alluded to some of this. Let's unpack how it actually works. So you mentioned there were three components. What is the architecture? What do those components do?
A So I think like our kind of very broad vision is that we are going to have a conjecture. We're going to have a prover. And then there is knowledge base. In a way, your knowledge base is like, let me take a metaphor. I think for the, cause this is like quite, I think like in the niche area of like very like subfield of, of AI, but, uh, suppose you are like selling on an ocean, right? And like, where do you know where to go? And sort of your ship, um, that basically decides where to navigate. That's your conjecture. And then you like, you know, sail and tour one direction. Um, and then you land at this island. Okay, well, like, do you know if you have been on this island before? You don't necessarily know. Basically, you need to look up your knowledge base. You want to make sure that this is indeed uncharted territory. And then once you realize that it is uncharted territory, like, how do you know if it's like, say, India or West Indy, right? Like, so is it, I don't know, um, is it gonna have some rare metal? Um, that's kind of where your prover starts coming in to basically prove this New conjecture that is not in the knowledge base that is mathematically correct and has merit. And so that's kind of the, and then there's all the formalization, which is the ability to, um, recent across informal and formal space, kind of weaving all these.
AI assessment note: “we are going to have a conjecture. We're going to have a prover. And then there is knowledge base.”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q PhD in math and JD in law at Stanford. All of this been fascinating, not just in terms of, um, achievement, but in terms of like range. And, um, I'm, I'm just, uh, curious, like how you were able to do all of this and whether there's any lesson for Anybody else? I mean, clearly there's an element of, like, of the chart IQ, but there must be something else.
A I think there's a lot of things, actually. My junior year, I mean, my last year at MIT, I'm like, I kind of grew up with a very sole focus and goal to do math Olympiad and then do math research, and the very next step is to probably go to grad school directly and probably not even do the Rhodes Scholarship, I'm not sure, because a lot of the Roe scholars are like, you know, politicians or aspiring lawyers, judges. I'm like, I just want to have some fun intellectually. At the end of the neuroscience, I was like, okay, I want to go to Stanford to start my math PhD, but also there's this incredible opportunity of Stanford Law School, one of the two, I think, first ranking law schools in the country, and they have really good, um, IP professors who marry, like, AI and copyright law. They have really good Professor, uh, Mitchell Polinsky in Law and Economics, where you basically are doing differential equations, but you are, like, analyzing, like, deference and, like, retribution, uh, the ratio of each sort of criminal law measure, and there are a lot of, like, other things, cool, um, methods to apply textualism to constitutional law that's very similar to looking up definitions in mass textbooks, and so I was like, okay, that's very cool, so I did my, I mean, the JD-PhD is, like, you have to spend one full resident year in law school, so that was my first year. I spent, like, One y…
AI assessment note: “I'm like, I just want to have some fun intellectually.”
Redirected raw tape
D 2 · C 4 · P 4 · Cm 4 3.40
Q Okay, so there is an LLM. So how does that work then? What creates the conjecture? Is that front-based, um, like, what goes into it?
A Yeah, I think I would say here, like, the conjecturing part is still the underdevelopment part. Like, we have been in the last, I think, seven months very focused on Prover and also made a lot of progress on the knowledge base. So, so in the PNM exam, right, you don't need to conjecture. You have 12 problems. They're incredibly hard and they are like, you know, basically test for your prover. So action prover, prover, uh, tried on PNM exam, got perfect score. The, the kind of, you know, underlying system is a system of ensemble of models. And there's also a set of deterministic tools. And also there's like a proprietary data set that's very large. So, um, kind of a combination of these three things that led to that success specifically. Um, I think that for the deterministic tooling, it's quite interesting because these are actually written for Lean in the language of Lean. So, um, a bit like metaprogramming, uh, and that's, that's very interesting. We are actually gonna release them, uh, on a public, like, API on these, all these dozen of pools. Uh, very, very soon. Beginning of March. Great. Um, so.
AI assessment note: “the conjecturing part is still the underdevelopment part. Like, we have been... focused on Prover”
Answered raw tape
D 4 · C 3 · P 3 · Cm 3 3.30
Q Yeah. Do you have a sense for where that threshold is? So if you have math on one side and English on the other side, effectively with your approach, you're going to be able to cover kind of like all of science, and the second you start getting into non-scientific fields, then the approach doesn't work anymore, or you don't know yet and you're about to explore? Yeah.
A I think we want to try to figure out, like, what are things that can be done in software. One is, I think, math and code, they really complement each other very well. I mean, there are a lot of great code generation companies. We can provide approval guarantee and code verification. There are a lot of other, like, domains where just that sort of verified generation capability is incredibly valuable, like hardware. And then, I think, if you can have a lot of theory, and you can have partners who are really good at reward testing, Then that is, you know, AI for science, and I think that's also incredibly promising. I feel like this is a generational effort. It's gonna take, like, a long time. Um, we're gonna see the DNA of the company remains math, and we're gonna see best first market, maybe verification. Best second market, I don't know what that is. Could be optimization. A lot of the things I think are waiting to be explored, um, but just the generation verification loop I think itself is going to have a large time. And then I think, you know, there are a lot of things that we're also learning together with the potential customers.
AI assessment note: “A lot of the things I think are waiting to be explored”
Not addressed raw tape
D 2 · C 2 · P 3 · Cm 2 2.25
Q How far do you think we are from that world where we have armies of it?
A Like, we have to move extremely fast. Like, there's so much to do. I think Axiom is a very, very young company, and we are at, like, you know, the very, very beginning tip, and, and we are already, I personally feel some sort of shock and emotional response, and I know some of my mathematician friends, Scott Commoners, who's a Harvard, uh, microeconomics professor, also a Morgan Price winner, we're good friends, and We all have this sort of emotional response when Axiom Prover proved false conjecture, proved that almost all primes are partially regular, partial Vandiver conjecture, which is one part of the original Vandiver conjecture that had been open for 90 years, the parity of differentials, um, for surfaces of genius zero and one by algebraic geometry paper. We are really just leaping across a point. I mean, Pundit marked, I think, The end of AI trying on Maths Olympiae. We are very glad that we got a perfect score. It's a really good period point. Pundum 20 25 is by a lot of sort of experts grading harder than IMO 20 25, so it's the hardest reward Maths Olympiae test, and now we are leaping. We are leaping to research, and I think I'm going to have another similar emotional response if it really does solve one of those Breakthrough mathematics problems. Interestingly, I think there are a lot of experts in domains that are currently overlooked by AI development. So, so if …
AI assessment note: “we are at, like, you know, the very, very beginning tip”
Not addressed raw tape
D 1 · C 3 · P 3 · Cm 2 2.25
Q Great. People may have heard, uh, of both OpenAI and Google DeepMind sort of winning IMO, the International Math Olympiad, and other very hard to crack. Kind of like math problems. How does their approach differ from what it is that you guys are doing?
A Yeah, I think the sort of concept formal theorem proving actually, uh, existed, like automated theorem proving as a field existed before, before deep learning. So I think there are a lot of researchers, a lot of them in Europe in 20, say, 20 18 and even, even before then were doing sort of like automated, uh, reasoning without, without the LM component. And we actually have some of these people on our team, like, they were the authors of ATP Boost, um, and, uh, it's a very interesting time. In 2019, um, François Charton and Guilhem Lampo, uh, the co-founder of Mistral, they have a paper which tries to put transformer on sort of symbolic integration and realize that it can beat, like, computer algebra systems such as MATLAB or Mathematica. Uh, I think Ilya was actually the reviewer of the paper, and he actually tweeted about it. There is this Other fields medalist, um, Tim Gower said it was, uh, either amusing or game changing. Cause it was, it was an open review. People don't know if it's correct or not. And that was not amusing. That was the beginning of AI for math. And Francois is now also at Axiom. Uh, there's like a long history of what people are trying to do with it. And I think, uh, Google, you know, started the alpha geometry effort in, uh, and that was very exciting effort. They realized that if you convert The figures and lines, triangles, circles, intersection point…
AI assessment note: “I think the sort of concept formal theorem proving actually, uh, existed”
Redirected raw tape
D 2 · C 2 · P 2 · Cm 2 2.00
Q How do you select people? How do you evaluate someone's taste? Is that their prior work?
A Right. So, so one is, basically that means I have to learn every day because I need to have a certain ground, you know, a basic amount of taste, and also the other people who are senior at the, the company, uh, need to have, research scientists have need to have that sort of amount of taste, and what makes us excited, and When we are excited, we just go after that person. Um, we have had a lot of, uh, I think, extraordinary hires from amazing and interesting backgrounds, and we, we, we like, we really like this team, and I think that we raise the bar for hiring, actually, um, continuously. So, uh, we, we recently have a uptick in the number of people who would like to join us and our candidates that we excited about. We, we take recruiting very, very seriously.
AI assessment note: “When we are excited, we just go after that person.”