Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q but I came away with a pretty, I think like a sober view of the, the path there. Okay. So fast forward a year and you left Google DeepMind, started a company to do that amongst other things. So I guess the first question that I have for you is, What changed in the last 12 months to make you, to give you conviction that, like, now is the time?
A Great question. So, when we talked, I was doing research in the field of computational material science and machine learning. You know, specifically, we were using graph neural networks, we were using density functional theory, and we were trying to discover materials. One thing that changed since our discussion was, uh, the LLMs have Improved even further. Um, so at the time, I wasn't using LLMs much at all. Um, but I think right around when we were talking, the O-one came out, right? The reasoning models started showing up. And that was a huge update for me, because you might remember that one of my big concerns is machine learning works best on the training set distribution. But in science and technology, we almost only care about auto-domain generalization, right? So what O-one showed is if you spend test time compute, you can get better results. So that was very exciting to me because there was one way of investing resources that was beyond the training set.
AI assessment note: “One thing that changed since our discussion was, uh, the LLMs have Improved even further.”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q two. I'm curious in practice, like how you imagine that, that feedback loop working. So is it a traditional, you develop a theory, you run an experiment, you You generate data from that experiment, but in this case, you feed the experiment back into your customized LLM as an additional set of training data, and then that's the way that the loop works, or is it more complicated than that?
A Yeah, exactly. I mean, it's pretty simple, I think, as you said. So the LLM can propose, for example, synthesis recipes, or it can propose simulations to run. And because the LLMs are pretty good at tool use, it can actually do it itself. And then you get some results back. So the results from experiment could be some characterization data. Results from the simulation can be some, you know, trace or some, uh, simulation you did. And now the LLM can go through it with the context of its previous training, maybe the context of relevant papers, textbooks, but also now the results that it just got that no one else have ever seen. And then now it can Kind of tweak the experiment, tweak the simulation for the next step.
AI assessment note: “Yeah, exactly. I mean, it's pretty simple, I think, as you said.”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q being true because you, again, you just don't have the same corpus of data. You can't build a ten billion parameter model right now because the data isn't there to do it. And so instead that cost is going to go more toward the robotic lab and all that kind of stuff. Like how should I be thinking about how much compute you'll use and where that costs come from?
A Yeah. So honestly, compute is very expensive and, um, We are going to train LLMs, we are going to use GPUs to run simulations, so that does end up being a large part of the cost. Um, yeah, it's funny, like, before, you know, if you asked me this question 10 years ago, I would have thought that the biggest part of the cost must come from the labs, because, like, physical is real, you know, you're building this lab, you're buying instruments, but turns out the GPUs are so expensive, and training LLMs is so expensive, so when we were thinking about how much to raise, we kind of Laid it out in terms of the GPU cost, the lab cost, and this was kind of a minimum number we felt like was viable. Um, and yeah, we'll see the GPUs have been getting more expensive recently. I guess we'll see how the, uh, market dynamics continue.
AI assessment note: “We are going to train LLMs, we are going to use GPUs to run simulations”
Answered produced feed
D 4 · C 5 · P 4 · Cm 4 4.30
Q trying to discover novel materials, and it was thousands of data points, not tens of billions or whatever. And so that presumably hasn't changed, at least yet, but you're saying that the reasoning models have gotten good enough that they are able to sort of Get, get around that challenge via reasoning or possibly generating their own synthetic data. Like what is it that allows them to, to break that?
A So I'm not saying that they're good enough already, but that was one step in the positive direction. Um, and another thing, you know, we've seen is they've gotten really good at math. So, you know, since last time we talked, they started winning gold medals in math Olympiads. And, uh, they're doing similarly well on, um, the coding, really well on physics olympiads, and I have to say there's a difference between those things and scientific discovery. Like, you can practice for math olympiads by studying previous year's problems. You can't really practice how to discover the next big theory, but it did show you that they were getting better at Uh, reasoning on complex problems. So then what else do we need? So I think the biggest thing we need is to have our own lab, because once you have a very intelligent reasoning LLM, you still can't discover things unless you make trials, right? Like, just like humans, the LLMs will be wrong often when they try to predict things outside of their training set. But you try many things, and then at some point, You get a really cool discovery, and this is, you know, as we talked about history last time, this is quite common in solid state chemistry, solid state physics, where a lot of discoveries happen somewhat by accident, but of course with a lot of background understanding of the physical system and a lot of trial and error.
AI assessment note: “I'm not saying that they're good enough already, but that was one step”
Answered produced feed
D 4 · C 4 · P 4 · Cm 4 4.00
Q temperature superconductors is on the roadmap. So, Why, first of all, and then, and then second of all, like this question of Um, do you think you have a path to the truly breakthrough? What would the path be to a truly breakthrough discovery as opposed to finding something that is a material that is superconductive at a ever so slightly higher temperature than the best that we've got today?
A Yeah. So to answer your first question, I think it's still true that it would be difficult to just reason your way into a much better superconductor. Um, I, I actually would guess that there's a Law out there that we haven't discovered yet that says that you can't just look at your training set that's different than what you're trying to discover and just predict it. Um, you know, there's been rules that we discovered from 1800 on where like you connect energy to work. So thermodynamics is the first example. There's more recently land hours limit, which shows that you have to spend a certain amount of energy to delete information, um, which can be used to describe Maxwell's demon, um, Contradiction. I bet there's something similar for how hard it is to discover things that's outside of your training domain. Okay, so I don't think that's been fixed since last time we talked, but because we have a lab internally, We can just try things, and try them at high, large scale, and often, and hopefully, as intelligently as possible. So, even though we won't reason our way into a much better superconductor, we'll be able to push our trials in the direction that's most promising, or, you know, most promising for us, given our training set at the time. So, yeah, I think that hasn't changed, and I think there's reason to be hopeful, because, you know, in the big scheme of things, it's a pre…
AI assessment note: “because we have a lab internally, We can just try things, and try them”
Answered produced feed
D 4 · C 4 · P 4 · Cm 4 4.00
Q But how do you set the reward function for your model? What are you, what are you optimizing it for then?
A I mean, I think that's an empirical question. I think one thing I should say is it's quite nice because it's a bit, it's hard to reward hack. You know, one of these issues with RL and training LLMs is you might worry about reward hacking. And in simulations, again, reward hacking can be a problem, even in DFT. Um, but for real life experimental measurement of TC, it's much harder to reward hack, which we love. So, like, if our reward was increasing TC, like, that just seems like a, uh, Nicer, unhackable reward. But in terms of, like, what specifically will get us there, you know, we're not sure yet. I mean, it's an empirical question. We can probably try all of them. I'll list the things you propose, and we'll try all of them.
AI assessment note: “if our reward was increasing TC, like, that just seems like a, uh, Nicer, unhackable reward.”
Answered produced feed
D 4 · C 4 · P 4 · Cm 4 4.00
Q this domain? I think the, or another way to ask this question is, um, you know, if you fast forward three years, you're fully up and running and, and you're operating, how much of the, uh, of the valuable insight you will generate, do you think will come from the, the physical data coming out of your lab? Versus this synthetic data that the LLMs create on top of that?
A Yeah, that's a great question. Um, I obviously don't know the answer, but it's great to, uh, uh, brainstorm about that. Um, because on, on the one hand, um, the lab data will be kind of our additional data that other LLMs may not have. And you might think then the lab data will only be as valuable as the results in it. But on the other hand, um, what's interesting about scientific data is it's not just a few bits or numbers, right? Like, for example, there are certain experiments you can run where the result you get from it is just, say, three floating point numbers. But the implications of those could be tremendous, right? It's not just going to be, like, a few bytes. It will actually be potentially an incredible amount of understanding just from a few experiments. And this has been how it is in human history, right? Like, there are certain experiments that told us so much about how we understand about the universe. And the way to do this with synthetic data can, of course, be, you know, you run simulations that relate to that experiment. And when you get the experimental result, that actually validates or Refuse so much of the simulations you ran, and then that is a lot of information in itself. So, you know, it's a very interesting question, and I think there are some, actually, differences about how you think about synthetic data when it comes to an LLM that's good as scien…
AI assessment note: “when you get the experimental result, that actually validates or Refuse so much of the simulations”
Answered produced feed
D 3 · C 4 · P 4 · Cm 3 3.55
Q uh, developing a theory of something, and that gets fed through the model, and you get the results, and you feed it back in, you see whether it's a promising category. Like, is the germ of the original idea of what to look for coming from a human, or is it coming predominantly from the model, and then the humans have to interpret and send it off in various directions?
A Yeah, I mean, that's a great question. You know, we, we're not really prioritizing full automation anyway, so if we get better results with humans doing part of it, that's great. Um, like, this is also actually a question for our lab, right? Like, do we want to automate every single aspect of the lab? Um, at some point, you end up needing humanoids for that, and I think that's not, like, um, Liam, my co-founder and I, we are trying to be very pragmatic about it. Like, our goal is to Uh, get the best result possible on the things we care about, and, you know, how much of the automation comes from the ML, how much of it comes from more traditional tools, and how much of it gets done by humans, I think that's kind of, again, an empirical question. So, yeah, we're not, like, I think, as you said, it does seem like today there are things that ML, AI is better than humans, but one of those things is not hypothesis generation. So, um, the, I mean, there are two options that we either have to improve these lens on hypothesis generation, which is possible. Or the other option is we have humans providing some of the hypotheses and then AI doing the execution.
AI assessment note: “one of those things is not hypothesis generation”