why aren't all 31 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Brady: Prompt Engineering Won't Be a Durable Differentiating Skill
“I don't think that prompt engineering is going to be a kind of durable differential skill that people will hold. I do think that the way that you set up the ML problem to kind of ask the right questions, if you see what I mean, rather than the specific phrasin…”
Insight
Byun: Supervising step-by-step AI reasoning makes models far easier to evaluate
“The importance of supervising the process of AI systems, not just the outcomes. And so a big part of how, then, like, how Elicit is built is, We're very intentional about not just throwing a ton of data into a model and training it and then saying, cool, here'…”
Disclosure
Brady: Elicit does not run an ML-focused interview for AI engineers
“We don't have an ML-focused interview for the AI engineer role at all, actually.”
Opinion
Brady: Centralized AI gateways restrict Elicit's model and prompt experimentation
“For illicit where really the secret source of the real secret source is which models we're using, how we're using them, how we're combining them, how we're thinking about the user problem, how we're thinking about all these pieces coming together. You really n…”
Opinion
Brady: Early-twenties engineers today show strikingly higher capability and maturity
“The maturity and capabilities and just kind of general put togetherness of people at that age now is strikingly different to where I was then.”
Disclosure
Brady: Elicit uses checked exceptions in Python to handle edge cases
“We use checked exceptions inside our Python code base, which means that we can use the type system to make sure we are handling, properly handling, all of the various things that could be going wrong, all the different exceptions that could be getting raised. …”
Insight
Byun: Foundational models will not commoditize Elicit due to deep workflow specialization
“I think about this a lot in the context of moats. People are like, oh, what's your moat? What happens if GPT-V comes out? It's like, if GPT-V comes out, there's still like all of this other space that we can go into. And so I think being really obsessed with t…”
Opinion
Byun: GPT-3 was a qualitative shift, while GPT-4 was an extension
“I think GPT-III was a big change because it kind of said, oh, now is the time to build to you that we can use AI to build these tools. And then GPT-IV was maybe a little bit more of an extension of GPT-III. It felt less like a level, GPT-III over GPT-II was li…”
Assertion Not checkable as stated
Byun: LLM self-reported uncertainty is reasonably well-calibrated in production
“We found it to be pretty calibrated. There varies on the model.”
Prediction Not checkable as stated
Stuhlmüller: Generalist AI research platforms will be a winner-take-all market
“So I think there will be, at least within research, I think there will be, like, one best platform, more or less for this type of generalist research. I think there may still be, like, some particular tools, like, for genomics, like, particular types of module…”
Insight
Brady: Handling LLM Non-Determinism and Latency at Scale Is Unsolved
“There isn't some kind of industry-wide accepted way of handling that at massive scale. There are definitely patterns and anti-patterns and tools and whatnot, but it's not like this is a solved problem. So I would expect that it's not going to go down easily as…”
Insight
Brady: Effective AI engineers need SWE fundamentals, ML curiosity, and a fault-first mindset
“The three things that we say are most important for a highly effective AI engineer first of all, conventional software engineering skills, which is kind of a given, but definitely worth mentioning. The second thing is a curiosity and enthusiasm for machine lea…”
Disclosure
Brady: Elicit hires through 80,000 Hours and aligns with EA movement
“We're definitely affiliate affiliated with the safety, effective altruists kind of movement. We've gone to a few EA globals and have hired people effectively through the 80,000 hours list as well.”
Disclosure
Brady: Elicit structures interviews as real work simulations, rejecting arbitrary questions
“I really have a strong dislike and distaste for interview questions, which are arbitrary and kind of strip away all the context from what it really is to do the work. We try to make the interview process that's illicit A simulation of working together.”
Insight
Brady: ML apps require distributed systems engineering skills from day one
“The kind of person that is deep in the guts of some kind of distributed systems, really high, high scale backend kind of a problem would probably naturally have these kinds of skills, but you'll find them on, on day one if you're building a, you know, an ML po…”
Disclosure
Stuhlmüller: Elicit builds scaffolding rather than training foundation models
“The way we are building Elicit is not let's train a foundation model to do more stuff. It's like let's build a scaffolding such that we can deploy powerful models to good ends.”
Prediction Not checkable as stated
Stuhlmüller: In 10-20 years, today's scientific methods will look incredibly unsystematic
“Probably, yeah, I, I'd guess in like, 1020 years, we'll look back and it will be incredible how unsystematic science was back in the day.”
Disclosure
Stuhlmüller: Seed VCs urged Elicit to build legal AI over research
“We did encounter, I guess talking to VCs for our seed round. A lot of VCs were like, you know, researchers, they don't have any money. Why don't you build a legal assistant?”
Assertion Not checkable as stated
Byun: Anthropic's Constitutional AI slashed Elicit's query costs tenfold in days
“At the start of twenty-twenty-three, Anthropik kind of launched their constitutional AI paper and within a few days, I think four days, he had basically implemented that in production, and then we had it in-app, like, a week or so after that, and he has since …”
Insight
Byun: Early LLMs prioritized answering questions over faithfulness to source text
“At the time, the models hadn't been trained at all to be faithful to a text. So they were just generating. So then when you ask them a question, they tried too hard to ask, answer the question, and didn't try hard enough to answer the question given the text o…”
Insight
Stuhlmüller: AI orchestration resembles software engineering far more than ML research
“I think a lot of this looks more like traditional software engineering than it does look like machine learning research, and I think the people who are, like, really good at building good abstractions building applications that can kind of survive even if some…”
Insight
Stuhlmüller: Agentic search ideally balances parametric memory with retrieved context documents
“I think probably the ideal thing looks a bit more like agent control where the model can issue a query that then is intended to surface documents that substantiate its hunch. So I would, that's maybe a reasonable middle ground between model just telling you an…”
Insight
Byun: Highly structured, reproducible research workflows are uniquely amenable to automation
“Because it's so structured and designed to be reproducible, it's really amenable to automation. So that's kind of the one, the workflow that we want to automate first.”
Insight
Stuhlmüller: Separate evaluator models yield better uncertainty estimates than self-evaluation
“I think in some cases we also use the different models for the uncertainty estimates. Yes, then, for the question answering. So, one model would say, here's my chain of thought, here's my answer, and then a different type of model. Let's say the first model is…”
Disclosure
Stuhlmüller: Closed-source models consume most of Elicit's compute budget
“I'd say, like, in terms of number of careers, it's maybe similar. In terms of, like, cost and compute, I think the closed models make, make up more of the budget, since the main cases where you want to use closed models are cases where they're just smarter, wh…”
Opinion
Stuhlmüller: Claude Haiku offers an optimal balance of cost and accuracy
“Specifically, I think Cloud Haiku is like a good point on the kind of Pareto frontier, so I think it's like, it's not the, it's neither the cheapest model nor is it the most accurate, most high quality model, but it's just like a really good trade-off between …”
Assertion Not checkable as stated
Brady: Elicit routinely sees 10x p90 latency variation when prompting LLMs
“We do often normally, in fact, see a 10 X variation in P-ninety latency over the course of half an hour. When we're prompting these models, which is way higher than if you're working with a, you know, a more, more kind of conventional conventionally backed API…”
Assertion Not checkable as stated
Stuhlmüller: Elicit continues to use T5-based models
“We do also use, like, T-Five-based models, even, even now but started, yeah, started with GPT-II.”
Assertion Supported
Byun: Scientific meta-analysis typically takes five people over a year
“Lisset was very much inspired by this workflow in literature called systematic reviews or meta-analysis, which is basically the human state of the art for summarizing scientific literature. It typically involves like five people working together for over a yea…”
Disclosure
Brady: Elicit generates TypeScript types from Python OpenAPI specs instead of GraphQL
“We don't use GraphQL. So we've got the types defined in Python. That's the source of truth. And we go from the open API spec and there's a tool that you can use to generate types dynamically, like TypeScript types from those opening API definitions.”
Assertion Not checkable as stated
Stuhlmüller: Elicit co-founders wrote a 50-page mutual evaluation document before starting
“We also did a pretty lengthy mutual evaluation process where we had a Google Doc where we had all kinds of questions for each other, and I think it ended up being around 50 pages or so of, like, various, like, questions and back and forth.”