Jungwon Byun reflects on what Elicit learned when attempting to build probabilistic forecasting tools before pivoting to academic research.
Insight
Byun: Supervising step-by-step AI reasoning makes models far easier to evaluate
“The importance of supervising the process of AI systems, not just the outcomes. And so a big part of how, then, like, how Elicit is built is, We're very intentional about not just throwing a ton of data into a model and training it and then saying, cool, here'…”
Opinion
Byun: GPT-3 was a qualitative shift, while GPT-4 was an extension
“I think GPT-III was a big change because it kind of said, oh, now is the time to build to you that we can use AI to build these tools. And then GPT-IV was maybe a little bit more of an extension of GPT-III. It felt less like a level, GPT-III over GPT-II was li…”
Insight
Byun: Foundational models will not commoditize Elicit due to deep workflow specialization
“I think about this a lot in the context of moats. People are like, oh, what's your moat? What happens if GPT-V comes out? It's like, if GPT-V comes out, there's still like all of this other space that we can go into. And so I think being really obsessed with t…”
Assertion Not checkable as stated
Byun: LLM self-reported uncertainty is reasonably well-calibrated in production
“We found it to be pretty calibrated. There varies on the model.”
Insight
Byun: Highly structured, reproducible research workflows are uniquely amenable to automation
“Because it's so structured and designed to be reproducible, it's really amenable to automation. So that's kind of the one, the workflow that we want to automate first.”
Assertion Not checkable as stated
Byun: Anthropic's Constitutional AI slashed Elicit's query costs tenfold in days
“At the start of twenty-twenty-three, Anthropik kind of launched their constitutional AI paper and within a few days, I think four days, he had basically implemented that in production, and then we had it in-app, like, a week or so after that, and he has since …”