Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6

Ian Fisher · The Powerful Alternative To Fine-Tuning · Y Combinator · Feb 27, 2026 · at 7:00

Ian Fischer is the co-founder of Poetiq, discussing benchmark results on Humanity's Last Exam (HLE) compared to Anthropic's frontier model.

0:00 / 0:19exact quote · 19.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ian Fisher

Assertion Not checkable as stated
Poetiq achieves faster, cheaper recursive self-improvement than existing methods
“The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this.”
Ian Fisher Feb 27, 2026 ▶ 1:18 The Powerful Alternative To Fine-Tuning · Y Combinator
Assertion Not checkable as stated
Poetiq's agentic harness outperforms new base models without code changes
“With poetic what we end up giving you is a you know, people are calling these things harnesses now, but you know, or agentic system or whatever you want to call it, that sits on top of one or more language models, and it just performs better than them. And whe…”
Ian Fisher Feb 27, 2026 ▶ 4:18 The Powerful Alternative To Fine-Tuning · Y Combinator
Disclosure
Poetiq's meta-system generates reasoning systems for problems GPT-5 cannot reliably solve
“And so the core technology that we've developed at Poetic is recursive self-improvement. So we have a recursively self-improving system, which we call the Poetic meta system. The output of that system is systems that solve hard problems where a hard problem is…”
Ian Fisher Feb 27, 2026 ▶ 9:08 The Powerful Alternative To Fine-Tuning · Y Combinator
Insight
AI is replacing human engineers for dataset understanding and failure-mode detection
“Historically in machine learning, you always, you know, it's like the rule was you have to know your data set really well. But now we're kind of outsourcing that to the AI itself, where the AI is the, it's the AI's job to understand the dataset and figure out …”
Ian Fisher Feb 27, 2026 ▶ 12:52 The Powerful Alternative To Fine-Tuning · Y Combinator
Assertion Supported
Poetiq outperformed Gemini 3 Deep Think on ARC-AGI-2 at half the cost
“Yeah, so the interesting thing is that we were half the cost of Gemini Three Deep Think because we were building on top of Gemini Three Pro, which is a much cheaper model. But we still got in the end, a nine percentage point improvement on the official verific…”
Ian Fisher Feb 27, 2026 ▶ 6:14 The Powerful Alternative To Fine-Tuning · Y Combinator
Assertion Not checkable as stated
Poetiq's Humanity's Last Exam optimization run cost less than $100,000
“We didn't publish any cost for this, but I can say that the optimization costs us less than a hundred K, yeah.”
Ian Fisher Feb 27, 2026 ▶ 7:31 The Powerful Alternative To Fine-Tuning · Y Combinator
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.