Feb 27, 2026 · 19m · y-combinator

The Powerful Alternative To Fine-Tuning · Y Combinator

Ian Fisher · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Y Combinator's The Light Cone, Poetiq Co-founder Ian Fischer discusses how recursively self-improving reasoning harnesses provide a cost-effective alternative to model fine-tuning while sharing practical insights for developers.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 3.7 Guest teaching 4.3 Guest disagreement 0.4 The partners pushing back 0.6
05100:0010:001:01–3:06 · The partners as informed peer 4/10 Defining Poetiq and Recursive Self-Improvement The host demonstrates strong familiarity with the startup ecosystem's dilemma regarding fine-tuning vs. frontier model upgrades. Ian explains recursive self-improvement without requiring costly model retraining from scratch in an agreeable, collaborative dialogue.3:06–6:37 · The partners as informed peer 4/10 The Poetiq Harness vs. Traditional Fine-Tuning The host invokes the Bitter Lesson and benchmark tracking on Arc AGI V-II. Ian breaks down the exact benchmark score and cost savings achieved by Poetiq over Gemini 3 Deep Think.6:37–8:39 · The partners as informed peer 4/10 Record-Breaking Performance on Humanity's Last Exam The co-host highlights model routing behaviors among founders and notes the efficiency of small teams. Ian explains how a seven-person team achieved SOTA on Humanity's Last Exam for under six figures.8:39–11:31 · The partners as informed peer 4/10 Automated Self-Improvement and Custom Agent Optimization The co-host connects Poetiq's approach to RNN paradigms versus RL S-curves. Ian reframes the concept by explaining that Poetiq's meta-system and underlying models form compounding S-curves.11:31–14:50 · The partners as informed peer 5/10 Outsourcing Context and Prompt Engineering to AI The host probes the exact mechanics of whether gains come from prompt optimization or harness architecture. Ian educates the hosts with concrete DeepMind experimental data showing prompts only yielded 5% while reasoning strategies in code drove 95%.14:50–18:16 · The partners as informed peer 3/10 Partnering with Poetiq and Early Access The interview transitions into call-to-action details for early access and Ian's biographical journey from Apportable to Google Robotics and DeepMind research.18:16–19:19 · The partners as informed peer 2/10 Actionable Advice for AI Builders and Engineers A brief concluding segment where the guest shares practical advice encouraging engineers to build daily with AI tools.1:01–3:06 · Guest teaching 4/10 Defining Poetiq and Recursive Self-Improvement The host demonstrates strong familiarity with the startup ecosystem's dilemma regarding fine-tuning vs. frontier model upgrades. Ian explains recursive self-improvement without requiring costly model retraining from scratch in an agreeable, collaborative dialogue.3:06–6:37 · Guest teaching 5/10 The Poetiq Harness vs. Traditional Fine-Tuning The host invokes the Bitter Lesson and benchmark tracking on Arc AGI V-II. Ian breaks down the exact benchmark score and cost savings achieved by Poetiq over Gemini 3 Deep Think.6:37–8:39 · Guest teaching 5/10 Record-Breaking Performance on Humanity's Last Exam The co-host highlights model routing behaviors among founders and notes the efficiency of small teams. Ian explains how a seven-person team achieved SOTA on Humanity's Last Exam for under six figures.8:39–11:31 · Guest teaching 5/10 Automated Self-Improvement and Custom Agent Optimization The co-host connects Poetiq's approach to RNN paradigms versus RL S-curves. Ian reframes the concept by explaining that Poetiq's meta-system and underlying models form compounding S-curves.11:31–14:50 · Guest teaching 6/10 Outsourcing Context and Prompt Engineering to AI The host probes the exact mechanics of whether gains come from prompt optimization or harness architecture. Ian educates the hosts with concrete DeepMind experimental data showing prompts only yielded 5% while reasoning strategies in code drove 95%.14:50–18:16 · Guest teaching 3/10 Partnering with Poetiq and Early Access The interview transitions into call-to-action details for early access and Ian's biographical journey from Apportable to Google Robotics and DeepMind research.18:16–19:19 · Guest teaching 2/10 Actionable Advice for AI Builders and Engineers A brief concluding segment where the guest shares practical advice encouraging engineers to build daily with AI tools.1:01–3:06 · Guest disagreement 1/10 Defining Poetiq and Recursive Self-Improvement The host demonstrates strong familiarity with the startup ecosystem's dilemma regarding fine-tuning vs. frontier model upgrades. Ian explains recursive self-improvement without requiring costly model retraining from scratch in an agreeable, collaborative dialogue.3:06–6:37 · Guest disagreement 0/10 The Poetiq Harness vs. Traditional Fine-Tuning The host invokes the Bitter Lesson and benchmark tracking on Arc AGI V-II. Ian breaks down the exact benchmark score and cost savings achieved by Poetiq over Gemini 3 Deep Think.6:37–8:39 · Guest disagreement 0/10 Record-Breaking Performance on Humanity's Last Exam The co-host highlights model routing behaviors among founders and notes the efficiency of small teams. Ian explains how a seven-person team achieved SOTA on Humanity's Last Exam for under six figures.8:39–11:31 · Guest disagreement 1/10 Automated Self-Improvement and Custom Agent Optimization The co-host connects Poetiq's approach to RNN paradigms versus RL S-curves. Ian reframes the concept by explaining that Poetiq's meta-system and underlying models form compounding S-curves.11:31–14:50 · Guest disagreement 1/10 Outsourcing Context and Prompt Engineering to AI The host probes the exact mechanics of whether gains come from prompt optimization or harness architecture. Ian educates the hosts with concrete DeepMind experimental data showing prompts only yielded 5% while reasoning strategies in code drove 95%.14:50–18:16 · Guest disagreement 0/10 Partnering with Poetiq and Early Access The interview transitions into call-to-action details for early access and Ian's biographical journey from Apportable to Google Robotics and DeepMind research.18:16–19:19 · Guest disagreement 0/10 Actionable Advice for AI Builders and Engineers A brief concluding segment where the guest shares practical advice encouraging engineers to build daily with AI tools.1:01–3:06 · The partners pushing back 1/10 Defining Poetiq and Recursive Self-Improvement The host demonstrates strong familiarity with the startup ecosystem's dilemma regarding fine-tuning vs. frontier model upgrades. Ian explains recursive self-improvement without requiring costly model retraining from scratch in an agreeable, collaborative dialogue.3:06–6:37 · The partners pushing back 0/10 The Poetiq Harness vs. Traditional Fine-Tuning The host invokes the Bitter Lesson and benchmark tracking on Arc AGI V-II. Ian breaks down the exact benchmark score and cost savings achieved by Poetiq over Gemini 3 Deep Think.6:37–8:39 · The partners pushing back 0/10 Record-Breaking Performance on Humanity's Last Exam The co-host highlights model routing behaviors among founders and notes the efficiency of small teams. Ian explains how a seven-person team achieved SOTA on Humanity's Last Exam for under six figures.8:39–11:31 · The partners pushing back 1/10 Automated Self-Improvement and Custom Agent Optimization The co-host connects Poetiq's approach to RNN paradigms versus RL S-curves. Ian reframes the concept by explaining that Poetiq's meta-system and underlying models form compounding S-curves.11:31–14:50 · The partners pushing back 2/10 Outsourcing Context and Prompt Engineering to AI The host probes the exact mechanics of whether gains come from prompt optimization or harness architecture. Ian educates the hosts with concrete DeepMind experimental data showing prompts only yielded 5% while reasoning strategies in code drove 95%.14:50–18:16 · The partners pushing back 0/10 Partnering with Poetiq and Early Access The interview transitions into call-to-action details for early access and Ian's biographical journey from Apportable to Google Robotics and DeepMind research.18:16–19:19 · The partners pushing back 0/10 Actionable Advice for AI Builders and Engineers A brief concluding segment where the guest shares practical advice encouraging engineers to build daily with AI tools.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 10:47 Reframing the RNN and RL comparison

Ian politely pushes past the co-host's comparison to RNNs, clarifying how the meta-system creates its own distinct, shifting S-curve on top of underlying LLM models.

Hardest push from the partners ▶ 13:19 Drilling into prompt vs harness mechanics

The host presses Ian on whether Poetiq's magic is just superior prompt crafting or structural harness logic such as summarizing and re-ranking.

Biggest teaching moment ▶ 13:40 Prompt optimization limits vs code reasoning

Ian illustrates the limitation of popular prompt tuning methods like JEPA, citing empirical DeepMind data where prompts yielded 5% versus 95% achieved via coded reasoning strategies.

The partners hold their own ▶ 2:07 Synthesizing the startup fine-tuning trap

Garry Tan demonstrates keen industry domain insight by articulating why spending millions fine-tuning models is rendered obsolete by next-generation frontier releases.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Defining Poetiq and Recursive Self-Improvement 4411 The host demonstrates strong familiarity with the startup ecosystem's dilemma regarding fine-tuning vs. frontier model upgrades. Ian explains recursive self-improvement without requiring costly model retraining from scratch in an agreeable, collaborative dialogue.
The Poetiq Harness vs. Traditional Fine-Tuning 4500 The host invokes the Bitter Lesson and benchmark tracking on Arc AGI V-II. Ian breaks down the exact benchmark score and cost savings achieved by Poetiq over Gemini 3 Deep Think.
Record-Breaking Performance on Humanity's Last Exam 4500 The co-host highlights model routing behaviors among founders and notes the efficiency of small teams. Ian explains how a seven-person team achieved SOTA on Humanity's Last Exam for under six figures.
Automated Self-Improvement and Custom Agent Optimization 4511 The co-host connects Poetiq's approach to RNN paradigms versus RL S-curves. Ian reframes the concept by explaining that Poetiq's meta-system and underlying models form compounding S-curves.
Outsourcing Context and Prompt Engineering to AI 5612 The host probes the exact mechanics of whether gains come from prompt optimization or harness architecture. Ian educates the hosts with concrete DeepMind experimental data showing prompts only yielded 5% while reasoning strategies in code drove 95%.
Partnering with Poetiq and Early Access 3300 The interview transitions into call-to-action details for early access and Ian's biographical journey from Apportable to Google Robotics and DeepMind research.
Actionable Advice for AI Builders and Engineers 2200 A brief concluding segment where the guest shares practical advice encouraging engineers to build daily with AI tools.

Statements from this episode (12)

Assertion Not checkable as stated
Poetiq achieves faster, cheaper recursive self-improvement than existing methods
“The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this.”
Ian Fisher Feb 27, 2026 ▶ 1:18
Assertion Not checkable as stated
Frontier labs achieve recursive self-improvement by retraining models at each step
“And, you know, of course, Anthropic and OpenAI and Google, they're exploring recursive self-improvement, but Typically at that level of having the, you know, having to train a new model for every step of self-improvement that they do.”
Ian Fisher Feb 27, 2026 ▶ 1:55
Assertion Not checkable as stated
Poetiq's agentic harness outperforms new base models without code changes
“With poetic what we end up giving you is a you know, people are calling these things harnesses now, but you know, or agentic system or whatever you want to call it, that sits on top of one or more language models, and it just performs better than them. And whe…”
Ian Fisher Feb 27, 2026 ▶ 4:18
Assertion Supported
Poetiq outperformed Gemini 3 Deep Think on ARC-AGI-2 at half the cost
“Yeah, so the interesting thing is that we were half the cost of Gemini Three Deep Think because we were building on top of Gemini Three Pro, which is a much cheaper model. But we still got in the end, a nine percentage point improvement on the official verific…”
Ian Fisher Feb 27, 2026 ▶ 6:14
Assertion Supported
Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”
Ian Fisher Feb 27, 2026 ▶ 7:00
Assertion Not checkable as stated
Poetiq's Humanity's Last Exam optimization run cost less than $100,000
“We didn't publish any cost for this, but I can say that the optimization costs us less than a hundred K, yeah.”
Ian Fisher Feb 27, 2026 ▶ 7:31
Disclosure
Poetiq's meta-system generates reasoning systems for problems GPT-5 cannot reliably solve
“And so the core technology that we've developed at Poetic is recursive self-improvement. So we have a recursively self-improving system, which we call the Poetic meta system. The output of that system is systems that solve hard problems where a hard problem is…”
Ian Fisher Feb 27, 2026 ▶ 9:08
Insight
Fischer predicts Poetiq and base model S-curves will compound toward AGI
“Effectively you could say, like, each model or each set of models that we're working with will have their own S-curve. The poetic system, the poetic meta system itself, is also going to have its own S-curve. And so as the poetic meta system gets better, and as…”
Ian Fisher Feb 27, 2026 ▶ 10:56
Assertion Not checkable as stated
Poetiq's autonomous prompt generation system produced unexpected, non-human prompt structures
“It was pretty interesting to look at the prompt outputs in particular, I'd say, for ArcGi in that you know, I think you can read those and say, well, that's not what a human would have written. Pretty clearly. And it's, you know, there's some unexpected stuff …”
Ian Fisher Feb 27, 2026 ▶ 12:20
Insight
AI is replacing human engineers for dataset understanding and failure-mode detection
“Historically in machine learning, you always, you know, it's like the rule was you have to know your data set really well. But now we're kind of outsourcing that to the AI itself, where the AI is the, it's the AI's job to understand the dataset and figure out …”
Ian Fisher Feb 27, 2026 ▶ 12:52
Assertion Not publicly verifiable
Code-based reasoning strategies boosted Gemini 1.5 Flash performance from 5% to 95%
“In this particular case, you know, the hardest task we were working on, we got like to five percent performance with Gemini, 1.5 flash. This was a while ago. And then when we added on the reasoning strategies, we went from five percent to 95%.”
Ian Fisher Feb 27, 2026 ▶ 14:02
Insight
Programmatic reasoning strategies in code vastly outperform automated prompt optimization
“That will get you some performance improvements, but it's very far from everything that you can get. If you actually think about these reasoning strategies that are really going to be written in code rather than in, in just better prompts.”
Ian Fisher Feb 27, 2026 ▶ 14:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.