Assertion Not checkable as stated
Poetiq achieves faster, cheaper recursive self-improvement than existing methods
“The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this.”
Assertion Not checkable as stated
Poetiq's agentic harness outperforms new base models without code changes
“With poetic what we end up giving you is a you know, people are calling these things harnesses now, but you know, or agentic system or whatever you want to call it, that sits on top of one or more language models, and it just performs better than them. And whe…”
Disclosure
Poetiq's meta-system generates reasoning systems for problems GPT-5 cannot reliably solve
“And so the core technology that we've developed at Poetic is recursive self-improvement. So we have a recursively self-improving system, which we call the Poetic meta system. The output of that system is systems that solve hard problems where a hard problem is…”
Insight
AI is replacing human engineers for dataset understanding and failure-mode detection
“Historically in machine learning, you always, you know, it's like the rule was you have to know your data set really well. But now we're kind of outsourcing that to the AI itself, where the AI is the, it's the AI's job to understand the dataset and figure out …”
Assertion Supported
Poetiq outperformed Gemini 3 Deep Think on ARC-AGI-2 at half the cost
“Yeah, so the interesting thing is that we were half the cost of Gemini Three Deep Think because we were building on top of Gemini Three Pro, which is a much cheaper model. But we still got in the end, a nine percentage point improvement on the official verific…”
Assertion Supported
Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”
Assertion Not checkable as stated
Poetiq's Humanity's Last Exam optimization run cost less than $100,000
“We didn't publish any cost for this, but I can say that the optimization costs us less than a hundred K, yeah.”
Insight
Fischer predicts Poetiq and base model S-curves will compound toward AGI
“Effectively you could say, like, each model or each set of models that we're working with will have their own S-curve. The poetic system, the poetic meta system itself, is also going to have its own S-curve. And so as the poetic meta system gets better, and as…”
Assertion Not publicly verifiable
Code-based reasoning strategies boosted Gemini 1.5 Flash performance from 5% to 95%
“In this particular case, you know, the hardest task we were working on, we got like to five percent performance with Gemini, 1.5 flash.
This was a while ago.
And then when we added on the reasoning strategies, we went from five percent to 95%.”
Insight
Programmatic reasoning strategies in code vastly outperform automated prompt optimization
“That will get you some performance improvements, but it's very far from everything that you can get. If you actually think about these reasoning strategies that are really going to be written in code rather than in, in just better prompts.”
Assertion Not checkable as stated
Frontier labs achieve recursive self-improvement by retraining models at each step
“And, you know, of course, Anthropic and OpenAI and Google, they're exploring recursive self-improvement, but Typically at that level of having the, you know, having to train a new model for every step of self-improvement that they do.”
Assertion Not checkable as stated
Poetiq's autonomous prompt generation system produced unexpected, non-human prompt structures
“It was pretty interesting to look at the prompt outputs in particular, I'd say, for ArcGi in that you know, I think you can read those and say, well, that's not what a human would have written. Pretty clearly. And it's, you know, there's some unexpected stuff …”