OpenAI

includes OpenAI O1, OpenAI Codex, OpenAI O3, OpenAI DevDay, OpenAI Realtime API, OpenAI Assistants API, OpenAI Deep Research, OpenAI Frontier, OpenAI API, OpenAI Responses API, OpenAI O1 Preview, OpenAI O1 Mini and 25 more

476 statements across 115 episodes · 254 bullish · 50 bearish · 113 people on the record · first statement Jun 20, 2023 by George Hotz · said 2,340 times in 230 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (467), Alessio Fanelli (151), Nathan Lambert (98), Michelle Pokrass (44), Ryan Lopopolo (39), Noam Brown (31), Brian Fioca (30), Dylan Patel (29)

tap a year for its mentions
00750751,5001502023202420252026episodesmentions
0751502023202420252026episodes it came up in
007.575151502023202420252026episodesmentions per episode
2026 483 mentions in 58 episodes 8 per episode
2025 1,113 mentions in 104 episodes 11 per episode
2024 657 mentions in 57 episodes 12 per episode
2023 87 mentions in 11 episodes 8 per episode

every mention, scene by scene, with the transcript →

Everything said about OpenAI, oldest first

Jun 20, 2023 neutral
Assertion Not publicly verifiable
Hotz: GPT-4 is an 8-way mixture model with 220B parameters per head
“GPT-IV is two hundred twenty billion in each head, and then it's an eight-way mixture model.”
George Hotz Jun 20, 2023 ▶ 49:48 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Jun 20, 2023 negative
Opinion
Hotz: Meta attracts researchers who want to publish while OpenAI keeps ideologues
“OpenAI can keep ideologues who, you know, believe ideological stuff, and Facebook can keep every researcher who's like, dude, I just want to build AI and publish it.”
George Hotz Jun 20, 2023 ▶ 55:16 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Aug 3, 2023 bullish
Prediction Not checkable as stated
LLaMA 2 will shift developers from closed APIs to self-hosting
“And I do see that's going to shift the balance of it. More and more folks are going to be using let's say derivatives of Lama two. More folks are going to Fine-tune and serve their own model instead of calling an API.”
Tri Dao Aug 3, 2023 ▶ 54:14 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 31, 2023 negative
Assertion Not checkable as stated
Most Open-Source LLMs Fall Flat Outside English-Speaking Nations
“Beyond OpenAI's model, and beyond ChatGPT and Claudia, the two big models, right? Outside of the English speaking nations, right? A lot of the open source models really fall flat.”
Eugene Cheah Aug 31, 2023 ▶ 23:00 RWKV: Reinventing RNNs for the Transformer Era
Sep 20, 2023 neutral
Opinion
Rizk: OpenAI's Moat Is Merely Paying the Training Bill
“OpenAI's moat is that they paid for the training bill. So they just have a good model.”
Youssef Rizk Sep 20, 2023 ▶ 53:24 Generating your AI Media Empire - with Youssef Rizk of Wondercraft.ai
Oct 20, 2023 positive
Insight
Howard: Rapid LLM race created massive technical debt and optimization opportunities
“There's a whole lot of technical debt everywhere, you know, nobody's really figured this stuff out because everybody's been so busy building what we know works as quickly as possible. So, yeah, I think there's a huge amount of opportunity to, you know, I think…”
Jeremy Howard Oct 20, 2023 ▶ 1:14:48 The End of Finetuning — with Jeremy Howard of Fast.ai
Oct 20, 2023 positive
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41 The End of Finetuning — with Jeremy Howard of Fast.ai
Oct 27, 2023 bearish
Prediction Not checkable as stated
Keydunov: AI will not make embedded analytics easier to monetize
“I don't think AI really going to change that just because it's using model, you just pay to open AI and that's it. Like everyone can do that, right? It's not much of a Competitive advantage. So it's going to be more like a commodity features that a lot of like…”
Artem Keydunov Oct 27, 2023 ▶ 38:46 Powering your Copilot for Data - with Artem Keydunov from Cube.dev
Nov 3, 2023 bearish
Prediction Not checkable as stated
Royzen: The leap from GPT-4 to GPT-5 will be smaller
“I think that GPT-IV, my hypothesis is that the jump from four to 4.5, or four to five, will be smaller than the jump from Three to four.”
Michael Royzen Nov 3, 2023 ▶ 37:38 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Nov 3, 2023 neutral
Assertion Supported
Swix: OpenAI is adopting the 'Applied AI Engineer' title
“Well, for what it's worth, OpenAI is adopting Applied AI Engineer.”
Shawn Wang Nov 3, 2023 ▶ 1:11:42 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Nov 3, 2023 negative
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Dec 5, 2023 bearish
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Dylan Patel Dec 5, 2023 ▶ 46:55 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 positive
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Dylan Patel Dec 5, 2023 ▶ 1:06:37 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 bullish
Opinion
Patel: Hardware vendors and developers are coalescing around OpenAI's Triton
“Likewise, there's OpenAI's Triton, like, what they're trying to do there, and like you know, everyone's really coalescing around Triton. You know, people, you know, third-party hardware vendors”
Dylan Patel Dec 5, 2023 ▶ 13:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Jan 5, 2024 bullish
Disclosure
Ruiz: tldraw is discussing a potential ChatGPT integration with OpenAI
“Hey, I'm talking to the good folks over at OpenAI tomorrow. Fingers crossed. Maybe we maybe we get it in, inside of ChatGPT or something.”
Steve Ruiz Jan 5, 2024 ▶ 1:26:40 The Accidental AI Canvas - with Steve Ruiz of tldraw
Jan 11, 2024 negative
Insight
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Nathan Lambert Jan 11, 2024 ▶ 29:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Prediction Not checkable as stated
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Nathan Lambert Jan 11, 2024 ▶ 50:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 bearish
Opinion
Lambert: OpenAI's rumored Q* was likely just a moderate benchmark bump
“They probably just got like a moderate bump on one of their benchmarks, and then everyone lost their minds, so it doesn't really matter.”
Nathan Lambert Jan 11, 2024 ▶ 3:34 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Insight
Lambert: Anthropic Constitutional AI and OpenAI Superalignment share intellectual roots
“The constitutional AI and the super alignment is, like, very conceptually linked. It's like a group of people that has, like, a very similar intellectual upbringing, and they work together for a long time, like, coming to the same conclusions in different ways…”
Nathan Lambert Jan 11, 2024 ▶ 1:11:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Nathan Lambert Jan 11, 2024 ▶ 14:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Assertion Supported
Lambert: GPT-4 Turbo Showed a Noticeable Jump on LMSYS Chatbot Arena
“GPT-IV Turbo is also notably ahead of the other GPT-IVs, which it kind of showed up immediately once they added it to the leaderboard, or to the arena, and I was like, all the GPT-IV memes aside, it seems like this is effectively a bump in the model.”
Nathan Lambert Jan 11, 2024 ▶ 1:25:01 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Supported
Lambert: Anthropic and OpenAI reward model loss functions are mathematically identical
“Fun fact is that these loss functions Look different and anthropic in opening eyes papers, but they're just literally just log transform. So if you start like expantiating both sides and taking the log of both sides, you'll like converge on one of the two, the…”
Nathan Lambert Jan 11, 2024 ▶ 54:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Feb 7, 2024 bullish
Prediction Not checkable as stated
Future Llama models will match frontier GPT models if progress slows
“As AI progress slows down, so if we get like Llama-IV, Llama-V for example, maybe it's a comparable at that point, like GPT-V or GPT-VI, like, It made it to the point where it was like, look, I just want to use Lama. Like, it's, you know, safe for me to, you k…”
David Hsu Feb 7, 2024 ▶ 50:30 The State of AI in production — with David Hsu of Retool
Feb 7, 2024 bullish
Opinion
OpenAI would decisively defeat any AI startup competitors
“I think if I was competing in startups, OpenAI would win for sure. Like, at this point, OpenAI is so far ahead from both a model and a pricing perspective that, like, there is no reason for it to go really, I think, in my opinion, at least, a startup model.”
David Hsu Feb 7, 2024 ▶ 49:48 The State of AI in production — with David Hsu of Retool
Feb 19, 2024 neutral
Disclosure
Modal Will Not Try to Compete Directly With OpenAI's API
“So many people get started with APIs and that's just, you know, they're just dominating a space in particular, open AI. Right. And that's not necessarily like a place where we aim to compete. I mean, maybe at some point, but like, it's just not like a core foc…”
Erik Bernhardsson Feb 19, 2024 ▶ 41:02 Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
Mar 6, 2024 bullish
Prediction Not checkable as stated
Chintala: Centralized feedback could trigger open source runaway over OpenAI
“If that central sinkhole is there, who's gonna go coordinate all of this integration across all of these, like, open source frontends? But I think if we do that, if that actually happens, I think that probably has a real chance of the open source models having…”
Soumith Chintala Mar 6, 2024 ▶ 1:14:45 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Mar 27, 2024 neutral
Assertion Not checkable as stated
Brockman and Sutskever handed Luan their direct reports to do IC work
“I think second or second or third day of my time at OpenAI Greg and Ilya pulled me in a room and were like, hey you know, the you should take over our directs and we'll go mostly do IC work.”
David Luan Mar 27, 2024 ▶ 1:57 Why Google failed to make GPT-3 -- with David Luan of Adept
Mar 27, 2024
Assertion Not checkable as stated
Altman and Luan pitched Microsoft's top leadership before the OpenAI investment
“The last meeting we did with Microsoft, Before Microsoft invested in OpenAI, Sam Altman, myself, and our CFO flew up to Seattle to do the final pitch meeting. And I'd been a founder before, so I always had like a tremendous amount of anxiety about partner meet…”
David Luan Mar 27, 2024 ▶ 14:18 Why Google failed to make GPT-3 -- with David Luan of Adept
Mar 27, 2024
What-if
Google would have crushed OpenAI by giving Noam Shazeer half its TPUs
“That muscle did not exist during my time at Google. And I think had they had it, what they would have done would be say, hey, Noam Shazir, you're a brilliant guy. You know how to scale these things up? Like, here's half of all of our TPUs. And then I think the…”
David Luan Mar 27, 2024 ▶ 8:29 Why Google failed to make GPT-3 -- with David Luan of Adept
Apr 6, 2024
Assertion Not checkable as stated
Murphy: Azure OpenAI significantly beats OpenAI's hosted API 400-600ms latency.
“And then GPD, 3.5 turbo or four, you probably get, you know, 400, maybe 600 milliseconds of latency in their hosted API. And if you go into Azure and you use their services, you can get that down a lot lower.”
Damien Murphy Apr 6, 2024 ▶ 10:17 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Apr 24, 2024 bullish
Opinion
Liu: Claude 3 Haiku outperforms OpenAI models at function calling
“Overall, I'm like super happy with the anthropic models compared to the OpenAM models. Like, Sonnet is very cost effective. Haiku is, in function calling, it's actually better.”
Jason Liu Apr 24, 2024 ▶ 21:57 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Apr 27, 2024 positive
Opinion
Claude 3 can bypass its assistant persona to expose the underlying simulator
“Instead of having this entity, like GPT-IV, that's an assistant that just pops up in your face that you have to kind of, like, punch your way through and continue to have to deal with as a headache, instead, there's ways to kindly coax Claude into having the a…”
Karan Malhotra Apr 27, 2024 ▶ 8:37 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Jun 25, 2024 neutral
Assertion Supported
Frankle: OpenAI, Google, Meta, and Apple have data deals with Shutterstock
“And you know, I, at least I've heard in the news, like opening, I Google, Meta Apple have all called Shutterstock and made those deals.”
Jonathan Frankle Jun 25, 2024 ▶ 3:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jul 5, 2024 neutral
Opinion
Tay: ChatGPT's release made task-specific academic NLP research obsolete
“The big thing about the ChatGPT moment of, like, twenty-twenty-two, the thing that changed drastically is, like, it completely, like, it was, like, this sharp, like, make all this work, like, kind of, like, obsolete”
Yi Tay Jul 5, 2024 ▶ 6:26 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Jul 5, 2024 neutral
Assertion Not checkable as stated
Tay: Google and OpenAI built general models three years before academia
“Places like Google and Meta, OpenAI, we will be working on things, like, Three years ahead of everybody else, and then suddenly, like, then Academia would be, like, still working on, like, these task-specific things.”
Yi Tay Jul 5, 2024 ▶ 7:07 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Jul 23, 2024 neutral
Opinion
Scialom: OpenAI likely understood scaling laws before Chinchilla was published
“To be fair, I think OpenAI knew that at the time of Chinchilla paper.”
Thomas Scialom Jul 23, 2024 ▶ 10:55 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Aug 2, 2024 positive
Assertion Supported
Reddit makes over 200 million dollars in AI data licensing deals
“Yeah, the, I guess the winner in all of this is Reddit, which is making over two hundred million just in data licensing to OpenAI and some of the other AI providers.”
Alessio Fanelli Aug 2, 2024 ▶ 36:09 The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
Aug 17, 2024 negative
Insight
Howard: Non-profit boards cannot control commercial entities with equity-compensated staff
“This didn't make sense to have like a so-called non-profit where then there are people working at a commercial company that's owned by or controlled nominally by the non-profit where the people in the company are being given the equivalent of stock options. Li…”
Jeremy Howard Aug 17, 2024 ▶ 7:55 Answer.ai & AI Magic with Jeremy Howard
Aug 22, 2024
Assertion Supported
Y Combinator companies had a direct Slack channel to OpenAI
“Which is a little bit, you know, for a year or so, YC companies had like a direct Slack channel to open AI.”
Shawn Wang Aug 22, 2024 ▶ 13:08 Is finetuning GPT4o worth it?
Aug 22, 2024 positive
Disclosure
Access to GPT-4 Turbo experimental fine-tuning enabled Cosine Genie's creation
“Eventually we were able to get on the experimental access program and we got access to four turbo fine tuning. As soon as we did that, because in the entire run up to that, we'd built the data pipeline. We already had all that set up. So we're like, right, we …”
Alistair Pullen Aug 22, 2024 ▶ 14:14 Is finetuning GPT4o worth it?
Aug 22, 2024 neutral
Disclosure
Cosine receives larger OpenAI LoRA adapters than public tiers due to volume
“Actually we use models that are larger than what's publicly available, something publicly available yet, but when this goes out, it will be, but we have larger law adapters available to us, just because the amount of data that we're pumping through it”
Alistair Pullen Aug 22, 2024 ▶ 43:41 Is finetuning GPT4o worth it?
Aug 22, 2024 negative
Assertion Not publicly verifiable
Genie's SWE-bench success rate drops to roughly 50% past 60k tokens
“Performance of Jeannie over the length of the context window degrades fairly linearly. So actually, I actually broke it down by probability of solving a SWE bench issue. Given the number of tokens of the context window at 60 K, it's basically .5. So if you go …”
Alistair Pullen Aug 22, 2024 ▶ 36:26 Is finetuning GPT4o worth it?
Aug 28, 2024
Assertion Supported
Carlini extracted production models from Google and OpenAI with legal permission
“We ran the attack that let us, yeah, stole several of OpenAI's models. With their permission... We notified everyone who was vulnerable to this attack. Some Google models were vulnerable. Some open AM models were vulnerable. There were one or two other people …”
Nicholas Carlini Aug 28, 2024 ▶ 57:21 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Aug 28, 2024 negative
Insight
Carlini: GPT-4 would exist identically without adversarial machine learning research
“Nothing about GPT-IV would be at all different if the field of, like the entire field of Everson machine learning disappeared. Like everything to do with Everson examples, like all of the, like for the most part, like GPT-IV would exist identically.”
Nicholas Carlini Aug 28, 2024 ▶ 58:56 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Sep 17, 2024 neutral
Assertion Not checkable as stated
OpenAI built Structured Outputs because developers were hacking function calling
“A lot of people were hacking function calling to get the response format they needed, and so this is why we shipped kind of this new response format. So you can get exactly what you want, and you get kind of more of the models verbosity.”
Michelle Pokrass Sep 17, 2024 ▶ 14:24 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Assertion Supported
OpenAI's Assistants API increased its file limit to 10,000 files
“Before, we only supported, I think, like, 20 files per assistant, and the way we used those files was, like, less effective. Basically, the model would decide, based on the file name, whether to search a file, and there's, like, not a ton of information in the…”
Michelle Pokrass Sep 17, 2024 ▶ 51:45 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Insight
Building custom evals is high leverage for AI application developers
“I think for customers, and we work with a lot of customers, really developing their own evals is super high leverage. Because then you can upgrade really quickly when we have a new model, you can experiment with these things with confidence.”
Michelle Pokrass Sep 17, 2024 ▶ 37:26 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Disclosure
Structured Outputs does not yet support parallel function calling
“All our models support it, or all of our newer models support it, but we don't support it with structured outputs right now.”
Michelle Pokrass Sep 17, 2024 ▶ 25:36 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024
Assertion Not checkable as stated
OpenAI's Applied team now exceeds its entire pre-ChatGPT headcount
“Applied now is bigger than the company when I joined.”
Michelle Pokrass Sep 17, 2024 ▶ 8:51 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 bullish
Prediction Open · timeframe Sep 2027
OpenAI likely to release streaming video API with frame sampling
“Yeah, I think it's very possible that we'll have an API where you stream video in, and maybe, you know, to start, we'll do the frame sampling for you.”
Michelle Pokrass Sep 17, 2024 ▶ 1:00:55 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Insight
Constrained decoding alone degrades output quality without model training
“And so it's not enough to just kind of constrain the model. I think of that as the engineering side, whereas basically you mask the available tokens that are produced every time to only fit the schema. And so you can do this engineering thing and you can force…”
Michelle Pokrass Sep 17, 2024 ▶ 11:07 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Disclosure
OpenAI combined model training with constrained decoding for Structured Outputs
“We trained a model which is significantly better than our past models at following formats. And we did the end work to serve like this constrained decoding concept at scale.”
Michelle Pokrass Sep 17, 2024 ▶ 11:41 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Assertion Not checkable as stated
Whisper v2 outperforms v3 at certain tasks, delaying its API rollout
“And so whisper V two is better at some things than whisper V three. And so it didn't seem that worthwhile to ship whisper V three compared to like the other things in our priorities. I think we still will at some point, but yeah, it's just, you know, there's a…”
Michelle Pokrass Sep 17, 2024 ▶ 1:02:23 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Disclosure
OpenAI aims to become an AI development platform, not just LLM-as-a-service
“We want to do more in this space and not just be an LLM as a service, but kind of AI development platform as a service.”
Michelle Pokrass Sep 17, 2024 ▶ 56:27 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Disclosure
OpenAI hires applied engineers without prior AI experience
“We've hired people with all kinds of backgrounds, people who have PhD in an ML or folks who have just done engineering like me, and we're really hiring for a lot of teams. We're hiring across the applied org, which is where I sit for engineering, and for a lot…”
Michelle Pokrass Sep 17, 2024 ▶ 1:11:04 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Assertion Supported
Structured response format is limited to GPT-4o and GPT-4o mini
“Actually, the new response format is only available on two models. It's Foro Mini and the new Foro. So the old Foro doesn't have the new response format. However, for function calling, we were able to enable it for all models that support function calling, and…”
Michelle Pokrass Sep 17, 2024 ▶ 30:23 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Insight
Fine-tuning requires only 100 to 1,000 high-quality examples
“It's actually a lot easier to get started than a lot of people expect. I think they might need Tens of thousands of examples, but even a hundred really high quality ones or a thousand is enough to get going.”
Michelle Pokrass Sep 17, 2024 ▶ 41:57 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Assertion Supported
The API was OpenAI's first commercial product
“The API is actually OpenAI's first product, and the first idea for commercialization, that predates me as well.”
Michelle Pokrass Sep 17, 2024 ▶ 48:41 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Opinion
OpenAI's API is its broadest vehicle for distributing AGI
“So I believe that the API is kind of our broadest vehicle for distributing AGI. You know, we're building some first party products, but they'll never reach every niche in the world and kind of every corner in community.”
Michelle Pokrass Sep 17, 2024 ▶ 47:41 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024
Assertion Not checkable as stated
OpenAI's API team had just five engineers before ChatGPT launched
“I would say the applied team was maybe like 30 or 40 people, and yeah, probably closer to 30, and there was maybe like five-ish total working on the API at most.”
Michelle Pokrass Sep 17, 2024 ▶ 8:37 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024
Assertion Not checkable as stated
OpenAI built its constrained decoding engine from scratch
“Yeah, we didn't use any kind of Other stuff. We kind of built, you know, our solution from scratch to meet our specific needs.”
Michelle Pokrass Sep 17, 2024 ▶ 16:45 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Assertion Supported
Structured Outputs supports recursive schemas for generating dynamic UIs
“The schemas, we support recursive schemas, and this allows you to do really cool stuff, like, you know, every UI is a nested tree that has children, and so I thought that was super cool. You can use one schema and generate, like, tons of UIs.”
Michelle Pokrass Sep 17, 2024 ▶ 31:20 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Assertion Supported
OpenAI Structured Outputs enforces schemas in one shot without retries
“We are not retrying, you know, we're doing it in one shot and this is how you save on latency and cost.”
Michelle Pokrass Sep 17, 2024 ▶ 38:32 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Assertion Supported
OpenAI's seed parameter is best-effort and not fully deterministic
“Yeah, the seed parameter is not fully deterministic, and it's kind of a best effort thing. So you'll notice there's more determinism in the first few tokens. That's kind of the current implementation.”
Michelle Pokrass Sep 17, 2024 ▶ 53:38 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Opinion
Standard request-response APIs will not work for speech-to-speech models
“I think just the regular request response probably isn't going to be the right solution.”
Michelle Pokrass Sep 17, 2024 ▶ 1:04:53 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Assertion Not checkable as stated
OpenAI's function calling originated from early Code Interpreter prototypes
“The history here is we started with function calling and function calling, you know, came from the idea of like, let's give the model access to tools and let's see what it does. And we basically had these internal prototypes of what code interpreter is now. An…”
Michelle Pokrass Sep 17, 2024 ▶ 13:27 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 neutral
Disclosure
OpenAI rejected pre-registering schema IDs to avoid developer complexity
“The alternative design space that we explored It's like pre-registering your schema, so like a totally different endpoint, and then passing in like a schema ID. But we thought, you know, that was a lot of overhead, and like another endpoint to maintain, and ju…”
Michelle Pokrass Sep 17, 2024 ▶ 27:00 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024
Assertion Not checkable as stated
ChatGPT's launch caused Postgres bottlenecks by reusing API developer accounts
“Surprisingly there were a lot of Postgres issues when ChatGPT came out because the accounts for like ChatGPT were tied to the accounts in the API. And so you're basically creating a developer account to log into ChatGPT at the time, because it's just what we h…”
Michelle Pokrass Sep 17, 2024 ▶ 9:20 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 positive
Assertion Supported
Structured Outputs prevent models from hallucinating invalid tools
“The model before was able to hallucinate a tool, but now it's it can't when you're using structured outputs.”
Michelle Pokrass Sep 17, 2024 ▶ 22:42 Building AGI with OpenAI's Structured Outputs API
Sep 19, 2024 neutral
Assertion Not checkable as stated
Jamil: OpenAI and Cohere Overlap Prefill and Generation to Maximize GPU Utilization
“Token generation is memory bound means that the limitation is only given by how much your KVCache can hold. So the memory can hold in terms of KVCache. While prefilling is compute bound, so to maximize the GPU utilization, whenever you work with OpenAI or Cohe…”
Umar Jamil Sep 19, 2024 ▶ 43:05 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Sep 20, 2024 neutral
Assertion Not checkable as stated
Schulhoff: GPT-4 Fails to Output Reasoning on 1 in 100 to 1,000 Prompts
“I remember I did a lot of experiments with GPT-IV, and especially when you look at it at scale, so I'll run thousands of prompts against it through the API, and I'll see, you know, every one in a hundred, every one in a thousand outputs no reasoning whatsoever…”
Sander Schulhoff Sep 20, 2024 ▶ 33:03 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Sep 20, 2024 neutral
Assertion Supported
Schulhoff: Preamble Discovered Prompt Injection Before Riley Goodside
“Preamble is the company that first discovered Prompt Injection, even before Riley, and they, like, responsibly disclosed it, kind of, internally to OpenAI”
Sander Schulhoff Sep 20, 2024 ▶ 4:46 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Sep 27, 2024 positive
Assertion Supported
Harrison Chase says OpenAI recommends adding a thought field to tool schemas
“I think open AI even recommended, like when you're doing tool calling, it's sometimes helpful to put like a thought field in the tool along with all the actual acquired arguments and then have that one first. So it fills out that first and then, and that's, th…”
Harrison Chase Sep 27, 2024 ▶ 12:10 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Sep 27, 2024
Assertion Not checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu Yao Sep 27, 2024 ▶ 2:12 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Oct 4, 2024 bullish
Prediction Not checkable as stated
Altman: AI will dynamically render custom real-time interfaces for any request
“At some point in not that many years in the future, you'll walk up to a piece of glass, you will say whatever you want they will have, like, there will be incredible reasoning models, agents connected to everything, there will be a video model streaming back t…”
Sam Altman Oct 4, 2024 ▶ 2:08:17 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 positive
Assertion Not checkable as stated
Altman: o1 is OpenAI's most aligned model ever by a lot
“And O-one is obviously our most capable model ever, but it's also our most aligned model ever by a lot.”
Sam Altman Oct 4, 2024 ▶ 1:33:35 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 neutral
Assertion Supported
Weil: ChatGPT supports over 200 million weekly active users
“As we, you know, we support over two hundred million people every week on ChatGPT.”
Kevin Weil Oct 4, 2024 ▶ 1:51:49 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 positive
Assertion Supported
Huet: Over 3 million developers build on OpenAI
“I'm sure we talked about this before, but there's now more than three million developers building on OpenAI, so it's pretty exciting to see all of that energy into creating new things.”
Romain Huet Oct 4, 2024 ▶ 39:27 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 positive
Disclosure
Godement: OpenAI plans API knobs for developers to adjust safety thresholds
“And so I think the direction where we'll go here is that, basically, there will always be, like, you know, a set of behavior that will, you know, just, like, forbid, frankly, because they're illegal against our terms of services, but then there will be, like, …”
Olivier Godement Oct 4, 2024 ▶ 34:59 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 neutral
Opinion
Pokras: OpenAI Assistants API requires too many initial API requests
“Some of the things that are good in the assistance API is hosted tools. People really like posted tools and especially RAG. And then some things that are, you know, less intuitive is just how many API requests you need to get going with the Assistant's API.”
Michelle Pokrass Oct 4, 2024 ▶ 1:05:05 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 positive
Assertion Not checkable as stated
Godement: Vision fine-tuning shows higher performance uplift than text
“We've been alpha testing, like, the vision fine tuning, like, for several weeks at that point. We are seeing, like, even higher performance uplift compared to text fine tuning.”
Olivier Godement Oct 4, 2024 ▶ 25:34 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 positive
Assertion Not checkable as stated
Weil: OpenAI customer support team is 20% expected size thanks to AI
“There are things that get closer to that, I mean, there, like, customer service, we have bots internally that do what's fun about answering external questions and fielding internal people's questions on Slack and so on, and our customer success, our customer s…”
Kevin Weil Oct 4, 2024 ▶ 1:57:42 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 bullish
Disclosure
Pokras: OpenAI will ship raw audio in Chat Completions API
“We're actually going to be shipping audio capabilities in chat completions. So this is like the lowest level capability. So you supply in audio and you can get back raw audio and it works at the request response layer.”
Michelle Pokrass Oct 4, 2024 ▶ 59:41 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 neutral
Disclosure
Godement: OpenAI will only build tools closest to the model
“There is no freaking way that OpenAI can build everything. Like, there is just too much to build, frankly. And so, my philosophy is, essentially, we'll focus on, like, the tools which are, like, the closest to the model itself. So that's why you see us, like, …”
Olivier Godement Oct 4, 2024 ▶ 31:05 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 bullish
Assertion Not checkable as stated
Altman: OpenAI reached Level 2 AGI with o1
“I think we clearly got to level two, or we clearly got to level two with O-one.”
Sam Altman Oct 4, 2024 ▶ 1:24:50 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 bullish
Assertion Supported
Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit
“Yeah, I sat in the distillation session just now, and they showed how they distilled from four to four mini, and it was like only like a two percent hit in the performance, and 15 X cheaper.”
Shawn Wang Oct 4, 2024 ▶ 48:14 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 positive
Disclosure
Weil: OpenAI o1 will support function calling by end of 2024
“I'm really excited to see things like system prompts, And structured outputs, and function calling, make it into a one, we will be there by the end of the year.”
Kevin Weil Oct 4, 2024 ▶ 1:48:29 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 bullish
Opinion
Pokras: Vision fine-tuning is the most underrated release for bespoke OCR
“Vision fine-tuning is so underrated. For the past, like, two months, whenever I talk to founders, they tell me this is the thing they need most. A lot of people are doing, like, OCR on, on very bespoke formats, like government documents, and vision fine-tuning…”
Michelle Pokrass Oct 4, 2024 ▶ 56:21 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024 bullish
Prediction Held up
Biggio: Real-time video/vision is likely the next Realtime API feature
“To use ChatGPT's voice mode as an example, like we've demoed the video, right? Like real-time image, right? So I'm not actually sure what timelines are, but I would expect, if I had to guess, that like that is probably the next thing that we're going to be mak…”
Ilan Biggio Oct 4, 2024 ▶ 16:15 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 4, 2024
Insight
Godement: Founders should build apps slightly too difficult for current models
“As a developer, as a founder, you basically want to build an app, which is a bit too difficult for the model today, right? Like what you think is right. It's like sort of working, sometimes not working. And that way, you know, that basically gives us like a go…”
Olivier Godement Oct 4, 2024 ▶ 37:14 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 11, 2024 bullish
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Ankur Goyal Oct 11, 2024 ▶ 1:27:25 Production AI Engineering starts with Evals
Oct 11, 2024 bearish
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Ankur Goyal Oct 11, 2024 ▶ 1:34:00 Production AI Engineering starts with Evals
Oct 11, 2024 neutral
Assertion Not checkable as stated
Goyal: OpenAI dominates production while Anthropic Sonnet leads side projects
“We still see an overwhelming majority of customers using OpenAI, but almost everyone is using Anthropic for the, and Sonnet specifically for their side projects, whether it's You know, via cursor or prototypes or whatever.”
Ankur Goyal Oct 11, 2024 ▶ 1:26:58 Production AI Engineering starts with Evals
Oct 11, 2024 bullish
Assertion Not checkable as stated
Goyal: Braintrust saw nearly 100% OpenAI market share pre-Claude 3
“Pre-Claude III, it was close to a hundred percent OpenAI.”
Ankur Goyal Oct 11, 2024 ▶ 1:24:59 Production AI Engineering starts with Evals
Oct 11, 2024 bearish
Assertion Not checkable as stated
Goyal: Public clouds fail to match direct OpenAI endpoint experience and capacity
“It hasn't been a smooth journey for people to get the capacity on public clouds that they're able to get through, you know, OpenAI directly. I mean, I think a lot of this is changing, catching up, et cetera. But it hasn't been perfectly smooth. And I think the…”
Ankur Goyal Oct 11, 2024 ▶ 1:28:43 Production AI Engineering starts with Evals
Oct 11, 2024 negative
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24 Production AI Engineering starts with Evals
Oct 19, 2024 positive
Assertion Supported
Hu: OpenAI o1-preview achieves bronze medals in 17% of MLE-bench competitions
“Their final results with a one preview and this a scaffolding from a different company was that they got a bronze medal. I don't think I've ever achieved once but I haven't competed that much in. 17% of competitions.”
Jesse Hu Oct 19, 2024 ▶ 45:06 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Oct 19, 2024 positive
Assertion Supported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Jesse Hu Oct 19, 2024 ▶ 47:29 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Oct 19, 2024 positive
Assertion Supported
Wang: OpenAI is opening an office in Singapore
“You know, OpenAI is opening an office in Singapore, and how we can grow more AI engineers in Singapore as well, because I think, I do think that that is something that people are interested in, whether or not it's for their own careers or to hire out in Singap…”
Shawn Wang Oct 19, 2024 ▶ 8:09 Singapore: the AI Engineer Nation — with Minister Josephine Teo
Oct 25, 2024 negative
Opinion
Swix: OpenAI's Realtime API does not handle interruptions well
“OpenAI has just launched a real-time chat. It's a very hot topic. I would say one of the toughest AI engineering disciplines out there, because even their API doesn't do interruptions that well, to be honest.”
Shawn Wang Oct 25, 2024 ▶ 1:08:09 How NotebookLM Was Made
Nov 1, 2024 bullish
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 2, 2024 positive
Opinion
Isolating Tangent Function Instability is OpenAI's Core Contribution in sCM
“And this is, I, in my opinion, the meat of the paper. So you have this, part of the, you have this tangent function that I had called out.”
RJ Honicky Nov 2, 2024 ▶ 31:47 [Paper Club] Intro to Diffusion Models and OpenAI sCM: Simple, Stable, Scalable Consistency Models
Nov 2, 2024 positive
Assertion Supported
Stabilization Techniques Help Continuous Consistency Models Outperform Discrete Models
“And so like when you stack all of these things together, then you're able to train much more effectively and continuous time does much better. Then these discrete, this n is the number of discrete steps that your model is taking, and, you know, maybe one inter…”
RJ Honicky Nov 2, 2024 ▶ 39:25 [Paper Club] Intro to Diffusion Models and OpenAI sCM: Simple, Stable, Scalable Consistency Models
Nov 2, 2024
Insight
Consistency Models Map Any Trajectory Point Directly to Original Data
“What a consistency model does is it says that everything should be on the same trajectory, right? So I'm gonna, if I estimate it, I can I'm gonna learn how to map from any point on this trajectory to the to this point in the data.”
RJ Honicky Nov 2, 2024 ▶ 19:46 [Paper Club] Intro to Diffusion Models and OpenAI sCM: Simple, Stable, Scalable Consistency Models
Nov 11, 2024 positive
Opinion
Polu: Sam Altman mastered technical ML details within two years at OpenAI
“One thing about Sam Altman, he really impressed me, because when I joined, he had joined not that long ago. And it felt like he was kind of a very high level CEO. And I was mind blown by how deep he was able to go into the subjects within a year or something, …”
Stanislas Polu Nov 11, 2024 ▶ 14:44 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024
Insight
Polu: OpenAI Managed Research Priorities Directly Through Compute Allocation
“In that space, there's a managing tool that is great, which is computer location. Basically, by managing the computer location, you can message the team of where you think the priority should go. And so it was really a question of you were free as a researcher…”
Stanislas Polu Nov 11, 2024 ▶ 10:31 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024 positive
Insight
Polu: Combining LLMs with formal math pairs creativity with proof verification
“Transformers are very creative, but yet they do mistakes. And formal math systems are the ability to verify a proof. And the tactics they can use to solve problems are very mechanical. So you miss the creativity. And so the idea was to try to explore both toge…”
Stanislas Polu Nov 11, 2024 ▶ 6:30 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024
Assertion Not checkable as stated
Polu: OpenAI Believed in Transformer Scaling Pre-Kaplan Paper
“Before that, there really was a strong belief in, in scale. I think it was just the belief that the transformer was a generic enough architecture that you could learn anything, and that this was just a question of scaling.”
Stanislas Polu Nov 11, 2024 ▶ 14:17 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024 positive
Opinion
Polu: GPT-4 Turbo performs better than GPT-4o on function calling
“I personally don't have proof, but I know many people, and I'm probably part of them, to think that GPT-IV Turbo is still better than GPT-IV on function calling.”
Stanislas Polu Nov 11, 2024 ▶ 42:04 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024
Assertion Partly supported
Polu: GPT-4 was ready internally at OpenAI months before September 2022
“I had seen GPT-IV internally at the time. It was September, 20, 22. So it was pre-chat GPT, but GPT-IV was ready since, I mean, I'd been ready for a few months internally.”
Stanislas Polu Nov 11, 2024 ▶ 19:16 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024
Assertion Not checkable as stated
Polu: OpenAI's GPT-3 Was Internally Codenamed Project Nest
“Most of the compute was going to a product called Nest, which was basically GPT-free.”
Stanislas Polu Nov 11, 2024 ▶ 10:05 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024 positive
Assertion Not checkable as stated
Polu: Ilya Sutskever Spent Surprising Amount of Time Communicating OpenAI Vision
“I think he was really focused on building the vision and communicating the vision within the company, which was extremely It's extremely useful. I was personally surprised that he spent so much time, you know, working on communicating that vision and getting t…”
Stanislas Polu Nov 11, 2024 ▶ 13:22 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024 neutral
Opinion
Polu: Anthropic split was driven by disagreement over OpenAI's API commercialization
“What I understood of it is that there was a disagreement of the commercialization of that technology. I think the focal point of the disagreement was the fact that we started working on the API and wanted to make those models available through an API. Is that …”
Stanislas Polu Nov 11, 2024 ▶ 17:03 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 25, 2024 positive
Disclosure
Fireworks AI will release a reasoning model inspired by OpenAI's o1
“So another announcement is we will also announce a, our next Declarative system is going to be appear as a model that has extremely high quality, and this model is inspired by O-one announcement from OpenAI. You should see that by the time we announce this o…”
Lin Qiao Nov 25, 2024 ▶ 29:26 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 28, 2024 neutral
Assertion Supported
Schluntz: SWE-bench Verified was created in partnership with OpenAI
“SweetBench Verified was actually made in partnership with OpenAI, and they hired humans to go review all these tasks and pick out a subset to try to remove any obstacle like this that would make the tasks impossible.”
Erik Schluntz Nov 28, 2024 ▶ 10:03 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Dec 21, 2024 neutral
Assertion Supported
Reddy: Flagship OpenAI API costs fell 80% to 85% in roughly 18 months
“This is a graph of flagship OpenAI model costs, where the cost of the API has come down roughly 80, 85%, and call it the last year, year and a half which is pretty remarkable.”
Pranav Reddy Dec 21, 2024 ▶ 6:56 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Dec 21, 2024 bullish
Prediction Not checkable as stated
Reddy: Test-time scaling currently works best for verifiable domains like math
“It seems like OpenAI has cracked a version of this that works, and we think A, Foundation Model Labs will come up with better ways of doing this, and B, so far it largely works for very verifiable domains, things that look like math and physics and maybe secon…”
Pranav Reddy Dec 21, 2024 ▶ 10:33 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Dec 21, 2024 bearish
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Pranav Reddy Dec 21, 2024 ▶ 5:08 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Dec 23, 2024 negative
Assertion Supported
Soldani: Content owners blanket block crawling due to closed AI models
“What they found is, as a reaction to, like, the close like, of the existence of closed models, like OpenAI or Cloud GPT or Cloud a lot of content owners have blanket blocked any type of crawling to their website.”
Luca Soldani Dec 23, 2024 ▶ 18:34 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Dec 23, 2024
Insight
Soldani: Replicating OpenAI's o1 requires roughly 10,000 GPUs
“If you're interested in you know, your, Open replication of what OpenAI's O-one is you're gonna be on the 10 K spectrum of our GPUs.”
Luca Soldani Dec 23, 2024 ▶ 12:08 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Dec 25, 2024 positive
Opinion
Neubig: GPT Loops On Errors While Claude Tries New Approaches
“So, like, GPT doesn't have very good air recovery ability. And so, because of this, it will go into loops and do the same thing over and over and over again, whereas Claude does not do this.”
Graham Neubig Dec 25, 2024 ▶ 14:25 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Jan 1, 2025 bullish
Prediction Open · timeframe Jan 2028
Swyx predicts OpenAI will launch a $2,000 per month ChatGPT tier
“I think that 2000 dollars ChatGPT will come.”
Shawn Wang Jan 1, 2025 ▶ 18:02 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 1, 2025 bullish
Prediction Not checkable as stated
Swyx: Diff mode will become the norm for AI code tools in 2025
“Canvas has incorporated the diff mode that both Anthropic and OpenAI and Fireworks has now shipped that I think is going to be the norm for next year, that everyone Need some kind of diff mode code interpreter thing.”
Shawn Wang Jan 1, 2025 ▶ 59:25 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 1, 2025 positive
Opinion
Swyx: DeepMind has roughly a four-year advantage over OpenAI in world modeling
“So like they have maybe four years advantage on world modeling that OpenAI does not have. Cause OpenAI basically only started Diffusion Transformers last year when they hired Build Peebles. So DeepMind has a bit of advantage here.”
Shawn Wang Jan 1, 2025 ▶ 51:53 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 1, 2025 neutral
Assertion Partly supported
Swyx: OpenAI production market share dropped from 95% to 50–75%
“Basically over the course of 23, going into 24, OpenAI has gone from 95 market share to reasonably somewhere between 50 to 75 market share.”
Shawn Wang Jan 1, 2025 ▶ 12:20 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 2, 2025 bullish
Opinion
Lambert: OpenAI's o1 uses token streams as intermediate state compute
“Why oh, one is exciting is because it's a new type of language models that are going to maximize on this view of reasoning, which is that chain of thought in kind of a forward stream of tokens can actually do a lot to achieve better outcomes when you're doing …”
Nathan Lambert Jan 2, 2025 ▶ 4:42 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 2, 2025 positive
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 2, 2025 bullish
Prediction Not checkable as stated
Lambert: Open community will eventually match OpenAI's large-scale RL infrastructure
“And this is something that these early relative models are not going to be doing because we don't like, no one has this infrastructure like open AI does. It'll take a while to do that, but people will make it.”
Nathan Lambert Jan 2, 2025 ▶ 8:34 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 2, 2025 positive
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Nathan Lambert Jan 2, 2025 ▶ 5:23 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 2, 2025 negative
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 10, 2025 bullish
Disclosure
Bryk: Exa applies OpenAI's o1 variable compute paradigm to web search
“One way of thinking about what we built is like O-one for search because, Oh, one is all about like, you know, some questions require more compute than others, and we'll put as much compute into the question as we need to solve it. So similarly with our search…”
Will Bryk Jan 10, 2025 ▶ 13:15 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Jan 10, 2025 positive
Opinion
Bryk: Exa is the 'OpenAI of search' building AGI for retrieval
“I often say we're the OpenAI of search because we're a research company, we're a research startup that does like fundamental research into making like AGI for search in a way. And then we have all these like business products that come out of that.”
Will Bryk Jan 10, 2025 ▶ 4:54 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Jan 17, 2025 negative
Insight
Hylak: OpenAI o1 struggles to match personal tone and writing styles
“I think that I've had a very hard time getting it to actually write stuff. I know that I've heard of people using it for writing where it's like processing diffs, more like providing critiques or feedback, but at least for myself, I haven't found a good way to…”
Ben Hillock Jan 17, 2025 ▶ 10:24 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 neutral
Insight
Swyx: Place dynamic context at the end of prompts for efficient caching
“Anyone who's doing, who's working with a ton of context has to put them at the end because they're swapping them out. That's just how it is. I don't see any way around it.”
Shawn Wang Jan 17, 2025 ▶ 30:16 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 bullish
Opinion
McAteer: OpenAI o1 is the first model that grows more impressive over time
“O-one is actually the first model where I'm getting more impressed by it the more I use it. So, like, when ChatGPT first came out, right, I think it was GPT-III. And at first it seemed like, oh wow, this is amazing, it can actually create text that sounds like…”
Dan McAteer Jan 17, 2025 ▶ 2:10 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 neutral
Insight
Hylak: Hidden reasoning tokens create an information asymmetry between OpenAI and developers
“I think that what makes a one even trickier than other models is that there is actually an asymmetric miss to how well open AI understands the model and how well we, for example, the fact that like reasoning tokens are hidden, right? So there's all this stuff …”
Ben Hillock Jan 17, 2025 ▶ 12:27 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 neutral
Insight
Hylak: Users willing to wait five minutes for AI will wait an hour
“I think that it's like the number of tasks that you're willing to wait, you know, like 3:05 minutes for it. It's like probably pretty similar to the number of tasks you're willing to wait like an hour for, which is interesting.”
Ben Hillock Jan 17, 2025 ▶ 16:40 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 positive
Insight
Fanelli: OpenAI o1 is a goal-based reasoning model, not a chatbot
“Like O-one is not a chat model. And I think this is both from a usage perspective, but also ties back to some of the training stuff and post training that we already talked about. Like the previous models were so focused on early chat based on, especially on c…”
Alessio Fanelli Jan 17, 2025 ▶ 4:10 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 bullish
Assertion Not checkable as stated
McAteer: o1 is the first model to achieve one-shot codebase implementation
“Using O-one was, it was the first time where I would connect it to my IDE. I would provide the full context of my code base. I'll just create a file that concatenates all my files into one just simple text file, give it to O-one, and then say, hey, based on th…”
Dan McAteer Jan 17, 2025 ▶ 6:58 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 positive
Opinion
Hylak: OpenAI o1 is OpenAI's most capable yet hardest model to use
“We're finding that like, oh, one is the most capable model. I think that opening eye has made. And it's also, I think the hardest to use.”
Ben Hillock Jan 17, 2025 ▶ 15:00 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 bullish
Prediction Open · timeframe Jan 2028
Swyx: Anthropic and OpenAI will launch automated API model routing
“And so I pick and OpenAI have both keys that model routing on APIs already. And so I think they'll launch them, especially at some point where you can sort of prioritize the three tradeoffs that are in model routing, cost, speed, intelligence.”
Shawn Wang Jan 17, 2025 ▶ 23:35 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 17, 2025 positive
Insight
Swyx: Diff critiques steer LLM style better than few-shot examples
“And actually I found that that is a better way of doing this than when doing, you know, XML bracket, good example, close bracket, bad example, close bracket. Those examples tend to meet. There's an issue of prompts leaking, example leaking. Where there's a few…”
Shawn Wang Jan 17, 2025 ▶ 19:33 OpenAI o1 isn’t a chat model (and that’s the point)
Jan 24, 2025 neutral
Opinion
OpenAI likely reaches frontier capabilities via search, then distills into mini models
“The only way you reach the frontier with the full size models of O-one and O-three is with that stuff. And then you can distill to the minis, the O-one mini, O-three mini. So in my writeup, I said like, maybe this is the formula for O-one mini, O-three mini. T…”
Shawn Wang Jan 24, 2025 ▶ 11:05 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Jan 26, 2025 bullish
Insight
Beauchamp: AI intelligence is generative LLMs combined with tree search
“I think if you want to talk about what would intelligence look like, it looks much more like tree search. Combining the generative nature of these LLMs with a really good tree search. And that's what opening I've done with O-one and O-three.”
William Beauchamp Jan 26, 2025 ▶ 1:10:38 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Feb 1, 2025
Assertion Not checkable as stated
Nguyen: OpenAI developed ChatGPT Tasks in under two months
“And actually, tasks was developed less than, like, two months. So if Canvas took, like, I don't know, four months, then tasks took, like, two months.”
Karina Nguyen Feb 1, 2025 ▶ 42:35 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 neutral
Opinion
Karina Nguyen: Verification difficulty makes alignment crucial for reasoning models
“The question of like alignment is actually more important for this like complex reasoning models to like, how do we help humans to like verify the outputs of these models is quite important.”
Karina Nguyen Feb 1, 2025 ▶ 21:41 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025
Disclosure
Karina Nguyen: OpenAI Retrained GPT-4o to Handle Canvas Edge Cases
“The only way to like fix some of the edge cases is actually through post training. So we actually, what we did was actually retrain the entire full O plus our canvas stuff.”
Karina Nguyen Feb 1, 2025 ▶ 30:17 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 positive
Insight
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Karina Nguyen Feb 1, 2025 ▶ 18:13 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 bullish
Prediction Open · timeframe Feb 2028
Nguyen: ChatGPT Will Evolve Into an Interface That Morphs Based on User Intent
“Chat CPT evolves into this Blank interface, which can morph itself in whatever you trying, like the model should try to like derive your true intent and then modify the interface based on your intent. And then if you like writing, it should become like the mos…”
Karina Nguyen Feb 1, 2025 ▶ 37:53 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 bearish
Opinion
Swix: Bearish on computer-use AI agents due to cost, speed, and accuracy
“I have been very bearish in computer use because they're slow. They're expensive. They're imprecise. Like the accuracy is horrible. Still, even with Anthropix new stuff, I'm really waiting to see what opening I might do to change my opinions.”
Shawn Wang Feb 1, 2025 ▶ 55:02 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 bullish
Prediction Not checkable as stated
Nguyen: Canvas and Tasks Will Evolve ChatGPT into Something Completely New
“There are different types of like. Features like Canvas, tasks, but all those components that go, they compose together to evolve ChatGPT into something completely new, I think, in the new year.”
Karina Nguyen Feb 1, 2025 ▶ 1:58 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 neutral
Opinion
Nguyen: OpenAI takes bigger product risks while Anthropic focuses on enterprise
“OpenAI and Anthropik is different in terms of like more like maybe like product mindset. Maybe OpenAI is much more willing to take some of the product risks and explore different bets. And I think Anthropik is much more focused and they have, I think it's fine…”
Karina Nguyen Feb 1, 2025 ▶ 1:01:06 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 1, 2025 neutral
Assertion Not checkable as stated
Nguyen: Writing and Coding Are Most Common Use Cases for Canvas
“So for Canvas, for example, one of the most common use cases is basically writing and coding”
Karina Nguyen Feb 1, 2025 ▶ 1:21 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 6, 2025 positive
Assertion Supported
Colvin: AI providers are centralizing around OpenAI's API standard
“I think the truth is that everyone is centralizing around OpenAI's SD API as the one to do. So DeepSeek support that. Grok with a K support that. Olama also does it. Well, I mean, if there is that library right now, it's more or less the OpenAI SDK.”
Samuel Colvin Feb 6, 2025 ▶ 30:32 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Feb 6, 2025 positive
Assertion Not checkable as stated
Colvin: OpenAI Said Pydantic AI Resembles Production-Ready Swarms
“OpenAI have got in touch with me, and basically, maybe I'm not supposed to say this, but basically said that Pydantic AI looks like what Swarms would become if it was production-ready.”
Samuel Colvin Feb 6, 2025 ▶ 21:26 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Feb 11, 2025 neutral
Opinion
Bret Taylor: AI infrastructure capital needs exceed all prior expectations
“The capital requirements for infrastructure are well beyond what anyone would have predicted two years ago, let alone whenever the Microsoft relationship started.”
Bret Taylor Feb 11, 2025 ▶ 1:22:20 The AI Architect: Bret Taylor
Feb 11, 2025 neutral
Insight
AGI Real-World Impact Will Be Bottlenecked by Physical Processes, Not Intelligence
“It seems likely to me that, you know, at first and something that is AGI will be good in digital domains you know, because it's software. So if you think about something like AI discovering a new say like pharmaceutical therapy, the barrier to that is probably…”
Bret Taylor Feb 11, 2025 ▶ 59:44 The AI Architect: Bret Taylor
Feb 11, 2025 neutral
Disclosure
Former OpenAI Board Chair Bret Taylor Confirms He Holds No Equity
“I still don't own any equity in OpenAI.”
Bret Taylor Feb 11, 2025 ▶ 1:18:39 The AI Architect: Bret Taylor
Feb 26, 2025 neutral
Assertion Supported
OpenAI API throws errors when users exceed a 128-tool limit
“Which OpenAI would be figured out at some point is, like, oh, there is an upper limit of, like, a 128 Tools you can have. And then the API basically gives you an error. So we ran into some of those issues as well.”
Thomas Paul Mann Feb 26, 2025 ▶ 12:09 Raycast: Your AI Automation Assistant
Feb 26, 2025 positive
Opinion
OpenAI models remain the industry best for chained function calling
“And so we find like the open AI models function calling wise for our use case. So for the best ones and like, yeah, basically verifying the others. We had a beta group and testing different models. So basically from all providers and yeah, find basically the b…”
Thomas Paul Mann Feb 26, 2025 ▶ 10:21 Raycast: Your AI Automation Assistant
Feb 26, 2025 neutral
Insight
Mann: OpenAI-Compatible APIs Still Have Subtle Behavioral Differences
“Like, even though all models say they support open AI compatible APIs, they're always in nuances which are slightly different and then put you off when you see it for the first time.”
Thomas Paul Mann Feb 26, 2025 ▶ 13:16 Raycast: Your AI Automation Assistant
Feb 28, 2025 bearish
Opinion
Klein: OpenAI's brand constraints will prevent them from offering CAPTCHA solving
“I think it's going to be really hard for a company like OpenAI to do things like support CAPTCHA solving or like have proxies. Like, I think it's hard for them structurally. Imagine this New York Times headline, OpenAI CAPTCHA solving. Like, that would be a pr…”
Paul Klein Feb 28, 2025 ▶ 40:37 Browserbase: Browser Infrastructure For Your AI Agents
Mar 11, 2025 positive
Insight
Handa: Vector stores require metadata filtering above 5,000 records
“Metadata filtering was like the main thing people were asking us for a while, and that's the one I'm super excited about. I mean, it's just so critical. Once you're like vector store size goes over, you know, more than like, you know, five, 10,000 records, you…”
Nikunj Handa Mar 11, 2025 ▶ 16:15 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Disclosure
Huet: OpenAI will continue to maintain and support Chat Completions API
“Chat completion is definitely, like, here to stay. You know, it's a bare metal API we've had for quite some time, lots of tools built around it, so we want to make sure that it's maintained and people can confidently keep on building on it.”
Romain Huet Mar 11, 2025 ▶ 2:42 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 bullish
Disclosure
OpenAI launches Agents SDK with built-in dashboard tracing
“Actually, the last thing we're launching is the agent's SDK. We launched this thing called swarm last year, where You know, it was an experimental SDK for people to do multi-agent orchestration and stuff like that. It was supposed to be like educational, exper…”
Nikunj Handa Mar 11, 2025 ▶ 1:45 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Disclosure
OpenAI Agents SDK supports any Chat Completions-compatible API provider
“We also, like, made this pretty flexible, so you can pick any API from any provider that supports the chat completions API format. So it supports responses by default, but you can, like, easily plug it into Anyone that uses the Chat Completions API.”
Nikunj Handa Mar 11, 2025 ▶ 22:23 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 neutral
Insight
OpenAI's Handa: Current computer use models are at the GPT-1 or GPT-2 stage
“The cool thing about computer use is that we're just so, so early. It's like the GPT-II of computer use or maybe GPT-I of computer use right now.”
Nikunj Handa Mar 11, 2025 ▶ 18:35 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 neutral
Assertion Supported
Swyx: OpenAI Assistants API target sunset is H1 2026
“And assistance API we've has a target sunset date of first half of 26.”
Shawn Wang Mar 11, 2025 ▶ 3:27 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Disclosure
Handa: OpenAI plans to merge preview models into core mainline models
“I think in the early days, research teams that open AI like operate with like fine-tuned models. And then once the thing gets like more stable, we sort of merge it into the mainline. So that's definitely the vision, like going out of preview as we get more com…”
Nikunj Handa Mar 11, 2025 ▶ 20:50 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Disclosure
Handa: Responses API will support all Chat Completions and Assistants features
“So, so the responses API is going to support everything that the chat, it's at launch going to support everything that chat completion supports. And then over time, it's going to support everything that assistance supports.”
Nikunj Handa Mar 11, 2025 ▶ 5:23 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 bullish
Disclosure
OpenAI launches Responses API to power future agentic products
“And to support all of these tools, we're going to have a new API. So, you know, we launched ChatCompletions, like, I think March, 23 or so. It's been a while. So, so we're looking for an update over here to support all the new things that the models can do. An…”
Nikunj Handa Mar 11, 2025 ▶ 1:17 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Assertion Supported
Handa: OpenAI Responses API stores conversation state for 30 days free
“Yeah, it's free. We store your state for 30 days. You can turn it off. But yeah, it's free.”
Nikunj Handa Mar 11, 2025 ▶ 6:44 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Nikunj Handa Mar 11, 2025 ▶ 9:26 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 bullish
Disclosure
OpenAI releases its Operator computer use tool to API developers
“And then we're also launching our computer use tool. So this is the tool behind the operator product in ChatGPT. So that's coming to developers today.”
Nikunj Handa Mar 11, 2025 ▶ 1:08 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Assertion Partly supported
Alessio Fanelli: GPT-4o Search jumps to 90% accuracy on simple QA
“On simple QA, GPT four O is 30% accuracy. Four O search is 90%.”
Alessio Fanelli Mar 11, 2025 ▶ 8:12 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 11, 2025 positive
Disclosure
OpenAI plans to connect agent traces to evals and reinforcement fine-tuning
“Like you got to tie the traces to the evals product so that you can generate good evals. Once you have good evals and graders and tasks, You can use that to do reinforcement fine tuning and you know, lots of details to be figured out over here, but that's the …”
Nikunj Handa Mar 11, 2025 ▶ 24:44 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Mar 14, 2025
Assertion Supported
Swix: Google's entire SigLIP vision team left to join OpenAI
“I think the most recent notable move, I think the entire vision team from Google Lucas Beyer and all the other authors of Siglip left Google to join OpenAI”
Shawn Wang Mar 14, 2025 ▶ 3:24 Snipd: The AI Podcast App for Learning — with CEO Kevin Ben-Smith
Mar 14, 2025 negative
Opinion
Swix: Descript still sucks despite major funding and OpenAI backing
“Descript is so much funding. They had OpenAI invested in them, and they still suck.”
Shawn Wang Mar 14, 2025 ▶ 39:59 Snipd: The AI Podcast App for Learning — with CEO Kevin Ben-Smith
Mar 23, 2025 negative
Opinion
Swyx: Frontier Models Exist Primarily to Distill Smaller, Usable Models
“Even GPT 4.5 is too expensive. Normally it's really gonna use it in, in any reasonable quantity. Like, you know, Claude 3.5 Opus, like if it does exist, still not like, you know, the thing that we actually use is Sonnet, right? So like, it's almost like a depl…”
Shawn Wang Mar 23, 2025 ▶ 3:58 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 28, 2025 bullish
Prediction Not checkable as stated
Shah: OpenAI will inevitably build a custom AI agent marketplace
“I'm an investor, but no inside information. Is because it makes too much sense for them not to like, and they, they've taken multiple passes at it, right? They did the plugins back in the day, then the custom GPTs, and then the GPT store, because, you know, be…”
Dharmesh Shah Mar 28, 2025 ▶ 56:26 The Agent Network — Dharmesh Shah, Agent.ai + CTO of HubSpot
Mar 28, 2025 positive
Disclosure
Shah: Investor in OpenAI, Perplexity, LangGraph, CrewAI, and Limitless
“Investor in OpenAI, perplexity, lane graph, crew AI, limitless, a bunch of them.”
Dharmesh Shah Mar 28, 2025 ▶ 1:19:35 The Agent Network — Dharmesh Shah, Agent.ai + CTO of HubSpot
Apr 11, 2025 neutral
Assertion Partly supported
Swyx: Microsoft and OpenAI account for 77% of CoreWeave revenue
“Which are together, 77% of the revenue of CoreWeave.”
Michael Swix (Swyx) Apr 11, 2025 ▶ 9:45 SF Compute: Commoditizing Compute
Apr 11, 2025 positive
Insight
Swyx: At $5B+ training runs, designing custom chips makes economic sense
“When you get the five billion dollar runs, when you get the fifty billion dollar runs it is actually makes sense to build your own chips, like to, for OpenAI to get into chip design”
Michael Swix (Swyx) Apr 11, 2025 ▶ 19:53 SF Compute: Commoditizing Compute
Apr 11, 2025 neutral
Assertion Not checkable as stated
Conrad: OpenRouter open-source traffic required only around 10 H100 nodes
“The entirety of Open Router that was not Anthropic or Google like, or Gemini or OpenAI or something. It was like, 10 H 100 nodes or something like that. It's just, like, not that much. It's like, not that many GPUs, actually, to service that entire demand.”
Evan Conrad Apr 11, 2025 ▶ 32:58 SF Compute: Commoditizing Compute
Apr 15, 2025 negative
Insight
Open-source AI benchmarks omit critical tasks because they are hard to grade
“And these are useful instructions, but we find that many of the really interesting instructions are actually challenging to grade. And so the open source evals often don't have them.”
Michelle Pokrass Apr 15, 2025 ▶ 18:56 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Assertion Not checkable as stated
GPT-4.1 significantly improves chain-of-thought planning over previous non-reasoning models
“We have found that 4.1 is a lot better at doing planning and thinking through its steps in COT when prompted than our previous non-reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 27:19 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Disclosure
OpenAI has no current plans to add GPT-4.1 to Realtime API
“I don't think we don't have any current plans to release 4.1 in the real time API, but you know, things, things may change.”
Michelle Pokrass Apr 15, 2025 ▶ 6:27 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Assertion Supported
OpenAI stealth-tested GPT-4.1 models on OpenRouter before official release
“Yeah yeah, we really wanted to get as much developer feedback as possible on this model to make sure it worked well in the real world, and so we tested it kind of through Open Router and it was super cool to see people latch on to the names and get the theorie…”
Michelle Pokrass Apr 15, 2025 ▶ 2:20 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Insight
Pokrass: Pair reasoning models for planning with smaller models for execution
“I do think reasoning models for planning and using kind of more targeted models to execute is definitely a good architecture.”
Michelle Pokrass Apr 15, 2025 ▶ 28:48 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Opinion
Pokrass: Developers are sleeping on preference fine-tuning for model style steering
“One thing I will say is that I think people have slept on the preference fine tuning offering or the, I think that's what we call the product. So SFT is, people know it pretty well. It's the original fine tuning we had, whereas this preference fine tuning is s…”
Michelle Pokrass Apr 15, 2025 ▶ 39:04 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Assertion Supported
OpenAI increases prompt caching discount from 50% to 75% on GPT-4.1
“We've increased our prompt caching discount from 50% to 75% on these models.”
Michelle Pokrass Apr 15, 2025 ▶ 43:27 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Disclosure
OpenAI uses its own models to categorize anonymized developer prompts
“Well, I will say we do use our own products internally where we can, and so we're not manually by hand reading every prompt. After they're like anonymized, we scrub them with any identifying data, then we use our models to take passes to categorize them.”
Michelle Pokrass Apr 15, 2025 ▶ 19:55 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Michelle Pokrass Apr 15, 2025 ▶ 23:43 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Disclosure
OpenAI is working to inject GPT-4.5's humor and nuance into future models
“We're working on incorporating kind of those improvements into the models more generally. People loved about 4.5 is like the humor, the green text, the nuance. So we've heard that feedback and I know, yeah, there's lots of folks working on that and trying to b…”
Michelle Pokrass Apr 15, 2025 ▶ 41:01 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Assertion Supported
OpenAI launches GPT-4.1 model lineup featuring 1M-token context window
“Yeah, I'll just say we released three new models today, GPT-Fort.one, GPT-Fort.one mini, and GPT-Fort.one data, and the real focus on these were just making the models that were great for developers so we improved instruction following, coding, and shipped our…”
Michelle Pokrass Apr 15, 2025 ▶ 1:27 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 bullish
Insight
Pokrass: AI model gains now driven by post-training, not larger pre-trains
“We find that actually a significant amount of the gains come from new post-training techniques. So I think in the past the narrative is that you need to pre-train these larger and larger models to get better performance, and we're finding that we're able to sq…”
Michelle Pokrass Apr 15, 2025 ▶ 7:57 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 negative
Insight
Pokrass warns against close collaboration between AI evaluators and model developers
“Honestly, I think it's best when eval authors and model developers don't collab too much because you want things you know, as objective as possible, not trying to game any evals.”
Michelle Pokrass Apr 15, 2025 ▶ 17:38 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Disclosure
GPT-4.1 powers OpenAI API while enhanced memory remains ChatGPT-exclusive
“So, 4.1 is powering the API, whereas the enhanced memory is, is ChatGPT only.”
Michelle Pokrass Apr 15, 2025 ▶ 16:11 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Insight
Pokrass: Prototype with GPT-4.1, then downscale for latency or upscale for reasoning
“I think the answer is always going to be the fastest model that accomplishes your task, right? So maybe you start prompting 4.1 as a starting point if it does your task super well, Then maybe you could drop down a 4.1 mini and save latency, or even nano. Where…”
Michelle Pokrass Apr 15, 2025 ▶ 27:58 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Assertion Supported
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 39:32 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025
Assertion Open · timeframe Apr 2028
GPT-4.1 Nano and Mini are new pre-trains; base 4.1 is mid-train
“Nano is obviously a new pre-train. We also have a new pre-train for Mini, and then, ah, the larger version is, ah, a new mid-train.”
Michelle Pokrass Apr 15, 2025 ▶ 7:46 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025
Assertion Not checkable as stated
GPT-4.1's multimodal vision improvements stem from pre-training, not post-training
“We talked about like coding instruction following long context, a lot of gains coming from post training, but in particular multimodal, like basically everything you're seeing, the gains are there from pre-training.”
Michelle Pokrass Apr 15, 2025 ▶ 35:09 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 bearish
Prediction Not checkable as stated
Pokrass predicts developers will abandon RAG vector stores for direct long-context
“So we do expect a lot of developers to start, you know, uploading their full context more directly to the model. So for smaller tasks, you maybe don't need The whole vector store.”
Michelle Pokrass Apr 15, 2025 ▶ 15:21 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Insight
Pokrass: Use XML for structuring LLM inputs and JSON for parsing outputs
“I do think XML is very helpful for structuring prompts, whereas for parsing outputs maybe the story is a bit different. Like sometimes it's really useful to get outputs in JSON, so you can plug them directly into your application. But I do think the models wor…”
Michelle Pokrass Apr 15, 2025 ▶ 24:38 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 bullish
Assertion Not checkable as stated
OpenAI researcher uses GPT-4.1 for 49 of 50 commits on massive PR
“I was actually just talking to one of the researchers on the team who worked on something over the weekend. And he said that this model, GBT, 4.1 was able to like get 49 out of 50 of his commits on this massive PR done.”
Michelle Pokrass Apr 15, 2025 ▶ 33:51 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Assertion Supported
OpenAI will cover inference costs for developers who share custom evaluation data
“You can upload an eval and opt in such that we'll pay for the inference, inference costs if we can also use the eval.”
Michelle Pokrass Apr 15, 2025 ▶ 42:04 GPT 4.1: The New OpenAI Workhorse
Apr 23, 2025 positive
Opinion
Claude is far better than OpenAI at slang and Gen Z tone
“Whenever we need to do stuff that's a little more conversational or like a little more like a little better at slang, like Claude is way better at slang. Like whenever you ask like Claude to generate something that's like, that sounds like human or like sounds…”
Sid Bendre Apr 23, 2025 ▶ 32:47 Tiny Teams: $6m ARR, 5m users with 4 employees — Sid Bendre, Oleve (Quizard AI/Unstuck AI)
Apr 24, 2025 neutral
Assertion Supported
Fanelli: Anthropic scrapes 6,000 pages per referral, compared to OpenAI's 250
“Google would be a two to one crawl to referral ratio, so for every two pages, they will read, they will send you one visitor. He said OpenAI is 250 to one, so they'll read 250 of your pages and send you one person. And Anthropic was like 6000 to one. So they'l…”
Alessio Fanelli Apr 24, 2025 ▶ 50:01 Why Every Agent needs Open Source Cloud Sandboxes
May 9, 2025 negative
Assertion Not checkable as stated
No open-source model currently matches OpenAI's general agent capabilities
“There really isn't currently an open source model that behaves in the way that these models do. We have things like R-one, which are great at kind of the single-term math and code reasoning problems. But they are not in the general purpose agents world yet.”
Will Brown May 9, 2025 ▶ 0:38 ⚡️Open Questions in Agentic RL — Will Brown (Prime Intellect)
May 16, 2025 neutral
Opinion
Josh Ma: Software engineering stays human-centric due to review and deployment
“As long as you see humans and AI writing it, like maybe there's a world where it's only AI's maintaining a code base and the assumptions change, but the moment you start to break that fourth wall and a human's coming in, doing code review, deploying the code, …”
Josh Ma May 16, 2025 ▶ 21:40 ChatGPT Codex: The Missing Manual
May 16, 2025 negative
Insight
Embiricos: Scaffolding-heavy AI agents are limited by developers' mental capacity
“A lot of, like, agents that I see are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper …”
Alexander Embiricos May 16, 2025 ▶ 30:37 ChatGPT Codex: The Missing Manual
May 16, 2025 positive
Disclosure
Embiricos: OpenAI open-sourced Codex CLI to standardize agent safety
“Part of why we made the Codex CLI open source is, like, a lot of problems, like, safety issues that you need to figure out for how to deploy these things safely, and no one should have to figure these out, like, more than once. So, that's why we went for, like…”
Alexander Embiricos May 16, 2025 ▶ 24:45 ChatGPT Codex: The Missing Manual
May 16, 2025 positive
Insight
Embiricos: The best Codex users spend 30 seconds max prompt crafting
“The way we see people who, like, love Codex the most using it is they don't, they think for, like, maybe 30 seconds max about their prompt. It's just like, oh, I have this idea, like, boom. Oh, like, there's this thing I wanna do, like, boom. Oh, like, I just …”
Alexander Embiricos May 16, 2025 ▶ 39:48 ChatGPT Codex: The Missing Manual
May 16, 2025 neutral
Disclosure
Ma: ChatGPT Codex enforces a hard one-hour runtime limit per task
“Our hard call is an hour right now, although don't hold us to that. It may change over time.”
Josh Ma May 16, 2025 ▶ 37:07 ChatGPT Codex: The Missing Manual
May 16, 2025 bullish
Disclosure
Embiricos: OpenAI Prioritizing Multimodal Inputs and Tool Integration for Codex
“Some of the items that are top of mind for me are, like, multimodal inputs. You know, we've talked, yeah, I know you find that, right? Yeah, like, another, another example would be, like, you know, just giving it a little bit more access to the world. You know…”
Alexander Embiricos May 16, 2025 ▶ 47:42 ChatGPT Codex: The Missing Manual
May 16, 2025 positive
Insight
Embiricos: Specialized domain training yields outsized returns in general models
“If you can, like, build, do something very specific for, like, a specific purpose, actually, when you bring that and you bring it into the generalized model, like, you might even get outsized returns on that. Because there's, like, transfer from all these diff…”
Alexander Embiricos May 16, 2025 ▶ 36:15 ChatGPT Codex: The Missing Manual
May 16, 2025 neutral
Disclosure
Josh Ma: ChatGPT Codex cuts off internet access during agent execution
“Once the agent starts running, right what we actually do today, and we're hoping to like evolve on this, is we'll cut off internet access because we still don't fully understand what letting loose an agent in his own environment is going to do.”
Josh Ma May 16, 2025 ▶ 43:04 ChatGPT Codex: The Missing Manual
May 16, 2025 positive
Disclosure
Josh Ma: Codex focuses on pushing single-shot autonomous software engineering
“I think what we see as, like, the role of Codex here is to really push Frontier on that sort of single-shot autonomous software engineering.”
Josh Ma May 16, 2025 ▶ 46:00 ChatGPT Codex: The Missing Manual
May 16, 2025 bullish
Prediction Not checkable as stated
Ma: Industry will build an agentic software engineer within two years
“Whether or not I was involved in the next two years, I think we were going, we are going to build an agentic software engineer.”
Josh Ma May 16, 2025 ▶ 6:59 ChatGPT Codex: The Missing Manual
Jun 19, 2025 bullish
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Noam Brown Jun 19, 2025 ▶ 38:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 bearish
Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Noam Brown Jun 19, 2025 ▶ 19:00 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 bullish
Disclosure
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Noam Brown Jun 19, 2025 ▶ 41:59 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 bullish
Prediction Not checkable as stated
Brown: Reasoning models will progress rapidly into agentic behavior
“I think that we're going to continue to see, as I said before, that we're going to see this paradigm continue to progress rapidly. And I think that that's true even today, that we saw that with like going from O-one preview to O-one to O-three, consistent prog…”
Noam Brown Jun 19, 2025 ▶ 6:07 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 neutral
Assertion Not checkable as stated
Brown: OpenAI's o3 Gets 'Not Very Far' Playing Pokémon Unharnessed
“How far does O three get without any harness? How far does it get playing Pokemon? And the answer is like, not very far, you know?”
Noam Brown Jun 19, 2025 ▶ 14:33 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 bullish
Insight
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Noam Brown Jun 19, 2025 ▶ 21:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 positive
Assertion Not checkable as stated
Brown: OpenAI saw conclusive proof of its reasoning paradigm in late 2023
“I think it was around, like, November, twenty-twenty-three, or October, twenty-twenty-three, when I think I was convinced that we had, like, very conclusive signs of life, that, like, oh, this was going to be, this is the paradigm, and it's going to be a big d…”
Noam Brown Jun 19, 2025 ▶ 25:17 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 neutral
Opinion
Noam Brown: Closing the human data efficiency gap is a top unsolved problem
“I think it's a fair statement to say that these models are less data efficient than humans. And I think that that's an unsolved research question and probably one of the most important unsolved research questions.”
Noam Brown Jun 19, 2025 ▶ 30:21 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 bearish
Prediction Not checkable as stated
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Noam Brown Jun 19, 2025 ▶ 23:20 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 positive
Assertion Not checkable as stated
Noam Brown: OpenAI succeeded early by betting on scaling over small experiments
“One of OpenAI's big success was betting on the scaling paradigm. It is just kind of odd because, you know, they were not the biggest lab, you know, it was, like, difficult for them to scale. Back then, it was much more common to do, like, a lot of small experi…”
Noam Brown Jun 19, 2025 ▶ 32:07 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jun 19, 2025 positive
Disclosure
Brown: OpenAI models undergo mid-training and post-training before release
“For open AI models, like, they go through a mid-training step, and then they go through a post-training step, and then they're released, and they're a lot more useful. Like, frankly, if you interacted with the only pre-trained model, it would be super difficul…”
Noam Brown Jun 19, 2025 ▶ 1:11:38 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jul 11, 2025 bullish
Assertion Not checkable as stated
Hsu: Whisper outperformed human listeners on accented Korean English clips
“There were four of us in the room, we all closed our eyes, and none of us had any idea, and the model got it right. So, I mean, superhuman.”
Andrew Hsu Jul 11, 2025 ▶ 24:23 Personalized AI Language Education — with Andrew Hsu, Speak
Jul 18, 2025 positive
Assertion Supported
Marimo Surpasses 300K Monthly PyPI Downloads and Jupyter's GitHub Stars
“I think last I checked, over 300,000 monthly downloads on PyPy. More GitHub stars than Jupyter Notebook for whatever that's worth. And we're used at companies like OpenAI, Hugging Face, Cloudflare, BlackRock, universities like Stanford and Berkeley.”
Akshay Agrawal Jul 18, 2025 ▶ 2:07 ⚡️The Future of Notebooks - with Akshay Agrawal of Marimo
Jul 18, 2025 neutral
Prediction Open · timeframe Dec 2026
Kamradt Predicts OpenAI Will Delay o5 Release Until 2026
“O five is not coming out this year is my guess, you know, it's going to be coming out next year.”
Greg Kamradt Jul 18, 2025 ▶ 28:26 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Jul 23, 2025
Assertion Supported
McCloy: ChatGPT personalization and custom preferences directly alter AI search retrieval and sources
“ChatGPT personalization, memories, just explicit preferences, if you set them up, do definitely affect the results you get. Now, Again, obviously, there's, they're still using traditional search, so the search index itself is not necessarily personalized, but …”
Robert McCloy Jul 23, 2025 ▶ 29:06 AI is Eating Search
Jul 23, 2025
Opinion
McCloy: Stopping LLMs broadly from consuming web content will be nearly impossible
“You can fight the battle, I think, of saying, OpenAI shouldn't consume your content without paying for it, or shouldn't consume it at all. But I think it's gonna be really tough to fight the battle of saying, like, LLMs writ large shouldn't consume my content,…”
Robert McCloy Jul 23, 2025 ▶ 13:27 AI is Eating Search
Jul 23, 2025 bullish
Opinion
McCloy: ChatGPT holds the most durable consumer presence among AI platforms
“ChatGPT, I would say, by far, in a way, is the thing that people think about as being an AI platform that has the strongest consumer presence and the most durable relationship with consumers. Right. So everything else is kind of a distant second.”
Robert McCloy Jul 23, 2025 ▶ 51:53 AI is Eating Search
Jul 24, 2025 neutral
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 24, 2025 positive
Assertion Supported
DeepMind and OpenAI used different reduction methods to solve IMO Problem 1
“Google has one method and then OpenAI AI also have another method but both kind of works. Basically you just reduced any n to three, and then you just do case by case analysis.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 8:48 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 24, 2025 positive
Assertion Not checkable as stated
Zhang confirms OpenAI's unverified IMO proofs are mathematically correct
“I read the proof that it's still correct, but it's just, like, less official, and that's why people kind of, like, kind of OpenAI received a few backlash over the weekend, and on Monday, DeepMind officially confirmed they have the gold medal and also fully ver…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:51 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 24, 2025 positive
Assertion Supported
DeepMind and OpenAI eliminated formal Lean translation for 2025 IMO solutions
“What surprised me is this time they don't use formal language, but instead they just use LM. And so last year when they tried to do the IMO, they need like a like a human to kind of translate the natural language. Problems to Lean, and then they use Lean to ki…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 6:10 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 28, 2025 bullish
Assertion Not checkable as stated
Hou: Windsurf is among Anthropic and OpenAI's largest consumers
“We've had immense success getting people onto the platform, and we've been very fortunate to have the issue of being some of Anthropic and OpenAI's largest consumers.”
Kevin Hou Jul 28, 2025 ▶ 2:49:15 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Jul 29, 2025 positive
Disclosure
Fortuna: Ambience uses OpenAI's platform for reinforcement fine-tuning
“Our foray into RFT has primarily been through the OpenAI kind of platform. So we're using those self-service, you know, APIs.”
Brendan Fortuna Jul 29, 2025 ▶ 6:09 ⚡️Using RFT to Build Clinical Superintelligence
Jul 29, 2025
Insight
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Brendan Fortuna Jul 29, 2025 ▶ 14:43 ⚡️Using RFT to Build Clinical Superintelligence
Jul 31, 2025 neutral
Opinion
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Nathan Lambert Jul 31, 2025 ▶ 7:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 positive
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Nathan Lambert Jul 31, 2025 ▶ 1:03:38 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 neutral
Prediction Open · timeframe Jul 2028
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Nathan Lambert Jul 31, 2025 ▶ 15:03 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 positive
Insight
Lambert: North star of reasoning models is dynamic token budget calibration
“I think that has to be the north star for most people working on reasoning, which is the model will just Spend the right amount of tokens on it.”
Nathan Lambert Jul 31, 2025 ▶ 20:44 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 neutral
Insight
Lambert: Inference scaling plots misleadingly suggest search is an easy control knob
“The core of that article is just, they're taking points from within training, or there's a natural variance, and then you line them up. And if you line them up, then you get this nice inference time-scaling behavior, which is, and now people, a lot of people h…”
Nathan Lambert Jul 31, 2025 ▶ 28:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 neutral
Prediction Open · timeframe Jul 2028
Lambert: Jony Ive and OpenAI hardware will run in the cloud
“I think that thing will run on the cloud. I don't think that'll run local anyways.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:54 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 bullish
Prediction Not checkable as stated
Lambert: OpenAI's open model will be best-in-class in its size category
“I expected. It'll be best in class for some size Category in some subset of tasks. That's like, OpenAI only does things like that.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:07 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 6, 2025 neutral
Assertion Partly supported
The Information: OpenAI hit $12B ARR as burn rose to $8B
“We had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something.”
Stephanie Palazzolo Aug 6, 2025 ▶ 40:26 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Aug 6, 2025 positive
Assertion Partly supported
Early GPT-5 testers report noticeable gains across coding, science, and writing
“And the story that we wrote, we kind of talked about how at least the people that we've talked to who tested it so far have been pretty impressed. They seem to think that it's been, you know, there's been improvements in a number of domains and both like scien…”
Stephanie Palazzolo Aug 6, 2025 ▶ 28:49 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Aug 15, 2025 positive
Insight
Brockman: As AI capability increases, the value of generated tokens scales dramatically
“One thing that Ilya used to say a lot that I think is, is, is very, very astute is that when the models are not very capable, right? That the value of a token that they generate is very low. When the models are extremely capable, the value of a token they gene…”
Greg Brockman Aug 15, 2025 ▶ 5:07 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 neutral
Prediction Held up
Brockman: Most AI compute will shift from training to inference
“We're going to move from a world where most of the compute is training the model as we've deployed these models more, you know, more of the compute goes to inferencing them and actually using them.”
Greg Brockman Aug 15, 2025 ▶ 14:13 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 bullish
Assertion Not checkable as stated
Brockman: Physicists say GPT-5 re-derived research insights taking months of work
“We've seen physicists starting to kick the tires on GPT-V and say that, like, hey, this thing was able to get, this model was able to re-derive an insight that took me many months worth of research to produce.”
Greg Brockman Aug 15, 2025 ▶ 21:50 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Insight
Brockman: Open-source AI creates tech stack dependency beneficial to OpenAI and the US
“Another thing at a very practical level that we've thought about with open source models is that people building on our open source model are kind of building on our tech stack, right? If you are relying on us to help improve the model that you're relying on u…”
Greg Brockman Aug 15, 2025 ▶ 50:29 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 neutral
Disclosure
Brockman: OpenAI is not yet deploying models that learn online continuously
“And now there's a next step of just having a model that as it goes, it's learning online. We're not quite doing that yet, but the future is not yet written.”
Greg Brockman Aug 15, 2025 ▶ 6:37 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 neutral
Assertion Not checkable as stated
Brockman: OpenAI is compute-limited, preventing further price cuts for now
“Right now we are extremely compute limited, and so I think that if we were to cut prices a lot, it wouldn't actually increase the amount that this model's used.”
Greg Brockman Aug 15, 2025 ▶ 45:04 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Assertion Supported
Brockman: OpenAI open-source models saw millions of downloads within days
“Now being used by, you know, there's been millions of downloads of that just over the past couple days.”
Greg Brockman Aug 15, 2025 ▶ 0:44 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Insight
Brockman: Instruction hierarchy orders trust by system, developer, and user
“With instruction hierarchy, you sort of indicate that, hey, there's this message is from the system. This message is from the developer. This message is from the user and that they should be trusted in that order. And so that way the model can know something t…”
Greg Brockman Aug 15, 2025 ▶ 30:20 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Assertion Not checkable as stated
Brockman: LLMs consistently generalize to untrained preferences
“In order to get them to be able to operate according to different preferences and values, we just need to show that to them during training, and they are able to sort of generalize to different preferences and values that we didn't actually train against, and …”
Greg Brockman Aug 15, 2025 ▶ 38:56 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Insight
Brockman: Best AI codebases use modular units with fast unit tests
“The thing I've seen be most successful is that you really build code bases Around the strengths and weaknesses of these models. And so what that means is more self-contained units have very good unit tests that run super quickly and that have good documentatio…”
Greg Brockman Aug 15, 2025 ▶ 54:26 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Assertion Not checkable as stated
Brockman: OpenAI's 80% o3 price cut yielded neutral or positive revenue
“And you can see it with O three, I think we did like an 80% price cut and actually the usage grew such that it was like, I think in the revenue, it either was neutral or positive.”
Greg Brockman Aug 15, 2025 ▶ 44:15 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Insight
Brockman: AI developer productivity gains will increase demand for engineers
“The productivity impacts of people being able to do more means we actually want more people, right? It's like we are so limited by the ability to produce software, so limited by the ability of our team to actually clean up tech debt and go and refactor things,…”
Greg Brockman Aug 15, 2025 ▶ 53:37 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Insight
Brockman: Math reasoning and proof capabilities transfer directly to competitive programming
“Learning how to solve hard math problems and write proofs turns out to actually transfer to writing program and competition problems.”
Greg Brockman Aug 15, 2025 ▶ 11:53 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 neutral
Disclosure
Brockman: OpenAI trained GPT-5 with feedback from interactive coding applications
“The second thing we did Was we really spent a long time seeing how are people using it in interactive coding applications? And just taking a ton of feedback and feeding that back into our training. And that was something we didn't try as hard in the past, righ…”
Greg Brockman Aug 15, 2025 ▶ 23:56 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025
Assertion Not checkable as stated
Brockman: OpenAI's core IMO team was only three people
“The core IMO team at OpenAI was actually three people.”
Greg Brockman Aug 15, 2025 ▶ 11:24 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025
Assertion Contradicted
Brockman: OpenAI's Dota AI used only 300 million parameters
“And by the way, Dota was like a three hundred million parameter neural net. Tiny, tiny little insect brain, right?”
Greg Brockman Aug 15, 2025 ▶ 15:20 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 bullish
Insight
Brockman: Neural networks learn biological language as naturally as human text
“Why should human language be any more natural to a neural net than biological language? And the answer is, they're not, right? That, that actually these things are literally the same hardware. Exactly, and so one of the amazing hypotheses is that it's like, we…”
Greg Brockman Aug 15, 2025 ▶ 17:21 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 bullish
Prediction Not checkable as stated
Brockman: AI models will soon get very good at CUDA kernels
“Things like Cuda kernels are a good example of a very self-contained problem that actually our models should get very good at very soon, but it's just difficult because it requires a lot of domain expertise, a lot of like real abstract thinking. But again, it'…”
Greg Brockman Aug 15, 2025 ▶ 52:14 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025
Assertion Supported
Brockman: OpenAI's robotics team pivoted to build GitHub Copilot
“And we've been through times where, for example, robotics was one in 2018, where we had a great result, but we kind of realized that actually, like, that we can move so much faster in a different domain, right? That, that actually, you know, we had this great …”
Greg Brockman Aug 15, 2025 ▶ 1:02:01 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025
Assertion Not checkable as stated
Brockman: GPT-4 handled multi-turn chat without being trained on it
“We actually did a instruction following post-train on it, so it was really just a data set that was, here's a query, here's what the model completion should be, and I remember that we were like, well, what happens if you just follow up with another query? And …”
Greg Brockman Aug 15, 2025 ▶ 1:32 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 bullish
Assertion Not checkable as stated
Brockman: GPT-5 is OpenAI's most personalizable model to date
“And GPT-V itself is extremely good at instruction following. And so it actually is the most personalizable model that we've ever produced. You can have it operate according to whatever you prefer, just by saying it, just by providing that instruction.”
Greg Brockman Aug 15, 2025 ▶ 36:13 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Opinion
Brockman: AGI will be a menagerie of models, not a single model
“The flip side, though, is that I think that the evidence has been away from having the final form factor, the AGI itself being a single model. But instead thinking about this menagerie of models that have different strengths and weaknesses.”
Greg Brockman Aug 15, 2025 ▶ 41:10 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025
Assertion Supported
Brockman: OpenAI Dota used pure RL without human demonstrations
“If you rewind to even 2017, we were working on Dota, which was all reinforcement learning, no behavioral cloning from human demonstrations or anything. It was just From a randomly initialized neural net, you'd get these amazingly complicated, very sophisticate…”
Greg Brockman Aug 15, 2025 ▶ 2:46 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 bullish
Assertion Not checkable as stated
Brockman: AI models are reaching parameter counts comparable to human synapses
“It's a hundred T synapses, which kind of corresponds to the weights of the neural net. And so there's some sort of equivalence there. And so we're starting to get to the right numbers. Let me just say that.”
Greg Brockman Aug 15, 2025 ▶ 16:11 Greg Brockman on OpenAI's Road to AGI
Aug 15, 2025 positive
Assertion Not checkable as stated
Brockman: Wet lab tests of o3 produced mid-tier journal-level work
“We have wet lab scientists who took models like O-three, ask it for some hypotheses of, here's an experimental setup, what should I do? They have five ideas, They tried these five ideas out, four of them don't work, but one of them does. And the kind of feedba…”
Greg Brockman Aug 15, 2025 ▶ 12:01 Greg Brockman on OpenAI's Road to AGI
Sep 8, 2025 bearish
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Gorkem Yurtseven Sep 8, 2025 ▶ 8:50 A Technical History of Generative Media
Sep 8, 2025 bullish
Opinion
Fal CEO: Newer Video Models Are Now Much Better Than OpenAI's Sora
“Now we have video models that are much better than Sora.”
Gorkem Yurtseven Sep 8, 2025 ▶ 31:10 A Technical History of Generative Media
Sep 11, 2025 bullish
Assertion Contradicted
Martin: OpenDeep Research is the top-ranked open-source Deep Research agent
“OpenDeep Research is a deep research agent that I've been working on for about a year, and it's now, according to Deep Research Spence, the best performing Deep Research agent at least on that particular benchmark. So it's pretty good. Listen, it's not as good…”
Lance Martin Sep 11, 2025 ▶ 8:32 Context Engineering for Agents - Lance Martin, LangChain
Sep 25, 2025 negative
Insight
Ball: Custom MCP tools fail if workflows diverge from frontier training
“If I give it this other custom-made MCP that we built internally, and our processes don't map to anything that OpenAI and Anthropic have seen or trained for, it won't be used, and you won't get good results.”
Thorsten Ball Sep 25, 2025 ▶ 49:15 Amp: The Emperor Has No Clothes
Oct 1, 2025 positive
Assertion Supported
Feldman: Sam Altman and Ilya Sutskever invested in Cerebras' early rounds
“In 2016, we met with Sam Altman and Ilya Suskovard at OpenAI and they were an idea and we were PowerPoint, right? That's amazing. And what AI was doing was identifying cats in pictures. And I think they ended up investing in us, both of them and many of their …”
Andrew Feldman Oct 1, 2025 ▶ 1:47 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Oct 7, 2025 positive
What-if
Huang: Building Visual Agent Builder in under 2 months required Codex
“For the Visual Agents Builder, we only started that Probably less than two months ago, and that, that wouldn't be possible without Codex.”
Christina Huang Oct 7, 2025 ▶ 41:12 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Supported
OpenAI adds third-party model support to its evals product
“One of the things that we launched today with evals too is ability to use, like, third-party models as well and kind of bring that into one place”
Christina Huang Oct 7, 2025 ▶ 16:38 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 bullish
Assertion Supported
OpenAI's API throughput has surpassed six billion tokens per minute
“We actually zoomed past that.”
Sherwin Wu Oct 7, 2025 ▶ 44:33 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 neutral
Insight
Wu: Cheaper inference does not cut developer spending due to surging demand
“What we realized is as we make it cheaper, you know, the demand for that goes up even more, and you end up, you know, still spending quite a bit”
Sherwin Wu Oct 7, 2025 ▶ 32:11 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Disclosure
OpenAI plans two-way code sync and code execution in Agent Builder
“Eventually, like, that's definitely what we want to do. Maybe you could start off in code. You could bring it in. We'll also probably have, like, ability to, you know, run code in, in the agent builder as well”
Christina Huang Oct 7, 2025 ▶ 13:42 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Supported
OpenAI integrates with OpenRouter for multi-provider evals
“We have a really cool setup with Open Router, where we're working with them, and then you can bring your Open Router setup. And then with that, you can actually, you know, you write your evals using our data sets tool, or use our data set tool to create a bunc…”
Sherwin Wu Oct 7, 2025 ▶ 17:02 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 bullish
Assertion Supported
OpenAI reports 4 million active developers at DevDay 2025
“Every year in Dev Day, you report the number of developers. This year is four million. I think last year was like three.”
Shawn Wang Oct 7, 2025 ▶ 1:24 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Supported
OpenAI's Nick Cooper sits on Anthropic's MCP steering committee
“We actually have a member of our team, Nick Cooper, who is sitting on kind of like that, that steering committee for MCP as well.”
Sherwin Wu Oct 7, 2025 ▶ 6:45 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Not checkable as stated
OpenAI Codex successfully one-shots entire features 30% to 40% of the time
“What a lot of the interns would do is just, like, full YOLO mode, like, trust it to, like, write the whole feature. And it, like, it doesn't work. It, like, doesn't work sometimes. But, like, I don't know, like, 30, 40% of the time it just, like, one-shots it.”
Sherwin Wu Oct 7, 2025 ▶ 40:00 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 neutral
Disclosure
Wu: Managing and serving fine-tuned snapshots is extremely difficult for OpenAI
“We have a fine-tuning API, and, like, it is extremely difficult for us to run, you know, and serve, like, all of these different snapshots... But like, man, it is like pretty difficult for us to like manage all of these different snapshots.”
Sherwin Wu Oct 7, 2025 ▶ 21:38 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 neutral
Insight
AI industry has completed only 10% of necessary agent evaluation progress
“I actually think agent evals is still a work in progress. So I think we've, like, made maybe 10% of the progress that we need here.”
Sherwin Wu Oct 7, 2025 ▶ 17:44 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Not publicly verifiable
Huang: OpenAI Customer Support Is Powered by AgentKit
“We use this internally and externally, like our customer support, help.open.au.com already powered on agent kit and then various other like internal use cases as well.”
Christina Huang Oct 7, 2025 ▶ 33:52 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 bullish
Insight
Prompt engineering has grown more entrenched despite predictions of its demise
“I feel like two years ago, people were like, oh, at some point, prop, like, prompting's gonna be dead. Like, you know, and it's like, you know... And if anything, it is, like, become more and more entrenched. And I think that, you know, there's this interestin…”
Sherwin Wu Oct 7, 2025 ▶ 20:46 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Insight
Huang: Chat Interfaces Should Be Standardized Drop-in Components Like Stripe Checkout
“Yeah, so it's very similar philosophically, right? So Stripe, you know, can build elements and check out, and not every business needs to rebuild, right, the pieces that are really common, and I think we see the same with chat. We see chat being built over and…”
Christina Huang Oct 7, 2025 ▶ 36:30 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Supported
OpenAI launches Agent Kit to build, deploy, and optimize agents
“We launched Agent Kit today. Full set of solutions to build, deploy, and optimize agents.”
Christina Huang Oct 7, 2025 ▶ 9:26 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 neutral
Disclosure
Wu: OpenAI has no current plans to become a generic IdP
“Direct answer is like no plans right now, of course but I actually think we currently have some version of this, which is our partnership with Apple because with Apple, you can actually sign in to your ChatGPT account, and some of that identity does carry with…”
Sherwin Wu Oct 7, 2025 ▶ 29:32 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 positive
Assertion Not checkable as stated
John Schulman developed the Tinker API concept across OpenAI and Anthropic
“Right when I joined OpenAI, like, this has actually been, I think, a passion project of John's. Like, he's been talking about doing something in this, like, in this shape for a while, which is, like, a truly, like, low-level research, like, fine-tuning library…”
Sherwin Wu Oct 7, 2025 ▶ 22:35 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 7, 2025 neutral
Assertion Supported
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Sherwin Wu Oct 7, 2025 ▶ 15:46 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Oct 11, 2025 neutral
Assertion Not checkable as stated
Lenz: GPT-4o is an AI orchestration system, not a raw model
“GPT-IV-O is already an AI system. It's not calling a model directly. It can do certain, you know, it can do tool calls. It can orchestrate this entire thing.”
Barak Lenz Oct 11, 2025 ▶ 30:09 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Oct 11, 2025 negative
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Barak Lenz Oct 11, 2025 ▶ 40:33 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Oct 16, 2025 bullish
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year. And my money is on, they have to do a coin. Like it's, I'm not a crypto guy at all, but like, y…”
Shawn Wang Oct 16, 2025 ▶ 49:13 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Oct 24, 2025 neutral
Insight
Webster: Enterprises do not want AI models to be maximally helpful
“OpenAI Anthropic, everyone else, they're all building models that are like maximally helpful. And in Actually, most cases in a corporate environment, you don't want that to be maximum. You don't want the model to be like helpful in every way possible.”
Ian Webster Oct 24, 2025 ▶ 8:52 Breaking AI to Fix It: Ian Webster's Journey from Discord's Clyde to Promptfoo's $18M Series A
Nov 2, 2025 bullish
Prediction Not checkable as stated
Swyx: Tuning reasoning activations could let Anthropic leapfrog OpenAI
“If Anthropic ever found The activations for reasoning and could break down the different kinds of reasoning and turn, tune them properly. I think that's the thing that takes Anthropic to leapfrog OpenAI.”
Shawn Wang Nov 2, 2025 ▶ 4:33 ⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs
Nov 3, 2025 neutral
Prediction Open · timeframe Nov 2030
Anthropic and OpenAI will never open-source their high-performance inference kernels
“The high performance inference kernels that sort of drive a lot of, you know, anthropic and open AI and stuff, their models, those aren't open source. They're not going to be open source.”
Quentin Anthony Nov 3, 2025 ▶ 40:52 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Nov 14, 2025 neutral
Insight
Smarter Consumer AI Models Do Not Increase User Growth or Retention
“Furthering the intelligence of models and chat GPT, a consumer product does not lead to more users or more retention. It only is really applicable to us the thin slice of users who care about very smart type queries, right?”
Deedy Das Nov 14, 2025 ▶ 30:50 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 bullish
Assertion Partly supported
OpenAI Plans to Scale Compute Power Capacity to 125 Gigawatts
“For OpenAI to go from like two gigawatts of compute this year to 30 with everything they've already announced, and then there's a plan for the next 125. Like, the United States uses 300.”
Shawn Wang Nov 14, 2025 ▶ 1:15:33 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 bullish
Opinion
OpenAI Has Effectively Won the Consumer AI Market
“Now that means we're at a point in consumer where, maybe this is too early to say, but OpenAI has kind of won, right? Like, How do you catch up to something where model quality is not going to be differentiated? You already have the users, you already have the…”
Deedy Das Nov 14, 2025 ▶ 31:26 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 bearish
Prediction Not checkable as stated
AI Labs Will Not Dedicate Engineering Talent to Deep Enterprise Search
“If you really want to go deep, I don't think you will ever dedicate the people to do it. And the last thing I'll say is you think about from an anthropic engineer's perspective, you joined a big AI lab to work on models, not to build Google drive connectors, r…”
Deedy Das Nov 14, 2025 ▶ 10:10 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 neutral
Assertion Contradicted
OpenAI Spent $7 Billion on Compute, With $5 Billion for R&D
“This year, OpenAI spent seven billion dollars on compute. Only two of that was for all of their inference. The remaining five was R&D. So all of ChatGPT, all eight hundred million users, all of Sora, all of like, all, all the sort of like API volume, two billi…”
Shawn Wang Nov 14, 2025 ▶ 1:17:40 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025
Assertion Supported
Glean Operates at a Several Hundred Million Dollar Revenue Scale
“Look at the revenue of Anthropic and OpenAI right now. These are billion dollar revenue scale businesses. Glean is several hundred million dollar revenue scale business.”
Deedy Das Nov 14, 2025 ▶ 9:16 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 bullish
Insight
AI Coding Intelligence Has an Uncapped Frontier That Drives Revenue
“But the interesting about Anthropic is if you look at coding, that's probably never going to be the case. Like there's always an increasing frontier of how you good you could be at a task like that. And we're nowhere close to that frontier. So it's more possib…”
Deedy Das Nov 14, 2025 ▶ 31:43 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 neutral
Insight
Enterprise LLM Churn Is Low Due to Long-Term Compute Commitments
“In terms of enterprises, often what will happen is they'll buy up large chunks of long-term compute and dedicated instances, in which case you just don't churn, right? Like this is what you use.”
Deedy Das Nov 14, 2025 ▶ 26:26 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Deedy Das Nov 14, 2025 ▶ 9:02 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 negative
Opinion
Corporate Venture Funds Perform Poorly by Prioritizing Product Usage Over Quality
“Typically, if you look at corporate venture funds in history, obviously, besides OpenAI as a notable exception, they tend to not be very good because all they prioritize is who uses my stuff the most.”
Deedy Das Nov 14, 2025 ▶ 43:33 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Dec 7, 2025 neutral
Assertion Not checkable as stated
Goyal: Commercial AI customers are reticent to give eval data to labs
“The interesting thing is that most customers, or actually I'd say a stronger statement, like all customers are quite afraid and reticent to just hand over the data that they use to do evals on to labs.”
Ankur Goyal Dec 7, 2025 ▶ 29:24 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 positive
Disclosure
Ubl: Vercel publishes evals to influence OpenAI and Anthropic models
“I'm Vercel and I publish at Eval. That I want OpenAI and Anthropic to use to make sure when they ship the next model that they're better at the stuff that I care about.”
Malte Ubl Dec 7, 2025 ▶ 28:12 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 11, 2025
Opinion
Superhuman views OpenAI and ChatGPT as direct competitors in productivity
“Oh, the answer is like ChatGPT, or like OpenAI and superhuman are competitors. Like this is what we fight against to some extent.”
Loïc Houssier Dec 11, 2025 ▶ 57:20 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Dec 11, 2025 positive
Opinion
Claude Sonnet outperforms OpenAI on agent handoffs and reducing laziness
“Sonnet was really great for, like agent head off. Like the laziness was really great. OpenAI version of it was not that good.”
Loïc Houssier Dec 11, 2025 ▶ 19:59 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Dec 26, 2025 positive
Disclosure
Fioca: OpenAI evaluates GPT-5 coding models on behavioral software engineering practices
“And so these are just best software engineering practices that turn out to be behavior characteristics, and we can measure the model's performance on those behaviors and grade it that way.”
Brian Fioca Dec 26, 2025 ▶ 4:00 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 bullish
Insight
Chen: AI software abstraction is shifting from raw models to packaged agents
“So we're actually shipping this Entirety, entire agent altogether, then you can actually build on top of that agent. That's one of the patterns that we're seeing here is rather than focusing on optimizing with every single model release, you're actually just b…”
Bill Chen Dec 26, 2025 ▶ 12:28 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 positive
Assertion Supported
Fioca: Codex Max can run continuously for 24 hours or more
“Max can run for a really long time. We can go 24 hours or more. I've actually, like, sort of had it gone for more than that”
Brian Fioca Dec 26, 2025 ▶ 1:45 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 bullish
Prediction Held up
Chen: AI agents will master GUI-based computer use by 2026
“And I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 20, 26.”
Bill Chen Dec 26, 2025 ▶ 25:40 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 bullish
Insight
Fioca: Coding agents enable self-customizing software by writing integrations at runtime
“So now if it doesn't have a tool, it can make a tool that it needs to solve a problem, right? So that's like another layer of abstraction and it's not just coding. You can write software that has an agent that can spin up a codex instance and write a custom pl…”
Brian Fioca Dec 26, 2025 ▶ 13:40 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 bullish
Disclosure
Fioca: Has not hand-written a single line of code in months
“I haven't written a single line of code by hand in months, because I know what I can trust it to do.”
Brian Fioca Dec 26, 2025 ▶ 16:06 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025
Insight
Fioca: AI models develop operational habits during training analogous to muscle memory
“This is one of the coolest things about, like, model training is literally, like, they develop habits. It's just like a person does. Like, if you're, like, working on some podcasting tool, right, you're really good at editing, and then somebody makes you use a…”
Brian Fioca Dec 26, 2025 ▶ 8:37 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 neutral
Assertion Not checkable as stated
Chen: Approximately 50% of OpenAI employees adopted Codex at launch
“Initially when Codex first launched, it was around 50% of folks that open AI started using it.”
Bill Chen Dec 26, 2025 ▶ 16:28 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 bullish
Assertion Supported
Fioca: Codex Max manages its own context window to run indefinitely
“Codex Max manages its own context window. And so it can run basically forever without you having to worry about it while it's inside of the Codex harness.”
Brian Fioca Dec 26, 2025 ▶ 14:54 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 neutral
Assertion Supported
Fioca: GPT-5.1 allows disabling preambles, unlike the reasoning-dependent Codex model
“So Five One, you can turn that off, you can prompt it not to do that, but the Codex model can't actually do that, and it relies on the reasoning summarizer to give you that update.”
Brian Fioca Dec 26, 2025 ▶ 11:26 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025
Insight
Fioca: Raw frontier models resemble new PhD hires needing explicit job prompts
“I like to think of it as like we have, I mean, people say it's a PhD in, in an API, right? But you, if you know, you hire a PhD student, they don't know how to do the job. You have to give them a job description. Okay. That's a prompt, right? So now you have y…”
Brian Fioca Dec 26, 2025 ▶ 18:30 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 positive
Insight
Chen: Naming custom tools identically to terminal tools boosts Codex performance
“We found some, like, partners of ours, like, they discovered that what you can do is that you can actually still have a lot of the tools just named in the same way as the terminal tools, as well as having the same input and output. And all of a sudden, the too…”
Bill Chen Dec 26, 2025 ▶ 8:05 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 positive
Disclosure
Fioca: OpenAI trains models to flexibly adapt across varied developer toolsets
“Initially, you know, our models are trained the way they were trained to use tools, and that kind of bakes in a habit, and so we've been getting the models better at using different types of tools.”
Brian Fioca Dec 26, 2025 ▶ 4:50 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025
Assertion Not checkable as stated
Fioca: Codex is OpenAI's frontier coding model optimized for its harness
“Codex is, just to be clear, Codex is the frontier coding model that we have that is optimized for its harness.”
Brian Fioca Dec 26, 2025 ▶ 5:20 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 positive
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Brian Fioca Dec 26, 2025 ▶ 10:39 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 26, 2025 bullish
Assertion Not checkable as stated
Yegge: OpenAI sees a 10x productivity gap between AI adopters and non-adopters
“Anecdotally, they're sharing that performance, the performance differences like 10 X By any way that you measure it. So lines of code, commits, business impact, whatever. And it's so stark and pronounced that the people who aren't adopting it are now 10 times …”
Steve Yegge Dec 26, 2025 ▶ 3:02 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Dec 26, 2025 negative
Opinion
Yegge: Google, Anthropic, and OpenAI are unbelievably chaotic internally
“All three of those companies, Google, Anthropic, and OpenAI are unbelievably chaotic internally right now.”
Steve Yegge Dec 26, 2025 ▶ 29:48 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Dec 28, 2025 positive
Disclosure
OpenAI is contributing AgentsMD to the new Agentic AI Foundation ecosystem
“In a similar way, the agentic foundation is, well, it's a foundation, but also it's like the starting point where I really look forward to other contributions, like starting with Goose, our own AgentsMD, where we're really open for like a lot of technical cont…”
Nick Cooper / Brad Dec 28, 2025 ▶ 1:07:13 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 28, 2025 positive
Assertion Supported
Anthropic and OpenAI are collaborating to build a unified AI UI standard
“And now one thing we just announced three weeks ago on the MCP blog is that we're actually working with all, all two of them together to build like a common standard.”
David Soria Parra Dec 28, 2025 ▶ 53:04 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 28, 2025 positive
Insight
Agent ecosystems need common standards but diverse, competing concrete agent implementations
“Like, agents.md is an example, which is open up any GitHub repository, it has this file, it works the same way. Like, if everyone sort of did their own thing there, That's very low value, potentially damaging in a way. But so there's commonality value, whereas…”
Nick Cooper / Brad Dec 28, 2025 ▶ 1:16:55 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 28, 2025 bullish
Assertion Not yet assessed · timeframe Dec 2025
Altman, Nadella, and Pichai publicly committed to adopting MCP around April
“And then like, you had this like inflection point around April with like Sam Altman and Satya and Sundar and all posting about like MCP and that they're going to adopt MCP at Microsoft, at Google. At OpenAI and that was really like the big inflection point.”
David Soria Parra Dec 28, 2025 ▶ 1:44 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 28, 2025 positive
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
David Soria Parra Dec 28, 2025 ▶ 1:35:53 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 28, 2025 neutral
Assertion Not checkable as stated
AAIF began when Block asked Anthropic about donating the Goose agent
“We got approached by our friends at block to discuss because they were looking into like donating goose, I think at the time. And so there was a question around doing something together. And then we approached open AI and they were very, very welcoming and lik…”
David Soria Parra Dec 28, 2025 ▶ 1:05:44 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 30, 2025 positive
Opinion
Nair: Public corporate boards may govern AI more democratically than non-profit boards
“When the blip happened, one of my reactions was like, well, you know, this nonprofit board stuff, like, actually, if it takes such somewhat, like, surprising, Maybe erratic actions, like maybe you'd rather just have, like, you know, a thing like the Microsoft …”
Ashvin Nair Dec 30, 2025 ▶ 20:27 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025
Disclosure
Nair: Internal Slack posts replaced reading external papers at OpenAI
“Unfortunately I've like, kind of gotten the habit, especially at OpenAI, of like, not reading that much external work, and just like reading people's like, Slack posts internally. That's like the main, like, way to like, you know like, learn new stuff.”
Ashvin Nair Dec 30, 2025 ▶ 40:39 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 positive
Opinion
Nair: Sutskever and Pachocki Drove OpenAI's First-Principles Research Conviction
“I think in general, OpenAI is really good about, like, having conviction in something, and just, like, really, like, from first principles, like, going after it, and I think, like, the people who are kind of most responsible for that is probably, like, Ilya Se…”
Ashvin Nair Dec 30, 2025 ▶ 22:28 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 bullish
Assertion Not checkable as stated
Nair: OpenAI already possessed a superior model during the DeepSeek release
“The feeling in OpenAI is that like, well, I think we had a better model already at the time, right?”
Ashvin Nair Dec 30, 2025 ▶ 31:21 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 neutral
Insight
Nair: OpenAI Progress Feels Smooth Internally, Not Like Sudden Leaps
“It seems like externally people are kind of very, like, oh like, research seems to come in these, like, big leaps. But I think internally at OpenAI, it feels very smooth.”
Ashvin Nair Dec 30, 2025 ▶ 25:48 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 neutral
Opinion
Nair: OpenAI model splits happen because it ships its org chart
“OpenAI has a tendency to ship the org chart, basically.”
Ashvin Nair Dec 30, 2025 ▶ 17:45 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 bullish
Assertion Not checkable as stated
Nair: OpenAI internal models surpassed forecasters' 2027 benchmark targets before o1 launch
“And their estimates were, like, oh, we'll be at, like, 10, 20% in, like, 20, 27, and I think at the time, there was, like, you know, models internally that were, like, already better than their estimates, so that, like, it's, like, off by, like, you know, two …”
Ashvin Nair Dec 30, 2025 ▶ 28:28 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025
What-if
Nair: Previously Believed IOI Gold Would Mean AI Was Solved
“If you told me that we could have gotten IOI Gold then, I would have just assumed that we could all just go on vacation, like, you know, it's all over, like, AI is solved, like, no point in working anymore.”
Ashvin Nair Dec 30, 2025 ▶ 7:09 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 neutral
Insight
Nair: RL on LLMs is peaky and fails to generalize beyond training
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like …”
Ashvin Nair Dec 30, 2025 ▶ 12:26 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 31, 2025
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 neutral
Disclosure
McGrath: OpenAI Continues to Release Non-Thinking Models for Specific APIs
“No, we're still, we still are releasing non-thinking models but that one was the one that we did that was like API-specific non-thinking so, you know, focus has shifted a little.”
Josh McGrath Dec 31, 2025 ▶ 0:47 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 positive
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Josh McGrath Dec 31, 2025 ▶ 13:58 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 positive
Prediction Not checkable as stated
McGrath: Specialized and Frontier Reasoning AI Models Will Eventually Converge
“You know, I think if you look at like deep research, the original one and GPT-Five thinking on like high reasoning today, I think you'll see that like eventually the models all sort of converge in their capabilities.”
Josh McGrath Dec 31, 2025 ▶ 6:13 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 bullish
Prediction Not checkable as stated
McGrath: AGI will be a single tool that decides its own thinking time
“Yeah, I think, like, eventually, you know, we'll have AGI, and like, you're not gonna have to worry too much about how hard to think directly. It'll just, you know, we'll have a one tool that you always go to, and it knows how long to think for, and things lik…”
Josh McGrath Dec 31, 2025 ▶ 15:23 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 positive
Assertion Partly supported
McGrath: OpenAI 10xed Effective Context Window for GPT-4.1
“I worked on long context, that was why I was on last, was for 4.1, where we, you know, I think, tenxed the effective context window for 4.1”
Josh McGrath Dec 31, 2025 ▶ 16:40 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 positive
Assertion Supported
McGrath: GPT-5 Thinking Matches or Beats Deep Research on Published Evals
“I mean, I think if you look at our published evals, they're, they look, like, basically on par if it's not better, so, like, I mean, that's personally what I do.”
Josh McGrath Dec 31, 2025 ▶ 6:46 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 positive
Insight
McGrath: Design specs let Codex complete hours of coding in 15 minutes
“If I spend, like, you know, 30, 40 minutes writing something that looks like a design doc or something, Codex can do more work than I can do in a few hours in, like, 15 minutes.”
Josh McGrath Dec 31, 2025 ▶ 3:50 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 neutral
Insight
McGrath: AI Frontier Is Bottlenecked by Shortage of Hybrid Systems-ML Talent
“I think we're still having trouble not at OpenAI, but I think as a whole, producing lots of people that do lot, want to do lots of both systems work and ML work. And I think if you're trying to push the frontier, you don't know which Place is currently bottlen…”
Josh McGrath Dec 31, 2025 ▶ 21:36 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025
Insight
McGrath: RL runs have far more infrastructure failure points than pre-training
“The issue with RL is, like, you're doing tasks, and each task could have, like, a different grading setup, and each one of those different grading setups, that's, like, more infrastructure, and so, You know, when I'm staying up late trying to figure out what's…”
Josh McGrath Dec 31, 2025 ▶ 2:12 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 negative
Assertion Contradicted
All major US AI labs stopped publishing research after OpenAI closed
“Whereas in the United States, since OpenAI closed their doors and stopped publishing, so did all the other labs.”
Andy Konwinski Dec 31, 2025 ▶ 19:25 [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
Jan 9, 2026 neutral
Assertion Supported
Hill-Smith: OpenAI Was Untouchable for Well Over a Year
“If we go back even a little bit before then, we're in the era where, when you look at this chart, like, OpenAI was untouchable for well over a year.”
Micah Hill-Smith Jan 9, 2026 ▶ 23:52 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 27, 2026 bullish
Prediction Not checkable as stated
Powell: LaTeX code will recede as researchers interact directly with papers
“To your point about voice though, I do think maybe over time the code kind of might recede into the background more as you're just really interact, you're interacting with the paper.”
Victor Powell Jan 27, 2026 ▶ 12:43 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026 bullish
Disclosure
Weil: OpenAI wants 100 external scientists to win Nobels with its AI
“Our goal is not To win a Nobel Prize ourselves, it is for a hundred scientists to win Nobel Prizes using our technology.”
Kevin Weil Jan 27, 2026 ▶ 32:10 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026 bullish
Prediction Not checkable as stated
Powell: AI tools will progress to running multi-day autonomous scientific analyses
“I do think that's sort of the progression where it's like, Doing, doing maybe work for a few seconds versus maybe we're already at a point where it's doing work for a few minutes, eventually doing work for hours, days, coming back with very complicated analysi…”
Victor Powell Jan 27, 2026 ▶ 19:56 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026 positive
Disclosure
Weil: OpenAI will mostly partner in science rather than build verticals
“By and large, we're going to partner because the surface area of science is massive. And we want to accelerate all of science.”
Kevin Weil Jan 27, 2026 ▶ 32:38 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026 bullish
Prediction Not checkable as stated
Weil: 2026 for AI in science will mirror 2025 software engineering
“I think, 2026 for ai and science is going to look a lot like what 2025 looked like for soft ai and software engineering yeah where if you go back to the beginning of 2025 if you were using ai heavily to write your code you were sort of an early adopter and lik…”
Kevin Weil Jan 27, 2026 ▶ 23:50 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026 neutral
Prediction Not checkable as stated
Weil: Editor UIs will converge as AI interaction replaces direct document editing
“The UI probably changes for all of these things, right? You don't need your document front and center because you're actually not looking at your document as much. You're, that's sort of your backup and your interaction with your AI is primary. And as that hap…”
Kevin Weil Jan 27, 2026 ▶ 18:54 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026
Disclosure
OpenAI launches Prism, a free AI-native LaTeX editor
“So we're launching Prism, which is a free AI native LaTeX editor.”
Kevin Weil Jan 27, 2026 ▶ 0:34 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 27, 2026 positive
Assertion Supported
Weil: OpenAI's Prism allows unlimited collaborators for free
“I think most other tools in the space have hard limits and charge you money and other things. In Prism, it's as many collaborators as you want for free.”
Kevin Weil Jan 27, 2026 ▶ 16:04 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Jan 28, 2026 neutral
Disclosure
Andrew White Began Red-Teaming GPT-4 in August 2022
“And so I was a red teamer for GPT-IV and I was using it like nine months with me for release with August. So GPT-IV came out in March and I was using it in August.”
Andrew White Jan 28, 2026 ▶ 9:13 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Jan 28, 2026 neutral
Assertion Supported
White: OpenAI reached out to red team new models after reading his chemistry paper
“And then opening eye, some people there Lama was there. She saw this paper, and they reached out, like, hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemis…”
Andrew White Jan 28, 2026 ▶ 8:58 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Feb 10, 2026
Assertion Supported
Pre-December 2024 OpenAI models did not exhibit seahorse emoji self-correction loops
“And so I like ran the OpenAI API across like models released from 23 to 25, and you would see like all the models until twenty-twenty-four December had very Terce and short responses to the question. Is there a seahorse emoji? They would either say that there …”
Pratyush Maini Feb 10, 2026 ▶ 9:04 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Feb 10, 2026 neutral
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Pratyush Maini Feb 10, 2026 ▶ 15:26 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Feb 19, 2026 bullish
Prediction Didn’t hold up
Swyx: OpenAI will always release both general and Codex model variants
“I'm pretty, like, have pretty high confidence that basically OpenAI will always release a GPT-V and a GPT-V codex.”
Shawn Wang Feb 19, 2026 ▶ 36:15 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Feb 23, 2026 neutral
Disclosure
Watkins: OpenAI will probably not release proprietary AI research coding benchmarks
“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make evals for that, like, we're probably not gonna release them. And it's harder for people in the field to make evals that kind of measure, like, is this a realist…”
Olivia Watkins Feb 23, 2026 ▶ 20:04 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 neutral
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 bearish
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 negative
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026
Assertion Supported
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Olivia Watkins Feb 23, 2026 ▶ 2:56 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 negative
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia Watkins Feb 23, 2026 ▶ 11:54 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Mar 14, 2026 neutral
Disclosure
Pydantic built a VIP issue scorer after closing an OpenAI founder's ticket
“Basically this started off because one of the OpenAI co-founders created an issue on Pydantic. And we just closed it and said it was wrong. And so we have this that, like, injects itself and tries to summarize someone and it gives them a, like, brutal score of…”
Samuel Colvin Mar 14, 2026 ▶ 13:45 ⚡️Monty: the ultrafast Python interpreter by Agents for Agents — Samuel Colvin, Pydantic
Mar 20, 2026 positive
Disclosure
Dreamer uses continuous evaluations to dynamically route tasks across AI models
“Dreamer actually uses all of the state-of-the-art models. As a user, you don't have to think about, should I be using, you know, Opus four six, or should I be using the five four model from OpenAI? We are continually doing evals and so forth to make sure that …”
David Singleton Mar 20, 2026 ▶ 30:31 Dreamer: the Agent OS for Everyone — David Singleton
Mar 30, 2026 positive
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 26:34 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Apr 2, 2026 bearish
Opinion
Manning: OpenAI's Sora cannot produce compelling gameplay or persistent mechanics
“Don't think you can take Sora and produce compelling gameplay, right? If you want to have a world that you can wander around in a bit, you're good, but what are your abilities to have gameplay mechanics implemented the way you'd like them to be, and to have th…”
Chris Manning Apr 2, 2026 ▶ 52:02 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Apr 7, 2026 positive
Assertion Not checkable as stated
Lopopolo: OpenAI Engineers Face No Internal Rate Limits for Development
“It certainly helps that we have no rate limits internally and I can go, like you said, full send at this thing.”
Ryan Lopopolo Apr 7, 2026 ▶ 3:26 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Standardizing codebase structure and skills maximizes AI agent effectiveness
“I do think that there is leverage to be had in making the code and the processes as much the same as possible. If you think that code is context, code is prompts, it's better from the agent behavior perspective to be able to look in a package in directory XYZ …”
Ryan Lopopolo Apr 7, 2026 ▶ 42:35 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: BEAM process supervision gives agent orchestration free concurrency
“The process supervision and the gen servers are super amenable to the type of process orchestration that we're doing here, right? You are essentially spinning up little daemons for every task that is in execution and driving it to completion, which means the m…”
Ryan Lopopolo Apr 7, 2026 ▶ 34:37 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Disclosure
Lopopolo: OpenAI Frontier team operates with post-merge or zero human code review
“You know, we, we've moved beyond even the humans reviewing the code as well. Most of the human review is post merge at this point, but it's not even reviewed.”
Ryan Lopopolo Apr 7, 2026 ▶ 9:21 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026
Insight
Lopopolo: When coding agents fail, decompose tasks into smaller reusable building blocks
“Whenever the model just cannot, you always pop open the task, double click into it and build smaller building blocks that then you can reassemble into the broader objective.”
Ryan Lopopolo Apr 7, 2026 ▶ 4:58 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: AI-amplified small teams require extreme package decomposition and strict boundaries
“The structure of the repository is like, 500 NPM packages. It's like architecture to the access for what you would consider, I think, normal for a seven person team. But if every person is actually, like, 10 to 50. Then the, like, numbers on, like, being super…”
Ryan Lopopolo Apr 7, 2026 ▶ 39:56 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 bullish
Assertion Not checkable as stated
Lopopolo: Codex App Hits 2M WAUs, Growing 25% Week-Over-Week
“We just passed two million weekly active users growing at a phenomenally fast rate, 25% week over week.”
Ryan Lopopolo Apr 7, 2026 ▶ 1:15:41 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 bearish
Opinion
Lopopolo: Bearish on MCP due to forced token injection and compaction issues
“MCPs I'm pretty bearish on because the harness forcibly injects all those tokens in the context and I don't really get a say over it. They mess with auto compaction. The agent can forget how to use the tool. There's probably only like, what, three calls in Pla…”
Ryan Lopopolo Apr 7, 2026 ▶ 38:37 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 bullish
Opinion
Lopopolo: Coding models have largely solved all tasks except hard and new
“And I think things that are hard and new is still something that the models need humans. Yeah. Drive. Yeah. But I think those other quadrants are largely solved, given the right scaffold and the right thing that's going to drive the agent to completion.”
Ryan Lopopolo Apr 7, 2026 ▶ 33:46 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Disclosure
Lopopolo: OpenAI feeds 12-month business vision and customer context to agents
“One thing that's in core beliefs.md is like, Who's on the team, what product we're building, who our end customers are, who our pilot customers are, what the full vision of what we want to achieve over the next 12 months is. Like these are all bits of context …”
Ryan Lopopolo Apr 7, 2026 ▶ 1:07:38 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Assertion Not checkable as stated
Lopopolo: Better AI models propose their own code abstractions
“As the models have gotten better, they have gotten better at proposing these abstractions to unblock themselves, which again, lets me move higher and higher up the stack to look deeper into the future on what ultimately blocked the team from shipping.”
Ryan Lopopolo Apr 7, 2026 ▶ 21:45 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026
Disclosure
Lopopolo: OpenAI Frontier codebase operates on approximately six core skills
“So like in our code base, we have, I think six skills. That's it. And if some part of the software development loop is not being covered, Our first attempt is to encode it in one of the existing setup skills, which means that we can change the agent behavior m…”
Ryan Lopopolo Apr 7, 2026 ▶ 43:13 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Fast Spark models excel at prototyping, docs, and lint healing
“It's very great for spiking out prototypes, exploring ideas quickly, doing those documentation updates. It, Is fantastic for us in taking that feedback and transforming it into a lint where we already have good infrastructure for ES lints in the code base. The…”
Ryan Lopopolo Apr 7, 2026 ▶ 59:09 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 neutral
Assertion Not checkable as stated
Lopopolo: PR review agents initially caused non-convergence by bullying author agents
“Initially the codex driving the code author was willing to be bullied by the PR reviewer, which meant you could kind of end up in a situation where things were not converging.”
Ryan Lopopolo Apr 7, 2026 ▶ 15:30 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Harness engineering codifies implicit engineering standards into agent context
“The whole meta of the thing is to basically tease out of the heads of all the engineers on my team, what they think good looks like, what they would do by default. Or what they would coach a new hire on the team to do, to get things to merge. And that's why we…”
Ryan Lopopolo Apr 7, 2026 ▶ 26:21 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Converting UI Images to ASCII Art Improves AI Agent Layout Perception
“If we want to actually, like, make it see the layout, it's almost easier to rasterize that image to ASCII arc and feed it in to the agent.”
Ryan Lopopolo Apr 7, 2026 ▶ 47:57 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 bullish
Opinion
Lopopolo: Coding models and harnesses are now isomorphic to human engineering capability
“The models are there enough. The harnesses are there enough where they're isomorphic to me and capability and the ability to do the job.”
Ryan Lopopolo Apr 7, 2026 ▶ 4:02 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 bullish
Disclosure
Lopopolo: Codex authors Grafana dashboards and handles on-call incident paging
“Like the dashboard thing you mentioned, we have Codex authoring the JSON for the Grafana dashboards and publishing them, and also responding to the pages, which means when it gets the page, it knows exactly which dashboards are defined and what alerts. What al…”
Ryan Lopopolo Apr 7, 2026 ▶ 19:00 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026
Assertion Not checkable as stated
Lopopolo: OpenAI used iterative Codex loops to generate Symphony specs
“Like we have taken all the scaffolding that has existed in our proprietary repo, spun up a new one. Ask codex with our repo as a reference. Write the spec. We tell it, spin up a tmux, spawn a disconnected codex to implement the spec. Wait for it to be done. Sp…”
Ryan Lopopolo Apr 7, 2026 ▶ 32:27 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 bullish
Insight
Lopopolo: AI models can in-house 2,000-line dependencies in an afternoon
“The level of complexity of the dependencies that we can internalize is I would say low medium right now, right? Just based on model capability. What is medium? I would say like a couple thousand line dependency is a thing that we could in house no problem in a…”
Ryan Lopopolo Apr 7, 2026 ▶ 28:25 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 neutral
Disclosure
Lopopolo: Symphony discards failed PRs entirely to regenerate from scratch
“In Symfony, there's this like rework state where once the PR is proposed and it's escalated to the human for review, it should be a cheap review, right? It is either mergeable or it is not. And if it's not, you move it to rework. The Elixir service will comple…”
Ryan Lopopolo Apr 7, 2026 ▶ 36:53 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Disclosure
Lopopolo: OpenAI Frontier Is OpenAI's Platform for Enterprise Agent Deployment
“I work on frontier product exploration, new product development in the space of open AI frontier, which is our enterprise platform for deploying agents safely at scale with good governance in any business.”
Ryan Lopopolo Apr 7, 2026 ▶ 2:21 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 negative
Disclosure
Lopopolo: GPT-5.3 Spark burned three compactions before coding complex tasks
“I was adapting it to the same sorts of tasks I would use X high reasoning for, and it would blow through three compactions before writing a line of code.”
Ryan Lopopolo Apr 7, 2026 ▶ 58:41 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Assertion Supported
Lopopolo: GPT OSS Safeguard supports custom enterprise safety specs
“The GPT OSS Safeguard model, for example. One thing that's really cool about it is it ships the ability to interface with a safety spec. Safety specs are things that are bespoke to enterprises. We owe it to these folks to figure out ways for them to instrument…”
Ryan Lopopolo Apr 7, 2026 ▶ 1:03:38 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Disclosure
Lopopolo: OpenAI uses an automated landing skill to delegate PR merges to Codex
“We invoke a dollar land skill and that coaches codecs to push the PR, wait for human and agent reviewers, wait for CI to be green, fix the flakes if there are any Merge upstream if the PR comes into conflict, wait for everything to pass, put it in the merge qu…”
Ryan Lopopolo Apr 7, 2026 ▶ 24:55 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Assertion Supported
Lopopolo: Codex can run background builds while concurrently reviewing code
“It basically means that Codex is able to spawn commands in the background and then go continue to work while it waits for them to finish. So it can spawn an expensive build and then continue reviewing the code, for example.”
Ryan Lopopolo Apr 7, 2026 ▶ 7:19 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 negative
Insight
Lopopolo: Optimizing AI debugging workflows for human legibility is wrong
“Optimizing for human legibility of that debugging process was wrong. It kept him in the loop unnecessarily, when instead he could have just like codex cooked for five minutes and gotten the same.”
Ryan Lopopolo Apr 7, 2026 ▶ 31:18 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Assertion Not checkable as stated
Lopopolo: Zero-code harness was 10x slower initially before outperforming any single engineer
“Honestly, the first month and a half was 10 times slower than I would be. But because we paid that cost, we ended up getting to something much more productive than any one engineer could be, because we built the tools, the assembly station for the agent to do …”
Ryan Lopopolo Apr 7, 2026 ▶ 5:17 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Disclosure
Lopopolo: OpenAI Agents SDK provides an out-of-the-box agent harness
“Agents SDK is a core part of this to enable both Startup builders as well as enterprise builders to have a works by default harness that is able to use all the best features of our models from the shell tool down to the codex harness with file attachments and …”
Ryan Lopopolo Apr 7, 2026 ▶ 1:03:05 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Disclosure
Lopopolo: OpenAI uses incident pages to update repository reliability rules via Codex
“When we get a page because we're missing a timeout, for example, I can just add codecs in Slack on that page and say, I'm going to fix this by adding a timeout. Please update our reliability documentation to require that all network calls have timeouts. So I h…”
Ryan Lopopolo Apr 7, 2026 ▶ 14:04 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Codex reviews internalized dependencies with less friction than upstream patching
“When we deploy Codex security on the repo, it is able to deeply review and change The internalized dependencies in a much lower friction way than it would be to like push patches upstream, wait for them to be released, pull them down, make sure that's compatib…”
Ryan Lopopolo Apr 7, 2026 ▶ 29:07 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Autonomous Coding Removes Human Language Familiarity Constraints
“No humans in the loop here. So like my, Own personal ability to write or not write Elixir doesn't really have to bias us away from using the right tool for the job, which is just wild.”
Ryan Lopopolo Apr 7, 2026 ▶ 47:38 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026
Disclosure
Lopopolo: OpenAI inverts harnesses by having Codex spawn dev environments
“One neat thing here is we have tried to invert things as much as possible, which is instead of setting up an environment to spawn the coding agent into, instead we spawn the coding agent, like that's the entry point, just codex, and then we give codex via skil…”
Ryan Lopopolo Apr 7, 2026 ▶ 11:32 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 positive
Insight
Lopopolo: Coding agents should summarize proof instead of requiring full oversight
“I would expect you to do what you think you need to do to convince me that the code is good and mergeable and compress that full trajectory in a way that is legible to me, the reviewer.”
Ryan Lopopolo Apr 7, 2026 ▶ 56:43 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 neutral
Insight
Lopopolo: Managing coding agents resembles tech leading a 500-person organization
“The mindset is very much that I'm removed from the process, right? I can't really have Deep code level opinions about things. It's as if I'm group tech leading a 500 person organization. Like, yeah, like it's not appropriate for me to be in the weeds on every …”
Ryan Lopopolo Apr 7, 2026 ▶ 20:32 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 7, 2026 neutral
Disclosure
Lopopolo: OpenAI Frontier requires human-approved smoke tests before distribution
“So because we are building a native application here, we're not doing continuous deploy. Right. So there's still a human in the loop for cutting the release branch. We require a blessed human approved smoke test of the app before we promote it to distribution,…”
Ryan Lopopolo Apr 7, 2026 ▶ 17:17 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Apr 15, 2026
Disclosure
Notion Partnered With Anthropic and OpenAI to Build 30% Pass Rate Evals
“And then what we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate. And that's actually been a effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because we actually hit a…”
Sarah Sachs Apr 15, 2026 ▶ 26:28 Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
Apr 15, 2026 negative
Assertion Not checkable as stated
Claude Sonnet Crashed on Duplicate Tool Names While OpenAI Handled the Error
“Sonic couldn't handle two tools with the same name in OpenAI, GPT, 5.2. It was like, ah, I can figure this out. So that was an interesting one that we learned by accident through a SEV.”
Sarah Sachs Apr 15, 2026 ▶ 55:35 Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
Apr 22, 2026 bullish
What-if
Parakhin: Liquid AI could beat frontier models with equal compute
“I think if they if they had similar level of compute, they would be very competitive and maybe even beat the largest models, at least from what I've seen.”
Mikhail Parakhin Apr 22, 2026 ▶ 1:06:01 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
Apr 22, 2026
Assertion Supported
Parakhin: Bing Sydney first launched in India using Megatron, not OpenAI
“The funny thing, I mean, the most interesting anecdote is that Sydney was first shipped in India for and it was not noticed for a long time. And first implementation of Sydney didn't even have open AI model under it. It was during Megatron. Microsoft and the N…”
Mikhail Parakhin Apr 22, 2026 ▶ 1:10:53 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
May 2, 2026 neutral
Assertion Not checkable as stated
Open-source AI demand spiked on hype before reverting to frontier labs
“Like all the open source models, I think what happened was they got like very hyped and people were very interested in using them. But I think like over time, like there was a spike in usage for these models. And then it goes back to open AI, Anthropic and Goo…”
Yasser Elsaid May 2, 2026 ▶ 16:45 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
May 2, 2026
Assertion Contradicted
No commercial products augmented GPT models with custom data pre-ChatGPT
“Like I saw some people doing demos, but like in like a CLI or something like that, but there was no product doing like this model, but with additional data on top of it.”
Yasser Elsaid May 2, 2026 ▶ 2:59 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
May 2, 2026
Disclosure
Chatbase customer model usage is split 50% OpenAI, 50% Anthropic and Google
“Maybe 50% is still on Okunai. Yeah. Yeah. And then 50 on everything else. Yeah. But everything else is like mainly Anthropic and Google.”
Yasser Elsaid May 2, 2026 ▶ 13:30 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
May 5, 2026 positive
Assertion Not checkable as stated
ChatGPT Pro Generated Lupsasca's Exact Top Three Follow-Up Physics Questions
“You can take this page of this paper and you can feed it to ChatGPT Pro, say, like the best model we have out right now, and you can ask it, what should I do next? Give me the top three follow-up questions to ask based on this paper. I've done this experiment …”
Alex Lupsasca May 5, 2026 ▶ 1:18:37 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
May 5, 2026 bullish
Assertion Not checkable as stated
GPT-5 Reproduced Lupsasca's Best Physics Paper in 30 Minutes
“Then when GPT-V came out. It was able to reproduce one of my best papers that took me a very long time to come up with, in like, 30 minutes.”
Alex Lupsasca May 5, 2026 ▶ 2:32 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
May 5, 2026 bullish
Assertion Not checkable as stated
Internal OpenAI Model Proved Gluon Amplitude Formula in 12 Hours
“We had this Internal model that could think for a very long time and was extra strong in physics. So we gave it the whole problem from scratch without actually giving it this. We just formulated the problem in a very sharp way and asked the model to solve, to …”
Alex Lupsasca May 5, 2026 ▶ 35:56 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
May 5, 2026 negative
Assertion Not checkable as stated
Terry Tao Says AI Math Proofs Merely Cite Obscure References
“I talked to Terry Tao a couple of weeks ago at UCLA. We had an OpenAI event with IPAM, which is this Institute of Mathematics there. And I talked to Terry Tao and he said that in his view, all of the proofs that he's seen AI come up with in math, even the ones…”
Alex Lupsasca May 5, 2026 ▶ 1:09:35 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
May 5, 2026 bullish
Assertion Not checkable as stated
ChatGPT Solved Open Physics Problem Before Collaborator's Flight Landed
“We decided to start working on it using AI a little bit before Andy was scheduled to come, like the week before. And in fact, using ChatGPT, we solved the problem before he even got off the plane.”
Alex Lupsasca May 5, 2026 ▶ 20:53 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Jun 3, 2026 positive
Opinion
Nadella confirms intelligence scaling logarithmically with compute still largely holds true
“This, you know, this crude way of saying it is intelligence is log of compute kind of works.”
Satya Nadella Jun 3, 2026 ▶ 5:45 Satya Nadella on AI: @NoPriorsPodcast x Latent Space Crossover Special at Microsoft Build 2026
Jun 3, 2026 negative
Assertion Supported
Hong: All OpenAI formal math researchers have left the company
“No, no, they all left.”
Carina Hong Jun 3, 2026 ▶ 1:05:45 Scaling Past Informal AI - Carina Hong, Axiom Math
Jun 4, 2026 negative
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Jun 25, 2026 negative
Assertion Not checkable as stated
OpenAI's Mark Chen: AI is in an evals crisis with saturated benchmarks
“So I think beyond that, the other scary thing in the field is the number of canonical gold standard benchmarks is low. And we really are kind of in an evals crisis, right? Where all the really great evals that we all know, like growing up, like taking the SAT …”
Mark Chen Jun 25, 2026 ▶ 20:23 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 positive
Opinion
Chen: Meta poaching has calmed down and OpenAI came out on top
“I think that met us calmed down a little bit. I think we came out on top”
Mark Chen Jun 25, 2026 ▶ 0:43 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 bullish
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Mark Chen Jun 25, 2026 ▶ 9:58 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 bullish
Opinion
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
Mark Chen Jun 25, 2026 ▶ 38:01 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 neutral
Insight
OpenAI's Chen: Reinforcement learning struggles in subjective, hard-to-grade fields
“RLs traditionally had headwinds when it's come to fields that, you know, it's more kind of, Subjective than objective. So if you kind of think of, you know, one kind of, you know example of this is creative writing, where, you know, you could take two pieces o…”
Mark Chen Jun 25, 2026 ▶ 5:54 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 positive
Disclosure
Mark Chen brought soup to researchers to counter Zuckerberg's talent poaching
“Oh, you know, it's absolutely a true story. And I have brought soup to our own researchers.”
Mark Chen Jun 25, 2026 ▶ 0:39 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026
Disclosure
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”
Mark Chen Jun 25, 2026 ▶ 14:03 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 bullish
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Mark Chen Jun 25, 2026 ▶ 7:27 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 positive
Assertion Not checkable as stated
Jakob Pachocki and Ilya Sutskever overcame internal inertia to build OpenAI o1
“Even at a company like OpenAI, you would have people ask naturally, why do something when you have a machine that works? And fundamentally, you know, it's to the credit of, you know, Jakob, Ilya, many of the people who really had conviction and vision in this …”
Mark Chen Jun 25, 2026 ▶ 10:46 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 bullish
Disclosure
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Mark Chen Jun 25, 2026 ▶ 34:15 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 positive
Disclosure
Chen: OpenAI favors unifying modalities in as few architectures as possible
“For a research lab, I think there are a lot of advantages for it to being under one. So you just have to maintain one infrastructure stack, for instance. I think the cost to, like, maintaining and scaling many infrastructure stacks at once I think that's somet…”
Mark Chen Jun 25, 2026 ▶ 31:58 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 positive
Insight
Chen: Eval teams and model optimization teams must remain separate
“Yeah, I think there's a kind of interesting philosophy of separate the teams that are creating the evals from the teams that are optimizing the models themselves, because that way you don't, like, co-incentivize them, right? Like, the way the evals theme can w…”
Mark Chen Jun 25, 2026 ▶ 22:34 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026 positive
Insight
Mark Chen: A PhD is not necessary to excel in AI research
“There are a lot of researchers who just started out without formal training in machine learning or AI research. We've very much believed in training people up to do this. I think the real hard thing is the ability to creatively solve problems and think outside…”
Mark Chen Jun 25, 2026 ▶ 2:14 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jul 22, 2026 neutral
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Jul 28, 2026
Disclosure
Nathan: I would never write a performance review solely via AI
“I would never write something via, like, solely via AI and, like, present it as, like, a review for someone. What I was talking about is more, like, gathering context.”
Akshay Nathan Jul 28, 2026 ▶ 40:26 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 bullish
Opinion
Nathan: Untapped enterprise agent market is 10x to 100x larger
“So now like we're seeing with agents, like there is Probably a contingent of like early adopters still who, you know, truly get it. We're like, you know, you can do anything. You just have to make sure the right context is there. It's connected to the right to…”
Akshay Nathan Jul 28, 2026 ▶ 6:23 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 positive
Assertion Not checkable as stated
Nathan: OpenAI saw an inflection in Codex adoption among non-developers
“When we release codex or even internally add codex. Like it was really surprising to us. I think we recently put out some stats on this, but there was this like real inflection of like adoption among non-developers at OpenAI.”
Akshay Nathan Jul 28, 2026 ▶ 7:24 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 neutral
Insight
Nathan: Open-ended AI chat interfaces confuse enterprise users without guided use cases
“Using these models and these products, you have this box and you can say anything to it, which is the magic, but it's on the flip side. It also means that, like, you don't know what to do with it, and in enterprise, I think a big part of that is, like, actuall…”
Akshay Nathan Jul 28, 2026 ▶ 5:05 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026
Disclosure
Nathan: OpenAI limits sub-agent UI visibility to prevent user overwhelm
“There's another, you know, iteration of this where like you can see exactly what they're doing and things like that, which I think is like, you know, could verge on like overwhelming with information. And so this is like the deliberate trade-off that we've mad…”
Akshay Nathan Jul 28, 2026 ▶ 52:50 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 bullish
Disclosure
Nathan: OpenAI sequencing agents from developers to knowledge workers to everyone
“The vision is like bring useful agents to everyone. We started with like developers Historically are like early adopters that are willing to put up with more friction, set things up, et cetera. Like that's where, you know, Codex started. I think the next oppor…”
Akshay Nathan Jul 28, 2026 ▶ 35:57 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026
Assertion Not checkable as stated
Nathan: OpenAI built its model slider fully within a site artifact
“Even the model slider that you guys were referencing earlier, like that was developed almost fully in a site. Like, you know, the collaboration between design and engineering and product on that was like on a site where we play with, you know, the affordance a…”
Akshay Nathan Jul 28, 2026 ▶ 26:25 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 neutral
Insight
Nathan: Traditional software productivity proxies are falling apart in AI era
“I think with AI now, those proxies starting to fall apart, like, you know, and the number of tokens you use or the number of pull requests you make are like no longer like maybe as hyper correlated with that. Is your team able to hit the goal or are they on tr…”
Akshay Nathan Jul 28, 2026 ▶ 1:08:10 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 positive
Disclosure
Nathan: OpenAI plans to keep developing Codex specifically for developers
“Like, I think we fully intend to like, you know, treat developer, like developers have been, you know, a core market for us for so long. And like, there's so much more that we can do to make Codex great specifically for software development, and we'll continue…”
Akshay Nathan Jul 28, 2026 ▶ 45:14 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 bullish
Opinion
Nathan: AI models are experiencing another step-function capability jump
“The models were getting infinitely more capable. That's happening again. I think it's like another step function jump now.”
Akshay Nathan Jul 28, 2026 ▶ 18:42 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 positive
Assertion Not checkable as stated
Nathan: Artifact quality improved dramatically over GPT-5.4 and GPT-5.5
“One of the big, like, pushes that we made for this launch was, like, artifacts, right? Like, both on the model side, like, I think if you compare this with 5.5 and 5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of the…”
Akshay Nathan Jul 28, 2026 ▶ 22:30 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 positive
Assertion Supported
Nathan: ChatGPT Work includes persistent computer environments across sessions
“In ChatGPT work in web and mobile, like, you get access to those, like, persistent computer environment where, you know, you can store files, and those files stay around between sessions.”
Akshay Nathan Jul 28, 2026 ▶ 46:59 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 neutral
Insight
Akshay Nathan: Software creation bottleneck has shifted to ideas and taste
“I think the bottleneck some becomes like sort of like ideas and taste, I guess. I think because anyone can build now, I think it really is the era of like bottoms up ambition. And because there's so much to be built, like you're always going to be bottlenecked…”
Akshay Nathan Jul 28, 2026 ▶ 1:04:44 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026
Assertion Not checkable as stated
Nathan: Codex and ChatGPT Work share the same underlying agent harness
“So the harness is the same. The harness is shared. On, In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plugins or computer use or artifacts. You get that power regardless of what your…”
Akshay Nathan Jul 28, 2026 ▶ 11:18 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 neutral
Insight
Nathan: AI tooling makes motion easier, risking conflation with true progress
“I think maybe the trap is like conflating motion and progress. I think motion is much easier now than ever before because of the tooling that we have, but progress requires you to be like very prescriptive and deliberate about like what you're actually trying …”
Akshay Nathan Jul 28, 2026 ▶ 1:09:47 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Jul 28, 2026 bullish
Prediction Not checkable as stated
Akshay Nathan: AI will make tech workers into T-shaped generalists
“My suspicion is that there's everything, everyone will be, like, T-shaped in a way, and that, like, AI will enable everyone to become a generalist. Like, you know, things that, like, I never would be able to, like, come up with a design before, and, like, even…”
Akshay Nathan Jul 28, 2026 ▶ 1:03:59 OpenAI’s Vision for the AI Super App — Akshay Nathan, OpenAI
Aug 11, 2026 neutral
Assertion Supported
McPartlon: OpenAI co-led Chai Discovery's seed funding round
“Actually OpenAI co-led our seed round.”
Matt McPartlon Aug 11, 2026 ▶ 20:01 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Sep 2, 2026 neutral
Disclosure
Sean Lie: OpenAI is Cerebras's biggest customer
“OpenAI is our biggest customer.”
Sean Lie Sep 2, 2026 ▶ 18:26 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 positive
Assertion Not checkable as stated
Lie: OpenAI uses Cerebras hardware internally for incident response and research
“So right now internally, they're using it for a lot of really critical use cases where the speed really, really matters. Like they're using it in like, Their incidents response teams, right? When there's an outage in their service, for example, every single se…”
Sean Lie Sep 2, 2026 ▶ 13:29 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 bullish
Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 positive
Disclosure
Sean Lie: Cerebras uses OpenAI's internal AI tools for chip design
“We're also collaborating very closely with OpenAI, right, to use their tools to help us also continue to push what's possible in our chip design, in our software, and all that.”
Sean Lie Sep 2, 2026 ▶ 19:59 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 bullish
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.