McGrath: OpenAI Continues to Release Non-Thinking Models for Specific APIs
“No, we're still, we still are releasing non-thinking models but that one was the one that we did that was like API-specific non-thinking so, you know, focus has shifted a little.”
McGrath: OpenAI 10xed Effective Context Window for GPT-4.1
“I worked on long context, that was why I was on last, was for 4.1, where we, you know, I think, tenxed the effective context window for 4.1”
OpenAI launches GPT-4.1 model lineup featuring 1M-token context window
“Yeah, I'll just say we released three new models today, GPT-Fort.one, GPT-Fort.one mini, and GPT-Fort.one data, and the real focus on these were just making the models that were great for developers so we improved instruction following, coding, and shipped our…”
OpenAI stealth-tested GPT-4.1 models on OpenRouter before official release
“Yeah yeah, we really wanted to get as much developer feedback as possible on this model to make sure it worked well in the real world, and so we tested it kind of through Open Router and it was super cool to see people latch on to the names and get the theorie…”
OpenAI has no current plans to add GPT-4.1 to Realtime API
“I don't think we don't have any current plans to release 4.1 in the real time API, but you know, things, things may change.”
GPT-4.1 Nano and Mini are new pre-trains; base 4.1 is mid-train
“Nano is obviously a new pre-train. We also have a new pre-train for Mini, and then, ah, the larger version is, ah, a new mid-train.”
GPT-4.1 powers OpenAI API while enhanced memory remains ChatGPT-exclusive
“So, 4.1 is powering the API, whereas the enhanced memory is, is ChatGPT only.”
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
GPT-4.1 significantly improves chain-of-thought planning over previous non-reasoning models
“We have found that 4.1 is a lot better at doing planning and thinking through its steps in COT when prompted than our previous non-reasoning models.”
Pokrass: Prototype with GPT-4.1, then downscale for latency or upscale for reasoning
“I think the answer is always going to be the fastest model that accomplishes your task, right? So maybe you start prompting 4.1 as a starting point if it does your task super well, Then maybe you could drop down a 4.1 mini and save latency, or even nano. Where…”
GPT-4.1 excels at exploring repositories, while reasoning models dominate targeted file changes
“Basically, where GPT, 4.1, can it kind of explore, go through a repo? It's been trained to do that particularly well. Whereas you know, to just get some code and produce a change, a reasoning model might do better because it can kind of reason over the entire …”
OpenAI researcher uses GPT-4.1 for 49 of 50 commits on massive PR
“I was actually just talking to one of the researchers on the team who worked on something over the weekend. And he said that this model, GBT, 4.1 was able to like get 49 out of 50 of his commits on this massive PR done.”
GPT-4.1's multimodal vision improvements stem from pre-training, not post-training
“We talked about like coding instruction following long context, a lot of gains coming from post training, but in particular multimodal, like basically everything you're seeing, the gains are there from pre-training.”
OpenAI increases prompt caching discount from 50% to 75% on GPT-4.1
“We've increased our prompt caching discount from 50% to 75% on these models.”