Apr 15, 2025 · 45m · latent-space

GPT 4.1: The New OpenAI Workhorse

Shawn Wang · 14m spoken Michelle Pokrass · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI researchers Michelle Pokrass and Josh join the Latent Space podcast to detail the GPT-4.1 model family, highlighting advances in post-training, software engineering, one-million-token context reasoning, and developer tooling. They provide actionable guidance on prompt engineering, benchmark evaluations, model selection, and cost optimization for production AI workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 36.2% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 3.4 Guest disagreement 0.8 The hosts pushing back 1.7
05100:0015:0030:0045:000:58–3:39 · The hosts as informed peer 4/10 Unveiling the GPT-4.1 Model Family and Launch Lore Swyx opens with conversational context regarding OpenAI's pre-release code names on OpenRouter and the Waterloo engineering connection. Michelle and Josh explain the high-level release details and the inclusion of the nano tier. The dynamic is friendly and collaborative.3:40–8:12 · The hosts as informed peer 5/10 Decoupling Model Names and Post-Training Research Architectural Insights Swyx probes the naming logic, model size comparisons with GPT-4.5, and Omni architecture lineage. Josh and Michelle clarify that versioning reflects user capability improvements and post-training gains rather than linear pre-training size scaling.8:15–17:12 · The hosts as informed peer 6/10 Achieving One Million Context Windows Through Graph Reasoning Benchmarks The hosts and guests discuss 1M context evaluation, where Josh details synthetic graph walk tasks. Swyx cites previous research literature on graph traversals for agent planning (CogEval at NeurIPS) to frame the discussion.17:12–20:22 · The hosts as informed peer 5/10 Real-World Developer Instruction Following and API Data Evaluation Swyx asks about Shen Yu's work and internal API instruction-following evals. Michelle explains the limitations of open-source evals with programmatic grading versus real-world developer prompts with complex negative constraints.20:22–26:49 · The hosts as informed peer 6/10 Effective Prompt Engineering Guidelines and Agentic Workflow Persistence Swyx challenges the prompt guide recommendation of placing instructions at both the top and bottom of the context because of conflicts with prompt caching. Josh and Michelle clarify how caching can still work depending on prompt structure and dispute the idea that JSON is deprecated in favor of XML.26:51–28:59 · The hosts as informed peer 4/10 Navigating Model Selection Between GPT-4.1 and Reasoning Models Alessio and Swyx ask how developers should choose between GPT-4.1 chain-of-thought prompting and dedicated reasoning models like o1. Michelle outlines a tiered heuristic based on latency, cost, and planning horizon.29:00–34:25 · The hosts as informed peer 5/10 State-of-the-Art Coding Capabilities and Internal OpenAI Engineering Workflows Swyx asks about SWE-bench benchmarks and breakdowns across diff generation and full-repo exploration. Michelle shares internal developer metrics and anecdotes about massive PR commits completed by 4.1.34:27–38:31 · The hosts as informed peer 5/10 Multimodal Vision Advancements and Infrastructure Optimization Swyx questions the premise that deprecating GPT-4.5 reclaims GPUs when OpenAI supports old models concurrently for months. Michelle explains the balance between compute reclamation and developer API reliability commitments.38:32–41:23 · The hosts as informed peer 4/10 Day-One Fine-Tuning Offerings and Upcoming Model Horizons Michelle highlights preference fine-tuning for style steering, which Swyx confuses with reinforcement fine-tuning (RFT). Michelle and Josh politely correct him, clarifying that RFT is strictly for reasoning models.41:23–44:53 · The hosts as informed peer 4/10 Developer Ecosystem Feedback, Caching Pricing Reductions, and Concluding Remarks Swyx asks about blended pricing and assumes all models received a price drop. Michelle corrects the pricing assumption regarding mini and highlights the newly increased 75% prompt caching discount.0:58–3:39 · Guest teaching 1/10 Unveiling the GPT-4.1 Model Family and Launch Lore Swyx opens with conversational context regarding OpenAI's pre-release code names on OpenRouter and the Waterloo engineering connection. Michelle and Josh explain the high-level release details and the inclusion of the nano tier. The dynamic is friendly and collaborative.3:40–8:12 · Guest teaching 4/10 Decoupling Model Names and Post-Training Research Architectural Insights Swyx probes the naming logic, model size comparisons with GPT-4.5, and Omni architecture lineage. Josh and Michelle clarify that versioning reflects user capability improvements and post-training gains rather than linear pre-training size scaling.8:15–17:12 · Guest teaching 3/10 Achieving One Million Context Windows Through Graph Reasoning Benchmarks The hosts and guests discuss 1M context evaluation, where Josh details synthetic graph walk tasks. Swyx cites previous research literature on graph traversals for agent planning (CogEval at NeurIPS) to frame the discussion.17:12–20:22 · Guest teaching 3/10 Real-World Developer Instruction Following and API Data Evaluation Swyx asks about Shen Yu's work and internal API instruction-following evals. Michelle explains the limitations of open-source evals with programmatic grading versus real-world developer prompts with complex negative constraints.20:22–26:49 · Guest teaching 4/10 Effective Prompt Engineering Guidelines and Agentic Workflow Persistence Swyx challenges the prompt guide recommendation of placing instructions at both the top and bottom of the context because of conflicts with prompt caching. Josh and Michelle clarify how caching can still work depending on prompt structure and dispute the idea that JSON is deprecated in favor of XML.26:51–28:59 · Guest teaching 3/10 Navigating Model Selection Between GPT-4.1 and Reasoning Models Alessio and Swyx ask how developers should choose between GPT-4.1 chain-of-thought prompting and dedicated reasoning models like o1. Michelle outlines a tiered heuristic based on latency, cost, and planning horizon.29:00–34:25 · Guest teaching 2/10 State-of-the-Art Coding Capabilities and Internal OpenAI Engineering Workflows Swyx asks about SWE-bench benchmarks and breakdowns across diff generation and full-repo exploration. Michelle shares internal developer metrics and anecdotes about massive PR commits completed by 4.1.34:27–38:31 · Guest teaching 3/10 Multimodal Vision Advancements and Infrastructure Optimization Swyx questions the premise that deprecating GPT-4.5 reclaims GPUs when OpenAI supports old models concurrently for months. Michelle explains the balance between compute reclamation and developer API reliability commitments.38:32–41:23 · Guest teaching 6/10 Day-One Fine-Tuning Offerings and Upcoming Model Horizons Michelle highlights preference fine-tuning for style steering, which Swyx confuses with reinforcement fine-tuning (RFT). Michelle and Josh politely correct him, clarifying that RFT is strictly for reasoning models.41:23–44:53 · Guest teaching 5/10 Developer Ecosystem Feedback, Caching Pricing Reductions, and Concluding Remarks Swyx asks about blended pricing and assumes all models received a price drop. Michelle corrects the pricing assumption regarding mini and highlights the newly increased 75% prompt caching discount.0:58–3:39 · Guest disagreement 0/10 Unveiling the GPT-4.1 Model Family and Launch Lore Swyx opens with conversational context regarding OpenAI's pre-release code names on OpenRouter and the Waterloo engineering connection. Michelle and Josh explain the high-level release details and the inclusion of the nano tier. The dynamic is friendly and collaborative.3:40–8:12 · Guest disagreement 1/10 Decoupling Model Names and Post-Training Research Architectural Insights Swyx probes the naming logic, model size comparisons with GPT-4.5, and Omni architecture lineage. Josh and Michelle clarify that versioning reflects user capability improvements and post-training gains rather than linear pre-training size scaling.8:15–17:12 · Guest disagreement 1/10 Achieving One Million Context Windows Through Graph Reasoning Benchmarks The hosts and guests discuss 1M context evaluation, where Josh details synthetic graph walk tasks. Swyx cites previous research literature on graph traversals for agent planning (CogEval at NeurIPS) to frame the discussion.17:12–20:22 · Guest disagreement 0/10 Real-World Developer Instruction Following and API Data Evaluation Swyx asks about Shen Yu's work and internal API instruction-following evals. Michelle explains the limitations of open-source evals with programmatic grading versus real-world developer prompts with complex negative constraints.20:22–26:49 · Guest disagreement 2/10 Effective Prompt Engineering Guidelines and Agentic Workflow Persistence Swyx challenges the prompt guide recommendation of placing instructions at both the top and bottom of the context because of conflicts with prompt caching. Josh and Michelle clarify how caching can still work depending on prompt structure and dispute the idea that JSON is deprecated in favor of XML.26:51–28:59 · Guest disagreement 0/10 Navigating Model Selection Between GPT-4.1 and Reasoning Models Alessio and Swyx ask how developers should choose between GPT-4.1 chain-of-thought prompting and dedicated reasoning models like o1. Michelle outlines a tiered heuristic based on latency, cost, and planning horizon.29:00–34:25 · Guest disagreement 1/10 State-of-the-Art Coding Capabilities and Internal OpenAI Engineering Workflows Swyx asks about SWE-bench benchmarks and breakdowns across diff generation and full-repo exploration. Michelle shares internal developer metrics and anecdotes about massive PR commits completed by 4.1.34:27–38:31 · Guest disagreement 1/10 Multimodal Vision Advancements and Infrastructure Optimization Swyx questions the premise that deprecating GPT-4.5 reclaims GPUs when OpenAI supports old models concurrently for months. Michelle explains the balance between compute reclamation and developer API reliability commitments.38:32–41:23 · Guest disagreement 1/10 Day-One Fine-Tuning Offerings and Upcoming Model Horizons Michelle highlights preference fine-tuning for style steering, which Swyx confuses with reinforcement fine-tuning (RFT). Michelle and Josh politely correct him, clarifying that RFT is strictly for reasoning models.41:23–44:53 · Guest disagreement 1/10 Developer Ecosystem Feedback, Caching Pricing Reductions, and Concluding Remarks Swyx asks about blended pricing and assumes all models received a price drop. Michelle corrects the pricing assumption regarding mini and highlights the newly increased 75% prompt caching discount.0:58–3:39 · The hosts pushing back 1/10 Unveiling the GPT-4.1 Model Family and Launch Lore Swyx opens with conversational context regarding OpenAI's pre-release code names on OpenRouter and the Waterloo engineering connection. Michelle and Josh explain the high-level release details and the inclusion of the nano tier. The dynamic is friendly and collaborative.3:40–8:12 · The hosts pushing back 2/10 Decoupling Model Names and Post-Training Research Architectural Insights Swyx probes the naming logic, model size comparisons with GPT-4.5, and Omni architecture lineage. Josh and Michelle clarify that versioning reflects user capability improvements and post-training gains rather than linear pre-training size scaling.8:15–17:12 · The hosts pushing back 2/10 Achieving One Million Context Windows Through Graph Reasoning Benchmarks The hosts and guests discuss 1M context evaluation, where Josh details synthetic graph walk tasks. Swyx cites previous research literature on graph traversals for agent planning (CogEval at NeurIPS) to frame the discussion.17:12–20:22 · The hosts pushing back 1/10 Real-World Developer Instruction Following and API Data Evaluation Swyx asks about Shen Yu's work and internal API instruction-following evals. Michelle explains the limitations of open-source evals with programmatic grading versus real-world developer prompts with complex negative constraints.20:22–26:49 · The hosts pushing back 3/10 Effective Prompt Engineering Guidelines and Agentic Workflow Persistence Swyx challenges the prompt guide recommendation of placing instructions at both the top and bottom of the context because of conflicts with prompt caching. Josh and Michelle clarify how caching can still work depending on prompt structure and dispute the idea that JSON is deprecated in favor of XML.26:51–28:59 · The hosts pushing back 1/10 Navigating Model Selection Between GPT-4.1 and Reasoning Models Alessio and Swyx ask how developers should choose between GPT-4.1 chain-of-thought prompting and dedicated reasoning models like o1. Michelle outlines a tiered heuristic based on latency, cost, and planning horizon.29:00–34:25 · The hosts pushing back 1/10 State-of-the-Art Coding Capabilities and Internal OpenAI Engineering Workflows Swyx asks about SWE-bench benchmarks and breakdowns across diff generation and full-repo exploration. Michelle shares internal developer metrics and anecdotes about massive PR commits completed by 4.1.34:27–38:31 · The hosts pushing back 3/10 Multimodal Vision Advancements and Infrastructure Optimization Swyx questions the premise that deprecating GPT-4.5 reclaims GPUs when OpenAI supports old models concurrently for months. Michelle explains the balance between compute reclamation and developer API reliability commitments.38:32–41:23 · The hosts pushing back 1/10 Day-One Fine-Tuning Offerings and Upcoming Model Horizons Michelle highlights preference fine-tuning for style steering, which Swyx confuses with reinforcement fine-tuning (RFT). Michelle and Josh politely correct him, clarifying that RFT is strictly for reasoning models.41:23–44:53 · The hosts pushing back 2/10 Developer Ecosystem Feedback, Caching Pricing Reductions, and Concluding Remarks Swyx asks about blended pricing and assumes all models received a price drop. Michelle corrects the pricing assumption regarding mini and highlights the newly increased 75% prompt caching discount.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 49.1% · guest 50.9%0:00 · the hosts 49.1% · guest 50.9%3:00 · the hosts 29.8% · guest 70.2%3:00 · the hosts 29.8% · guest 70.2%6:00 · the hosts 29.9% · guest 70.1%6:00 · the hosts 29.9% · guest 70.1%9:00 · the hosts 20.1% · guest 79.9%9:00 · the hosts 20.1% · guest 79.9%12:00 · the hosts 36.2% · guest 63.8%12:00 · the hosts 36.2% · guest 63.8%15:00 · the hosts 46.2% · guest 53.8%15:00 · the hosts 46.2% · guest 53.8%18:00 · the hosts 30.4% · guest 69.6%18:00 · the hosts 30.4% · guest 69.6%21:00 · the hosts 19.9% · guest 80.1%21:00 · the hosts 19.9% · guest 80.1%24:00 · the hosts 49.1% · guest 50.9%24:00 · the hosts 49.1% · guest 50.9%27:00 · the hosts 22.4% · guest 77.6%27:00 · the hosts 22.4% · guest 77.6%30:00 · the hosts 33.4% · guest 66.6%30:00 · the hosts 33.4% · guest 66.6%33:00 · the hosts 32.6% · guest 67.4%33:00 · the hosts 32.6% · guest 67.4%36:00 · the hosts 52.3% · guest 47.7%36:00 · the hosts 52.3% · guest 47.7%39:00 · the hosts 35.9% · guest 64.1%39:00 · the hosts 35.9% · guest 64.1%42:00 · the hosts 59.5% · guest 40.5%42:00 · the hosts 59.5% · guest 40.5%45:00 · the hosts 0% · guest 0%45:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 24:20 Dismissing the claim that JSON is deprecated for XML

Michelle immediately rejects Swyx's playful assertion that JSON is now bad and everyone must switch to XML, clarifying the distinct separation between prompt structuring inputs and parsed outputs.

Hardest push from the hosts ▶ 37:41 Challenging the GPU reclamation rationale

Swyx directly challenges OpenAI executive messaging about reclaiming GPUs via deprecations, pointing out that concurrent model hosting over three-month windows actually increases immediate resource usage.

Biggest teaching moment ▶ 39:28 Differentiating RFT from preference fine-tuning

Michelle and Josh correct Swyx's mistaken assumption that preference fine-tuning is restricted to reasoning models, clearly explaining the difference between paired preference tuning and reinforcement fine-tuning.

The host holds their own ▶ 14:13 Referencing CogEval literature on graph walks

Swyx demonstrates deep domain knowledge by connecting OpenAI's newly released synthetic graph benchmarks to prior research in the CogEval paper at NeurIPS regarding agent planning.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Unveiling the GPT-4.1 Model Family and Launch Lore 4101 Swyx opens with conversational context regarding OpenAI's pre-release code names on OpenRouter and the Waterloo engineering connection. Michelle and Josh explain the high-level release details and the inclusion of the nano tier. The dynamic is friendly and collaborative.
Decoupling Model Names and Post-Training Research Architectural Insights 5412 Swyx probes the naming logic, model size comparisons with GPT-4.5, and Omni architecture lineage. Josh and Michelle clarify that versioning reflects user capability improvements and post-training gains rather than linear pre-training size scaling.
Achieving One Million Context Windows Through Graph Reasoning Benchmarks 6312 The hosts and guests discuss 1M context evaluation, where Josh details synthetic graph walk tasks. Swyx cites previous research literature on graph traversals for agent planning (CogEval at NeurIPS) to frame the discussion.
Real-World Developer Instruction Following and API Data Evaluation 5301 Swyx asks about Shen Yu's work and internal API instruction-following evals. Michelle explains the limitations of open-source evals with programmatic grading versus real-world developer prompts with complex negative constraints.
Effective Prompt Engineering Guidelines and Agentic Workflow Persistence 6423 Swyx challenges the prompt guide recommendation of placing instructions at both the top and bottom of the context because of conflicts with prompt caching. Josh and Michelle clarify how caching can still work depending on prompt structure and dispute the idea that JSON is deprecated in favor of XML.
Navigating Model Selection Between GPT-4.1 and Reasoning Models 4301 Alessio and Swyx ask how developers should choose between GPT-4.1 chain-of-thought prompting and dedicated reasoning models like o1. Michelle outlines a tiered heuristic based on latency, cost, and planning horizon.
State-of-the-Art Coding Capabilities and Internal OpenAI Engineering Workflows 5211 Swyx asks about SWE-bench benchmarks and breakdowns across diff generation and full-repo exploration. Michelle shares internal developer metrics and anecdotes about massive PR commits completed by 4.1.
Multimodal Vision Advancements and Infrastructure Optimization 5313 Swyx questions the premise that deprecating GPT-4.5 reclaims GPUs when OpenAI supports old models concurrently for months. Michelle explains the balance between compute reclamation and developer API reliability commitments.
Day-One Fine-Tuning Offerings and Upcoming Model Horizons 4611 Michelle highlights preference fine-tuning for style steering, which Swyx confuses with reinforcement fine-tuning (RFT). Michelle and Josh politely correct him, clarifying that RFT is strictly for reasoning models.
Developer Ecosystem Feedback, Caching Pricing Reductions, and Concluding Remarks 4512 Swyx asks about blended pricing and assumes all models received a price drop. Michelle corrects the pricing assumption regarding mini and highlights the newly increased 75% prompt caching discount.

Statements from this episode (24)

Assertion Supported
OpenAI launches GPT-4.1 model lineup featuring 1M-token context window
“Yeah, I'll just say we released three new models today, GPT-Fort.one, GPT-Fort.one mini, and GPT-Fort.one data, and the real focus on these were just making the models that were great for developers so we improved instruction following, coding, and shipped our…”
Michelle Pokrass Apr 15, 2025 ▶ 1:27
Assertion Supported
OpenAI stealth-tested GPT-4.1 models on OpenRouter before official release
“Yeah yeah, we really wanted to get as much developer feedback as possible on this model to make sure it worked well in the real world, and so we tested it kind of through Open Router and it was super cool to see people latch on to the names and get the theorie…”
Michelle Pokrass Apr 15, 2025 ▶ 2:20
Disclosure
OpenAI has no current plans to add GPT-4.1 to Realtime API
“I don't think we don't have any current plans to release 4.1 in the real time API, but you know, things, things may change.”
Michelle Pokrass Apr 15, 2025 ▶ 6:27
Assertion Open · timeframe Apr 2028
GPT-4.1 Nano and Mini are new pre-trains; base 4.1 is mid-train
“Nano is obviously a new pre-train. We also have a new pre-train for Mini, and then, ah, the larger version is, ah, a new mid-train.”
Michelle Pokrass Apr 15, 2025 ▶ 7:46
Insight
Pokrass: AI model gains now driven by post-training, not larger pre-trains
“We find that actually a significant amount of the gains come from new post-training techniques. So I think in the past the narrative is that you need to pre-train these larger and larger models to get better performance, and we're finding that we're able to sq…”
Michelle Pokrass Apr 15, 2025 ▶ 7:57
Prediction Not checkable as stated
Pokrass predicts developers will abandon RAG vector stores for direct long-context
“So we do expect a lot of developers to start, you know, uploading their full context more directly to the model. So for smaller tasks, you maybe don't need The whole vector store.”
Michelle Pokrass Apr 15, 2025 ▶ 15:21
Disclosure
GPT-4.1 powers OpenAI API while enhanced memory remains ChatGPT-exclusive
“So, 4.1 is powering the API, whereas the enhanced memory is, is ChatGPT only.”
Michelle Pokrass Apr 15, 2025 ▶ 16:11
Insight
Pokrass warns against close collaboration between AI evaluators and model developers
“Honestly, I think it's best when eval authors and model developers don't collab too much because you want things you know, as objective as possible, not trying to game any evals.”
Michelle Pokrass Apr 15, 2025 ▶ 17:38
Insight
Open-source AI benchmarks omit critical tasks because they are hard to grade
“And these are useful instructions, but we find that many of the really interesting instructions are actually challenging to grade. And so the open source evals often don't have them.”
Michelle Pokrass Apr 15, 2025 ▶ 18:56
Disclosure
OpenAI uses its own models to categorize anonymized developer prompts
“Well, I will say we do use our own products internally where we can, and so we're not manually by hand reading every prompt. After they're like anonymized, we scrub them with any identifying data, then we use our models to take passes to categorize them.”
Michelle Pokrass Apr 15, 2025 ▶ 19:55
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Michelle Pokrass Apr 15, 2025 ▶ 23:43
Insight
Pokrass: Use XML for structuring LLM inputs and JSON for parsing outputs
“I do think XML is very helpful for structuring prompts, whereas for parsing outputs maybe the story is a bit different. Like sometimes it's really useful to get outputs in JSON, so you can plug them directly into your application. But I do think the models wor…”
Michelle Pokrass Apr 15, 2025 ▶ 24:38
Assertion Not checkable as stated
GPT-4.1 significantly improves chain-of-thought planning over previous non-reasoning models
“We have found that 4.1 is a lot better at doing planning and thinking through its steps in COT when prompted than our previous non-reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 27:19
Insight
Pokrass: Prototype with GPT-4.1, then downscale for latency or upscale for reasoning
“I think the answer is always going to be the fastest model that accomplishes your task, right? So maybe you start prompting 4.1 as a starting point if it does your task super well, Then maybe you could drop down a 4.1 mini and save latency, or even nano. Where…”
Michelle Pokrass Apr 15, 2025 ▶ 27:58
Insight
Pokrass: Pair reasoning models for planning with smaller models for execution
“I do think reasoning models for planning and using kind of more targeted models to execute is definitely a good architecture.”
Michelle Pokrass Apr 15, 2025 ▶ 28:48
Insight
GPT-4.1 excels at exploring repositories, while reasoning models dominate targeted file changes
“Basically, where GPT, 4.1, can it kind of explore, go through a repo? It's been trained to do that particularly well. Whereas you know, to just get some code and produce a change, a reasoning model might do better because it can kind of reason over the entire …”
Michelle Pokrass Apr 15, 2025 ▶ 31:20
Assertion Not checkable as stated
GPT-4.1 Mini significantly outperforms 4o Mini, nearing original GPT-4o performance
“4.1 mini is actually quite significantly better than four o mini and not that far away from the old four o.”
Michelle Pokrass Apr 15, 2025 ▶ 32:15
Assertion Not checkable as stated
OpenAI researcher uses GPT-4.1 for 49 of 50 commits on massive PR
“I was actually just talking to one of the researchers on the team who worked on something over the weekend. And he said that this model, GBT, 4.1 was able to like get 49 out of 50 of his commits on this massive PR done.”
Michelle Pokrass Apr 15, 2025 ▶ 33:51
Assertion Not checkable as stated
GPT-4.1's multimodal vision improvements stem from pre-training, not post-training
“We talked about like coding instruction following long context, a lot of gains coming from post training, but in particular multimodal, like basically everything you're seeing, the gains are there from pre-training.”
Michelle Pokrass Apr 15, 2025 ▶ 35:09
Opinion
Pokrass: Developers are sleeping on preference fine-tuning for model style steering
“One thing I will say is that I think people have slept on the preference fine tuning offering or the, I think that's what we call the product. So SFT is, people know it pretty well. It's the original fine tuning we had, whereas this preference fine tuning is s…”
Michelle Pokrass Apr 15, 2025 ▶ 39:04
Assertion Supported
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 39:32
Disclosure
OpenAI is working to inject GPT-4.5's humor and nuance into future models
“We're working on incorporating kind of those improvements into the models more generally. People loved about 4.5 is like the humor, the green text, the nuance. So we've heard that feedback and I know, yeah, there's lots of folks working on that and trying to b…”
Michelle Pokrass Apr 15, 2025 ▶ 41:01
Assertion Supported
OpenAI will cover inference costs for developers who share custom evaluation data
“You can upload an eval and opt in such that we'll pay for the inference, inference costs if we can also use the eval.”
Michelle Pokrass Apr 15, 2025 ▶ 42:04
Assertion Supported
OpenAI increases prompt caching discount from 50% to 75% on GPT-4.1
“We've increased our prompt caching discount from 50% to 75% on these models.”
Michelle Pokrass Apr 15, 2025 ▶ 43:27
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.