Feb 6, 2026 · 48m · startup-ideas

Claude Opus 4.6 vs GPT-5.3 Codex

Morgan Linton · 32m spoken Greg Isenberg · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Greg Isenberg and guest Morgan Linton stage a live head-to-head coding challenge between Claude Opus 4.6 and GPT-5.3 Codex to rebuild the decentralized prediction platform Polymarket from scratch. Through direct comparison of setup, real-time code generation, interactive steering, and autonomous multi-agent orchestration, the episode reveals the distinct engineering strengths and trade-offs of both frontier AI models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Greg holds 21.4% of the talking time here. How this is scored →

Greg as informed peer 3.3 Guest teaching 4.9 Guest disagreement 0.8 Greg pushing back 0.7
05100:0015:0030:0045:001:12–8:23 · Greg as informed peer 2/10 Welcoming Morgan Linton and Framing the Matchup Morgan provides an extensive technical walkthrough on configuring Opus 4.6, updating CLI versions, editing settings.json for agent teams, and managing API adaptive thinking settings. Greg acts primarily as a receptive host validating the instructions.8:23–11:11 · Greg as informed peer 3/10 Philosophical Divergence Between Collaborative and Autonomous AI Morgan reads an insightful community breakdown outlining the philosophical split between human-in-the-loop pair programming and autonomous multi-agent delegation. Greg summarizes the takeaway as a matter of developer preference.11:12–15:41 · Greg as informed peer 3/10 Technical Comparison: Context Windows, Benchmarks, and Personas Morgan contrasts model specs, pointing out Opus 4.6's 1M context window and deep architectural comprehension versus GPT-5.3 Codex's ~200k context window and superior benchmark execution speed. Greg contributes relatable persona metaphors to frame the comparison.15:42–18:29 · Greg as informed peer 2/10 Configuring the Head-to-Head Polymarket Challenge The pair set up the benchmark competition to recreate Polymarket, with Morgan tailoring prompt phrasing to match each model's operational style. Greg jokingly affirms their strict impartiality.18:30–22:15 · Greg as informed peer 3/10 Live Build Execution: Multi-Agent Research vs Rapid Scaffolding Both models kick off the build; Opus spawns four parallel web research agents while Codex immediately scaffolds code and writes market math engines. Greg asks how beginner vibe-coders should choose between the two models.22:15–26:37 · Greg as informed peer 3/10 Reviewing and Interacting with Codex's Signal Market Codex finishes its build in under four minutes, and Morgan spins up the local Node server to test market creation. Greg creates a Bitcoin price prediction market to verify trade execution and order book functionality.26:40–32:52 · Greg as informed peer 5/10 Analyzing Opus Token Consumption and AI Business Models While Opus continues building, Greg and Morgan analyze token consumption, noting multi-agent architectures multiply token burns and align with Anthropic's monetization strategy. Greg estimates token dollar costs against monthly limits.32:55–39:32 · Greg as informed peer 5/10 Iterative Steering and Redesigning Codex's Interface Greg aggressively pushes Codex's UI design by instructing it to mimic Jack Dorsey's minimalist aesthetic, prompting Morgan to test mid-execution pause and resume steering. Both critique Codex's halting behavior and iterative refinements.39:33–43:54 · Greg as informed peer 4/10 Exploring Claude Opus's Completed Forecast Application Opus completes its build after using ~150k tokens, delivering a polished Next.js app with 96 automated tests, category filters, and detailed market cards. Greg and Morgan are visibly stunned by the UI polish and comprehensive architecture.43:54–48:54 · Greg as informed peer 3/10 Final Application Comparison and Declaring a Winner Greg and Morgan inspect Codex's updated design and declare Claude Opus the clear winner for full-stack build quality in this specific benchmark test. Morgan shares advice for enterprise engineering teams exploring agentic workflows.1:12–8:23 · Guest teaching 7/10 Welcoming Morgan Linton and Framing the Matchup Morgan provides an extensive technical walkthrough on configuring Opus 4.6, updating CLI versions, editing settings.json for agent teams, and managing API adaptive thinking settings. Greg acts primarily as a receptive host validating the instructions.8:23–11:11 · Guest teaching 5/10 Philosophical Divergence Between Collaborative and Autonomous AI Morgan reads an insightful community breakdown outlining the philosophical split between human-in-the-loop pair programming and autonomous multi-agent delegation. Greg summarizes the takeaway as a matter of developer preference.11:12–15:41 · Guest teaching 7/10 Technical Comparison: Context Windows, Benchmarks, and Personas Morgan contrasts model specs, pointing out Opus 4.6's 1M context window and deep architectural comprehension versus GPT-5.3 Codex's ~200k context window and superior benchmark execution speed. Greg contributes relatable persona metaphors to frame the comparison.15:42–18:29 · Guest teaching 4/10 Configuring the Head-to-Head Polymarket Challenge The pair set up the benchmark competition to recreate Polymarket, with Morgan tailoring prompt phrasing to match each model's operational style. Greg jokingly affirms their strict impartiality.18:30–22:15 · Guest teaching 6/10 Live Build Execution: Multi-Agent Research vs Rapid Scaffolding Both models kick off the build; Opus spawns four parallel web research agents while Codex immediately scaffolds code and writes market math engines. Greg asks how beginner vibe-coders should choose between the two models.22:15–26:37 · Guest teaching 4/10 Reviewing and Interacting with Codex's Signal Market Codex finishes its build in under four minutes, and Morgan spins up the local Node server to test market creation. Greg creates a Bitcoin price prediction market to verify trade execution and order book functionality.26:40–32:52 · Guest teaching 4/10 Analyzing Opus Token Consumption and AI Business Models While Opus continues building, Greg and Morgan analyze token consumption, noting multi-agent architectures multiply token burns and align with Anthropic's monetization strategy. Greg estimates token dollar costs against monthly limits.32:55–39:32 · Guest teaching 3/10 Iterative Steering and Redesigning Codex's Interface Greg aggressively pushes Codex's UI design by instructing it to mimic Jack Dorsey's minimalist aesthetic, prompting Morgan to test mid-execution pause and resume steering. Both critique Codex's halting behavior and iterative refinements.39:33–43:54 · Guest teaching 5/10 Exploring Claude Opus's Completed Forecast Application Opus completes its build after using ~150k tokens, delivering a polished Next.js app with 96 automated tests, category filters, and detailed market cards. Greg and Morgan are visibly stunned by the UI polish and comprehensive architecture.43:54–48:54 · Guest teaching 4/10 Final Application Comparison and Declaring a Winner Greg and Morgan inspect Codex's updated design and declare Claude Opus the clear winner for full-stack build quality in this specific benchmark test. Morgan shares advice for enterprise engineering teams exploring agentic workflows.1:12–8:23 · Guest disagreement 1/10 Welcoming Morgan Linton and Framing the Matchup Morgan provides an extensive technical walkthrough on configuring Opus 4.6, updating CLI versions, editing settings.json for agent teams, and managing API adaptive thinking settings. Greg acts primarily as a receptive host validating the instructions.8:23–11:11 · Guest disagreement 1/10 Philosophical Divergence Between Collaborative and Autonomous AI Morgan reads an insightful community breakdown outlining the philosophical split between human-in-the-loop pair programming and autonomous multi-agent delegation. Greg summarizes the takeaway as a matter of developer preference.11:12–15:41 · Guest disagreement 1/10 Technical Comparison: Context Windows, Benchmarks, and Personas Morgan contrasts model specs, pointing out Opus 4.6's 1M context window and deep architectural comprehension versus GPT-5.3 Codex's ~200k context window and superior benchmark execution speed. Greg contributes relatable persona metaphors to frame the comparison.15:42–18:29 · Guest disagreement 0/10 Configuring the Head-to-Head Polymarket Challenge The pair set up the benchmark competition to recreate Polymarket, with Morgan tailoring prompt phrasing to match each model's operational style. Greg jokingly affirms their strict impartiality.18:30–22:15 · Guest disagreement 1/10 Live Build Execution: Multi-Agent Research vs Rapid Scaffolding Both models kick off the build; Opus spawns four parallel web research agents while Codex immediately scaffolds code and writes market math engines. Greg asks how beginner vibe-coders should choose between the two models.22:15–26:37 · Guest disagreement 1/10 Reviewing and Interacting with Codex's Signal Market Codex finishes its build in under four minutes, and Morgan spins up the local Node server to test market creation. Greg creates a Bitcoin price prediction market to verify trade execution and order book functionality.26:40–32:52 · Guest disagreement 1/10 Analyzing Opus Token Consumption and AI Business Models While Opus continues building, Greg and Morgan analyze token consumption, noting multi-agent architectures multiply token burns and align with Anthropic's monetization strategy. Greg estimates token dollar costs against monthly limits.32:55–39:32 · Guest disagreement 2/10 Iterative Steering and Redesigning Codex's Interface Greg aggressively pushes Codex's UI design by instructing it to mimic Jack Dorsey's minimalist aesthetic, prompting Morgan to test mid-execution pause and resume steering. Both critique Codex's halting behavior and iterative refinements.39:33–43:54 · Guest disagreement 0/10 Exploring Claude Opus's Completed Forecast Application Opus completes its build after using ~150k tokens, delivering a polished Next.js app with 96 automated tests, category filters, and detailed market cards. Greg and Morgan are visibly stunned by the UI polish and comprehensive architecture.43:54–48:54 · Guest disagreement 0/10 Final Application Comparison and Declaring a Winner Greg and Morgan inspect Codex's updated design and declare Claude Opus the clear winner for full-stack build quality in this specific benchmark test. Morgan shares advice for enterprise engineering teams exploring agentic workflows.1:12–8:23 · Greg pushing back 0/10 Welcoming Morgan Linton and Framing the Matchup Morgan provides an extensive technical walkthrough on configuring Opus 4.6, updating CLI versions, editing settings.json for agent teams, and managing API adaptive thinking settings. Greg acts primarily as a receptive host validating the instructions.8:23–11:11 · Greg pushing back 1/10 Philosophical Divergence Between Collaborative and Autonomous AI Morgan reads an insightful community breakdown outlining the philosophical split between human-in-the-loop pair programming and autonomous multi-agent delegation. Greg summarizes the takeaway as a matter of developer preference.11:12–15:41 · Greg pushing back 0/10 Technical Comparison: Context Windows, Benchmarks, and Personas Morgan contrasts model specs, pointing out Opus 4.6's 1M context window and deep architectural comprehension versus GPT-5.3 Codex's ~200k context window and superior benchmark execution speed. Greg contributes relatable persona metaphors to frame the comparison.15:42–18:29 · Greg pushing back 0/10 Configuring the Head-to-Head Polymarket Challenge The pair set up the benchmark competition to recreate Polymarket, with Morgan tailoring prompt phrasing to match each model's operational style. Greg jokingly affirms their strict impartiality.18:30–22:15 · Greg pushing back 1/10 Live Build Execution: Multi-Agent Research vs Rapid Scaffolding Both models kick off the build; Opus spawns four parallel web research agents while Codex immediately scaffolds code and writes market math engines. Greg asks how beginner vibe-coders should choose between the two models.22:15–26:37 · Greg pushing back 0/10 Reviewing and Interacting with Codex's Signal Market Codex finishes its build in under four minutes, and Morgan spins up the local Node server to test market creation. Greg creates a Bitcoin price prediction market to verify trade execution and order book functionality.26:40–32:52 · Greg pushing back 1/10 Analyzing Opus Token Consumption and AI Business Models While Opus continues building, Greg and Morgan analyze token consumption, noting multi-agent architectures multiply token burns and align with Anthropic's monetization strategy. Greg estimates token dollar costs against monthly limits.32:55–39:32 · Greg pushing back 4/10 Iterative Steering and Redesigning Codex's Interface Greg aggressively pushes Codex's UI design by instructing it to mimic Jack Dorsey's minimalist aesthetic, prompting Morgan to test mid-execution pause and resume steering. Both critique Codex's halting behavior and iterative refinements.39:33–43:54 · Greg pushing back 0/10 Exploring Claude Opus's Completed Forecast Application Opus completes its build after using ~150k tokens, delivering a polished Next.js app with 96 automated tests, category filters, and detailed market cards. Greg and Morgan are visibly stunned by the UI polish and comprehensive architecture.43:54–48:54 · Greg pushing back 0/10 Final Application Comparison and Declaring a Winner Greg and Morgan inspect Codex's updated design and declare Claude Opus the clear winner for full-stack build quality in this specific benchmark test. Morgan shares advice for enterprise engineering teams exploring agentic workflows.

speaking balance: gold is Greg, purple is the guest (3 minute bins)

0:00 · Greg 65.6% · guest 34.4%0:00 · Greg 65.6% · guest 34.4%3:00 · Greg 4.3% · guest 95.7%3:00 · Greg 4.3% · guest 95.7%6:00 · Greg 0.1% · guest 99.9%6:00 · Greg 0.1% · guest 99.9%9:00 · Greg 10.7% · guest 89.3%9:00 · Greg 10.7% · guest 89.3%12:00 · Greg 4.4% · guest 95.6%12:00 · Greg 4.4% · guest 95.6%15:00 · Greg 2.7% · guest 97.3%15:00 · Greg 2.7% · guest 97.3%18:00 · Greg 11.6% · guest 88.4%18:00 · Greg 11.6% · guest 88.4%21:00 · Greg 11.3% · guest 88.7%21:00 · Greg 11.3% · guest 88.7%24:00 · Greg 19.5% · guest 80.5%24:00 · Greg 19.5% · guest 80.5%27:00 · Greg 37.7% · guest 62.3%27:00 · Greg 37.7% · guest 62.3%30:00 · Greg 29% · guest 71%30:00 · Greg 29% · guest 71%33:00 · Greg 42.2% · guest 57.8%33:00 · Greg 42.2% · guest 57.8%36:00 · Greg 35.7% · guest 64.3%36:00 · Greg 35.7% · guest 64.3%39:00 · Greg 15.6% · guest 84.4%39:00 · Greg 15.6% · guest 84.4%42:00 · Greg 21.5% · guest 78.5%42:00 · Greg 21.5% · guest 78.5%45:00 · Greg 31.2% · guest 68.8%45:00 · Greg 31.2% · guest 68.8%48:00 · Greg 31.7% · guest 68.3%48:00 · Greg 31.7% · guest 68.3%
Sharpest disagreement ▶ 37:05 Morgan dismisses Codex's literal interpretation of user prompts

Morgan mocks Codex for acting like Data from Star Trek when it offered a dry dictionary explanation of credits rather than handling the design brief naturally.

Hardest push from Greg ▶ 34:30 Greg rejects Codex's minor UI tweaks and demands a major overhaul

Greg explicitly rejects Codex's incremental styling changes, demanding an aggressive caps-lock redesign modeled on Jack Dorsey's product sensibilities.

Biggest teaching moment ▶ 5:05 Morgan explains configuring experimental agent teams and API thinking depth

Morgan educates Greg and the audience on setting the CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS environment variable and configuring max effort parameters in API calls.

Greg holds their own ▶ 27:48 Greg decodes Anthropic's business incentive behind multi-agent architecture

Greg demonstrates strong business acumen by connecting token-hungry multi-agent systems to Anthropic's anti-ad stance and monetization model.

the scores for every segment, with the reasoning behind each
ChapterTopicGreg as informed peerGuest teachingGuest disagreementGreg pushing backWhy
Welcoming Morgan Linton and Framing the Matchup 2710 Morgan provides an extensive technical walkthrough on configuring Opus 4.6, updating CLI versions, editing settings.json for agent teams, and managing API adaptive thinking settings. Greg acts primarily as a receptive host validating the instructions.
Philosophical Divergence Between Collaborative and Autonomous AI 3511 Morgan reads an insightful community breakdown outlining the philosophical split between human-in-the-loop pair programming and autonomous multi-agent delegation. Greg summarizes the takeaway as a matter of developer preference.
Technical Comparison: Context Windows, Benchmarks, and Personas 3710 Morgan contrasts model specs, pointing out Opus 4.6's 1M context window and deep architectural comprehension versus GPT-5.3 Codex's ~200k context window and superior benchmark execution speed. Greg contributes relatable persona metaphors to frame the comparison.
Configuring the Head-to-Head Polymarket Challenge 2400 The pair set up the benchmark competition to recreate Polymarket, with Morgan tailoring prompt phrasing to match each model's operational style. Greg jokingly affirms their strict impartiality.
Live Build Execution: Multi-Agent Research vs Rapid Scaffolding 3611 Both models kick off the build; Opus spawns four parallel web research agents while Codex immediately scaffolds code and writes market math engines. Greg asks how beginner vibe-coders should choose between the two models.
Reviewing and Interacting with Codex's Signal Market 3410 Codex finishes its build in under four minutes, and Morgan spins up the local Node server to test market creation. Greg creates a Bitcoin price prediction market to verify trade execution and order book functionality.
Analyzing Opus Token Consumption and AI Business Models 5411 While Opus continues building, Greg and Morgan analyze token consumption, noting multi-agent architectures multiply token burns and align with Anthropic's monetization strategy. Greg estimates token dollar costs against monthly limits.
Iterative Steering and Redesigning Codex's Interface 5324 Greg aggressively pushes Codex's UI design by instructing it to mimic Jack Dorsey's minimalist aesthetic, prompting Morgan to test mid-execution pause and resume steering. Both critique Codex's halting behavior and iterative refinements.
Exploring Claude Opus's Completed Forecast Application 4500 Opus completes its build after using ~150k tokens, delivering a polished Next.js app with 96 automated tests, category filters, and detailed market cards. Greg and Morgan are visibly stunned by the UI polish and comprehensive architecture.
Final Application Comparison and Declaring a Winner 3400 Greg and Morgan inspect Codex's updated design and declare Claude Opus the clear winner for full-stack build quality in this specific benchmark test. Morgan shares advice for enterprise engineering teams exploring agentic workflows.

Statements from this episode (11)

Assertion Supported
Altman announced GPT-5.3 Codex 18 minutes after Anthropic released Opus 4.6
“Opus four, six came out and then Sam Altman put together a quick tweet. I want to say like maybe 18 minutes later announcing GPT five, three codex”
Morgan Linton Feb 6, 2026 ▶ 1:43
Prediction Not checkable as stated
Morgan Linton predicts engineering teams will use both Codex and Opus 4.6
“Where I think you're going to see a lot of teams using both, because Codex really is your collaborator, and what they've added with Five-Three is, like, really good, like, mid-execution steering, whereas with Opus Four-Six, it's probably the best of the best n…”
Morgan Linton Feb 6, 2026 ▶ 10:09
Assertion Supported
Linton: Claude Opus 4.6 features a 1-million-token context window
“So, with Opus Four Six, ah, much bigger context window, so you have a million token context window here.”
Morgan Linton Feb 6, 2026 ▶ 11:39
Assertion Contradicted
Linton: GPT-5.3 Codex has a context window of roughly 200k tokens
“Five, three, they talk about large context, but it's not a headline feature, and I actually went back and forth with it to get it to actually give me a number, and the number's around 200,000 tokens, which is not that impressive.”
Morgan Linton Feb 6, 2026 ▶ 11:55
Insight
Claude excels at system comprehension while GPT-5.3 rules rapid pair programming
“So, high level, what that means is Claude is better when the task is understand everything first and then decide. GPT Thrive Five Three Codex is probably better when the task is decide fast, act, iterate, more of that, you know, pair programming, you know, mid…”
Morgan Linton Feb 6, 2026 ▶ 12:26
Assertion Supported
Linton: GPT-5.3 Codex beat Opus 4.6 on SWE-bench Pro and TerminalBench
“Five, five, three codecs did win on SWD bench pro terminal bench overall. It's like scored better on coding benchmarks. So probably better end end app generation.”
Morgan Linton Feb 6, 2026 ▶ 13:28
Assertion Not checkable as stated
Claude Opus 4.6 creates agent teams while Codex requires direct reasoning prompts
“When I'm talking to Opus, I can tell Opus, build me a team, and here's what I want each member of the team to do. When I'm talking to Codex, I can't really tell it to build me a team, but I can tell it to think about stuff.”
Morgan Linton Feb 6, 2026 ▶ 15:57
Assertion Not checkable as stated
GPT-5.3 Codex generated a functional Polymarket competitor in under four minutes
“So Codex built a competitor to Polymarket in three minutes and 47 seconds.”
Morgan Linton Feb 6, 2026 ▶ 22:17
Insight
Autonomous AI agents will massively multiply token consumption and revenue for Anthropic
“I think that's one of the very good things for like, Investors in Anthropic, right, is with agents and agents now being, I think probably the new killer feature in Opus. You're going to take whatever token usage and multiply it by the number of agents.”
Morgan Linton Feb 6, 2026 ▶ 28:11
Assertion Supported
Claude Opus 4.6 created 96 automated tests versus GPT-5.3 Codex's 10
“Codex created 10 tests, right? Opus created 96 tests.”
Morgan Linton Feb 6, 2026 ▶ 40:27
Opinion
Claude Opus 4.6 beat GPT-5.3 Codex in a head-to-head coding test
“I mean, I would say, you know, like I said, I'm not going to say which one is, it's not that Opus is better than Codex or vice versa, but I would say in this test, Opus won.”
Morgan Linton Feb 6, 2026 ▶ 45:07
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.