Jun 6, 2026 · 40m · latent-space

⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai

Ahmad Awais · 31m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Ahmad Awais discusses how CommandCode utilizes deterministic tool repair logic and an automated neurosymbolic Taste engine to elevate open-source models like DeepSeek v4 to frontier-level coding performance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 2.8 Guest teaching 4.3 Guest disagreement 1.5 The hosts pushing back 0.7
05100:0015:0030:001:12–4:32 · The hosts as informed peer 1/10 From Corona CLI to CommandCode and Taste Host invites Ahmad to share his background leading into CommandCode. Ahmad delivers an extensive overview tracing his work from early GPT-3 access and LangBase to creating the Taste engine.4:32–9:18 · The hosts as informed peer 4/10 Identifying the Tool Confusion Problem in DeepSeek Host frames the tool-calling issue and asks why models fail to listen to Zod schema errors like Instructor prompts do. Ahmad explains that open models exhibit stubborn behavior due to synthetic distillation.9:19–13:27 · The hosts as informed peer 1/10 Deterministic Repair Logic for Agent Tool Calling Ahmad screen-shares and explains deterministic repair logic, drawing analogies to database migrations and driving lessons to show how intercepted tool errors guide models without looping.13:28–16:02 · The hosts as informed peer 3/10 Scaling Deterministic Repairs Across Billions of Open Model Tokens Host asks whether tool confusion is unique to DeepSeek. Ahmad explains how the pattern generalized across Kimi and Minimax models over hundreds of billions of tokens.16:03–23:10 · The hosts as informed peer 5/10 Eliminating AI Design Slop with Deterministic Composition Frameworks Ahmad discusses AI design slop and pattern-first composition frameworks. Host contributes domain knowledge by referencing Mario Zechner and the rationale behind CSS OKLCH color models.23:10–26:52 · The hosts as informed peer 3/10 Demonstrating CommandCode's Design Skill and Security Extensions Host validates Ahmad's design and engineering credibility while Ahmad demonstrates the design skill generating UI from raw data and notes extensions into automated security patching.26:52–31:49 · The hosts as informed peer 1/10 Continuous Preference Learning via the Taste Engine Ahmad breaks down continuous preference learning, explaining why static rules files rot and how CommandCode automatically learns micro-decisions and git habits directly into repository markdown.31:50–36:56 · The hosts as informed peer 5/10 Taste vs. Skills Architecture and Multi-Tier Agent Workflows Host presses for architectural clarity regarding whether Taste is a backend model or portable repository files. Ahmad clarifies the distinction between static skills and the dynamic Taste engine.36:57–39:10 · The hosts as informed peer 2/10 CommandCode Open-Source Roadmap and Architectural Philosophy Ahmad outlines open-sourcing CommandCode, contrasting its curated Apple-like philosophy with the broad Windows and Linux approaches of alternative harnesses. Host briefly clarifies the repository's age.39:10–40:35 · The hosts as informed peer 3/10 Concluding Thoughts on Open Ecosystems and Collaborative AI Progress Host notes DeepSeek's recruitment for dedicated coding tooling, and both agree that shared open-source improvements benefit all developer harnesses across the ecosystem.1:12–4:32 · Guest teaching 3/10 From Corona CLI to CommandCode and Taste Host invites Ahmad to share his background leading into CommandCode. Ahmad delivers an extensive overview tracing his work from early GPT-3 access and LangBase to creating the Taste engine.4:32–9:18 · Guest teaching 5/10 Identifying the Tool Confusion Problem in DeepSeek Host frames the tool-calling issue and asks why models fail to listen to Zod schema errors like Instructor prompts do. Ahmad explains that open models exhibit stubborn behavior due to synthetic distillation.9:19–13:27 · Guest teaching 6/10 Deterministic Repair Logic for Agent Tool Calling Ahmad screen-shares and explains deterministic repair logic, drawing analogies to database migrations and driving lessons to show how intercepted tool errors guide models without looping.13:28–16:02 · Guest teaching 5/10 Scaling Deterministic Repairs Across Billions of Open Model Tokens Host asks whether tool confusion is unique to DeepSeek. Ahmad explains how the pattern generalized across Kimi and Minimax models over hundreds of billions of tokens.16:03–23:10 · Guest teaching 4/10 Eliminating AI Design Slop with Deterministic Composition Frameworks Ahmad discusses AI design slop and pattern-first composition frameworks. Host contributes domain knowledge by referencing Mario Zechner and the rationale behind CSS OKLCH color models.23:10–26:52 · Guest teaching 4/10 Demonstrating CommandCode's Design Skill and Security Extensions Host validates Ahmad's design and engineering credibility while Ahmad demonstrates the design skill generating UI from raw data and notes extensions into automated security patching.26:52–31:49 · Guest teaching 6/10 Continuous Preference Learning via the Taste Engine Ahmad breaks down continuous preference learning, explaining why static rules files rot and how CommandCode automatically learns micro-decisions and git habits directly into repository markdown.31:50–36:56 · Guest teaching 5/10 Taste vs. Skills Architecture and Multi-Tier Agent Workflows Host presses for architectural clarity regarding whether Taste is a backend model or portable repository files. Ahmad clarifies the distinction between static skills and the dynamic Taste engine.36:57–39:10 · Guest teaching 3/10 CommandCode Open-Source Roadmap and Architectural Philosophy Ahmad outlines open-sourcing CommandCode, contrasting its curated Apple-like philosophy with the broad Windows and Linux approaches of alternative harnesses. Host briefly clarifies the repository's age.39:10–40:35 · Guest teaching 2/10 Concluding Thoughts on Open Ecosystems and Collaborative AI Progress Host notes DeepSeek's recruitment for dedicated coding tooling, and both agree that shared open-source improvements benefit all developer harnesses across the ecosystem.1:12–4:32 · Guest disagreement 1/10 From Corona CLI to CommandCode and Taste Host invites Ahmad to share his background leading into CommandCode. Ahmad delivers an extensive overview tracing his work from early GPT-3 access and LangBase to creating the Taste engine.4:32–9:18 · Guest disagreement 3/10 Identifying the Tool Confusion Problem in DeepSeek Host frames the tool-calling issue and asks why models fail to listen to Zod schema errors like Instructor prompts do. Ahmad explains that open models exhibit stubborn behavior due to synthetic distillation.9:19–13:27 · Guest disagreement 2/10 Deterministic Repair Logic for Agent Tool Calling Ahmad screen-shares and explains deterministic repair logic, drawing analogies to database migrations and driving lessons to show how intercepted tool errors guide models without looping.13:28–16:02 · Guest disagreement 1/10 Scaling Deterministic Repairs Across Billions of Open Model Tokens Host asks whether tool confusion is unique to DeepSeek. Ahmad explains how the pattern generalized across Kimi and Minimax models over hundreds of billions of tokens.16:03–23:10 · Guest disagreement 2/10 Eliminating AI Design Slop with Deterministic Composition Frameworks Ahmad discusses AI design slop and pattern-first composition frameworks. Host contributes domain knowledge by referencing Mario Zechner and the rationale behind CSS OKLCH color models.23:10–26:52 · Guest disagreement 1/10 Demonstrating CommandCode's Design Skill and Security Extensions Host validates Ahmad's design and engineering credibility while Ahmad demonstrates the design skill generating UI from raw data and notes extensions into automated security patching.26:52–31:49 · Guest disagreement 2/10 Continuous Preference Learning via the Taste Engine Ahmad breaks down continuous preference learning, explaining why static rules files rot and how CommandCode automatically learns micro-decisions and git habits directly into repository markdown.31:50–36:56 · Guest disagreement 1/10 Taste vs. Skills Architecture and Multi-Tier Agent Workflows Host presses for architectural clarity regarding whether Taste is a backend model or portable repository files. Ahmad clarifies the distinction between static skills and the dynamic Taste engine.36:57–39:10 · Guest disagreement 2/10 CommandCode Open-Source Roadmap and Architectural Philosophy Ahmad outlines open-sourcing CommandCode, contrasting its curated Apple-like philosophy with the broad Windows and Linux approaches of alternative harnesses. Host briefly clarifies the repository's age.39:10–40:35 · Guest disagreement 0/10 Concluding Thoughts on Open Ecosystems and Collaborative AI Progress Host notes DeepSeek's recruitment for dedicated coding tooling, and both agree that shared open-source improvements benefit all developer harnesses across the ecosystem.1:12–4:32 · The hosts pushing back 0/10 From Corona CLI to CommandCode and Taste Host invites Ahmad to share his background leading into CommandCode. Ahmad delivers an extensive overview tracing his work from early GPT-3 access and LangBase to creating the Taste engine.4:32–9:18 · The hosts pushing back 2/10 Identifying the Tool Confusion Problem in DeepSeek Host frames the tool-calling issue and asks why models fail to listen to Zod schema errors like Instructor prompts do. Ahmad explains that open models exhibit stubborn behavior due to synthetic distillation.9:19–13:27 · The hosts pushing back 0/10 Deterministic Repair Logic for Agent Tool Calling Ahmad screen-shares and explains deterministic repair logic, drawing analogies to database migrations and driving lessons to show how intercepted tool errors guide models without looping.13:28–16:02 · The hosts pushing back 0/10 Scaling Deterministic Repairs Across Billions of Open Model Tokens Host asks whether tool confusion is unique to DeepSeek. Ahmad explains how the pattern generalized across Kimi and Minimax models over hundreds of billions of tokens.16:03–23:10 · The hosts pushing back 1/10 Eliminating AI Design Slop with Deterministic Composition Frameworks Ahmad discusses AI design slop and pattern-first composition frameworks. Host contributes domain knowledge by referencing Mario Zechner and the rationale behind CSS OKLCH color models.23:10–26:52 · The hosts pushing back 1/10 Demonstrating CommandCode's Design Skill and Security Extensions Host validates Ahmad's design and engineering credibility while Ahmad demonstrates the design skill generating UI from raw data and notes extensions into automated security patching.26:52–31:49 · The hosts pushing back 0/10 Continuous Preference Learning via the Taste Engine Ahmad breaks down continuous preference learning, explaining why static rules files rot and how CommandCode automatically learns micro-decisions and git habits directly into repository markdown.31:50–36:56 · The hosts pushing back 2/10 Taste vs. Skills Architecture and Multi-Tier Agent Workflows Host presses for architectural clarity regarding whether Taste is a backend model or portable repository files. Ahmad clarifies the distinction between static skills and the dynamic Taste engine.36:57–39:10 · The hosts pushing back 1/10 CommandCode Open-Source Roadmap and Architectural Philosophy Ahmad outlines open-sourcing CommandCode, contrasting its curated Apple-like philosophy with the broad Windows and Linux approaches of alternative harnesses. Host briefly clarifies the repository's age.39:10–40:35 · The hosts pushing back 0/10 Concluding Thoughts on Open Ecosystems and Collaborative AI Progress Host notes DeepSeek's recruitment for dedicated coding tooling, and both agree that shared open-source improvements benefit all developer harnesses across the ecosystem.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 8:00 Rejecting conventional error return paradigms for open models

Ahmad forcefully dismisses standard error feedback loops, arguing that DeepSeek refuses to learn from returned schema errors due to its stubborn synthetic pretraining.

Hardest push from the hosts ▶ 31:47 Host demands clear boundary between model, memory, and skills

Alessio cuts through marketing terminology to demand whether Taste is an active model or a system of markdown files, pressing on how it differs from Anthropic-style skills.

Biggest teaching moment ▶ 9:30 Educating on deterministic repair hints over raw schema failures

Ahmad demonstrates how harness-level error recovery and repair hints stop models from looping into 50+ failed attempts per session.

The host holds their own ▶ 21:23 Host provides underlying CSS color science context

Alessio demonstrates technical depth by explaining the exact historical color space advancements in CSS that necessitated the adoption of OKLCH over HSL.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
From Corona CLI to CommandCode and Taste 1310 Host invites Ahmad to share his background leading into CommandCode. Ahmad delivers an extensive overview tracing his work from early GPT-3 access and LangBase to creating the Taste engine.
Identifying the Tool Confusion Problem in DeepSeek 4532 Host frames the tool-calling issue and asks why models fail to listen to Zod schema errors like Instructor prompts do. Ahmad explains that open models exhibit stubborn behavior due to synthetic distillation.
Deterministic Repair Logic for Agent Tool Calling 1620 Ahmad screen-shares and explains deterministic repair logic, drawing analogies to database migrations and driving lessons to show how intercepted tool errors guide models without looping.
Scaling Deterministic Repairs Across Billions of Open Model Tokens 3510 Host asks whether tool confusion is unique to DeepSeek. Ahmad explains how the pattern generalized across Kimi and Minimax models over hundreds of billions of tokens.
Eliminating AI Design Slop with Deterministic Composition Frameworks 5421 Ahmad discusses AI design slop and pattern-first composition frameworks. Host contributes domain knowledge by referencing Mario Zechner and the rationale behind CSS OKLCH color models.
Demonstrating CommandCode's Design Skill and Security Extensions 3411 Host validates Ahmad's design and engineering credibility while Ahmad demonstrates the design skill generating UI from raw data and notes extensions into automated security patching.
Continuous Preference Learning via the Taste Engine 1620 Ahmad breaks down continuous preference learning, explaining why static rules files rot and how CommandCode automatically learns micro-decisions and git habits directly into repository markdown.
Taste vs. Skills Architecture and Multi-Tier Agent Workflows 5512 Host presses for architectural clarity regarding whether Taste is a backend model or portable repository files. Ahmad clarifies the distinction between static skills and the dynamic Taste engine.
CommandCode Open-Source Roadmap and Architectural Philosophy 2321 Ahmad outlines open-sourcing CommandCode, contrasting its curated Apple-like philosophy with the broad Windows and Linux approaches of alternative harnesses. Host briefly clarifies the repository's age.
Concluding Thoughts on Open Ecosystems and Collaborative AI Progress 3200 Host notes DeepSeek's recruitment for dedicated coding tooling, and both agree that shared open-source improvements benefit all developer harnesses across the ecosystem.

Statements from this episode (19)

Assertion Not checkable as stated
Awais: LangBase grew to 1.2 billion agent runs per month
“It was called LangBase. It grew quite big 1.2 billion agent runs a month ended up building a memory infrastructure or whatnot.”
Ahmad Awais Jun 6, 2026 ▶ 2:10
Insight
Awais: There is only one type of AI agent, the coding agent
“There's only one type of agent, and that is a coding agent. It can do it all, right?”
Ahmad Awais Jun 6, 2026 ▶ 2:18
Assertion Not checkable as stated
Awais: Deterministic tool repair with hints fixes model tool-calling loops
“What we saw is the moment you send the result with the repair logic, right after that, the third tool call is fixed. Instead of, you know, it all of a sudden becomes super smart. It understands like, okay, I got the result, what I was looking for, and I'm gonn…”
Ahmad Awais Jun 6, 2026 ▶ 11:04
Assertion Not checkable as stated
Awais: Claude Code hides 50+ tool-call failures per session on DeepSeek
“In CloudCode you know, they hide a lot of the errors behind control O, right? So you don't even know that, you know, you have like 50 plus tool call failures plus per session. You're just sitting there and you're like, oh, why is DeepSeq so slow?”
Ahmad Awais Jun 6, 2026 ▶ 12:21
Insight
Awais: Open LLM coding failures are harness issues, not model issues
“So I feel like this always ends up being a tool call, a hardness issue. Then, you know, an actual model issue.”
Ahmad Awais Jun 6, 2026 ▶ 12:58
Disclosure
Awais: CommandCode has 16,000 tool repair variations across 600B tokens
“And now we have like 16,000 different repair, you know, variations across hundreds of billions of tokens. We are doing anywhere from six hundred billion tokens right now.”
Ahmad Awais Jun 6, 2026 ▶ 14:01
Insight
Awais: Interactive permission prompts degrade coding agent model performance
“If you run any coding agent with permissions on, The models are actually number. And if you run them without, you know, the complete bypass of permissions, they do much better. Even if you like sit through those yes, yes, yes, accept or whatnot, you will see t…”
Ahmad Awais Jun 6, 2026 ▶ 14:52
Insight
Awais: Reducing tool call errors enables LLMs to sustain longer exploration
“Like if they are seeing a lot less tool call errors, they're much more creative. They are, they can explore a lot and they can continue a lot longer.”
Ahmad Awais Jun 6, 2026 ▶ 15:23
Disclosure
CommandCode offers 600 million DeepSeek tokens for $1 per month
“We launched a Go plan with just dollar one per month to, where you can do like, six hundred million tokens of DeepSQL for Pro in it, just to prove like, open models are actually really, really good, and they are catching up, right?”
Ahmad Awais Jun 6, 2026 ▶ 16:44
Insight
Awais: Forcing LLMs to use OKLCH significantly improves color palette control
“I personally don't use OKLCH, but apparently LLMs are really good at it. And if you see them using HSL or something, they are, they don't actually are able to control the lightness in HSL very quickly, but on to human eye, it's very, very easy to see like this…”
Ahmad Awais Jun 6, 2026 ▶ 20:43
Insight
Awais: AI design slop is a contract gap, not a model capability deficit
“Feels like you can fix 90% of a design slop, which is not a capability gap. It's more like a contract gap in what your harness is telling an LLM to do versus what your user is saying.”
Ahmad Awais Jun 6, 2026 ▶ 22:24
Insight
Awais: Design is now the key differentiator in software as building gets commoditized
“Like a lot of people are right now able to build just about anything. And we are now differentiating between good work and bad work based on their design.”
Ahmad Awais Jun 6, 2026 ▶ 23:37
Assertion Not checkable as stated
Awais: Claude tolerates tool errors and self-corrects, unlike open models
“Claude is actually really, really lenient for tool calls. So even if, you know, your coding agent harness messes up, it can figure out that, oh, I'm being sent this error and can fix itself. Not the case with you know open models”
Ahmad Awais Jun 6, 2026 ▶ 27:12
Insight
Awais: Skills files shouldn't duplicate knowledge LLMs already have
“If an LLM already knows about something, it should not end up in your, you know, skill or taste file. That is absolutely useless context, right?”
Ahmad Awais Jun 6, 2026 ▶ 28:15
Assertion Not checkable as stated
Awais: 70+ developer study showed CommandCode Taste cut edits and steering
“We ran a study with like 70 plus developers in the number of times that they had to go edit files because their LLM made a different, you know, the scene took a different turn, or steer their LLM like, yeah, don't do this, don't use this, don't use TRPC or som…”
Ahmad Awais Jun 6, 2026 ▶ 31:20
Assertion Not checkable as stated
Awais: Devs bootstrap Taste files with frontier models, then execute cheaply
“A lot of people, what they're doing is they're building one project with a really high quality LLM, like Opus or GPT, 5.5, right? They're building a taste file, and then they're using, you know, super cheap models to continuously build on that more.”
Ahmad Awais Jun 6, 2026 ▶ 35:59
Disclosure
Awais: CommandCode will open-source its codebase soon
“So we are going to open source command code very, very soon. I'm hoping we can announce that on the AI engineering conference, NSF.”
Ahmad Awais Jun 6, 2026 ▶ 37:04
Disclosure
Awais: Matt Mullenweg is an angel investor in CommandCode
“Matt is actually one of our angels now. When he heard that we are open sourcing command code, he reached out.”
Ahmad Awais Jun 6, 2026 ▶ 37:55
Assertion Not checkable as stated
Awais: Claude 3.7 Max is already CommandCode's second most used model
“But they will only be for deep seek when to 3.7 max is the second most used model on command code right now. It's just two or three days old.”
Ahmad Awais Jun 6, 2026 ▶ 39:43
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.