Oct 25, 2023 · 44m · mad

Diffblue’s AI Testing Paradigm Shift — CEO Mathew Lodge Explains How Code Writes Itself

Mathew Lodge · 29m spoken Matt Turck · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

On The MAD Podcast, Diffblue CEO Mathew Lodge explains how Reinforcement Learning enables autonomous unit test generation, revolutionizing enterprise legacy software modernization and developer productivity beyond traditional Large Language Models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 24.1% of the talking time here. How this is scored →

Matt as informed peer 3.9 Guest teaching 3.4 Guest disagreement 1.6 Matt pushing back 1.6
05100:0015:0030:000:00–9:53 · Matt as informed peer 3/10 Opening Titles and FirstMark Branding Sequence The host asks a broad intro question, allowing the guest to deliver an extended tutorial on LLM predictive completion vs Reinforcement Learning search spaces. The host interjects briefly to steer interview structure before asking a solid technical follow-up about RLHF and hybrid model ensembles.9:53–18:09 · Matt as informed peer 2/10 Understanding Unit Testing and Legacy QA Challenges The host deliberately asks basic ELI5 questions to clarify unit testing concepts for a general audience. The guest walks through shift-left practices, legacy QA bottlenecks, and the friction of developer testing.18:09–23:27 · Matt as informed peer 4/10 Autonomous Test Generation and Compute Efficiency The host demonstrates good tech intuition by mapping call graphs to visual mapping exercises and inquiring about CPU vs GPU compute overhead. The guest explains Diffblue's iterative test generation process.23:27–27:56 · Matt as informed peer 5/10 Enterprise Adoption, COBOL Migration, and Commercial Strategy The host displays strong domain knowledge by bringing up COBOL migration challenges in financial infrastructure and enterprise software buyer personas. The exchange is highly collaborative as the guest outlines Diffblue's enterprise GTM.27:56–33:25 · Matt as informed peer 4/10 Developer Experience, AI Autonomy, and Impact on Software Jobs The host raises potential developer friction regarding autonomous AI and quotes a Citibank executive regarding job impact. The guest reframes autonomy using a historical analogy of assembly language transitioning to compiled high-level languages.33:25–39:28 · Matt as informed peer 5/10 Industry Perspectives: Prompt Engineering, Model Size, and Open Source The host engages with guest's public Twitter commentary regarding AI trends. The guest delivers strong contrarian views, dismissing prompt engineering as trial-and-error and explaining why smaller open-source models will win.39:28–42:50 · Matt as informed peer 4/10 The UK AI Startup Ecosystem and Global Talent The host highlights notable UK tech startups, leading into a discussion on talent pools. The guest offers a candid critique of Brexit's impact on European tech talent acquisition.0:00–9:53 · Guest teaching 5/10 Opening Titles and FirstMark Branding Sequence The host asks a broad intro question, allowing the guest to deliver an extended tutorial on LLM predictive completion vs Reinforcement Learning search spaces. The host interjects briefly to steer interview structure before asking a solid technical follow-up about RLHF and hybrid model ensembles.9:53–18:09 · Guest teaching 4/10 Understanding Unit Testing and Legacy QA Challenges The host deliberately asks basic ELI5 questions to clarify unit testing concepts for a general audience. The guest walks through shift-left practices, legacy QA bottlenecks, and the friction of developer testing.18:09–23:27 · Guest teaching 3/10 Autonomous Test Generation and Compute Efficiency The host demonstrates good tech intuition by mapping call graphs to visual mapping exercises and inquiring about CPU vs GPU compute overhead. The guest explains Diffblue's iterative test generation process.23:27–27:56 · Guest teaching 2/10 Enterprise Adoption, COBOL Migration, and Commercial Strategy The host displays strong domain knowledge by bringing up COBOL migration challenges in financial infrastructure and enterprise software buyer personas. The exchange is highly collaborative as the guest outlines Diffblue's enterprise GTM.27:56–33:25 · Guest teaching 4/10 Developer Experience, AI Autonomy, and Impact on Software Jobs The host raises potential developer friction regarding autonomous AI and quotes a Citibank executive regarding job impact. The guest reframes autonomy using a historical analogy of assembly language transitioning to compiled high-level languages.33:25–39:28 · Guest teaching 4/10 Industry Perspectives: Prompt Engineering, Model Size, and Open Source The host engages with guest's public Twitter commentary regarding AI trends. The guest delivers strong contrarian views, dismissing prompt engineering as trial-and-error and explaining why smaller open-source models will win.39:28–42:50 · Guest teaching 2/10 The UK AI Startup Ecosystem and Global Talent The host highlights notable UK tech startups, leading into a discussion on talent pools. The guest offers a candid critique of Brexit's impact on European tech talent acquisition.0:00–9:53 · Guest disagreement 1/10 Opening Titles and FirstMark Branding Sequence The host asks a broad intro question, allowing the guest to deliver an extended tutorial on LLM predictive completion vs Reinforcement Learning search spaces. The host interjects briefly to steer interview structure before asking a solid technical follow-up about RLHF and hybrid model ensembles.9:53–18:09 · Guest disagreement 1/10 Understanding Unit Testing and Legacy QA Challenges The host deliberately asks basic ELI5 questions to clarify unit testing concepts for a general audience. The guest walks through shift-left practices, legacy QA bottlenecks, and the friction of developer testing.18:09–23:27 · Guest disagreement 1/10 Autonomous Test Generation and Compute Efficiency The host demonstrates good tech intuition by mapping call graphs to visual mapping exercises and inquiring about CPU vs GPU compute overhead. The guest explains Diffblue's iterative test generation process.23:27–27:56 · Guest disagreement 1/10 Enterprise Adoption, COBOL Migration, and Commercial Strategy The host displays strong domain knowledge by bringing up COBOL migration challenges in financial infrastructure and enterprise software buyer personas. The exchange is highly collaborative as the guest outlines Diffblue's enterprise GTM.27:56–33:25 · Guest disagreement 2/10 Developer Experience, AI Autonomy, and Impact on Software Jobs The host raises potential developer friction regarding autonomous AI and quotes a Citibank executive regarding job impact. The guest reframes autonomy using a historical analogy of assembly language transitioning to compiled high-level languages.33:25–39:28 · Guest disagreement 3/10 Industry Perspectives: Prompt Engineering, Model Size, and Open Source The host engages with guest's public Twitter commentary regarding AI trends. The guest delivers strong contrarian views, dismissing prompt engineering as trial-and-error and explaining why smaller open-source models will win.39:28–42:50 · Guest disagreement 2/10 The UK AI Startup Ecosystem and Global Talent The host highlights notable UK tech startups, leading into a discussion on talent pools. The guest offers a candid critique of Brexit's impact on European tech talent acquisition.0:00–9:53 · Matt pushing back 3/10 Opening Titles and FirstMark Branding Sequence The host asks a broad intro question, allowing the guest to deliver an extended tutorial on LLM predictive completion vs Reinforcement Learning search spaces. The host interjects briefly to steer interview structure before asking a solid technical follow-up about RLHF and hybrid model ensembles.9:53–18:09 · Matt pushing back 1/10 Understanding Unit Testing and Legacy QA Challenges The host deliberately asks basic ELI5 questions to clarify unit testing concepts for a general audience. The guest walks through shift-left practices, legacy QA bottlenecks, and the friction of developer testing.18:09–23:27 · Matt pushing back 1/10 Autonomous Test Generation and Compute Efficiency The host demonstrates good tech intuition by mapping call graphs to visual mapping exercises and inquiring about CPU vs GPU compute overhead. The guest explains Diffblue's iterative test generation process.23:27–27:56 · Matt pushing back 1/10 Enterprise Adoption, COBOL Migration, and Commercial Strategy The host displays strong domain knowledge by bringing up COBOL migration challenges in financial infrastructure and enterprise software buyer personas. The exchange is highly collaborative as the guest outlines Diffblue's enterprise GTM.27:56–33:25 · Matt pushing back 2/10 Developer Experience, AI Autonomy, and Impact on Software Jobs The host raises potential developer friction regarding autonomous AI and quotes a Citibank executive regarding job impact. The guest reframes autonomy using a historical analogy of assembly language transitioning to compiled high-level languages.33:25–39:28 · Matt pushing back 2/10 Industry Perspectives: Prompt Engineering, Model Size, and Open Source The host engages with guest's public Twitter commentary regarding AI trends. The guest delivers strong contrarian views, dismissing prompt engineering as trial-and-error and explaining why smaller open-source models will win.39:28–42:50 · Matt pushing back 1/10 The UK AI Startup Ecosystem and Global Talent The host highlights notable UK tech startups, leading into a discussion on talent pools. The guest offers a candid critique of Brexit's impact on European tech talent acquisition.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 38.3% · guest 61.7%0:00 · Matt 38.3% · guest 61.7%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 22.4% · guest 77.6%6:00 · Matt 22.4% · guest 77.6%9:00 · Matt 25.8% · guest 74.2%9:00 · Matt 25.8% · guest 74.2%12:00 · Matt 11.8% · guest 88.2%12:00 · Matt 11.8% · guest 88.2%15:00 · Matt 9% · guest 91%15:00 · Matt 9% · guest 91%18:00 · Matt 6.6% · guest 93.4%18:00 · Matt 6.6% · guest 93.4%21:00 · Matt 45.9% · guest 54.1%21:00 · Matt 45.9% · guest 54.1%24:00 · Matt 36% · guest 64%24:00 · Matt 36% · guest 64%27:00 · Matt 24.7% · guest 75.3%27:00 · Matt 24.7% · guest 75.3%30:00 · Matt 22.2% · guest 77.8%30:00 · Matt 22.2% · guest 77.8%33:00 · Matt 27.6% · guest 72.4%33:00 · Matt 27.6% · guest 72.4%36:00 · Matt 16.5% · guest 83.5%36:00 · Matt 16.5% · guest 83.5%39:00 · Matt 46.3% · guest 53.7%39:00 · Matt 46.3% · guest 53.7%42:00 · Matt 32.4% · guest 67.6%42:00 · Matt 32.4% · guest 67.6%
Sharpest disagreement ▶ 33:52 Dismissal of Prompt Engineering

The guest forcefully rejects the term 'prompt engineering', arguing that applying traditional engineering logic to unpredictable LLMs is a complete fallacy.

Hardest push from Matt ▶ 6:53 Host halts premature deep dive

The host firmly cuts off the guest mid-sentence to hold off on unit testing details and keep the interview structure on track.

Biggest teaching moment ▶ 3:20 Search spaces in Chess vs Go

The guest educates the host on reinforcement learning fundamentals by contrasting chess brute-force search with Go's massive search space requiring probabilistic predictions.

Matt holds his own ▶ 23:44 COBOL legacy infrastructure context

The host demonstrates deep knowledge of enterprise tech debt by citing COBOL's persistent dominance in global banking backends.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Opening Titles and FirstMark Branding Sequence 3513 The host asks a broad intro question, allowing the guest to deliver an extended tutorial on LLM predictive completion vs Reinforcement Learning search spaces. The host interjects briefly to steer interview structure before asking a solid technical follow-up about RLHF and hybrid model ensembles.
Understanding Unit Testing and Legacy QA Challenges 2411 The host deliberately asks basic ELI5 questions to clarify unit testing concepts for a general audience. The guest walks through shift-left practices, legacy QA bottlenecks, and the friction of developer testing.
Autonomous Test Generation and Compute Efficiency 4311 The host demonstrates good tech intuition by mapping call graphs to visual mapping exercises and inquiring about CPU vs GPU compute overhead. The guest explains Diffblue's iterative test generation process.
Enterprise Adoption, COBOL Migration, and Commercial Strategy 5211 The host displays strong domain knowledge by bringing up COBOL migration challenges in financial infrastructure and enterprise software buyer personas. The exchange is highly collaborative as the guest outlines Diffblue's enterprise GTM.
Developer Experience, AI Autonomy, and Impact on Software Jobs 4422 The host raises potential developer friction regarding autonomous AI and quotes a Citibank executive regarding job impact. The guest reframes autonomy using a historical analogy of assembly language transitioning to compiled high-level languages.
Industry Perspectives: Prompt Engineering, Model Size, and Open Source 5432 The host engages with guest's public Twitter commentary regarding AI trends. The guest delivers strong contrarian views, dismissing prompt engineering as trial-and-error and explaining why smaller open-source models will win.
The UK AI Startup Ecosystem and Global Talent 4221 The host highlights notable UK tech startups, leading into a discussion on talent pools. The guest offers a candid critique of Brexit's impact on European tech talent acquisition.

Statements from this episode (12)

Assertion Partly supported
Lodge: DeepMind shaved 7% off C++ sort algorithm using reinforcement learning
“Google DeepMind has also applied this approach Do things like code optimization, where essentially they're searching for more efficient implementations of a particular algorithm. So they shaved about seven percent off the sort algorithm in the C++ library, whi…”
Mathew Lodge Oct 25, 2023 ▶ 7:37
Disclosure
Lodge: Diffblue uses ensemble of reinforcement learning and LLMs
“Yeah, we, so we use an ensemble approach. So we've got different aspects of the models and we've obviously been playing with large language models to see if we can improve on that process. So some of this is like, how do you make predictions? And you can use l…”
Mathew Lodge Oct 25, 2023 ▶ 8:36
Insight
Mathew Lodge: Developers dislike unit tests because incentives favor feature delivery
“Fundamentally they're not paid to write tests. They're paid to deliver the functionality, the new things in the application that fixes the application, the updates and tests are just there to, as part of the process to help them.”
Mathew Lodge Oct 25, 2023 ▶ 16:49
Assertion Not checkable as stated
Diffblue generated 3,100 unit tests in eight hours for a bank
“You know, for one of the banks, they had an application. We wrote like 3100 tests in about eight hours. Now they estimated that that's like a year's work for a single developer, 3000 tests. And we can do that in eight hours.”
Mathew Lodge Oct 25, 2023 ▶ 21:33
Assertion Partly supported
Private equity mainframe software acquisitions and price hikes drove companies to Java
“Because PE companies bought all of those software products and raised the price.”
Mathew Lodge Oct 25, 2023 ▶ 25:19
Insight
Lodge: AI code generation mirrors the shift from assembly to compiled languages
“The analogy I like to make is, it's kind of like, ah, the shift, you know, 30, ah, 35 years ago from sort of assembly language to compiled languages, right? So what you did as a developer really changed”
Mathew Lodge Oct 25, 2023 ▶ 31:30
Insight
Interactive AI code suggestions are strictly bottlenecked by human review speed
“You could really still only go as fast as a human can go, because a human has to review that, decide whether it's useful or not, if it is useful, make it, fit it into the code and finish the work. You're still basically going at the speed of a human, whereas w…”
Mathew Lodge Oct 25, 2023 ▶ 32:44
Opinion
Mathew Lodge argues prompt engineering is a fallacy and merely trial-and-error
“And so the idea that there's something called prompt engineering, engineering in the sense of you're applying an approach because it's going to give you the output that you desire. Is a complete fallacy. At best it's trial and error. It's prompt trial and erro…”
Mathew Lodge Oct 25, 2023 ▶ 34:49
Assertion Not checkable as stated
Diffblue generates unit tests in 1.5 seconds versus 40 seconds for LLMs
“We built a version of our product where we use the large language model to generate all the tests. So we took our reinforcement lending engine out and put in large language model, and we did a lot of work on prompts and so on. And you know, our product will wr…”
Mathew Lodge Oct 25, 2023 ▶ 36:13
Prediction Not checkable as stated
The AI market will shift toward smaller models running without GPUs
“It's one of the reasons I think that we're going to see more smaller models and you've got startups now emerging, which are essentially trying to, like, distill down, like from a large model into something much smaller that will run without a GPU, and they're …”
Mathew Lodge Oct 25, 2023 ▶ 37:02
Prediction Not checkable as stated
Lodge: Open-source AI will ultimately defeat proprietary models
“I think so. I think because it allows that level of innovation I just talked about, and I think also it gives people the ability to specialize these models to a particular task.”
Mathew Lodge Oct 25, 2023 ▶ 38:42
Opinion
Lodge: Brexit caused an exodus of UK tech talent
“We had, you know, we had a Brexit exodus of talent where people just said, you know what, I'm just going, I'm going back to France or I'm going back to, you know, wherever.”
Mathew Lodge Oct 25, 2023 ▶ 42:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.