Jan 22, 2020 · 21m · mad

Production AI: Lessons Learned the Hard Way // Adam Wenchel, Arthur.ai (FirstMark's Data Driven NYC)

Adam Wenchel · 17m spoken Matt Turck · 45s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At FirstMark's Data Driven NYC, Arthur AI CEO Adam Wenchel presents 'Production AI: Lessons Learned the Hard Way,' detailing why real-world machine learning models degrade after deployment and how enterprise monitoring and explainability guardrails bridge the gap between lab performance and production reliability.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.9% of the talking time here. How this is scored →

Matt as informed peer 0.7 Guest teaching 2.0 Guest disagreement 0.2 Matt pushing back 0.5
05100:0010:0020:000:14–3:19 · Matt as informed peer 0/10 Speaker Background and Audience Engagement Adam Wenchel opens with his background at DARPA and Capital One before introducing the topic. As this is a monologue presentation, host-side scores are zero.3:19–6:27 · Matt as informed peer 0/10 Bad Decisions, Model Degradation, and Operational Gaps Wenchel details how AI models degrade in production due to changing macroeconomic conditions and upstream data format changes. The segment remains a monologue, keeping host scores at zero.6:27–9:35 · Matt as informed peer 0/10 Inappropriate Decisions, Bias, and Regulatory Compliance Wenchel addresses issues of regulatory compliance, bias monitoring, and trust in deployed models. Host involvement is zero during this presentation section.9:35–12:50 · Matt as informed peer 0/10 Case Study: Dumbarton Oaks Harvard Research Wenchel presents a Harvard Dumbarton Oaks case study showing how computer vision explainability identified vegetation misclassifications in nave architecture. Host scores remain zero.12:50–15:13 · Matt as informed peer 0/10 Case Study: US Air Force Supply Chain Optimization Wenchel covers the US Air Force supply chain optimization case study and concludes his deck with a hiring pitch. The segment is entirely guest-led with no host interaction.15:13–21:19 · Matt as informed peer 4/10 Q&A: Platform Architecture and Model Agnosticism Host Matt Turck and audience members ask targeted Q&A questions regarding data ingestion, benchmark comparison, model agnosticism, and dataset access. Turck demonstrates strong technical understanding when questioning data access limits.0:14–3:19 · Guest teaching 1/10 Speaker Background and Audience Engagement Adam Wenchel opens with his background at DARPA and Capital One before introducing the topic. As this is a monologue presentation, host-side scores are zero.3:19–6:27 · Guest teaching 2/10 Bad Decisions, Model Degradation, and Operational Gaps Wenchel details how AI models degrade in production due to changing macroeconomic conditions and upstream data format changes. The segment remains a monologue, keeping host scores at zero.6:27–9:35 · Guest teaching 2/10 Inappropriate Decisions, Bias, and Regulatory Compliance Wenchel addresses issues of regulatory compliance, bias monitoring, and trust in deployed models. Host involvement is zero during this presentation section.9:35–12:50 · Guest teaching 2/10 Case Study: Dumbarton Oaks Harvard Research Wenchel presents a Harvard Dumbarton Oaks case study showing how computer vision explainability identified vegetation misclassifications in nave architecture. Host scores remain zero.12:50–15:13 · Guest teaching 2/10 Case Study: US Air Force Supply Chain Optimization Wenchel covers the US Air Force supply chain optimization case study and concludes his deck with a hiring pitch. The segment is entirely guest-led with no host interaction.15:13–21:19 · Guest teaching 3/10 Q&A: Platform Architecture and Model Agnosticism Host Matt Turck and audience members ask targeted Q&A questions regarding data ingestion, benchmark comparison, model agnosticism, and dataset access. Turck demonstrates strong technical understanding when questioning data access limits.0:14–3:19 · Guest disagreement 0/10 Speaker Background and Audience Engagement Adam Wenchel opens with his background at DARPA and Capital One before introducing the topic. As this is a monologue presentation, host-side scores are zero.3:19–6:27 · Guest disagreement 0/10 Bad Decisions, Model Degradation, and Operational Gaps Wenchel details how AI models degrade in production due to changing macroeconomic conditions and upstream data format changes. The segment remains a monologue, keeping host scores at zero.6:27–9:35 · Guest disagreement 0/10 Inappropriate Decisions, Bias, and Regulatory Compliance Wenchel addresses issues of regulatory compliance, bias monitoring, and trust in deployed models. Host involvement is zero during this presentation section.9:35–12:50 · Guest disagreement 0/10 Case Study: Dumbarton Oaks Harvard Research Wenchel presents a Harvard Dumbarton Oaks case study showing how computer vision explainability identified vegetation misclassifications in nave architecture. Host scores remain zero.12:50–15:13 · Guest disagreement 0/10 Case Study: US Air Force Supply Chain Optimization Wenchel covers the US Air Force supply chain optimization case study and concludes his deck with a hiring pitch. The segment is entirely guest-led with no host interaction.15:13–21:19 · Guest disagreement 1/10 Q&A: Platform Architecture and Model Agnosticism Host Matt Turck and audience members ask targeted Q&A questions regarding data ingestion, benchmark comparison, model agnosticism, and dataset access. Turck demonstrates strong technical understanding when questioning data access limits.0:14–3:19 · Matt pushing back 0/10 Speaker Background and Audience Engagement Adam Wenchel opens with his background at DARPA and Capital One before introducing the topic. As this is a monologue presentation, host-side scores are zero.3:19–6:27 · Matt pushing back 0/10 Bad Decisions, Model Degradation, and Operational Gaps Wenchel details how AI models degrade in production due to changing macroeconomic conditions and upstream data format changes. The segment remains a monologue, keeping host scores at zero.6:27–9:35 · Matt pushing back 0/10 Inappropriate Decisions, Bias, and Regulatory Compliance Wenchel addresses issues of regulatory compliance, bias monitoring, and trust in deployed models. Host involvement is zero during this presentation section.9:35–12:50 · Matt pushing back 0/10 Case Study: Dumbarton Oaks Harvard Research Wenchel presents a Harvard Dumbarton Oaks case study showing how computer vision explainability identified vegetation misclassifications in nave architecture. Host scores remain zero.12:50–15:13 · Matt pushing back 0/10 Case Study: US Air Force Supply Chain Optimization Wenchel covers the US Air Force supply chain optimization case study and concludes his deck with a hiring pitch. The segment is entirely guest-led with no host interaction.15:13–21:19 · Matt pushing back 3/10 Q&A: Platform Architecture and Model Agnosticism Host Matt Turck and audience members ask targeted Q&A questions regarding data ingestion, benchmark comparison, model agnosticism, and dataset access. Turck demonstrates strong technical understanding when questioning data access limits.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 27.6% · guest 72.4%15:00 · Matt 27.6% · guest 72.4%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 15.7% · guest 84.3%21:00 · Matt 15.7% · guest 84.3%
Sharpest disagreement ▶ 20:50 Wenchel Playfully Critiques Academic Open Source Code

Wenchel lightheartedly dismisses open-source explainability scripts like Lime as unscalable code written by PhDs purely to support academic papers.

Hardest push from Matt ▶ 17:10 Matt Turck Challenges Data Access Feasibility

Host Matt Turck pushes Wenchel on whether Arthur requires access to raw underlying data, forcing Wenchel to clarify proxy metrics versus ground truth monitoring.

Biggest teaching moment ▶ 17:31 Wenchel Differentiates Delayed Outcomes vs Instant Feedback

Wenchel educates the host and audience on how to monitor models like credit underwriting where ground truth outcomes take years to materialise.

Matt holds his own ▶ 15:20 Matt Turck Questions Model Agnosticism Across Frameworks

Host Matt Turck demonstrates domain knowledge by asking whether Arthur's architecture can handle diverse paradigms like deep learning, NLP, and tabular data.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Speaker Background and Audience Engagement 0100 Adam Wenchel opens with his background at DARPA and Capital One before introducing the topic. As this is a monologue presentation, host-side scores are zero.
Bad Decisions, Model Degradation, and Operational Gaps 0200 Wenchel details how AI models degrade in production due to changing macroeconomic conditions and upstream data format changes. The segment remains a monologue, keeping host scores at zero.
Inappropriate Decisions, Bias, and Regulatory Compliance 0200 Wenchel addresses issues of regulatory compliance, bias monitoring, and trust in deployed models. Host involvement is zero during this presentation section.
Case Study: Dumbarton Oaks Harvard Research 0200 Wenchel presents a Harvard Dumbarton Oaks case study showing how computer vision explainability identified vegetation misclassifications in nave architecture. Host scores remain zero.
Case Study: US Air Force Supply Chain Optimization 0200 Wenchel covers the US Air Force supply chain optimization case study and concludes his deck with a hiring pitch. The segment is entirely guest-led with no host interaction.
Q&A: Platform Architecture and Model Agnosticism 4313 Host Matt Turck and audience members ask targeted Q&A questions regarding data ingestion, benchmark comparison, model agnosticism, and dataset access. Turck demonstrates strong technical understanding when questioning data access limits.

Statements from this episode (12)

Insight
Real-world AI deployments introduce distinct failure modes beyond lab environments
“And it's not only hard to develop in the lab, but once you develop it and put it in the real world, there's a whole new set of categories of ways it can go wrong.”
Adam Wenchel Jan 22, 2020 ▶ 1:49
Assertion Supported
Historical anti-bias regulations apply directly to AI models
“There's a lot of historical regulation around anti-discrimination and bias and things like that that, ah, certainly applies just as much to AI models as it does to humans and more simple analytical models.”
Adam Wenchel Jan 22, 2020 ▶ 2:35
Disclosure
Lack of trust delays enterprise AI deployments for months
“We, you know, encounter this all the time, where in organizations, they have these big plans for AI but they're just, They're unsure about actually deploying them and turning them on, and things get held up for months and months and months because of that.”
Adam Wenchel Jan 22, 2020 ▶ 3:06
Assertion Not checkable as stated
Most companies using AI suffer from unpublicized model failures
“Every company that's doing anything substantive with AI probably suffers from any of these problems, it's just a few of them have actually, ah, made the headlines for it”
Adam Wenchel Jan 22, 2020 ▶ 3:27
Insight
Deployed AI models suffer immediate performance gaps and ongoing degradation
“The second you put it in the real world, models, ah, number one, there's a gap right from day one, and they get worse over time.”
Adam Wenchel Jan 22, 2020 ▶ 4:00
Insight
Not collecting protected class data does not prevent algorithmic bias
“What's happened a lot in the past is people have kind of like taken the head in the sand approach where they've sort of said like, oh you know, we're not even collecting protected classes, so we can't possibly be biased as far as we know, and that's no longer …”
Adam Wenchel Jan 22, 2020 ▶ 6:56
Insight
A fractional drop in AI performance can cost hundreds of millions
“Even if your model encounters some sort of issue that drops at a couple 10th of a percent in performance, that can literally be hundreds of millions of dollars over time, and so the ROI on having that kind of monitoring in place is, is huge.”
Adam Wenchel Jan 22, 2020 ▶ 8:46
Insight
Enterprise AI adoption fails without risk mitigation despite performance gains
“Especially large traditional enterprises there tend to be very consensus-driven cultures by nature, and so the, even if people, if a data scientist can demonstrate they have a model that, you know, generally, like, gets some huge five or 10% lift, which, you k…”
Adam Wenchel Jan 22, 2020 ▶ 9:00
Insight
Visual AI explainability helps build trust in skeptical academic communities
“Ah, and then the other thing is, you can imagine this world of humanities research has not changed a whole lot in the, like, the last 200 years of study, and so bringing this sort of innovation, like, we can automate this and computers can find patterns that w…”
Adam Wenchel Jan 22, 2020 ▶ 12:18
Insight
Early safety guardrails enable companies to pursue more aggressive AI strategies
“Like the more you can build these guardrails in from kind of day one of your AI projects, when you, it allows you to be more aggressive, right? Just like the safety systems on an F-one car allow you to lap faster. If you build this stuff in from the beginning …”
Adam Wenchel Jan 22, 2020 ▶ 14:23
Assertion Not checkable as stated
Credit underwriting AI feedback loops take three to four years
“There's other ones like, ah, underwriting credit cards, where you might not know for three or four years whether you should have given that person a credit card, right?”
Adam Wenchel Jan 22, 2020 ▶ 17:44
Assertion Not checkable as stated
Open-source explainable AI tools like LIME and SHAP fail enterprise scale
“If you look at the open source components they're not very scalable. They're not easy to deploy at scale. They're really, they're useful, like, if you're a data scientist, and you have your Jupyter notebook, and you're, you know, running an experiment locally,…”
Adam Wenchel Jan 22, 2020 ▶ 20:43
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.