Apr 26, 2024 · 38m · saastr

How Shopify Implements AI Across Sales and Product with the Head of AI at Shopify

Mike Tamir · 24m spoken Rudina Seseri · 10m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this fireside chat, Rudina Seseri of Glasswing Ventures and Mike Tamir, Head of AI at Shopify, discuss the strategic, architectural, and organizational frameworks required to successfully deploy artificial intelligence and machine learning at enterprise scale. They explore practical lessons from Shopify's e-commerce search pipelines, compute infrastructure economics, data curation strategies, and cross-functional alignment between engineering loss functions and commercial ROI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Jason as informed peer 5.0 Guest teaching 5.0 Guest disagreement 0.1 Jason pushing back 0.9
05100:0010:0020:0030:001:29–6:26 · Jason as informed peer 7/10 Enterprise AI Adoption and the Ambient AI Framework Rudina delivers a comprehensive framework outlining enterprise AI adoption, citing specific enterprise case studies like Klaviyo's $40M contribution and introducing the 'ambient AI' paradigm. Tamir readily agrees with her foundational pillars of data, infrastructure, and culture.6:28–9:14 · Jason as informed peer 2/10 Data Challenges, Iterative Modeling, and Overfitting Tamir details the mechanics of iterative model training, explaining loss minimization and the nuances of out-of-sample overfitting. The host remains in listening mode as the guest breaks down core machine learning concepts.9:15–12:01 · Jason as informed peer 4/10 Translating Business Goals into Machine Learning Metrics The host poses a practical question on why data processing takes weeks for P&L owners. Tamir educates on the gap between business objectives and mathematical loss functions as well as the necessity of curated training examples.12:02–16:18 · Jason as informed peer 6/10 AI Infrastructure Requirements and GPU Resource Management The host challenges Tamir with a curveball regarding cloud aggregators like AWS Bedrock versus direct model providers, framing cloud costs as a variable fixed line. Tamir explains latency tradeoffs and GPU reservation bottlenecks.16:19–18:34 · Jason as informed peer 3/10 Fostering a Culture of Rigorous ML Evaluation Tamir discusses the shift from manual architecture design to fine-tuning pre-trained models and the danger of non-experts deploying models without stratified evaluation. The host listens attentively.18:35–24:02 · Jason as informed peer 6/10 Bridging the Gap Between Technical Metrics and Business Value Rudina interjects to emphasize that metric significance varies wildly across business contexts. Tamir builds on this by detailing search relevance at Shopify, differentiating direct product matches from complementary cross-sell relevance.24:03–26:23 · Jason as informed peer 5/10 Measuring ROI and Cross-Functional ML Product Collaboration Rudina probes on measuring concrete ROI and navigating trade-offs between engineering resources and business results. Tamir describes breaking down departmental silos so product managers and ML scientists co-design learning objectives.26:24–29:33 · Jason as informed peer 6/10 Understanding Vectorization and Semantic Embeddings Tamir explains vector spaces and semantic embeddings using the contextual meaning of the word 'bank'. Rudina synthesizes this by explaining that embeddings constitute roughly 80% of the backbone of generative AI applications.29:33–34:42 · Jason as informed peer 6/10 Scaling Search Architectures and Dynamic Pricing Strategies Tamir gently reframes the pricing question from dynamic pricing to dynamic discounts and details two-tower architectures versus model distillation. Rudina presses for clarification on the boundary between distillation and fine-tuning.34:42–37:44 · Jason as informed peer 5/10 Data Curation, Hard Negatives, and Automated Data Prep Tamir unpacks hard negatives in search training (boots vs sandals vs Christmas trees) and balancing LLM synthetic data generation with human validation. Rudina concludes with the operational principle of 'trust and verify'.1:29–6:26 · Guest teaching 0/10 Enterprise AI Adoption and the Ambient AI Framework Rudina delivers a comprehensive framework outlining enterprise AI adoption, citing specific enterprise case studies like Klaviyo's $40M contribution and introducing the 'ambient AI' paradigm. Tamir readily agrees with her foundational pillars of data, infrastructure, and culture.6:28–9:14 · Guest teaching 6/10 Data Challenges, Iterative Modeling, and Overfitting Tamir details the mechanics of iterative model training, explaining loss minimization and the nuances of out-of-sample overfitting. The host remains in listening mode as the guest breaks down core machine learning concepts.9:15–12:01 · Guest teaching 6/10 Translating Business Goals into Machine Learning Metrics The host poses a practical question on why data processing takes weeks for P&L owners. Tamir educates on the gap between business objectives and mathematical loss functions as well as the necessity of curated training examples.12:02–16:18 · Guest teaching 5/10 AI Infrastructure Requirements and GPU Resource Management The host challenges Tamir with a curveball regarding cloud aggregators like AWS Bedrock versus direct model providers, framing cloud costs as a variable fixed line. Tamir explains latency tradeoffs and GPU reservation bottlenecks.16:19–18:34 · Guest teaching 6/10 Fostering a Culture of Rigorous ML Evaluation Tamir discusses the shift from manual architecture design to fine-tuning pre-trained models and the danger of non-experts deploying models without stratified evaluation. The host listens attentively.18:35–24:02 · Guest teaching 5/10 Bridging the Gap Between Technical Metrics and Business Value Rudina interjects to emphasize that metric significance varies wildly across business contexts. Tamir builds on this by detailing search relevance at Shopify, differentiating direct product matches from complementary cross-sell relevance.24:03–26:23 · Guest teaching 5/10 Measuring ROI and Cross-Functional ML Product Collaboration Rudina probes on measuring concrete ROI and navigating trade-offs between engineering resources and business results. Tamir describes breaking down departmental silos so product managers and ML scientists co-design learning objectives.26:24–29:33 · Guest teaching 5/10 Understanding Vectorization and Semantic Embeddings Tamir explains vector spaces and semantic embeddings using the contextual meaning of the word 'bank'. Rudina synthesizes this by explaining that embeddings constitute roughly 80% of the backbone of generative AI applications.29:33–34:42 · Guest teaching 6/10 Scaling Search Architectures and Dynamic Pricing Strategies Tamir gently reframes the pricing question from dynamic pricing to dynamic discounts and details two-tower architectures versus model distillation. Rudina presses for clarification on the boundary between distillation and fine-tuning.34:42–37:44 · Guest teaching 6/10 Data Curation, Hard Negatives, and Automated Data Prep Tamir unpacks hard negatives in search training (boots vs sandals vs Christmas trees) and balancing LLM synthetic data generation with human validation. Rudina concludes with the operational principle of 'trust and verify'.1:29–6:26 · Guest disagreement 0/10 Enterprise AI Adoption and the Ambient AI Framework Rudina delivers a comprehensive framework outlining enterprise AI adoption, citing specific enterprise case studies like Klaviyo's $40M contribution and introducing the 'ambient AI' paradigm. Tamir readily agrees with her foundational pillars of data, infrastructure, and culture.6:28–9:14 · Guest disagreement 0/10 Data Challenges, Iterative Modeling, and Overfitting Tamir details the mechanics of iterative model training, explaining loss minimization and the nuances of out-of-sample overfitting. The host remains in listening mode as the guest breaks down core machine learning concepts.9:15–12:01 · Guest disagreement 0/10 Translating Business Goals into Machine Learning Metrics The host poses a practical question on why data processing takes weeks for P&L owners. Tamir educates on the gap between business objectives and mathematical loss functions as well as the necessity of curated training examples.12:02–16:18 · Guest disagreement 0/10 AI Infrastructure Requirements and GPU Resource Management The host challenges Tamir with a curveball regarding cloud aggregators like AWS Bedrock versus direct model providers, framing cloud costs as a variable fixed line. Tamir explains latency tradeoffs and GPU reservation bottlenecks.16:19–18:34 · Guest disagreement 0/10 Fostering a Culture of Rigorous ML Evaluation Tamir discusses the shift from manual architecture design to fine-tuning pre-trained models and the danger of non-experts deploying models without stratified evaluation. The host listens attentively.18:35–24:02 · Guest disagreement 0/10 Bridging the Gap Between Technical Metrics and Business Value Rudina interjects to emphasize that metric significance varies wildly across business contexts. Tamir builds on this by detailing search relevance at Shopify, differentiating direct product matches from complementary cross-sell relevance.24:03–26:23 · Guest disagreement 0/10 Measuring ROI and Cross-Functional ML Product Collaboration Rudina probes on measuring concrete ROI and navigating trade-offs between engineering resources and business results. Tamir describes breaking down departmental silos so product managers and ML scientists co-design learning objectives.26:24–29:33 · Guest disagreement 0/10 Understanding Vectorization and Semantic Embeddings Tamir explains vector spaces and semantic embeddings using the contextual meaning of the word 'bank'. Rudina synthesizes this by explaining that embeddings constitute roughly 80% of the backbone of generative AI applications.29:33–34:42 · Guest disagreement 1/10 Scaling Search Architectures and Dynamic Pricing Strategies Tamir gently reframes the pricing question from dynamic pricing to dynamic discounts and details two-tower architectures versus model distillation. Rudina presses for clarification on the boundary between distillation and fine-tuning.34:42–37:44 · Guest disagreement 0/10 Data Curation, Hard Negatives, and Automated Data Prep Tamir unpacks hard negatives in search training (boots vs sandals vs Christmas trees) and balancing LLM synthetic data generation with human validation. Rudina concludes with the operational principle of 'trust and verify'.1:29–6:26 · Jason pushing back 0/10 Enterprise AI Adoption and the Ambient AI Framework Rudina delivers a comprehensive framework outlining enterprise AI adoption, citing specific enterprise case studies like Klaviyo's $40M contribution and introducing the 'ambient AI' paradigm. Tamir readily agrees with her foundational pillars of data, infrastructure, and culture.6:28–9:14 · Jason pushing back 0/10 Data Challenges, Iterative Modeling, and Overfitting Tamir details the mechanics of iterative model training, explaining loss minimization and the nuances of out-of-sample overfitting. The host remains in listening mode as the guest breaks down core machine learning concepts.9:15–12:01 · Jason pushing back 1/10 Translating Business Goals into Machine Learning Metrics The host poses a practical question on why data processing takes weeks for P&L owners. Tamir educates on the gap between business objectives and mathematical loss functions as well as the necessity of curated training examples.12:02–16:18 · Jason pushing back 3/10 AI Infrastructure Requirements and GPU Resource Management The host challenges Tamir with a curveball regarding cloud aggregators like AWS Bedrock versus direct model providers, framing cloud costs as a variable fixed line. Tamir explains latency tradeoffs and GPU reservation bottlenecks.16:19–18:34 · Jason pushing back 0/10 Fostering a Culture of Rigorous ML Evaluation Tamir discusses the shift from manual architecture design to fine-tuning pre-trained models and the danger of non-experts deploying models without stratified evaluation. The host listens attentively.18:35–24:02 · Jason pushing back 2/10 Bridging the Gap Between Technical Metrics and Business Value Rudina interjects to emphasize that metric significance varies wildly across business contexts. Tamir builds on this by detailing search relevance at Shopify, differentiating direct product matches from complementary cross-sell relevance.24:03–26:23 · Jason pushing back 1/10 Measuring ROI and Cross-Functional ML Product Collaboration Rudina probes on measuring concrete ROI and navigating trade-offs between engineering resources and business results. Tamir describes breaking down departmental silos so product managers and ML scientists co-design learning objectives.26:24–29:33 · Jason pushing back 0/10 Understanding Vectorization and Semantic Embeddings Tamir explains vector spaces and semantic embeddings using the contextual meaning of the word 'bank'. Rudina synthesizes this by explaining that embeddings constitute roughly 80% of the backbone of generative AI applications.29:33–34:42 · Jason pushing back 2/10 Scaling Search Architectures and Dynamic Pricing Strategies Tamir gently reframes the pricing question from dynamic pricing to dynamic discounts and details two-tower architectures versus model distillation. Rudina presses for clarification on the boundary between distillation and fine-tuning.34:42–37:44 · Jason pushing back 0/10 Data Curation, Hard Negatives, and Automated Data Prep Tamir unpacks hard negatives in search training (boots vs sandals vs Christmas trees) and balancing LLM synthetic data generation with human validation. Rudina concludes with the operational principle of 'trust and verify'.

speaking balance: gold is Jason, purple is the guest (3 minute bins)

0:00 · Jason 0% · guest 100%0:00 · Jason 0% · guest 100%3:00 · Jason 0% · guest 100%3:00 · Jason 0% · guest 100%6:00 · Jason 0% · guest 100%6:00 · Jason 0% · guest 100%9:00 · Jason 0% · guest 100%9:00 · Jason 0% · guest 100%12:00 · Jason 0% · guest 100%12:00 · Jason 0% · guest 100%15:00 · Jason 0% · guest 100%15:00 · Jason 0% · guest 100%18:00 · Jason 0% · guest 100%18:00 · Jason 0% · guest 100%21:00 · Jason 0% · guest 100%21:00 · Jason 0% · guest 100%24:00 · Jason 0% · guest 100%24:00 · Jason 0% · guest 100%27:00 · Jason 0% · guest 100%27:00 · Jason 0% · guest 100%30:00 · Jason 0% · guest 100%30:00 · Jason 0% · guest 100%33:00 · Jason 0% · guest 100%33:00 · Jason 0% · guest 100%36:00 · Jason 0% · guest 100%36:00 · Jason 0% · guest 100%
Sharpest disagreement ▶ 29:45 Dynamic pricing reframe

Tamir mildly rejects the premise of doing dynamic pricing, correcting the approach to 'dynamic discounts' based on product framing experience.

Hardest push from Jason ▶ 14:50 Challenging cloud aggregator performance parity

Rudina presses Tamir on whether cloud aggregator layers like AWS Bedrock introduce a persistent and severe performance delta compared to direct API access.

Biggest teaching moment ▶ 28:10 Contextual embeddings and semantic vector spaces

Tamir provides a clear, intuitive pedagogical breakdown of how vector spaces map multimodal concepts and resolve context-dependent meanings like financial banks versus river banks.

Jason holds their own ▶ 2:08 Klaviyo enterprise AI adoption proof point

Rudina demonstrates deep domain mastery by citing Klaviyo's $40M bottom-line impact and automation of 700 customer success roles as evidence of core AI integration.

the scores for every segment, with the reasoning behind each
ChapterTopicJason as informed peerGuest teachingGuest disagreementJason pushing backWhy
Enterprise AI Adoption and the Ambient AI Framework 7000 Rudina delivers a comprehensive framework outlining enterprise AI adoption, citing specific enterprise case studies like Klaviyo's $40M contribution and introducing the 'ambient AI' paradigm. Tamir readily agrees with her foundational pillars of data, infrastructure, and culture.
Data Challenges, Iterative Modeling, and Overfitting 2600 Tamir details the mechanics of iterative model training, explaining loss minimization and the nuances of out-of-sample overfitting. The host remains in listening mode as the guest breaks down core machine learning concepts.
Translating Business Goals into Machine Learning Metrics 4601 The host poses a practical question on why data processing takes weeks for P&L owners. Tamir educates on the gap between business objectives and mathematical loss functions as well as the necessity of curated training examples.
AI Infrastructure Requirements and GPU Resource Management 6503 The host challenges Tamir with a curveball regarding cloud aggregators like AWS Bedrock versus direct model providers, framing cloud costs as a variable fixed line. Tamir explains latency tradeoffs and GPU reservation bottlenecks.
Fostering a Culture of Rigorous ML Evaluation 3600 Tamir discusses the shift from manual architecture design to fine-tuning pre-trained models and the danger of non-experts deploying models without stratified evaluation. The host listens attentively.
Bridging the Gap Between Technical Metrics and Business Value 6502 Rudina interjects to emphasize that metric significance varies wildly across business contexts. Tamir builds on this by detailing search relevance at Shopify, differentiating direct product matches from complementary cross-sell relevance.
Measuring ROI and Cross-Functional ML Product Collaboration 5501 Rudina probes on measuring concrete ROI and navigating trade-offs between engineering resources and business results. Tamir describes breaking down departmental silos so product managers and ML scientists co-design learning objectives.
Understanding Vectorization and Semantic Embeddings 6500 Tamir explains vector spaces and semantic embeddings using the contextual meaning of the word 'bank'. Rudina synthesizes this by explaining that embeddings constitute roughly 80% of the backbone of generative AI applications.
Scaling Search Architectures and Dynamic Pricing Strategies 6612 Tamir gently reframes the pricing question from dynamic pricing to dynamic discounts and details two-tower architectures versus model distillation. Rudina presses for clarification on the boundary between distillation and fine-tuning.
Data Curation, Hard Negatives, and Automated Data Prep 5600 Tamir unpacks hard negatives in search training (boots vs sandals vs Christmas trees) and balancing LLM synthetic data generation with human validation. Rudina concludes with the operational principle of 'trust and verify'.

Statements from this episode (13)

Disclosure
Seseri: Glasswing advises AI startups not to lead pitches with AI
“In fact, we tell our AI native companies not to lead with AI when they're speaking to prospective customers, because it's such a top of mind topic that it's a superficial indication of interest.”
Rudina Seseri Apr 26, 2024 ▶ 1:54
Assertion Contradicted
Seseri: Klaviyo automating 700 roles with AI added $40M to bottom line
“We have seen a few green shoots, if you will, in this space, particularly, for example, what we saw with Klaviyo and the impact that they have had by bringing AI from a project siloed or customer interface type level usage to their core business, where they au…”
Rudina Seseri Apr 26, 2024 ▶ 2:22
Insight
Seseri: Data readiness, not technology, is typically the biggest enterprise AI bottleneck
“And oftentimes there is a almost naive assumption that the data will be easy to leverage. Oh, it's an asset and we'll use it and leverage it. And it turns out that oftentimes is the biggest problem.”
Rudina Seseri Apr 26, 2024 ▶ 3:58
Insight
Tamir: Gap between ML loss metrics and product goals determines AI success
“There, that gap between what I want to mathematically measure, which is how the machine is going to learn based on that error and what my actual product result is. Really matters.”
Mike Tamir Apr 26, 2024 ▶ 11:10
Insight
Tamir: On-premise infrastructure is cheaper for AI training if already established
“Maybe you buy your own infrastructure, which is also a cheaper answer. If you have, if you already have an infrastructure, an on-prem infrastructure system, that's really going to depend on what your AI training expectations are anticipated.”
Mike Tamir Apr 26, 2024 ▶ 13:36
Insight
Tamir: Direct OpenAI or Anthropic APIs outperform cloud intermediary platforms
“Using a first party solution with say OpenAI or with Anthropic is going to have different performance, different latency than if you do it via an intermediary, like a cloud. So you maybe get some simplicity by not having to onboard different vendors and differ…”
Mike Tamir Apr 26, 2024 ▶ 14:53
Assertion Not checkable as stated
Tamir: Serving 70B open-source LLMs forces teams to buy reserved GPUs
“What if you want to do something like a seventy billion Lama or minstrel or mixed or any of these, those are going to be very hard to get on an ad hoc basis. And you're going to end up having to pay the price for a reserve instance, if you want to be able to s…”
Mike Tamir Apr 26, 2024 ▶ 15:57
Insight
Tamir: Rigorous evaluation culture is critical as non-experts deploy pre-trained AI models
“Having that strong culture to back up the performance of your model before you go into release is probably one of the biggest issues in terms of having a strong ML culture when you're, as it gets easier and easier for non-experts to start leveraging these tool…”
Mike Tamir Apr 26, 2024 ▶ 18:08
Insight
Tamir: Search must balance direct and complementary relevance for optimal revenue
“And maybe what you would want to do is have one metric that shows for direct relevance and another that shows complementary relevance, right? And then it is somewhat a product decision of how much you want to balance these, but it's also something that you can…”
Mike Tamir Apr 26, 2024 ▶ 23:20
Insight
Tamir: Successful ML teams require serving engineers and embedded product managers
“You need to have engineering for serving. You need to have your ML scientists and engineers, and you need to have product. You need to have product every year in the trenches with your ML builders.”
Mike Tamir Apr 26, 2024 ▶ 25:42
Insight
Tamir: E-commerce platforms should use dynamic discounts, never dynamic pricing
“You never want to do dynamic pricing. You want to do dynamic discounts as a product framing bit of advice”
Mike Tamir Apr 26, 2024 ▶ 29:52
Insight
Tamir: Per-token commercial LLM costs are unscalable for every enterprise query
“And if you pay a commodity LLM, which right now does usually beat out open source LLMs, you're going to end up paying per token and you can't scalably, if you have the right size of a business, do this for every query or every kind of use case.”
Mike Tamir Apr 26, 2024 ▶ 32:23
Insight
Tamir: Search retrieval models learn most from near-miss negative training examples
“If I search for boots and I show a Christmas tree, that's not gonna, the model's not gonna learn much from saying, oh, that was wrong, right? It's gonna learn a lot more by showing its sneakers or by showing its sandals and saying that's because it's closer to…”
Mike Tamir Apr 26, 2024 ▶ 35:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.