Feb 20, 2025 · 1h 13m · mad

From Selfie to Studio: Captions CEO on AI Video for 10M Creators

Gaurav Misra · 53m spoken Matt Turck · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Captions CEO Gaurav Misra about building an AI-powered video creation studio that empowers over 10 million creators. They explore Captions' technical architecture, product evolution, model orchestration, and the broader future of democratized AI storytelling.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 20% of the talking time here. How this is scored →

Matt as informed peer 4.3 Guest teaching 4.5 Guest disagreement 1.1 Matt pushing back 1.1
05100:0015:0030:0045:001:00:001:27–3:40 · Matt as informed peer 2/10 Defining Captions as a Studio in Your Pocket Matt asks high-level introductory questions about Captions while Gaurav breaks down the framework of skipping keyframe editing complexity. Matt adds supportive summary comments on the magical user experience.3:40–6:46 · Matt as informed peer 3/10 Captions Origin Story and Co-Founders' Career Backgrounds Matt demonstrates familiarity with the Boston startup ecosystem and ZIRP valuation peaks when discussing Localytics and Snapchat. Gaurav shares his personal founder journey and career pivot.6:46–11:31 · Matt as informed peer 2/10 Macro Trends and Democratizing Hollywood Level Video Tools Gaurav educates Matt on his macro framework for identifying societal shifts and reveals an insider story about Snap running A/B tests on selling ad inventory to TikTok. Matt listens and reacts to the unexpected story.11:31–18:14 · Matt as informed peer 5/10 Captions Growth Scale, Funding Metrics, and Team Culture Matt probes differentiation against legacy software like Premiere and competitor products like CapCut using a software-versus-nocode analogy. Gaurav clarifies why Canva-style templates fail in video creation without AI.18:14–23:39 · Matt as informed peer 4/10 Small Business Target Demographics and Paid-Only Product Launch Matt tests whether small business users are a fickle target demographic. Gaurav explains how launch-time paywalls filtered noise and details the organic viral launch story.23:39–26:29 · Matt as informed peer 3/10 Product Architecture Breakdown of AI Creator and AI Edit Matt asks for a breakdown of the product suite, and Gaurav explains the distinct technical challenges of AI Creator generation versus AI Edit agentic workflows.26:29–28:37 · Matt as informed peer 4/10 Third-Party Model Stack and Provider Selection Strategy Matt asks detailed questions about provider choices across OpenAI and Anthropic. Gaurav details their rotation between Azure OpenAI and Claude 3.5 Sonnet for reliability and performance.28:37–30:50 · Matt as informed peer 6/10 Model Evaluation Frameworks and User Preference RLHF Matt accurately identifies Gaurav's model selection mechanism as an implicit RLHF flywheel using paying user preference signals inside the video workflow.30:50–34:02 · Matt as informed peer 6/10 Frontier AI Pace of Releases and Industry Shifts Matt references recent frontier AI releases like Mini o3 and Deep Research as well as industry talk from Replit's CEO. Gaurav breaks down why current text-to-video models focus on B-roll while Captions targets talking A-roll.34:02–37:15 · Matt as informed peer 6/10 Proprietary Data Sourcing and YouTube Crawling Compliance Matt directly challenges Gaurav on YouTube crawling compliance and legal liabilities. Gaurav acknowledges the distinction between technical capability and enterprise legal risk.37:15–42:10 · Matt as informed peer 5/10 Multilingual Transcription Engineering and Arabic Adaptation Matt inquires about transcription engineering and GPU rendering constraints. Gaurav highlights Whisper's deficiencies in non-English languages like Arabic and the timing loss in text encoders.42:10–45:08 · Matt as informed peer 6/10 Model Agnosticism, Latency Optimization, and Dynamic UX Matt presses Gaurav on what happens if frontier labs like OpenAI or Claude launch superior native video generation models. Gaurav defends Captions' moat around user relationship and workflow UI.45:08–47:56 · Matt as informed peer 4/10 Parallelized Cloud Video Generation and Multi-GPU Orchestration Gaurav provides a deep technical explanation of their video generation parallelization, splitting video into 25-frame segments with overlapping frames across GPU clusters to reduce latency.47:56–50:59 · Matt as informed peer 4/10 Multi-Cloud Stack and Custom Video Editing Pipelines Matt asks about the underlying cloud stack. Gaurav educates on why high-end H100s perform poorly at video encoding compared to older T4 GPUs, requiring multi-cloud GPU pipelines.50:59–53:55 · Matt as informed peer 5/10 AI Hallucinations as Features and System Prompting Matt raises the issue of AI hallucinations. Gaurav pushes back on the premise, asserting that hallucinations in creative video generation are a necessary feature that saves users from manual choices.53:55–58:22 · Matt as informed peer 4/10 Uncanny Valley and Global AI Video Localization Matt asks if video has surpassed the uncanny valley. Gaurav explains the shifting perception standards of audiences and reveals how AI localized dubbing converts surprisingly well across global markets.58:22–1:00:31 · Matt as informed peer 4/10 The Future of Professional Video Editing in the AI Era Matt asks about the impact on professional video editors. Gaurav draws a historical parallel to digital music production, arguing AI will elevate video editors rather than eliminate them.1:00:31–1:04:01 · Matt as informed peer 5/10 Key Technical Hurdles for AI Generated Feature Films Matt asks about timeline hurdles for full AI feature films and raises deepfake concerns. Gaurav introduces a clear conceptual framework separating documentation from storytelling.1:04:01–1:08:46 · Matt as informed peer 5/10 Captions Product Roadmap and Industry Competition Dynamic Matt asks about go-to-market shifts into B2B enterprise sales and inference unit economics. Gaurav points out the compounding 10x annual drop in GPU inference costs.1:08:46–1:12:53 · Matt as informed peer 4/10 Building an In-Person AI Startup in New York City Matt and Gaurav discuss building in New York City. Gaurav details why short commutes and apartment living make NYC ideal for mandatory in-person startup culture.1:27–3:40 · Guest teaching 4/10 Defining Captions as a Studio in Your Pocket Matt asks high-level introductory questions about Captions while Gaurav breaks down the framework of skipping keyframe editing complexity. Matt adds supportive summary comments on the magical user experience.3:40–6:46 · Guest teaching 2/10 Captions Origin Story and Co-Founders' Career Backgrounds Matt demonstrates familiarity with the Boston startup ecosystem and ZIRP valuation peaks when discussing Localytics and Snapchat. Gaurav shares his personal founder journey and career pivot.6:46–11:31 · Guest teaching 6/10 Macro Trends and Democratizing Hollywood Level Video Tools Gaurav educates Matt on his macro framework for identifying societal shifts and reveals an insider story about Snap running A/B tests on selling ad inventory to TikTok. Matt listens and reacts to the unexpected story.11:31–18:14 · Guest teaching 5/10 Captions Growth Scale, Funding Metrics, and Team Culture Matt probes differentiation against legacy software like Premiere and competitor products like CapCut using a software-versus-nocode analogy. Gaurav clarifies why Canva-style templates fail in video creation without AI.18:14–23:39 · Guest teaching 4/10 Small Business Target Demographics and Paid-Only Product Launch Matt tests whether small business users are a fickle target demographic. Gaurav explains how launch-time paywalls filtered noise and details the organic viral launch story.23:39–26:29 · Guest teaching 4/10 Product Architecture Breakdown of AI Creator and AI Edit Matt asks for a breakdown of the product suite, and Gaurav explains the distinct technical challenges of AI Creator generation versus AI Edit agentic workflows.26:29–28:37 · Guest teaching 3/10 Third-Party Model Stack and Provider Selection Strategy Matt asks detailed questions about provider choices across OpenAI and Anthropic. Gaurav details their rotation between Azure OpenAI and Claude 3.5 Sonnet for reliability and performance.28:37–30:50 · Guest teaching 4/10 Model Evaluation Frameworks and User Preference RLHF Matt accurately identifies Gaurav's model selection mechanism as an implicit RLHF flywheel using paying user preference signals inside the video workflow.30:50–34:02 · Guest teaching 4/10 Frontier AI Pace of Releases and Industry Shifts Matt references recent frontier AI releases like Mini o3 and Deep Research as well as industry talk from Replit's CEO. Gaurav breaks down why current text-to-video models focus on B-roll while Captions targets talking A-roll.34:02–37:15 · Guest teaching 3/10 Proprietary Data Sourcing and YouTube Crawling Compliance Matt directly challenges Gaurav on YouTube crawling compliance and legal liabilities. Gaurav acknowledges the distinction between technical capability and enterprise legal risk.37:15–42:10 · Guest teaching 5/10 Multilingual Transcription Engineering and Arabic Adaptation Matt inquires about transcription engineering and GPU rendering constraints. Gaurav highlights Whisper's deficiencies in non-English languages like Arabic and the timing loss in text encoders.42:10–45:08 · Guest teaching 4/10 Model Agnosticism, Latency Optimization, and Dynamic UX Matt presses Gaurav on what happens if frontier labs like OpenAI or Claude launch superior native video generation models. Gaurav defends Captions' moat around user relationship and workflow UI.45:08–47:56 · Guest teaching 6/10 Parallelized Cloud Video Generation and Multi-GPU Orchestration Gaurav provides a deep technical explanation of their video generation parallelization, splitting video into 25-frame segments with overlapping frames across GPU clusters to reduce latency.47:56–50:59 · Guest teaching 6/10 Multi-Cloud Stack and Custom Video Editing Pipelines Matt asks about the underlying cloud stack. Gaurav educates on why high-end H100s perform poorly at video encoding compared to older T4 GPUs, requiring multi-cloud GPU pipelines.50:59–53:55 · Guest teaching 6/10 AI Hallucinations as Features and System Prompting Matt raises the issue of AI hallucinations. Gaurav pushes back on the premise, asserting that hallucinations in creative video generation are a necessary feature that saves users from manual choices.53:55–58:22 · Guest teaching 5/10 Uncanny Valley and Global AI Video Localization Matt asks if video has surpassed the uncanny valley. Gaurav explains the shifting perception standards of audiences and reveals how AI localized dubbing converts surprisingly well across global markets.58:22–1:00:31 · Guest teaching 5/10 The Future of Professional Video Editing in the AI Era Matt asks about the impact on professional video editors. Gaurav draws a historical parallel to digital music production, arguing AI will elevate video editors rather than eliminate them.1:00:31–1:04:01 · Guest teaching 6/10 Key Technical Hurdles for AI Generated Feature Films Matt asks about timeline hurdles for full AI feature films and raises deepfake concerns. Gaurav introduces a clear conceptual framework separating documentation from storytelling.1:04:01–1:08:46 · Guest teaching 5/10 Captions Product Roadmap and Industry Competition Dynamic Matt asks about go-to-market shifts into B2B enterprise sales and inference unit economics. Gaurav points out the compounding 10x annual drop in GPU inference costs.1:08:46–1:12:53 · Guest teaching 3/10 Building an In-Person AI Startup in New York City Matt and Gaurav discuss building in New York City. Gaurav details why short commutes and apartment living make NYC ideal for mandatory in-person startup culture.1:27–3:40 · Guest disagreement 1/10 Defining Captions as a Studio in Your Pocket Matt asks high-level introductory questions about Captions while Gaurav breaks down the framework of skipping keyframe editing complexity. Matt adds supportive summary comments on the magical user experience.3:40–6:46 · Guest disagreement 0/10 Captions Origin Story and Co-Founders' Career Backgrounds Matt demonstrates familiarity with the Boston startup ecosystem and ZIRP valuation peaks when discussing Localytics and Snapchat. Gaurav shares his personal founder journey and career pivot.6:46–11:31 · Guest disagreement 1/10 Macro Trends and Democratizing Hollywood Level Video Tools Gaurav educates Matt on his macro framework for identifying societal shifts and reveals an insider story about Snap running A/B tests on selling ad inventory to TikTok. Matt listens and reacts to the unexpected story.11:31–18:14 · Guest disagreement 2/10 Captions Growth Scale, Funding Metrics, and Team Culture Matt probes differentiation against legacy software like Premiere and competitor products like CapCut using a software-versus-nocode analogy. Gaurav clarifies why Canva-style templates fail in video creation without AI.18:14–23:39 · Guest disagreement 1/10 Small Business Target Demographics and Paid-Only Product Launch Matt tests whether small business users are a fickle target demographic. Gaurav explains how launch-time paywalls filtered noise and details the organic viral launch story.23:39–26:29 · Guest disagreement 0/10 Product Architecture Breakdown of AI Creator and AI Edit Matt asks for a breakdown of the product suite, and Gaurav explains the distinct technical challenges of AI Creator generation versus AI Edit agentic workflows.26:29–28:37 · Guest disagreement 0/10 Third-Party Model Stack and Provider Selection Strategy Matt asks detailed questions about provider choices across OpenAI and Anthropic. Gaurav details their rotation between Azure OpenAI and Claude 3.5 Sonnet for reliability and performance.28:37–30:50 · Guest disagreement 0/10 Model Evaluation Frameworks and User Preference RLHF Matt accurately identifies Gaurav's model selection mechanism as an implicit RLHF flywheel using paying user preference signals inside the video workflow.30:50–34:02 · Guest disagreement 1/10 Frontier AI Pace of Releases and Industry Shifts Matt references recent frontier AI releases like Mini o3 and Deep Research as well as industry talk from Replit's CEO. Gaurav breaks down why current text-to-video models focus on B-roll while Captions targets talking A-roll.34:02–37:15 · Guest disagreement 2/10 Proprietary Data Sourcing and YouTube Crawling Compliance Matt directly challenges Gaurav on YouTube crawling compliance and legal liabilities. Gaurav acknowledges the distinction between technical capability and enterprise legal risk.37:15–42:10 · Guest disagreement 1/10 Multilingual Transcription Engineering and Arabic Adaptation Matt inquires about transcription engineering and GPU rendering constraints. Gaurav highlights Whisper's deficiencies in non-English languages like Arabic and the timing loss in text encoders.42:10–45:08 · Guest disagreement 3/10 Model Agnosticism, Latency Optimization, and Dynamic UX Matt presses Gaurav on what happens if frontier labs like OpenAI or Claude launch superior native video generation models. Gaurav defends Captions' moat around user relationship and workflow UI.45:08–47:56 · Guest disagreement 0/10 Parallelized Cloud Video Generation and Multi-GPU Orchestration Gaurav provides a deep technical explanation of their video generation parallelization, splitting video into 25-frame segments with overlapping frames across GPU clusters to reduce latency.47:56–50:59 · Guest disagreement 1/10 Multi-Cloud Stack and Custom Video Editing Pipelines Matt asks about the underlying cloud stack. Gaurav educates on why high-end H100s perform poorly at video encoding compared to older T4 GPUs, requiring multi-cloud GPU pipelines.50:59–53:55 · Guest disagreement 3/10 AI Hallucinations as Features and System Prompting Matt raises the issue of AI hallucinations. Gaurav pushes back on the premise, asserting that hallucinations in creative video generation are a necessary feature that saves users from manual choices.53:55–58:22 · Guest disagreement 1/10 Uncanny Valley and Global AI Video Localization Matt asks if video has surpassed the uncanny valley. Gaurav explains the shifting perception standards of audiences and reveals how AI localized dubbing converts surprisingly well across global markets.58:22–1:00:31 · Guest disagreement 1/10 The Future of Professional Video Editing in the AI Era Matt asks about the impact on professional video editors. Gaurav draws a historical parallel to digital music production, arguing AI will elevate video editors rather than eliminate them.1:00:31–1:04:01 · Guest disagreement 2/10 Key Technical Hurdles for AI Generated Feature Films Matt asks about timeline hurdles for full AI feature films and raises deepfake concerns. Gaurav introduces a clear conceptual framework separating documentation from storytelling.1:04:01–1:08:46 · Guest disagreement 2/10 Captions Product Roadmap and Industry Competition Dynamic Matt asks about go-to-market shifts into B2B enterprise sales and inference unit economics. Gaurav points out the compounding 10x annual drop in GPU inference costs.1:08:46–1:12:53 · Guest disagreement 1/10 Building an In-Person AI Startup in New York City Matt and Gaurav discuss building in New York City. Gaurav details why short commutes and apartment living make NYC ideal for mandatory in-person startup culture.1:27–3:40 · Matt pushing back 0/10 Defining Captions as a Studio in Your Pocket Matt asks high-level introductory questions about Captions while Gaurav breaks down the framework of skipping keyframe editing complexity. Matt adds supportive summary comments on the magical user experience.3:40–6:46 · Matt pushing back 0/10 Captions Origin Story and Co-Founders' Career Backgrounds Matt demonstrates familiarity with the Boston startup ecosystem and ZIRP valuation peaks when discussing Localytics and Snapchat. Gaurav shares his personal founder journey and career pivot.6:46–11:31 · Matt pushing back 0/10 Macro Trends and Democratizing Hollywood Level Video Tools Gaurav educates Matt on his macro framework for identifying societal shifts and reveals an insider story about Snap running A/B tests on selling ad inventory to TikTok. Matt listens and reacts to the unexpected story.11:31–18:14 · Matt pushing back 2/10 Captions Growth Scale, Funding Metrics, and Team Culture Matt probes differentiation against legacy software like Premiere and competitor products like CapCut using a software-versus-nocode analogy. Gaurav clarifies why Canva-style templates fail in video creation without AI.18:14–23:39 · Matt pushing back 2/10 Small Business Target Demographics and Paid-Only Product Launch Matt tests whether small business users are a fickle target demographic. Gaurav explains how launch-time paywalls filtered noise and details the organic viral launch story.23:39–26:29 · Matt pushing back 0/10 Product Architecture Breakdown of AI Creator and AI Edit Matt asks for a breakdown of the product suite, and Gaurav explains the distinct technical challenges of AI Creator generation versus AI Edit agentic workflows.26:29–28:37 · Matt pushing back 0/10 Third-Party Model Stack and Provider Selection Strategy Matt asks detailed questions about provider choices across OpenAI and Anthropic. Gaurav details their rotation between Azure OpenAI and Claude 3.5 Sonnet for reliability and performance.28:37–30:50 · Matt pushing back 0/10 Model Evaluation Frameworks and User Preference RLHF Matt accurately identifies Gaurav's model selection mechanism as an implicit RLHF flywheel using paying user preference signals inside the video workflow.30:50–34:02 · Matt pushing back 0/10 Frontier AI Pace of Releases and Industry Shifts Matt references recent frontier AI releases like Mini o3 and Deep Research as well as industry talk from Replit's CEO. Gaurav breaks down why current text-to-video models focus on B-roll while Captions targets talking A-roll.34:02–37:15 · Matt pushing back 5/10 Proprietary Data Sourcing and YouTube Crawling Compliance Matt directly challenges Gaurav on YouTube crawling compliance and legal liabilities. Gaurav acknowledges the distinction between technical capability and enterprise legal risk.37:15–42:10 · Matt pushing back 1/10 Multilingual Transcription Engineering and Arabic Adaptation Matt inquires about transcription engineering and GPU rendering constraints. Gaurav highlights Whisper's deficiencies in non-English languages like Arabic and the timing loss in text encoders.42:10–45:08 · Matt pushing back 5/10 Model Agnosticism, Latency Optimization, and Dynamic UX Matt presses Gaurav on what happens if frontier labs like OpenAI or Claude launch superior native video generation models. Gaurav defends Captions' moat around user relationship and workflow UI.45:08–47:56 · Matt pushing back 0/10 Parallelized Cloud Video Generation and Multi-GPU Orchestration Gaurav provides a deep technical explanation of their video generation parallelization, splitting video into 25-frame segments with overlapping frames across GPU clusters to reduce latency.47:56–50:59 · Matt pushing back 0/10 Multi-Cloud Stack and Custom Video Editing Pipelines Matt asks about the underlying cloud stack. Gaurav educates on why high-end H100s perform poorly at video encoding compared to older T4 GPUs, requiring multi-cloud GPU pipelines.50:59–53:55 · Matt pushing back 2/10 AI Hallucinations as Features and System Prompting Matt raises the issue of AI hallucinations. Gaurav pushes back on the premise, asserting that hallucinations in creative video generation are a necessary feature that saves users from manual choices.53:55–58:22 · Matt pushing back 1/10 Uncanny Valley and Global AI Video Localization Matt asks if video has surpassed the uncanny valley. Gaurav explains the shifting perception standards of audiences and reveals how AI localized dubbing converts surprisingly well across global markets.58:22–1:00:31 · Matt pushing back 0/10 The Future of Professional Video Editing in the AI Era Matt asks about the impact on professional video editors. Gaurav draws a historical parallel to digital music production, arguing AI will elevate video editors rather than eliminate them.1:00:31–1:04:01 · Matt pushing back 2/10 Key Technical Hurdles for AI Generated Feature Films Matt asks about timeline hurdles for full AI feature films and raises deepfake concerns. Gaurav introduces a clear conceptual framework separating documentation from storytelling.1:04:01–1:08:46 · Matt pushing back 1/10 Captions Product Roadmap and Industry Competition Dynamic Matt asks about go-to-market shifts into B2B enterprise sales and inference unit economics. Gaurav points out the compounding 10x annual drop in GPU inference costs.1:08:46–1:12:53 · Matt pushing back 0/10 Building an In-Person AI Startup in New York City Matt and Gaurav discuss building in New York City. Gaurav details why short commutes and apartment living make NYC ideal for mandatory in-person startup culture.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 38.6% · guest 61.4%0:00 · Matt 38.6% · guest 61.4%3:00 · Matt 14.9% · guest 85.1%3:00 · Matt 14.9% · guest 85.1%6:00 · Matt 15.1% · guest 84.9%6:00 · Matt 15.1% · guest 84.9%9:00 · Matt 8.9% · guest 91.1%9:00 · Matt 8.9% · guest 91.1%12:00 · Matt 32.5% · guest 67.5%12:00 · Matt 32.5% · guest 67.5%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 25.8% · guest 74.2%18:00 · Matt 25.8% · guest 74.2%21:00 · Matt 21.9% · guest 78.1%21:00 · Matt 21.9% · guest 78.1%24:00 · Matt 17.6% · guest 82.4%24:00 · Matt 17.6% · guest 82.4%27:00 · Matt 15.2% · guest 84.8%27:00 · Matt 15.2% · guest 84.8%30:00 · Matt 37.8% · guest 62.2%30:00 · Matt 37.8% · guest 62.2%33:00 · Matt 22.8% · guest 77.2%33:00 · Matt 22.8% · guest 77.2%36:00 · Matt 24.3% · guest 75.7%36:00 · Matt 24.3% · guest 75.7%39:00 · Matt 22% · guest 78%39:00 · Matt 22% · guest 78%42:00 · Matt 17.8% · guest 82.2%42:00 · Matt 17.8% · guest 82.2%45:00 · Matt 16% · guest 84%45:00 · Matt 16% · guest 84%48:00 · Matt 10.7% · guest 89.3%48:00 · Matt 10.7% · guest 89.3%51:00 · Matt 20.7% · guest 79.3%51:00 · Matt 20.7% · guest 79.3%54:00 · Matt 9.2% · guest 90.8%54:00 · Matt 9.2% · guest 90.8%57:00 · Matt 8.7% · guest 91.3%57:00 · Matt 8.7% · guest 91.3%1:00:00 · Matt 16% · guest 84%1:00:00 · Matt 16% · guest 84%1:03:00 · Matt 23.3% · guest 76.7%1:03:00 · Matt 23.3% · guest 76.7%1:06:00 · Matt 24.1% · guest 75.9%1:06:00 · Matt 24.1% · guest 75.9%1:09:00 · Matt 12.7% · guest 87.3%1:09:00 · Matt 12.7% · guest 87.3%1:12:00 · Matt 79.9% · guest 20.1%1:12:00 · Matt 79.9% · guest 20.1%
Sharpest disagreement ▶ 51:00 Re-framing Hallucinations as Features

Gaurav explicitly rejects Matt's premise that hallucinations are a bug in AI models, asserting 'I think it's a feature, actually' and arguing that users delegate creative choices to the system.

Hardest push from Matt ▶ 42:10 Pressing on Frontier Model Commoditization Risk

Matt directly challenges Gaurav on Captions' long-term defensibility and platform risk if OpenAI or Anthropic release native, perfect video generation models.

Biggest teaching moment ▶ 9:30 Behind-the-Scenes Snap vs TikTok Ad Dynamics

Gaurav reveals insider knowledge from his time leading Snap's TikTok response team, detailing how TikTok originally expanded in the US by spending heavily on Snapchat ads despite internal A/B test warnings.

Matt holds his own ▶ 29:40 Identifying User Workflow Preference as RLHF

Matt demonstrates sharp technical expertise by accurately identifying Gaurav's custom user preference evaluation harness as an implicit real-world RLHF flywheel.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Captions as a Studio in Your Pocket 2410 Matt asks high-level introductory questions about Captions while Gaurav breaks down the framework of skipping keyframe editing complexity. Matt adds supportive summary comments on the magical user experience.
Captions Origin Story and Co-Founders' Career Backgrounds 3200 Matt demonstrates familiarity with the Boston startup ecosystem and ZIRP valuation peaks when discussing Localytics and Snapchat. Gaurav shares his personal founder journey and career pivot.
Macro Trends and Democratizing Hollywood Level Video Tools 2610 Gaurav educates Matt on his macro framework for identifying societal shifts and reveals an insider story about Snap running A/B tests on selling ad inventory to TikTok. Matt listens and reacts to the unexpected story.
Captions Growth Scale, Funding Metrics, and Team Culture 5522 Matt probes differentiation against legacy software like Premiere and competitor products like CapCut using a software-versus-nocode analogy. Gaurav clarifies why Canva-style templates fail in video creation without AI.
Small Business Target Demographics and Paid-Only Product Launch 4412 Matt tests whether small business users are a fickle target demographic. Gaurav explains how launch-time paywalls filtered noise and details the organic viral launch story.
Product Architecture Breakdown of AI Creator and AI Edit 3400 Matt asks for a breakdown of the product suite, and Gaurav explains the distinct technical challenges of AI Creator generation versus AI Edit agentic workflows.
Third-Party Model Stack and Provider Selection Strategy 4300 Matt asks detailed questions about provider choices across OpenAI and Anthropic. Gaurav details their rotation between Azure OpenAI and Claude 3.5 Sonnet for reliability and performance.
Model Evaluation Frameworks and User Preference RLHF 6400 Matt accurately identifies Gaurav's model selection mechanism as an implicit RLHF flywheel using paying user preference signals inside the video workflow.
Frontier AI Pace of Releases and Industry Shifts 6410 Matt references recent frontier AI releases like Mini o3 and Deep Research as well as industry talk from Replit's CEO. Gaurav breaks down why current text-to-video models focus on B-roll while Captions targets talking A-roll.
Proprietary Data Sourcing and YouTube Crawling Compliance 6325 Matt directly challenges Gaurav on YouTube crawling compliance and legal liabilities. Gaurav acknowledges the distinction between technical capability and enterprise legal risk.
Multilingual Transcription Engineering and Arabic Adaptation 5511 Matt inquires about transcription engineering and GPU rendering constraints. Gaurav highlights Whisper's deficiencies in non-English languages like Arabic and the timing loss in text encoders.
Model Agnosticism, Latency Optimization, and Dynamic UX 6435 Matt presses Gaurav on what happens if frontier labs like OpenAI or Claude launch superior native video generation models. Gaurav defends Captions' moat around user relationship and workflow UI.
Parallelized Cloud Video Generation and Multi-GPU Orchestration 4600 Gaurav provides a deep technical explanation of their video generation parallelization, splitting video into 25-frame segments with overlapping frames across GPU clusters to reduce latency.
Multi-Cloud Stack and Custom Video Editing Pipelines 4610 Matt asks about the underlying cloud stack. Gaurav educates on why high-end H100s perform poorly at video encoding compared to older T4 GPUs, requiring multi-cloud GPU pipelines.
AI Hallucinations as Features and System Prompting 5632 Matt raises the issue of AI hallucinations. Gaurav pushes back on the premise, asserting that hallucinations in creative video generation are a necessary feature that saves users from manual choices.
Uncanny Valley and Global AI Video Localization 4511 Matt asks if video has surpassed the uncanny valley. Gaurav explains the shifting perception standards of audiences and reveals how AI localized dubbing converts surprisingly well across global markets.
The Future of Professional Video Editing in the AI Era 4510 Matt asks about the impact on professional video editors. Gaurav draws a historical parallel to digital music production, arguing AI will elevate video editors rather than eliminate them.
Key Technical Hurdles for AI Generated Feature Films 5622 Matt asks about timeline hurdles for full AI feature films and raises deepfake concerns. Gaurav introduces a clear conceptual framework separating documentation from storytelling.
Captions Product Roadmap and Industry Competition Dynamic 5521 Matt asks about go-to-market shifts into B2B enterprise sales and inference unit economics. Gaurav points out the compounding 10x annual drop in GPU inference costs.
Building an In-Person AI Startup in New York City 4310 Matt and Gaurav discuss building in New York City. Gaurav details why short commutes and apartment living make NYC ideal for mandatory in-person startup culture.

Statements from this episode (38)

Insight
Misra: Product design and machine learning is a hard-to-beat founder background
“I transitioned to product design, which is a rare transition, I think. But it actually, like, really helped me in my path to starting this company. Because think about it, product design plus machine learning, like, hard to beat as a combination.”
Gaurav Misra Feb 20, 2025 ▶ 4:47
Assertion Supported
Gaurav Misra: TikTok's US growth was heavily fueled by Snap ads
“TikTok spread in the US mostly through Snapchat, and this is like not well known out there today, but a lot of it was they were spending a very, very large amount of money on advertising on Snap, right?”
Gaurav Misra Feb 20, 2025 ▶ 9:35
Disclosure
Snap A/B tests showed TikTok ads initially had no effect on Snapchat engagement
“We ran A-B tests actually Within the advertising groups to see like, oh, like let's segment some users out, not show them TikTok ads, see if they behave differently in the long run, right? In terms of engagement on Snapchat and stuff like that. And there was n…”
Gaurav Misra Feb 20, 2025 ▶ 10:00
Insight
Misra: Muscle memory makes UI innovation difficult in professional creative software
“I think that's what makes it actually hard to innovate in these spaces is because people have so much built in muscle memory about how things have been done that if you come in and change a bunch of stuff and be like, here's a simpler video editor, people actu…”
Gaurav Misra Feb 20, 2025 ▶ 13:02
Insight
Misra: Video cannot be templated without AI content understanding
“Like you actually can't templateize video because the nature of it is completely, it's unknown, right? It could be saying anything. It could be doing anything. Anything would be happening. The, template just can't fit with, it can't understand what's happening…”
Gaurav Misra Feb 20, 2025 ▶ 16:55
Opinion
Misra: CapCut relies on simple UI and TikTok distribution rather than AI
“Whereas I think Capco's approach is much more traditional. They're actually just trying to do exactly what Canva did, which is, okay, let's make the UI simpler and add a few more buttons and essentially bootstrap it to TikTok and distribute it like crazy.”
Gaurav Misra Feb 20, 2025 ▶ 18:00
Disclosure
Misra: Captions stayed paid-only initially to filter out user noise
“One thing that we did well sort of in the first couple of years with the company is we actually kept the application paid only. And that actually. It just filtered out anybody who's not. Right. Exactly. Right. Right. So it really left us with the people who ar…”
Gaurav Misra Feb 20, 2025 ▶ 19:23
Assertion Not checkable as stated
Misra: Captions hit top 30 on App Store after two-day build
“The original app was made in a weekend, like literally in two days. And launched and it wasn't marketed at all. It was just put on the app store. And the next day I woke up and I just saw like all the numbers just blowing up completely. And we're like number, …”
Gaurav Misra Feb 20, 2025 ▶ 21:40
Prediction Not checkable as stated
Misra: AI video generation will eventually make cameras obsolete for creators
“So once, as AI Creator develops and becomes perfect, you will not use to use the, need to use the camera anymore, right? So you'll just be able to generate all the footage that you want.”
Gaurav Misra Feb 20, 2025 ▶ 24:13
Prediction Not checkable as stated
Misra: LLMs will eventually solve video script generation
“The script generation part we don't take on because we think the LLMs and stuff will solve this eventually. I think they're still not very good at video scripts, and they're getting better and better, but we do think that as these models improve, this will be …”
Gaurav Misra Feb 20, 2025 ▶ 26:06
Assertion Not checkable as stated
Misra: Microsoft Azure OpenAI outperformed direct OpenAI on uptime and latency
“Microsoft OpenAI did much better in terms of reliability. It just worked perfectly for latency and never went down, which is what you want.”
Gaurav Misra Feb 20, 2025 ▶ 28:14
Disclosure
Misra: Captions switched its script generation LLM provider to Claude
“And then, but we ended up switching to Claude. So now we're on Claude.”
Gaurav Misra Feb 20, 2025 ▶ 28:26
Disclosure
Gaurav Misra: Captions uses Ideogram for in-app image generation
“So we are using Ideogram for that one because it generates really good text.”
Gaurav Misra Feb 20, 2025 ▶ 32:01
Assertion Contradicted
Misra: Existing text-to-video AI models produce only silent B-roll
“If you look at all the video generation models today that are doing text to video, they're all silent videos, right? There's never anybody talking. It's all B-roll, right?”
Gaurav Misra Feb 20, 2025 ▶ 32:15
Disclosure
Misra: Captions trains AI models entirely on fully licensed proprietary data
“So we, the data that we're training on internally is all fully licensed, completed proprietary.”
Gaurav Misra Feb 20, 2025 ▶ 34:09
Insight
Misra: Crawling YouTube for AI training data creates enterprise sales liability
“If the goal is to sell to enterprise, it's like really hard to justify it if you're crawling YouTube”
Gaurav Misra Feb 20, 2025 ▶ 34:46
Assertion Not checkable as stated
Misra: Audio conditioning for AI video has not been solved at scale
“Audio conditioning is just, it's not a thing that really has been solved at scale, and nobody else has done it, right?”
Gaurav Misra Feb 20, 2025 ▶ 35:21
Assertion Supported
Misra: Whisper and standard transcription models perform poorly in Arabic
“Whisper doesn't work in Arabic that well. Right. So we had to figure out how to make something work in Arabic, but it can't be bad because Almost all transcription is bad in Arabic, right? There's just not enough training that's been done on that language.”
Gaurav Misra Feb 20, 2025 ▶ 38:16
Opinion
Misra: Mobile video rendering software is among the hardest technical engineering challenges
“Just building the ability to render all this stuff on video on your phone, right? Like the technical challenges involved, this is probably one of the hardest things you can do besides maybe building a game, right? Which would be maybe the hardest thing you can…”
Gaurav Misra Feb 20, 2025 ▶ 40:19
Prediction Not checkable as stated
Misra: Captions can build better video models than frontier AI labs
“We are making the unique bet to say that we can actually build a better model than they can because of our unique data sets.”
Gaurav Misra Feb 20, 2025 ▶ 42:40
Insight
Misra: AI video becomes useful only when average viewers cannot spot generation
“One of the main criteria for the ability to actually Use something like for it to be like actually practically useful versus entertainment and just like interesting is how photorealistic is it? Right. And it has to be past the level that the average person can…”
Gaurav Misra Feb 20, 2025 ▶ 43:53
Assertion Not checkable as stated
Misra: AI-generated marketing videos outperform human actors via infinite variation testing
“And very quickly for us, it became, wait, this actually performs better than if we were to, you know, hire someone to like make this video basically. Right. And the reason is because we can customize it to an unlimited ability, right? Like we can get all the v…”
Gaurav Misra Feb 20, 2025 ▶ 44:50
Assertion Not checkable as stated
Captions parallelizes video rendering via 25-frame segments and overlapping GPU clusters
“We were splitting the video into like 25 frame segments basically, and 25 is 25 FPS. So that's basically a second. And then we were generating overlaps on both sides, four frame overlaps with the next segment. And then we would check for differences, the overl…”
Gaurav Misra Feb 20, 2025 ▶ 46:29
Assertion Partly supported
Misra: Nvidia H100 and A100 GPUs encode video slower than T4s
“And by the way, like the H 100 and A 100 kind of suck at encoding and decoding video, right? They're actually slower than like a T four, for example, which has like media sort of like drivers and stuff that can actually like do it really quickly.”
Gaurav Misra Feb 20, 2025 ▶ 48:48
Insight
Misra: Video encoding often takes longer than AI processing in video pipelines
“Oftentimes we'found that the video encoding decoding takes longer than the AI processing, you know, which is like, oh my God, like we're waiting a minute for the video to encode.”
Gaurav Misra Feb 20, 2025 ▶ 50:48
Insight
Gaurav Misra: AI hallucinations are creative features in video generation
“I think hallucinations in text obviously can be a bad thing, potentially, because they're saying stuff that, you know, isn't true, but in a video, like, you are delegating a lot of decisions to the model, right?”
Gaurav Misra Feb 20, 2025 ▶ 51:08
Prediction Not checkable as stated
Gaurav Misra: Indistinguishable AI video generation is a couple years away
“I do think the point of, you know, complete perfection in terms of generation where there is absolutely no way to tell is still a couple of years out, especially in terms of this type of a role video because a lot of the core problems haven't even been fully s…”
Gaurav Misra Feb 20, 2025 ▶ 55:35
Assertion Not checkable as stated
Misra: Captions AI videos get tens of millions of views without viewers noticing AI
“We run these things in ads like on scale. We also run these things on social media. People, big creators use, you know, these models and they get tens of millions, sometimes many, many tens of millions of views on their videos and nobody can tell that it was g…”
Gaurav Misra Feb 20, 2025 ▶ 55:58
Insight
Misra: AI lip-dubbed foreign language video ads perform as well as English originals
“Actually just dubbing it literally with LibDub. And now it's like the same exact script, just in a different language delivered the exact same way. Right. Actually performs just as well as the English version. Right. Which is like, which was a crazy discovery …”
Gaurav Misra Feb 20, 2025 ▶ 58:02
Prediction Not checkable as stated
Gaurav Misra: AI will likely create more professional video editors
“Our goal actually isn't to disrupt professional video editing, right? It's actually to enable, and it actually probably creates more professional video editors because a lot more people can get into the profession, right?”
Gaurav Misra Feb 20, 2025 ▶ 58:46
Prediction Not checkable as stated
Gaurav Misra: AI could allow a solo editor to create theatrical movies
“I think for video editors, there's a small chance that it actually becomes the most important profession of them all. It actually is possibly the profession that rules them all because suddenly one video editor sitting in their basement can make A movie that g…”
Gaurav Misra Feb 20, 2025 ▶ 1:00:08
Assertion Supported
Misra: Current AI character consistency is limited to face consistency
“Everything from character consistency, which we, you know, everybody talks about quite a bit, but really what it is today is just face consistency, right?”
Gaurav Misra Feb 20, 2025 ▶ 1:00:53
Opinion
Misra: Video generation is probably the most dangerous AI unlock
“Because I think of all the possible sort of AI unlocks, like even compared to LLMs and stuff, there's probably nothing more dangerous than video. Generation in a way, right?”
Gaurav Misra Feb 20, 2025 ▶ 1:01:59
Opinion
Misra: AI video has no positive use cases in documentary footage
“But on the documentation side, there's no good thing that AI video can do. Not one good thing, right? Like there's, it's all bad. There's nothing good.”
Gaurav Misra Feb 20, 2025 ▶ 1:03:31
Assertion Not checkable as stated
Captions uses under 1% of proprietary data in current models
“So I think in our current biggest model, we're only using like less than one percent of our data that we have, our proprietary data.”
Gaurav Misra Feb 20, 2025 ▶ 1:04:09
Assertion Not checkable as stated
Misra: Captions' AI model and infrastructure costs decrease over 10x yearly
“So it started off being more expensive and it's only the price has gone down, right? It's literally only gone down and it's gone down more than 10 times every year that this You know that we continue down this journey.”
Gaurav Misra Feb 20, 2025 ▶ 1:07:51
Disclosure
Misra: Captions would be cash flow positive without long-term model training R&D
“Today, like, You know, for us, COGS is like the least of our concerns, basically. I mean, to give you an idea, like if we were to remove a lot of the sort of capital investment in you know, model training and things that are long-term value for us, like we wou…”
Gaurav Misra Feb 20, 2025 ▶ 1:08:27
Disclosure
Captions transitioned from remote work to a five-day in-person office culture
“We were a fully remote company. But as soon as COVID started exiting, we realized that to start a company, there's just nothing like being in the same room. And so we started doing the in-person thing. First start with one day, then two days, then three days. …”
Gaurav Misra Feb 20, 2025 ▶ 1:09:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.