May 1, 2024 · 41m · big-technology

AWS VP of AI and Data Swami Sivasubramanian — GenAI's Growth Potential

Swami Sivasubramanian · 28m spoken Alex Kantrowitz · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

AWS Vice President of AI and Data Swami Sivasubramanian joins Alex Kantrowitz to examine whether generative AI is approaching practical limits in compute, data, and power. Sivasubramanian outlines AWS's multi-layered strategy to sustain AI progress through custom Trainium silicon, efficient model architectures, Amazon Bedrock, and tailored enterprise solutions.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Alex holds 26.3% of the talking time here. How this is scored →

Alex as informed peer 4.9 Guest teaching 5.1 Guest disagreement 2.0 Alex pushing back 3.8
05100:0015:0030:004:55–8:02 · Alex as informed peer 5/10 Energy Constraints and AWS Data Center Sustainability Alex pushes Swami repeatedly to address Mark Zuckerberg's point on energy constraints rather than allowing him to pivot purely to compute chip innovation. Swami acknowledges energy is a real bottleneck while highlighting AWS's focus on data center sustainability and custom silicon.8:03–13:00 · Alex as informed peer 6/10 Overcoming Data Limits with Enterprise Customization and Bedrock Alex cites reporting on Meta running out of training data and considering buying book publishers to ask if traditional scaling hits a wall. Swami counters that the enterprise future lies not in ever-larger generic models, but in customizing smaller, highly efficient models through Bedrock.13:01–15:55 · Alex as informed peer 6/10 The Debate Over Foundational Model Evolution vs Enterprise Impact Alex openly disagrees with Swami's framing, arguing that enterprise tuning on mortgage data does not advance frontier AI capabilities. Swami defends his position, stating that step-change productivity gains for individual enterprises are equally transformative.15:56–18:23 · Alex as informed peer 4/10 Model Efficiency and Smaller High-Performance LLMs Alex asks how smaller models can outperform previous-generation larger models. Swami provides a technical explanation covering token density, reinforcement recipes, and emerging non-transformer architectures like state-space models.18:23–23:13 · Alex as informed peer 4/10 AWS AI Strategy and the Anthropic Partnership Alex asks about the strategic rationale of AWS building internal LLMs while heavily backing Anthropic. Swami explains AWS's long-standing customer choice philosophy dating back to specialized databases, emphasizing that complex AI applications require multiple specialized models.23:14–25:38 · Alex as informed peer 5/10 Custom AI Silicon: Trainium and Inferentia Chips Alex probes the competitive viability of Trainium against Nvidia's CUDA ecosystem and mentions Groq's fast inference. Swami outlines Trainium2 and Inferentia cost-performance specs, while politely declining to comment in depth on Groq.25:39–29:58 · Alex as informed peer 4/10 The Mechanics of AI Reasoning and Amazon Q Alex inquires about the mechanics of frontier model reasoning. Swami delivers an extensive breakdown of SWE-bench, step-by-step plan generation in Amazon Q, and why real-world reasoning relies on coupling LLMs with deterministic tools like compilers.30:00–33:36 · Alex as informed peer 5/10 Cloud vs On-Device LLMs and Hierarchical Inference Alex cites Apple's open-source OpenELM models to ask if on-device inference threatens the cloud computing model. Swami educates on hierarchical inference using Alexa as a historical precedent, showing how edge and cloud complement each other.4:55–8:02 · Guest teaching 4/10 Energy Constraints and AWS Data Center Sustainability Alex pushes Swami repeatedly to address Mark Zuckerberg's point on energy constraints rather than allowing him to pivot purely to compute chip innovation. Swami acknowledges energy is a real bottleneck while highlighting AWS's focus on data center sustainability and custom silicon.8:03–13:00 · Guest teaching 5/10 Overcoming Data Limits with Enterprise Customization and Bedrock Alex cites reporting on Meta running out of training data and considering buying book publishers to ask if traditional scaling hits a wall. Swami counters that the enterprise future lies not in ever-larger generic models, but in customizing smaller, highly efficient models through Bedrock.13:01–15:55 · Guest teaching 4/10 The Debate Over Foundational Model Evolution vs Enterprise Impact Alex openly disagrees with Swami's framing, arguing that enterprise tuning on mortgage data does not advance frontier AI capabilities. Swami defends his position, stating that step-change productivity gains for individual enterprises are equally transformative.15:56–18:23 · Guest teaching 6/10 Model Efficiency and Smaller High-Performance LLMs Alex asks how smaller models can outperform previous-generation larger models. Swami provides a technical explanation covering token density, reinforcement recipes, and emerging non-transformer architectures like state-space models.18:23–23:13 · Guest teaching 4/10 AWS AI Strategy and the Anthropic Partnership Alex asks about the strategic rationale of AWS building internal LLMs while heavily backing Anthropic. Swami explains AWS's long-standing customer choice philosophy dating back to specialized databases, emphasizing that complex AI applications require multiple specialized models.23:14–25:38 · Guest teaching 5/10 Custom AI Silicon: Trainium and Inferentia Chips Alex probes the competitive viability of Trainium against Nvidia's CUDA ecosystem and mentions Groq's fast inference. Swami outlines Trainium2 and Inferentia cost-performance specs, while politely declining to comment in depth on Groq.25:39–29:58 · Guest teaching 7/10 The Mechanics of AI Reasoning and Amazon Q Alex inquires about the mechanics of frontier model reasoning. Swami delivers an extensive breakdown of SWE-bench, step-by-step plan generation in Amazon Q, and why real-world reasoning relies on coupling LLMs with deterministic tools like compilers.30:00–33:36 · Guest teaching 6/10 Cloud vs On-Device LLMs and Hierarchical Inference Alex cites Apple's open-source OpenELM models to ask if on-device inference threatens the cloud computing model. Swami educates on hierarchical inference using Alexa as a historical precedent, showing how edge and cloud complement each other.4:55–8:02 · Guest disagreement 2/10 Energy Constraints and AWS Data Center Sustainability Alex pushes Swami repeatedly to address Mark Zuckerberg's point on energy constraints rather than allowing him to pivot purely to compute chip innovation. Swami acknowledges energy is a real bottleneck while highlighting AWS's focus on data center sustainability and custom silicon.8:03–13:00 · Guest disagreement 3/10 Overcoming Data Limits with Enterprise Customization and Bedrock Alex cites reporting on Meta running out of training data and considering buying book publishers to ask if traditional scaling hits a wall. Swami counters that the enterprise future lies not in ever-larger generic models, but in customizing smaller, highly efficient models through Bedrock.13:01–15:55 · Guest disagreement 4/10 The Debate Over Foundational Model Evolution vs Enterprise Impact Alex openly disagrees with Swami's framing, arguing that enterprise tuning on mortgage data does not advance frontier AI capabilities. Swami defends his position, stating that step-change productivity gains for individual enterprises are equally transformative.15:56–18:23 · Guest disagreement 1/10 Model Efficiency and Smaller High-Performance LLMs Alex asks how smaller models can outperform previous-generation larger models. Swami provides a technical explanation covering token density, reinforcement recipes, and emerging non-transformer architectures like state-space models.18:23–23:13 · Guest disagreement 1/10 AWS AI Strategy and the Anthropic Partnership Alex asks about the strategic rationale of AWS building internal LLMs while heavily backing Anthropic. Swami explains AWS's long-standing customer choice philosophy dating back to specialized databases, emphasizing that complex AI applications require multiple specialized models.23:14–25:38 · Guest disagreement 2/10 Custom AI Silicon: Trainium and Inferentia Chips Alex probes the competitive viability of Trainium against Nvidia's CUDA ecosystem and mentions Groq's fast inference. Swami outlines Trainium2 and Inferentia cost-performance specs, while politely declining to comment in depth on Groq.25:39–29:58 · Guest disagreement 1/10 The Mechanics of AI Reasoning and Amazon Q Alex inquires about the mechanics of frontier model reasoning. Swami delivers an extensive breakdown of SWE-bench, step-by-step plan generation in Amazon Q, and why real-world reasoning relies on coupling LLMs with deterministic tools like compilers.30:00–33:36 · Guest disagreement 2/10 Cloud vs On-Device LLMs and Hierarchical Inference Alex cites Apple's open-source OpenELM models to ask if on-device inference threatens the cloud computing model. Swami educates on hierarchical inference using Alexa as a historical precedent, showing how edge and cloud complement each other.4:55–8:02 · Alex pushing back 6/10 Energy Constraints and AWS Data Center Sustainability Alex pushes Swami repeatedly to address Mark Zuckerberg's point on energy constraints rather than allowing him to pivot purely to compute chip innovation. Swami acknowledges energy is a real bottleneck while highlighting AWS's focus on data center sustainability and custom silicon.8:03–13:00 · Alex pushing back 4/10 Overcoming Data Limits with Enterprise Customization and Bedrock Alex cites reporting on Meta running out of training data and considering buying book publishers to ask if traditional scaling hits a wall. Swami counters that the enterprise future lies not in ever-larger generic models, but in customizing smaller, highly efficient models through Bedrock.13:01–15:55 · Alex pushing back 7/10 The Debate Over Foundational Model Evolution vs Enterprise Impact Alex openly disagrees with Swami's framing, arguing that enterprise tuning on mortgage data does not advance frontier AI capabilities. Swami defends his position, stating that step-change productivity gains for individual enterprises are equally transformative.15:56–18:23 · Alex pushing back 2/10 Model Efficiency and Smaller High-Performance LLMs Alex asks how smaller models can outperform previous-generation larger models. Swami provides a technical explanation covering token density, reinforcement recipes, and emerging non-transformer architectures like state-space models.18:23–23:13 · Alex pushing back 2/10 AWS AI Strategy and the Anthropic Partnership Alex asks about the strategic rationale of AWS building internal LLMs while heavily backing Anthropic. Swami explains AWS's long-standing customer choice philosophy dating back to specialized databases, emphasizing that complex AI applications require multiple specialized models.23:14–25:38 · Alex pushing back 3/10 Custom AI Silicon: Trainium and Inferentia Chips Alex probes the competitive viability of Trainium against Nvidia's CUDA ecosystem and mentions Groq's fast inference. Swami outlines Trainium2 and Inferentia cost-performance specs, while politely declining to comment in depth on Groq.25:39–29:58 · Alex pushing back 1/10 The Mechanics of AI Reasoning and Amazon Q Alex inquires about the mechanics of frontier model reasoning. Swami delivers an extensive breakdown of SWE-bench, step-by-step plan generation in Amazon Q, and why real-world reasoning relies on coupling LLMs with deterministic tools like compilers.30:00–33:36 · Alex pushing back 5/10 Cloud vs On-Device LLMs and Hierarchical Inference Alex cites Apple's open-source OpenELM models to ask if on-device inference threatens the cloud computing model. Swami educates on hierarchical inference using Alexa as a historical precedent, showing how edge and cloud complement each other.

speaking balance: gold is Alex, purple is the guest (3 minute bins)

0:00 · Alex 67.6% · guest 32.4%0:00 · Alex 67.6% · guest 32.4%3:00 · Alex 26.9% · guest 73.1%3:00 · Alex 26.9% · guest 73.1%6:00 · Alex 42.2% · guest 57.8%6:00 · Alex 42.2% · guest 57.8%9:00 · Alex 31.1% · guest 68.9%9:00 · Alex 31.1% · guest 68.9%12:00 · Alex 20.8% · guest 79.2%12:00 · Alex 20.8% · guest 79.2%15:00 · Alex 13.2% · guest 86.8%15:00 · Alex 13.2% · guest 86.8%18:00 · Alex 14.7% · guest 85.3%18:00 · Alex 14.7% · guest 85.3%21:00 · Alex 24% · guest 76%21:00 · Alex 24% · guest 76%24:00 · Alex 23.4% · guest 76.6%24:00 · Alex 23.4% · guest 76.6%27:00 · Alex 0.6% · guest 99.4%27:00 · Alex 0.6% · guest 99.4%30:00 · Alex 29.9% · guest 70.1%30:00 · Alex 29.9% · guest 70.1%33:00 · Alex 18.2% · guest 81.8%33:00 · Alex 18.2% · guest 81.8%36:00 · Alex 31.3% · guest 68.7%36:00 · Alex 31.3% · guest 68.7%39:00 · Alex 24% · guest 76%39:00 · Alex 24% · guest 76%
Sharpest disagreement ▶ 13:20 Swami firmly rejects Kantrowitz's characterization

Swami immediately counters Alex's assertion that he is ignoring frontier AI evolution, clarifying that customer application customization is where actual enterprise value and step-change innovation occur.

Hardest push from Alex ▶ 13:01 Alex directly challenges the enterprise customization pivot

Alex openly tells Swami he disagrees with his premise, asserting that tuning on Rocket Mortgage data does not advance the frontier of AI capabilities.

Biggest teaching moment ▶ 28:20 Swami details compiler-augmented reasoning systems

Swami demystifies LLM reasoning by explaining that practical reasoning in tools like Amazon Q requires interleaving language models with external compilers and feedback loops rather than relying on pure LLM generation.

Alex holds their own ▶ 8:03 Alex details Llama 3 compute clusters and data exhaustion

Alex demonstrates deep reporting knowledge by citing Meta's massive cluster scale and internal considerations of purchasing book publishers to challenge the limits of traditional scaling.

the scores for every segment, with the reasoning behind each
ChapterTopicAlex as informed peerGuest teachingGuest disagreementAlex pushing backWhy
Energy Constraints and AWS Data Center Sustainability 5426 Alex pushes Swami repeatedly to address Mark Zuckerberg's point on energy constraints rather than allowing him to pivot purely to compute chip innovation. Swami acknowledges energy is a real bottleneck while highlighting AWS's focus on data center sustainability and custom silicon.
Overcoming Data Limits with Enterprise Customization and Bedrock 6534 Alex cites reporting on Meta running out of training data and considering buying book publishers to ask if traditional scaling hits a wall. Swami counters that the enterprise future lies not in ever-larger generic models, but in customizing smaller, highly efficient models through Bedrock.
The Debate Over Foundational Model Evolution vs Enterprise Impact 6447 Alex openly disagrees with Swami's framing, arguing that enterprise tuning on mortgage data does not advance frontier AI capabilities. Swami defends his position, stating that step-change productivity gains for individual enterprises are equally transformative.
Model Efficiency and Smaller High-Performance LLMs 4612 Alex asks how smaller models can outperform previous-generation larger models. Swami provides a technical explanation covering token density, reinforcement recipes, and emerging non-transformer architectures like state-space models.
AWS AI Strategy and the Anthropic Partnership 4412 Alex asks about the strategic rationale of AWS building internal LLMs while heavily backing Anthropic. Swami explains AWS's long-standing customer choice philosophy dating back to specialized databases, emphasizing that complex AI applications require multiple specialized models.
Custom AI Silicon: Trainium and Inferentia Chips 5523 Alex probes the competitive viability of Trainium against Nvidia's CUDA ecosystem and mentions Groq's fast inference. Swami outlines Trainium2 and Inferentia cost-performance specs, while politely declining to comment in depth on Groq.
The Mechanics of AI Reasoning and Amazon Q 4711 Alex inquires about the mechanics of frontier model reasoning. Swami delivers an extensive breakdown of SWE-bench, step-by-step plan generation in Amazon Q, and why real-world reasoning relies on coupling LLMs with deterministic tools like compilers.
Cloud vs On-Device LLMs and Hierarchical Inference 5625 Alex cites Apple's open-source OpenELM models to ask if on-device inference threatens the cloud computing model. Swami educates on hierarchical inference using Alexa as a historical precedent, showing how edge and cloud complement each other.

Statements from this episode (15)

Opinion
AWS's Sivasubramanian: AI scaling is not close to hitting a wall
“One, I don't think we are anywhere close to yet hitting that wall yet. So I do think we are going to actually see a lot of net new innovations when it comes to optimizing how to actually train these models in parallel and get better Utilization so that we can …”
Swami Sivasubramanian May 1, 2024 ▶ 3:24
Assertion Supported
Sivasubramanian: Trainium2 offers 4x faster training and 2x energy efficiency
“AWS actually, for instance, we have invested in technologies such as Trinium-II chips, which actually can deliver up to four X faster training. And, ah, also it is while improving energy efficiency up to two times, as an example.”
Swami Sivasubramanian May 1, 2024 ▶ 3:49
Prediction Not checkable as stated
AWS AI VP: New model architectures will reshape AI scaling laws
“And if you see there are starting to see new kinds of architectures that are going to change how these models scale in the future. And I suspect, ah, that is going to change, ah, how These scaling laws are going to be reshaped as well when it comes to that is …”
Swami Sivasubramanian May 1, 2024 ▶ 4:25
Prediction Not checkable as stated
AWS AI Chief: AI architectures beyond pure transformers are absolutely necessary
“I actually think, ah, new architectures are absolutely necessary in the future, ah, and, ah, there are already hybrid architectures evolving. I mean, ah, I mean in terms of state space models to actually connectors hybrid models between state space to transfor…”
Swami Sivasubramanian May 1, 2024 ▶ 9:58
Insight
AWS AI Chief: Enterprise AI needs smaller distilled models, not monolithic LLMs
“To put these LLMs to work to solve real world problems, you got to actually take some of these models and then customize it, and the end result is not the biggest model. It is actually a much more customized, smaller model or a distal model to solve specific b…”
Swami Sivasubramanian May 1, 2024 ▶ 12:33
Opinion
Kantrowitz: Training on enterprise data will not advance the AI field
“If you start changing with trading with Rocket Mortgage data, you might just get a better application for Rocket Mortgage. You're not gonna advance the field of AI by getting mortgage data in there.”
Alex Kantrowitz May 1, 2024 ▶ 14:22
Prediction Not checkable as stated
AWS's Sivasubramanian: No Single AI Model Will Rule the World
“And that's why I keep saying no one model and will rule the world in the future.”
Swami Sivasubramanian May 1, 2024 ▶ 18:18
Disclosure
Sivasubramanian: Anthropic will use AWS Trainium and Inferentia for future models
“Anthropic will also use AWS Tranium and Inferentia tips to build, train, and deploy their future models and as well.”
Swami Sivasubramanian May 1, 2024 ▶ 22:44
Assertion Supported
AWS AI Chief: Inferentia chips can cut AI computing costs by 40%
“So the EC two instances that are powered by our Inferentia chips deliver up to 50% better price performance per watt over comparable EC-II as well, and can reduce cost by up to 40% as well.”
Swami Sivasubramanian May 1, 2024 ▶ 24:36
Insight
AWS VP: Pure LLMs Won't Deliver Effective AI Reasoning Without Other Tools
“Reasoning capability is not, if we view it as purely like everything is going to be LLM driven, we as industry are going to be very, very not satisfied. That's why it has to be actually more iterative. And we work with LLM as one of the ingredients. It is not …”
Swami Sivasubramanian May 1, 2024 ▶ 29:39
Prediction Not checkable as stated
AWS VP: LLM architectures will shift to hierarchical edge-and-cloud inference
“So, I expect LLNs to evolve in the same way, where there are a lot of, ah, decisions, especially if you have a powerful computing device, ah, which, ah, many smartphones and others, ah, to be having that, You can actually run some of those, ah, simple LLM thin…”
Swami Sivasubramanian May 1, 2024 ▶ 31:30
Insight
AWS VP: GenAI apps will chain multiple models and data connectors
“The way these generative AI applications are going to be built is actually like series of things chained together. Some of it is model inference. Some of it is actually workflows and whatnot. And that's why that pattern is going to be important. And they are g…”
Swami Sivasubramanian May 1, 2024 ▶ 32:23
Assertion Partly supported
Sivasubramanian: Amazon Bedrock guardrails improve AI model accuracy up to 84%
“We see up to 40, 84 percent improvement in further accuracy because of those capability as well.”
Swami Sivasubramanian May 1, 2024 ▶ 36:22
Opinion
AWS AI Chief: AI's future is assisting humans, not autonomous AGI
“Intelligence is going to be constantly getting better, but it is going to be human assisted. Ah, ah, intelligence assisting humans is going to be the future, rest my guess.”
Swami Sivasubramanian May 1, 2024 ▶ 37:48
Assertion Supported
Sivasubramanian: Tens of thousands of companies use Amazon Bedrock
“Bedrock, which is used by tens of thousands of companies already are with things like Amazon Q and so forth.”
Swami Sivasubramanian May 1, 2024 ▶ 39:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.