Transformer Models

topic on 8 shows · 13 statements across 13 episodes

Acquired Cheeky Pint Latent Space the Neon Show No Priors the MAD Podcast Big Technology 20VC

13 statements about Transformer Models, every show

Kutylowski: Task-specific AI models outperform general LLMs due to dedicated parameter capacity
“When they're made for different purposes, they also lose a little bit of the capability that they had maybe initially when they've been made for translation only. This The set of parameters that is available there, this, which kind of determines quite often th…”
Jarek Kutylowski Jul 7, 2026 ▶ 4:06 Why Specialized AI Models Are Challenging the Frontier Labs — With DeepL CEO Jarek Kutylowski
CHEEKY PINT Assertion Not checkable as stated
Staniszewski: ElevenLabs applied transformer and diffusion concepts to voice
“And here credit to my co-founder, Piotr, who effectively came with that new idea of how you can now create voice models, which are both reliable, high quality, quick, where you would bring a lot of the ideas from transformer models, from diffusion models into …”
Mati Staniszewski Apr 14, 2026 ▶ 1:47 The world of voice AI, with Mati Staniszewski of ElevenLabs
CHEEKY PINT Prediction Not checkable as stated
Collison: Everything will pass through a transformer model before consumption
“Everything will pass through a transformer model before it's consumed.”
Patrick Collison Aug 27, 2025 ▶ 44:00 A Cheeky Pint with Cognition CEO Scott Wu
NEON SHOW Insight
Namasivayam: Transformer models miss edge-case opinions by defaulting to population averages
“Transformer models, OpenAI, is great at providing feedback on the average, but it doesn't cover the edge cases. Or the edge cases of your opinions. So that's the problem. So ChatGPT is good at persona feedback. It kind of averages it across the population. But…”
Vasanth Namasivayam Jun 27, 2025 ▶ 27:37 How NVIDIA, Meta & Dropbox Taught Me To Build Great Products | Vasanth, Founder - Featurely
LATENT SPACE Assertion Supported
Jamil: 'Writing in the Margins' works on any transformer without fine-tuning
“So it can be used with any transformer model without fine-tuning, just by doing it, just by doing this inference differently.”
Umar Jamil Sep 19, 2024 ▶ 13:32 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
MAD Disclosure
Moody's automated feature engineering using transformer models on text data
“Then we actually automated the feature engineering process, for example, using transformer models to extract features from massive, massive text data.”
Yimei Fan Dec 21, 2023 ▶ 6:03 How Moody’s Analytics Is Using AI to Transform Credit Risk | Cristina Pieretti & Yimei Fan
Ramaswamy: Google open-sourced Transformers in 2017 to prevent top researchers from leaving
“If you had told these people at that time that they could not publish, they would have gone and worked for universities, they would have gone and worked for Microsoft, they would have gone and worked for other people.”
Sridhar Ramaswamy Sep 6, 2023 ▶ 24:17 Google’s Weird Year + Neeva Goes to Snowflake — With Sridhar Ramaswamy
ACQUIRED Insight
Scaling compute and data created unpredicted emergent reasoning in language models
“We don't change anything about the structure, we just give it way more data and let it Run these models for a long time and make the parameters of the model way bigger, and like, no researchers expected them to reason about the world as well as they do, but it…”
Ben Gilbert Sep 6, 2023 ▶ 42:44 Nvidia Part III: The Dawn of the AI Era (2022-2023) (Audio) · Acquired
20VC Insight
Tobi Lütke: Tech leaders must understand AI models, not use black boxes
“I need to not treat transformer models as a black box. Like I need to understand those because otherwise I cannot show up to a meeting with my teams and have a good chance of reasoning with them about not just what is sort of a convenient next task, but what's…”
Tobi Lütke Jun 9, 2023 ▶ 14:10 20VC: What are the World's Tech Leaders Running From? Fear? Insecurity? Poverty? What Drives the Best with Orlando Bravo, Bill Ackman, Dara Khosrowshahi, Parker Conrad, Tobi Luttke, Brian Armstrong and more..
NO PRIORS Insight
Liang: Language models should use calculators instead of computing internally
“There are cases where you want to just map natural language into say people call it tool use. Like you ask some question that reverse calculation, you should just use a calculator rather than trying to sort of quote unquote do it in the transformers head.”
Dr. Percy Liang Apr 25, 2023 ▶ 16:09 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Marcus: Pure transformer models lack mechanisms for truth and inherently hallucinate
“If you look at a completely pure case of a transformer model trained on a bunch of data, It doesn't have any mechanisms for truth. Now, except the sort of accidental contingency, and there are inherent reasons why these systems hallucinate, and maybe I can, in…”
Gary Marcus Feb 23, 2023 ▶ 4:02 Blake Lemoine and Gary Marcus Debate AI Chatbots
MAD Assertion Supported
Pesenti: Facebook AI transformer models matched or beat Mathematica at math
“My team also applied transformer models to mathematics, you know, doing, like, things like partial derivation equation or integration and showing that, hey, it could work as well better than Mathematica or some, you know, hundred page long algorithm.”
Jerome Pesenti Jun 10, 2020 ▶ 40:49 Fireside Chat: Jerome Pesenti (Head of AI, Facebook) with Matt Turck (Partner, FirstMark)
MAD Assertion Supported
Delangue: Google called transformer adoption one of its biggest changes ever
“Google did a lot of PR a few months ago saying that moving to these new transformer models has been one of the most impactful changes that they'done over the last five years, if not since the beginning of Google.”
Clement Delangue Jan 22, 2020 ▶ 11:44 NLP—The Most Important Field of ML // Clement Delangue, Hugging Face (FirstMark's Data Driven NYC)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.