The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Yi Tay: Gap Between Closed AI Labs and Open-Source Is Increasing
“I think the gap is definitely increasing.”
Yi Tay: IR and RecSys research lags significantly behind NeurIPS and ICML
“Also the IR community and the retrieval community is also like always behind the mainstream. And then now it's just probably gotten even more worse because of ILM and stuff. So, okay, I'm getting into Hottick territory, but it's just, like, certain conferences…”
Tay: Frontier AI researchers cannot maintain standard nine-to-five work-life balance
“You cannot be, like, checking out on, like, Friday, Saturday, Sunday, and, like, work at, like, nine to five if you want to, like, Make progress, or like, some people are just so good at detaching, like, ok, like, you know, like, eight pm, I'm not going to, my…”
Tay: Long context architecture is the future of AI over RAG
“And, yeah, I mean, I think long context is definitely the future, rather than rec. But I mean, they could be used in conjunction, like,”
DeepMind Abandoned AlphaProof to Run Gemini End-to-End for IMO Math
“We wanted to try to, like, use, actually use Gemini as an end-to-end model. Basically, no, no second system with alpha proof. No second system. In, text out.”
Tay: Most specialized tools will be subsumed directly into model parameters
“Then the most I can see in the future is there'll be a model then that, that is, there's something that really cannot be subsumed by a model. Then you just use a tool or something, right? But my prediction is that I think most things can be subsumed by the mod…”
Yi Tay: Today's Models Likely Couldn't Invent the Transformer from Pre-2015 Data
“Even today's models, they might not even be able to invent the transformer. Like, if you freeze the time at a certain time, and even you bring the time, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage.”
Yi Tay Fixes ML Bugs Automatically Using AI Without Reading Error Traces
“I think AI coding has started to become the point where I run a job, I get a bug. I almost don't look at the bug. I paste it into, like, anti-gravity, and, like, I throw it, that will fix the bug for me. And then I relaunched the job. And, like, beyond, like, …”
Yi Tay: The Architecture That Achieves AGI Will Still Be a Transformer
“It will be a transformer, I think. Like people, it depends on what you call it, but I think unless the paradigm shifts completely, which is, I mean, as a scientist, you cannot like completely say no to like that, this would never happen. But my feeling is that…”
Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the to…”
Yi Tay: The AI Industry Is Stuck in a Transformer Local Minimum
“So now we are like in this local minima of like transformers, everything, everything, right? Maybe it's not easy to like get totally out Of this, because also a lot of people's investment optimization have been done. So the things that play well needs to play …”
Yi Tay: The 'Bitter Lesson' Is Overapplied; Architectural Ideas Fundamentally Matter
“The bitter lesson gets used too much in, like, too conveniently used around, but actually there's also a little bit of a, not a bit, there's also a sweet lesson where it's like, ideas matter.”
Yi Tay: AI Research Has Not Entered Diminishing Returns on Ideas
“The number of ideas that actually work is not decreasing compared to the last, like, we're not in the era of diminishing returns yet.”
Tay: Newly provisioned GPU clusters are highly unreliable and require node filtering
“Usually when, like, a provider, like, provisions new notes, or they would, like give us... Yeah, it's usually, like, bad, like, dog shit, like, at the start. And then it gets, like, better as you go through the process of, like, returning notes, like, and, you…”
Yi Tay: Llama 3 shows Meta may have caught up to Google
“So I think I don't really follow, like, fine much, but I think that, like, Lama Tree actually shows that, like, kind of, like, Meta got a pretty, like, a good stack around training these models you know, like, oh, and I've even started to feel like, oh, they a…”
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Tay: Mixture-of-Experts is fundamentally the right architecture for scaling
“Fundamentally, I just think that MOEs are just, like, the way to go in terms of, like, floppyram ratio, they bring the benefit from the scaling curve, if you do it right, if you, they bring the benefit from the scaling curve, right, and then, Like, that's, lik…”
Tay: Hugging Face's Open LLM Leaderboard is a major problem
“The open LM leaderboard is, like, probably, like, the, a big, like, Problem, to be honest.”
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”
Gemini's IMO Model Checkpoint Required Only One Week of Training
“The training process of this IMO model itself was, like, maybe a week or so.”
Tay: Present AI Milestones Would Have Been Viewed as AGI Five Years Ago
“If you just look at the AI progress now and five years ago, I think people would think that we already reached like AGI.”
Yi Tay: AI model thoughts do not need to resemble human thoughts
“Generally, I'm not really, I don't really believe that model thoughts have to be the same with human thoughts. I'm actually like, generally in ML, I'm more of the school of thought of let the model do whatever it wants.”
Yi Tay: AI Tools Act as a Team Productivity Aura Rather Than Replacing Engineers
“These things are not like going to replace one person as it is, but more like a passive aura that buffs everybody.”
Yi Tay: AI Model Laziness and Edge Flaws Will Disappear via General Scaling
“I don't think there's anything that to be done to specifically like focus fire. These things is more like general capability improvements. The models just get better over time and then these things will just like go away.”
Yi Tay: AI progress is driven by compounding small incremental changes
“I think that it's true that sometimes a lot of progress on the whole is just a series of small incremental changes that, yeah, that push. I think that's accurate. That's true. There's also, it also feels that there's also a lot of like small, like seemingly mi…”
Yi Tay: The term 'world models' is not well-defined in AI
“I don't think about world models that often. I think because world models are just not really well defined in the first place.”
Yi Tay: Data-Bound AI Must Spend More Compute Per Token to Scale
“So if you are, you come to a point where you are Very data bound, but not compute bound at all. You just find algorithms that spend a lot of compute on every token.”
Yi Tay: Top Southeast Asian AI talent requires prominent leadership to unlock
“I feel like the talent we can get from the region is really, really good, but it's only It's only because it's us, we can unlock this talent. Otherwise, might join some other place.”
Yi Tay: AI research benefits from escaping the Bay Area monoculture
“I do think that to some extent, if you want to do research in, you need a little bit of peace and quiet somewhere, right? So this island may be good for that, but then you can, you're still, like, able to, like, be connected, right?”
Yi Tay: ML and RL Knowledge Can Be Learned Easily by Engineers
“ML. ML can be learned easily. Our knowledge can be learned easily.”
Tay: Underlying AI research principles have not changed much beyond compute scale
“Fundamentally, I don't think, like, the, like, the stuff has actually, like, the underlying principles of research hasn't really changed that much, except for, like, compute.”
Tay: ChatGPT's release made task-specific academic NLP research obsolete
“The big thing about the ChatGPT moment of, like, twenty-twenty-two, the thing that changed drastically is, like, it completely, like, it was, like, this sharp, like, make all this work, like, kind of, like, obsolete”
Tay: Google and OpenAI built general models three years before academia
“Places like Google and Meta, OpenAI, we will be working on things, like, Three years ahead of everybody else, and then suddenly, like, then Academia would be, like, still working on, like, these task-specific things.”
Tay: Academic best paper awards at AI conferences are completely meaningless
“Does best paper awards, like, mean anything? Actually, it doesn't mean anything, right, like, but like, I think that was more of, like, my, Like, where my angst was coming from, right?”
Tay: Inflection AI is completely gone and effectively defunct
“I wouldn't have left, like, for inflection, or something like that. Like, I mean, inflection is gone now. RIP.”
Tay: The biggest green flag for GPU providers is sharing node failure costs
“If you do it like a, like, compute startup or anything, the biggest green flag would be to share the cause of node failures with, ah, with, ah, your customers, right”
Tay: Encoder-decoders provide 2x parameter capacity at matched FLOPs
“The only big benefit of encoder decoders
[4425] Yi Tay: Is that it has this thing called, like, I mean, what I like to call intrinsic sparsity.
[4430] Podcast Host: Ok.
[4431] Yi Tay: So basically, an encoder decoder with, like, n baramps is, like, basically, …”
Tay: Serious AI labs should never release their good evaluation benchmarks
“Serious LMS that create their own evals, and they, a good eval set is one that you don't release. A good eval set is the one that you, like, ok, you release some of it, but, like, it's like, you don't, like, you know, let it be contaminated by the community.”
Tay: Multimodal AI architectures will eventually move completely to early fusion
“As early fusion models get more traction, I think the themes will start to get more and more, like, it's a bit like how all the tasks like unify, like from Like, two zero one nine to, like, now it's like all the tasks are unifying, now it's like all the modali…”
Tay: Vision models will unify screen intelligence and natural imagery without bifurcating
“I think at the end of the day, like, the models would become, like, I don't see that there will be, like, screen agents and, like, natural images. Humans, like, you can read what's on a screen, you can go out and appreciate the scenery, right? You're not, like…”
Yi Tay: Meta's Llama is corporate open weights, not grassroots open source
“To me, Lama Tree is like... Meta has an org that is hypothetically very similar to Gemini or something but they just decide to release the weights It's open weights It's open weights and everything”
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Tay: Researchers can rely on the Twitter algorithm to surface important papers
“You actually don't have to follow anything. If the paper is important enough, the Twitter algorithm will give it to you.”
Tay: Singapore AI research prioritizes paper counts over real-world impact
“I think, to be honest, the research here is, like, in Singapore is just basically, like, they just care about publishing papers and stuff like that and then it's not, like, impact-driven. I think, at U.S., it's mostly focused on impact-driven, and the thing ne…”
Yi Tay: Governments cannot artificially manufacture top AI talent ecosystems
“I don't think there's actually much, like, the government can do to, like, influence, like, this kind of thing is, like, a natural, like, organic, natural thing, right? The worst thing to do is probably, like, to create, like, create a lot of artificial things…”
Yi Tay: Reinforcement learning is the primary AI modeling toolset today
“So I think RL is basically the main modeling tool set that we play around with these days.”
DeepMind Shipped Full Gemini IMO Config Only to Select Mathematicians
“So the inference time config was like the one serve to most people is different, but And the full IMO, like, inference config was presented, like, shipped to some mathematicians just because of the inference cost, right? But that was good enough to be a genera…”