Insight certainty 3/5 debate potential 2/5

Liang: Interconnected AI accepting external inputs risks cascading jailbreak exploits

Dr. Percy Liang · No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang · Apr 25, 2023 · at 46:38

Stanford Professor Dr. Percy Liang discusses emerging AI safety and security vulnerabilities as language models are deployed with external tool use and web access.

0:00 / 0:17exact quote · 17.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“If these models start interacting with the world and accepting external inputs, now you can not only just sort of jailbreak your own model, but you can jailbreak other people's model and get them to do various things. And then, so that could lead to sort of a cascade of errors.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dr. Percy Liang

Assertion Not checkable as stated
Liang: The AI industry is retreating from its open culture
“And what we're seeing now is sort of a retreat of that open culture where models are now being only accessible via APIs. We don't really know all the secret sauce that's going behind them, and there's sort of limited access.”
Dr. Percy Liang Apr 25, 2023 ▶ 5:15 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Insight
Liang: Big tech scale forced AI academia to focus on understanding models
“And now today I think it's the dynamic is, is quite different because it's no longer academia's job isn't just to get things to work because you can do that in other ways. There's a lot of resources going into big tech companies where there's if you have data …”
Dr. Percy Liang Apr 25, 2023 ▶ 7:58 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Insight
Liang: AI benchmarks should target superhuman reliability over human mimicry
“I think we're getting to a point where along many axes, it's a superhuman or should be superhuman. And I think we should maybe define more of an objective measure of like what we actually want. We want something that's very reliable, is grounded. You know, I o…”
Dr. Percy Liang Apr 25, 2023 ▶ 12:21 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Opinion
Liang: LLMs are not just memorizing because novel concept fusion requires invention
“You know, people say that sometimes all language models just memorize because they're so big and train on clearly a lot of texts, but these examples, I think really indicate that there's no way That these language models are just memorizing because this text j…”
Dr. Percy Liang Apr 25, 2023 ▶ 21:53 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Insight
Liang: Next-token prediction forces language models to build world models
“If you think about predicting the next word, It's, it seems very simple, but you have to really internalize a lot of what is going on in this context. What are the previous words? What's the syntax? What's who's saying them? And all of that information and con…”
Dr. Percy Liang Apr 25, 2023 ▶ 24:54 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Insight
Liang: Foundation Models Cause the Concept of an AI Task to Dissolve
“And this was a paradigm shift, in my opinion, because it changed the way that we conceptualize Machine learning and NLP systems from these bespoke systems where you're, it's trained to do question answering, to train to do this, to just a general substrate whe…”
Dr. Percy Liang Apr 25, 2023 ▶ 3:00 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.