Everything Shreya Rajpal said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Production LLM deployment challenges mirror autonomous vehicle development
“The trajectory of issues and concerns that people are running into are similar to, you know, like the path is similar to what it was in self-driving, which is, you know, How do I get like this, you know, runtime safety? How do I get like runtime constraints, e…”
Existing model risk frameworks fail for third-party AI models
“Existing, for example, like model risk management frameworks don't really apply when you haven't built the model yourself. You know, you didn't like curate the data that the model was trained on. And so you can't make any claims to that.”
Fine-tuning cannot eliminate LLM hallucinations
“At the end of the day, these models are like next token predictors, which is, you know, like they kind of look at like what they've predicted until now, and then, you know, figure out like what the next token they're on is. And from that, like, even with fine-…”
GenAI solves ML's first mile, but traditional tools handle the last
“LLMs and Generative AI really helps solve, like the first mile problem in ML, right? But the last mile problem, which is, like, how do you take this generic generalizable technology and make it work, like, specifically for your use cases and for your actual ap…”
Rajpal: Foundation models rarely generate toxic outputs without explicit jailbreaks
“Most of the stuff that the frameworks will recommend is actually stuff that the model providers are already working on. So toxicity, unless you're doing, unless somebody is very explicitly trying to jailbreak what you've built, you know, you won't run into the…”
Rajpal: AI simulations should prioritize product KPIs over generic safety metrics
“So I would actually say that like a lot of the things to simulate are more aligned with like product KPIs or product metrics that actually make Whatever AI system you're building very sticky, rather than, you know, focusing more on, like, traditional safety se…”
Rajpal: Fine-tuning open-source models on synthetic data closes proprietary capability gaps
“Not out of the box, but with a lot of that fine tuning and that the training, et cetera, you are able to kind of close the gap and even have better performance on metrics.”
Fine-tuned Llama 2 achieves performance comparable to GPT-3.5 and GPT-4
“Straight out of the bat, if you just use Lama tool directly, I don't think you could get like comparable performance, you know, with GPD, 3.5 or four today. But like with fine tuning, if you make that investment in curating your data set in running that fine t…”
RAG is the definitive way to build generative AI today
“Rag is the way to build you know, these models today”
Rajpal: AI agents resemble autonomous vehicle architectures with cascading ML units
“What patterns really worked well in self-driving cars, which is weirdly a very similar system to, you know, agents of today where you have like these cascading kind of like units that are all machine learning based and, you know, they all kind of like feed int…”
Rajpal: Simulations reveal which failure modes actually require runtime guardrails
“And then, you know, in simulation, figure out, you know, what is actually robust, what isn't, and then the stuff that isn't robust is the stuff that you need guardrails for.”
Rajpal: Product managers already act as AI persona engineers
“Interestingly, there are already persona engineers, and we call them like product managers, basically, you know. So your product managers are already thinking about, okay, I've built this, you know, model or this chatbot or this agent. Who are the personas? Wh…”
Rajpal: Pure ChatGPT test conversations lack diversity and realism
“Compared to, let's say, you were asking, like, ChatGPT to generate, you know, these, like, conversations for you. They all kind of have that ChatGPT vibe, and this ends up looking, you know, very diverse and very grounded in, like, your use case and your data.”
Rajpal: Historically conservative US banks are becoming very AI-forward
“Even, especially in the US, right, like, a lot of banks that you would think would be, like, historically maybe more conservative, like, maybe not the earliest technology adopters, like, they are very, like, AI forward and tech forward.”
Rajpal: Most AI audio applications are a 'voice sandwich' around text
“I think like most audio applications today are like a text sandwich, or sorry, a voice sandwich with like kind of text in the middle.”
Rajpal: General-Purpose AI Simulators Now Replace Manually Crafted Simulation Systems
“And with Snowglobe, a big kind of idea is that for the first time in history, we can actually have, you know, a general purpose simulation system, right? Like simulation systems have existed, but, and, you know, we saw them like extensively in self-driving and…”
Rajpal: Running massive AI simulations yields diminishing returns
“So I think I think there's definitely I guess a diminishing kind of like returns. You know phenomena with, like, running simulations that are absolutely massive.”
Open source turns abstract AI safety into an actionable engineering problem
“What the open source really ends up doing as you know, a framework is taking down this very abstract problem of what it means to do safe AI development, right? Like it's a very abstract problem. It's almost an academic problem to some degree, and it takes that…”
Rajpal: Waymo had 20 million real-world miles versus 20 billion in simulation
“Like Waymo had twenty million miles in the real World driving, but twenty billion miles in simulation.”
Rajpal: Simulation testing revealed an early partner's real failure was over-refusal
“Organizations that we're, we were, we had as our design partner, we you know, they were like, oh, we're very worried about toxicity, and we want toxicity guardrails, and we did all of this testing for them in production, and toxicity was actually not a real co…”
Rajpal: Reusable persona libraries are Snowglobe's top requested feature
“Today, all personas are net new, but this is our number one requested feature, which is I want to be able to, you know, like maybe this, maybe some product leader already has a set of like personas that they want to test again. So I want to bring those, be abl…”