Opinion certainty 3/5 debate potential 3/5

Adding sequential LLM calls or filters to catch errors fails in production

Sharon Zhou · Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini · Jul 25, 2024 · at 38:49

Sharon Zhou, co-founder and CEO of Lamini, discusses architectural pitfalls in designing enterprise AI agents and managing latency and error rates.

0:00 / 0:12exact quote · 12.2s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“It's both of those things, and I think people are addressing error today by adding more calls to the model of filtering. Out the requests. And I think I don't think that'll work for serious production use cases.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Sharon Zhou

Assertion Not checkable as stated
Memory tuning eliminates hallucinations and enables near-perfect task performance
“Been able with memory tuning, which is what I've been working on to remove those hallucinations, to remove that and actually get these models from, you know, not necessarily being general for everything. And instead of being pretty good at everything, but perf…”
Sharon Zhou Jul 25, 2024 ▶ 10:02 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Assertion Supported
The enterprise GPU shortage has eased at the company level
“Today, actually, I'm seeing the GPU shortage go away at the level, at the company level, meaning companies are able to procure enough compute enough is a strong word, but they're able to procure compute at some level to work with, to fine tune and run heavy in…”
Sharon Zhou Jul 25, 2024 ▶ 16:01 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Disclosure
In mid-2023, multi-billion dollar companies could not obtain AWS GPU nodes
“Last year was, at this time, was absolutely insane. That's why we threw up our own cloud, because there was just like, large companies with multi-billion revenue numbers could not get a node from AWS, despite their accounts being tens of millions or hundreds o…”
Sharon Zhou Jul 25, 2024 ▶ 17:06 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Assertion Supported
Re-engineering LLM decoders can guarantee absolute schema accuracy for structured outputs
“So that's something else we offer through our inference service to actually make it a hundred percent by re-engineering the decoder of any LLM.”
Sharon Zhou Jul 25, 2024 ▶ 25:22 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Prediction Not checkable as stated
Future AI models will deliver 100B parameter intelligence at 1B speeds
“I even think there's a future where these models can be a hundred billion parameters, but at, you know, have that intelligence of a hundred billion parameters, but then have the speed, latency, and cost of something that's still one billion or seven billion pa…”
Sharon Zhou Jul 25, 2024 ▶ 28:15 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Prediction Not checkable as stated
Combining MoE and LoRA will eliminate big versus small model trade-offs
“And I do think that's the future so we can get something that is incredibly smart, incredibly huge, but with the latency cost and speed of something, something tiny. So no more big model versus small model paradigm. It's potentially one in the same.”
Sharon Zhou Jul 25, 2024 ▶ 32:24 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.