Everything Sharon Zhou said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Memory tuning eliminates hallucinations and enables near-perfect task performance
“Been able with memory tuning, which is what I've been working on to remove those hallucinations, to remove that and actually get these models from, you know, not necessarily being general for everything. And instead of being pretty good at everything, but perf…”
The enterprise GPU shortage has eased at the company level
“Today, actually, I'm seeing the GPU shortage go away at the level, at the company level, meaning companies are able to procure enough compute enough is a strong word, but they're able to procure compute at some level to work with, to fine tune and run heavy in…”
In mid-2023, multi-billion dollar companies could not obtain AWS GPU nodes
“Last year was, at this time, was absolutely insane. That's why we threw up our own cloud, because there was just like, large companies with multi-billion revenue numbers could not get a node from AWS, despite their accounts being tens of millions or hundreds o…”
Re-engineering LLM decoders can guarantee absolute schema accuracy for structured outputs
“So that's something else we offer through our inference service to actually make it a hundred percent by re-engineering the decoder of any LLM.”
Future AI models will deliver 100B parameter intelligence at 1B speeds
“I even think there's a future where these models can be a hundred billion parameters, but at, you know, have that intelligence of a hundred billion parameters, but then have the speed, latency, and cost of something that's still one billion or seven billion pa…”
Combining MoE and LoRA will eliminate big versus small model trade-offs
“And I do think that's the future so we can get something that is incredibly smart, incredibly huge, but with the latency cost and speed of something, something tiny. So no more big model versus small model paradigm. It's potentially one in the same.”
Adding sequential LLM calls or filters to catch errors fails in production
“It's both of those things, and I think people are addressing error today by adding more calls to the model of filtering. Out the requests. And I think I don't think that'll work for serious production use cases.”
RAG and prompt engineering are just search techniques, not real AI
“Today when people are running RAG or prompt engineering those are search. That's not AI. It's like keeping the AI frozen and fixed.”
Zhou: Public training data for LLMs is running out
“Yeah, I think even just zooming back out at a technical level, I think actually public data is running out for all, all of what LLMs can take advantage of.”
Zhou: Domain experts, not AI researchers, will drive top models
“However, I believe that, and this is based on my experience training these models, it's actually the domain experts will be driving the best models out there. It won't be people like me who can actually do all the model training, et cetera.”
Zhou: Lamini cuts LLM fine-tuning time from months to milliseconds
“And by efficiency, I mean, you know, it's instead of something that might take weeks or even months that's bringing it down to even like the millisecond level.”
Zhou: Lamini is the only platform running LLMs on AMD GPUs
“We are the only folks who can actually run your language models on top of AMD AMD GPUs.”
Lamini's AMD support unlocks 20,000 GPUs, enough to train GPT-4
“So that unlocks, what that means is that unlocks about 20,000 GPUs readily available today for enterprises to be able to use, and to get a sense of what that means you can train GBD-IV.”
Zhou: Zero-shot LLMs can replace manual human labeling in RLHF workflows
“Which is that why can't it be another LLM or a pipeline of LLMs that can help with that feedback? I think manual labeling is very tedious, especially for our target user, which is a software engineer. And I don't think people should necessarily have to do all …”
Zhou: Outsourcing specialized medical AI data labeling to Scale AI failed
“We tried outsourcing actually to like scale AI, et cetera. None of that worked. It had to basically be me.”
Zhou: LLMs can reach 99% accuracy today with narrow scoping
“I think we can get to that performance today, but it's based on how you scope out the problem. So if it's a very narrow scope, of course you can get that.”
Lamini's hosted service ran exclusively on AMD GPUs for a year
“The Lamini hosted service over the past year has been running on AMD GPUs only. We haven't been running on NVIDIA chips.”
Lamini has achieved software parity on AMD GPUs with CUDA
“We have reached software parity with essentially CUDA.”
Lower LLM latency requires specialized hardware over GPU software platforms
“Unfortunately, I think what people don't realize is that the way to get better latency, like significantly better latency, is actually in the hardware. And that's why we see Grok, G-R-O-Q be able to exceed all these GPU-based inference platforms significantly,…”
NVIDIA A100s and AMD MI300s are readily available, but H100s remain scarce
“A 100 in particular are pretty available. Obviously the AMD chips that we also agnostically work with the MI 300 and MI two fifties, those are available. H 100 still kind of. A little bit harder to get, but you can get started very easily with any of those oth…”
Memory tuning embeds enterprise data to enable near-deterministic factual recall
“To be able to embed facts of your data into the model, so memory tune the model so that it can recall those facts almost deterministically within its probabilistic context.”
Zhou: Best LLMs of the next wave will be enterprise models
“So actually the next frontier for LLMs is in enterprises, and I believe the best LLMs for this next, next wave essentially will be enterprise LLMs.”
Zhou: ML benchmark accuracy fails in production without low API latency
“In machine learning, you know, a lot of AI people are like, yeah, we push the performance of this model, and we define performance as accuracy or, you know, accuracy along these, like, general benchmarks. Maybe it's sixth grade science questions or something. …”
Zhou: Prompting Midjourney images matters more than writing blog posts
“And so we found that just spending some time on mid journey was worth more than spending time on the blog posts. In any single way, like hands down, just like spend a few extra minutes here, prompt engineering or really just like generating an image of your ch…”