The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Memory tuning eliminates hallucinations and enables near-perfect task performance
“Been able with memory tuning, which is what I've been working on to remove those hallucinations, to remove that and actually get these models from, you know, not necessarily being general for everything. And instead of being pretty good at everything, but perf…”
The enterprise GPU shortage has eased at the company level
“Today, actually, I'm seeing the GPU shortage go away at the level, at the company level, meaning companies are able to procure enough compute enough is a strong word, but they're able to procure compute at some level to work with, to fine tune and run heavy in…”
Re-engineering LLM decoders can guarantee absolute schema accuracy for structured outputs
“So that's something else we offer through our inference service to actually make it a hundred percent by re-engineering the decoder of any LLM.”
Future AI models will deliver 100B parameter intelligence at 1B speeds
“I even think there's a future where these models can be a hundred billion parameters, but at, you know, have that intelligence of a hundred billion parameters, but then have the speed, latency, and cost of something that's still one billion or seven billion pa…”
Combining MoE and LoRA will eliminate big versus small model trade-offs
“And I do think that's the future so we can get something that is incredibly smart, incredibly huge, but with the latency cost and speed of something, something tiny. So no more big model versus small model paradigm. It's potentially one in the same.”
Zhou: Public training data for LLMs is running out
“Yeah, I think even just zooming back out at a technical level, I think actually public data is running out for all, all of what LLMs can take advantage of.”
Zhou: Domain experts, not AI researchers, will drive top models
“However, I believe that, and this is based on my experience training these models, it's actually the domain experts will be driving the best models out there. It won't be people like me who can actually do all the model training, et cetera.”
Zhou: Lamini cuts LLM fine-tuning time from months to milliseconds
“And by efficiency, I mean, you know, it's instead of something that might take weeks or even months that's bringing it down to even like the millisecond level.”
Zhou: Lamini is the only platform running LLMs on AMD GPUs
“We are the only folks who can actually run your language models on top of AMD AMD GPUs.”
Lamini's AMD support unlocks 20,000 GPUs, enough to train GPT-4
“So that unlocks, what that means is that unlocks about 20,000 GPUs readily available today for enterprises to be able to use, and to get a sense of what that means you can train GBD-IV.”
Zhou: Outsourcing specialized medical AI data labeling to Scale AI failed
“We tried outsourcing actually to like scale AI, et cetera. None of that worked. It had to basically be me.”
Zhou: LLMs can reach 99% accuracy today with narrow scoping
“I think we can get to that performance today, but it's based on how you scope out the problem. So if it's a very narrow scope, of course you can get that.”
Lamini has achieved software parity on AMD GPUs with CUDA
“We have reached software parity with essentially CUDA.”
NVIDIA A100s and AMD MI300s are readily available, but H100s remain scarce
“A 100 in particular are pretty available. Obviously the AMD chips that we also agnostically work with the MI 300 and MI two fifties, those are available. H 100 still kind of. A little bit harder to get, but you can get started very easily with any of those oth…”
Zhou: Best LLMs of the next wave will be enterprise models
“So actually the next frontier for LLMs is in enterprises, and I believe the best LLMs for this next, next wave essentially will be enterprise LLMs.”
Lamini switches across 1,000 fine-tuned models in three milliseconds
“With the technology that we've used with parameter efficient fine tuning and just like efficiency, different efficiency methods, that time to switch across a thousand models is three milliseconds”
Fine-tuning a GPT-3 class model with LoRA provides a 10,000x efficiency boost
“I think for something like a GPD three level model, it's a 10,000 X speed up in efficiency while losing nearly not much at all in accuracy.”
Future AI models will undergo continuous fine-tuning as easily as prompt engineering
“I believe in a future where we're continuously fine tuning these models where it's as easy as prompt engineering and you know, these models continually improve.”