Memory tuning eliminates hallucinations and enables near-perfect task performance
“Been able with memory tuning, which is what I've been working on to remove those hallucinations, to remove that and actually get these models from, you know, not necessarily being general for everything. And instead of being pretty good at everything, but perf…”
Zhou: Lamini deploys enterprise LLMs on-premise in air-gapped environments
“We're an integrated inference and fine tuning platform for enterprises to be able to run factual LLMs. So essentially LLMs that don't hallucinate on their proprietary data within their secure walls. So we can deploy on premise air gapped, no internet sites. So…”
In mid-2023, multi-billion dollar companies could not obtain AWS GPU nodes
“Last year was, at this time, was absolutely insane. That's why we threw up our own cloud, because there was just like, large companies with multi-billion revenue numbers could not get a node from AWS, despite their accounts being tens of millions or hundreds o…”
Re-engineering LLM decoders can guarantee absolute schema accuracy for structured outputs
“So that's something else we offer through our inference service to actually make it a hundred percent by re-engineering the decoder of any LLM.”
Future AI models will undergo continuous fine-tuning as easily as prompt engineering
“I believe in a future where we're continuously fine tuning these models where it's as easy as prompt engineering and you know, these models continually improve.”
Zhou: ML benchmark accuracy fails in production without low API latency
“In machine learning, you know, a lot of AI people are like, yeah, we push the performance of this model, and we define performance as accuracy or, you know, accuracy along these, like, general benchmarks. Maybe it's sixth grade science questions or something. …”
Zhou: Adding images drove a 10x increase in marketing engagement
“I think one thing that, you know, we were analyzing our data when it came to all our marketing blog posts or like tweets, everything going out, and there was a 10 X increase in anything with an image, right?”
Zhou: Lamini cuts LLM fine-tuning time from months to milliseconds
“And by efficiency, I mean, you know, it's instead of something that might take weeks or even months that's bringing it down to even like the millisecond level.”
Zhou: Lamini is the only platform running LLMs on AMD GPUs
“We are the only folks who can actually run your language models on top of AMD AMD GPUs.”
Lamini's AMD support unlocks 20,000 GPUs, enough to train GPT-4
“So that unlocks, what that means is that unlocks about 20,000 GPUs readily available today for enterprises to be able to use, and to get a sense of what that means you can train GBD-IV.”
Zhou: Lamini can train models up to 100 billion parameters
“We can train up to a hundred billion parameters.”
Lamini switches across 1,000 fine-tuned models in three milliseconds
“With the technology that we've used with parameter efficient fine tuning and just like efficiency, different efficiency methods, that time to switch across a thousand models is three milliseconds”
Lamini's hosted service ran exclusively on AMD GPUs for a year
“The Lamini hosted service over the past year has been running on AMD GPUs only. We haven't been running on NVIDIA chips.”
Lamini has achieved software parity on AMD GPUs with CUDA
“We have reached software parity with essentially CUDA.”