Kevin Mandich, Head of Machine Learning at Sprig, discusses the challenges of benchmarking LLM performance on multi-question qualitative survey synthesis.
“And that's just, that's something that's really hard to evaluate in an automated manner, at least right now, because it's a more complex task, because the output is, you know, non-deterministic and freeform. For now, it really just does require some manual evaluation.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Kevin Mandich
Opinion
Tech is likely near the peak of the current generative AI hype cycle
“I think what we're going through now is we're probably close to the peak of the current hype cycle that was introduced earlier this year with, you know, ChatGPT and everything.”
Kevin MandichSep 7, 2023▶ 21:21A guide to building product in a post-LLM world | Ryan Glasgow and Kevin Mandich from Sprig
Insight
Word clouds and topic modeling fail to capture nuanced customer feedback
“You know, people have done word clouds, word counts topic modeling, but, you know, none of those really capture the nuance of what people are saying. And you know, it doesn't account for the fact that people could be saying multiple things per response.”
Kevin MandichSep 7, 2023▶ 5:28A guide to building product in a post-LLM world | Ryan Glasgow and Kevin Mandich from Sprig
Insight
Machine learning engineering skills transfer broadly across vision, NLP, and audio domains
“One of the things about machine learning and AI in this industry is that there's a lot of crossover between these different data domains. You know, if you're able to solve a computer vision problem pretty well, a lot of those skills are transferable to natural…”
Kevin MandichSep 7, 2023▶ 6:27A guide to building product in a post-LLM world | Ryan Glasgow and Kevin Mandich from Sprig
Insight
Mandich: User research synthesis lacks an objective universal ground truth
“You can take two expert user researchers, give them the same list of, you know, 500 responses, tell them to distill them down into 10 actionable takeaways, and those 10 will be completely different, or even if they are the same 10 takeaways, the responses that…”
Kevin MandichSep 7, 2023▶ 10:52A guide to building product in a post-LLM world | Ryan Glasgow and Kevin Mandich from Sprig
Insight
Analyzing machine learning data in batch produces better results than real-time streaming
“From an ML point of view. I think it's advantageous to analyze as much data in batch as possible. You tend to get more information to work with.”
Kevin MandichSep 7, 2023▶ 12:45A guide to building product in a post-LLM world | Ryan Glasgow and Kevin Mandich from Sprig
“I think given our current switch to hosted LLMs and away from, you know, taking Google's Burt and fine tuning it ourselves, more towards, I guess we call like a full stack ML engineer, somebody who's able to, you know, help integrate this and actually implemen…”
Kevin MandichSep 7, 2023▶ 57:04A guide to building product in a post-LLM world | Ryan Glasgow and Kevin Mandich from Sprig
Made with StarZero
Turn any episode into a week of clips.
This entire site, about 80 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.