why aren't all 11 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Agentic AI workflows are technically feasible today but face reliability hurdles
“Based on what we've seen today in the boundaries of the LLMs today is that we could, based on what we have tested with the models, actually build that now. And we are seeing the sophistication, the capability of the models to build that now. I think the roadbl…”
Disclosure
Pitching a product as an automated researcher alienated early customers
“I remember at first the very early days of Sprig, there was some wording around your automated user researcher, you know, or it does all the text analysis for you. And that was actually very off putting for both user researchers and product teams.”
Insight
AI product teams must present concrete options rather than asking what to build
“You know, I think with working with customers with AI, you also have to develop world-class techniques of not asking the customer what to build, but instead giving them a set of options to provide feedback on.”
Disclosure
Three-worker Mechanical Turk consensus often failed expert data accuracy checks
“We originally hired Amazon Mechanical Turks, and we said, hey, we'll save money. We'll have three of them review every response. But we'd often see that even if all three of them gave the same output, we would go to an expert researcher or a customer or look a…”
Insight
Mandich: User research synthesis lacks an objective universal ground truth
“You can take two expert user researchers, give them the same list of, you know, 500 responses, tell them to distill them down into 10 actionable takeaways, and those 10 will be completely different, or even if they are the same 10 takeaways, the responses that…”
Prediction Not checkable as stated
AI models will likely never reach total accuracy for subjective feedback categorization
“Cause we will never, it's very unlikely we'll ever get to a hundred percent. And so we'll get, you know, 90, 95, 99. But we'll always make sure that they have the ability to then do that last mile of analysis on their own. And so that's something that we have …”
Insight
Complex non-deterministic LLM tasks still require manual human evaluation
“And that's just, that's something that's really hard to evaluate in an automated manner, at least right now, because it's a more complex task, because the output is, you know, non-deterministic and freeform. For now, it really just does require some manual eva…”
Insight
AI product development strictly requires an AI engineer from day one
“The uniqueness that we found for product development specific with AI is that You know, we define product best in class product development as a product manager and a minimum designer starting from the very beginning of ideation and building a product spec and…”
Insight
AI engineers must be embedded directly in product squads, not isolated
“And so having, AI engineers embedded with traditional engineers has been the big breakthrough for us and working together To develop and build these features and not seeing AI as, like Kevin said, an input output, you know, the model spits out a response that …”
Assertion Not checkable as stated
Large language models currently fail at basic math and quantitative reasoning
“Right now a lot of these models fall flat and for an application like ours and for a lot of applications out there, that's kind of a pretty glaring omission. You know, these models are fantastic at summarization. They're really good at generating text based on…”
Disclosure
Sprig hired a Yale faculty member specifically to review model outputs
“We brought on a world-class researcher who was actually on the faculty at Yale as one of our other You know, very early team members. She was actually the fourth person to join SPRIG, and her first role was reviewing.”