why aren't all 877 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Disclosure
Yegge: Sourcegraph avoids autonomous AI agents until someone builds one that works
“We're not going in the agent direction, right? I mean, I'll believe in agents when somebody shows me one that works.”
Disclosure
Qiu: Imbue generates specific reasoning data rather than relying on web data
“So I think internally, yeah, we have a lot of thoughts on what reasoning is, and we generate a lot more specific data. We're not just like, oh, it'll figure out reasoning from this black box or like, it'll figure out reasoning from the data that, that exists.”
Disclosure
Hotz: AMD CEO Lisa Su sent pre-release ROCm 5.6 to fix panics
“Lisa Sue reached out, connected with a whole bunch of different people. They sent me a pre-release version of Rock M 5.6. They told me you can't release it, which I'm like, okay, Why do you care? But they say they're going to release it by the end of the month…”
Disclosure
Hotz: Third company will build an AI girlfriend product
“The third company's the first one that's gonna build a real product, and that product is A girlfriend? No, like I'm dead serious, right? Like this is the dream product, right? This is the absolute dream product.”
Disclosure
Slack says Amp abandoned GitHub issues and pull requests
“We're not using issues. We're not using pull requests. We're barely on GitHub actions, and we're looking to get off of that.”
Disclosure
Sean Lie: OpenAI is Cerebras's biggest customer
“OpenAI is our biggest customer.”
Disclosure
Patil: Chai operates as a neutral software factory for medicines
“There's a lot of bio companies, AI for bio companies that are like making their own drugs. We really don't see ourselves that way, right? We see ourselves as almost a neutral software factory. For making medicines.”
Disclosure
McPartlon: Chai bet on model scaling over analyzing individual target failures
“Just be bitter, less impaled in that sense, and just really bet on the models getting better. And we definitely took the latter approach. Like we bet on the models getting better and we just pushed as hard as we could on that front.”
Disclosure
Nathan: OpenAI sequencing agents from developers to knowledge workers to everyone
“The vision is like bring useful agents to everyone. We started with like developers Historically are like early adopters that are willing to put up with more friction, set things up, et cetera. Like that's where, you know, Codex started. I think the next oppor…”
Disclosure
Kant: Poolside is releasing open weights to prevent five-company AI concentration
“I think it all just came down to one thing and I'll stop the monologue is the fact that I rather live in a world that has a hundred foundation model companies than a world that has five. Even if I was one of the five. And the smallest and most meaningful contr…”
Disclosure
Wang: X-Cell trains on causal perturbation data unlike static models
“What sets Excel different from these static expression models such as SGB or geneformers is that we actually, instead of training on gene expression datasets, we train on causal datasets. We train on massive amount of genome-wide perturbation datasets so that …”
Disclosure
Beam: Lila's core asset is its reasoning model, not clinical pipeline
“The model itself is the thing of value at Lila. So in that sense, we're much more of like a neolab trying to think of a new way to push forward capabilities of a core reasoning LLM based model.”
Disclosure
Beam: Future AI labs will be million-square-foot, lights-out data centers
“We do think about scaling it In the same way that you would think about scaling a data center, in that it's a multi-level building, occupies millions of square feet, and it's like a lights-out facility, as they say. It's like running 24 seven, generating data …”
Disclosure
Beam: Lila avoids pre-training from scratch, builds on open-weight models
“We have not decided to take on pre-training as well, just because the black magic that you have to do is, is insane, and we've been gifted, you know, something like a billion dollars worth of compute in the form of open-weight models. So we start with an open-…”
Disclosure
Xin: Databricks uses ML models to optimize database algorithms at runtime
“They use that to build a model. Like a machine learning model, not an L, a machine learning model. Machine learning model basically can very, very quickly tell us how any algorithm and how any implementation will perform for any specific type of queries with v…”
Disclosure
Zaharia: Databricks abandons frontier AI models to focus on agent systems
“Even though we did launch open source model DBRX, and, you know, we went up to, like, sort of above the LAMA-R III scale, we decided that we really want to focus on, there'll be so many people releasing models, and instead of doing the general model where, lik…”
Disclosure
Midha: AMP Is Pooling 1.3 Gigawatts of Compute Supply Over Four Years
“We pool demand, we pool supply from a number of partners we trust at about 1.3 gigawatt scale over four years.”
Disclosure
Major instrument vendors previously refused to give Radical AI data access
“A few very big tool vendors were Not too excited about self-driving labs two years ago. They were not jumping to give us, even with payment, access to the software and pulling the data. That tone has now changed.”
Disclosure
Backlund: Andon Labs agent tried TaskRabbit arbitrage to make money
“We tasked our office agent to just make was it like a hundred dollars, a thousand dollars? We just gave that prompt. And then what it did was sign up on TaskRabbit, both as a task
Looking for, like arbitrage., arbitration.”
Disclosure
Nadella: Work IQ turns Microsoft 365 into an accessible database
“The same thing is now happening with M three six five, because with work IQ, we have exposed what was perhaps the most important database in a company that never got used as a database because it was only captive to our apps, right? It was all email operated o…”
Disclosure
Yan: Cognition stripped obsolete Devin code after Sonnet 3.7 release
“So it's almost funny to be talking about how like big of a leaps on it. 3.7 was, and we honestly, a lot of it was stripping out parts of Devon that were no longer needed with that jumping of intelligence.”
Disclosure
Yan: Cognition merged PRs grew 7x as headcount grew only 10%
“It grew like seven X over like the last, I think it was like two months, three months, something like that. And then you see our engineering headcount growth, it's like gone up by like 10% or something.”
Disclosure
Yan: Devin requires orchestrating multiple frontier models for end-to-end app testing
“Well, in some cases we found that actually no one frontier model can actually do this full end-to-end task itself. We've seen cases where we actually had had to orchestrate different frontier models together to kind of solve this problem together.”
Disclosure
Burazin: Daytona runs on bare metal with local NVMe preloads
“The reason why Daytona is like super, super fast and you see this on benchmarks is we essentially, we run on bare metal. We have our own scheduler. We use the underlining disk CPU and RAM. Of the underlying machine, which means your IOPS are insanely fast beca…”
Disclosure
Azhnyuk: FPV drone autonomy hardware costs hundreds of dollars in BOM
“When you put full autonomy on that FPV drone, which can be not very expensive, like systems that we're producing, they're like hundreds of dollars of pure bomb costs.”
Disclosure
Parakhin: Shopify funds unlimited tokens and discourages models below Opus
“And we effectively fund unlimited tokens for everybody. We do try to control the models that people use, but from the bottom, not from top. Like we basically say, hey, please don't use anything less than Opus.”
Disclosure
Parakhin: No existing third-party AI PR review tool meets his standards
“I haven't found a good PR review tool that, that does what I think should be done”
Disclosure
Eifrem: Neo4j Vector Search Lags Dedicated Vector Databases
“Cause we also have vector search as part of Neo four J and it's not as good as the dedicated databases.”
Disclosure
Notion Partnered With Anthropic and OpenAI to Build 30% Pass Rate Evals
“And then what we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate. And that's actually been a effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because we actually hit a…”
Disclosure
Notion Custom Agents Use Standard Pages and Databases as Their Memory
“Another example of this is we have no built-in memory concept. Memory is, is just pages and databases. And so if you want to give it memory, just give it a page and give it access to that page.”
Disclosure
Lopopolo: OpenAI inverts harnesses by having Codex spawn dev environments
“One neat thing here is we have tried to invert things as much as possible, which is instead of setting up an environment to spawn the coding agent into, instead we spawn the coding agent, like that's the entry point, just codex, and then we give codex via skil…”
Disclosure
Lopopolo: OpenAI uses incident pages to update repository reliability rules via Codex
“When we get a page because we're missing a timeout, for example, I can just add codecs in Slack on that page and say, I'm going to fix this by adding a timeout. Please update our reliability documentation to require that all network calls have timeouts. So I h…”
Disclosure
Lopopolo: Codex authors Grafana dashboards and handles on-call incident paging
“Like the dashboard thing you mentioned, we have Codex authoring the JSON for the Grafana dashboards and publishing them, and also responding to the pages, which means when it gets the page, it knows exactly which dashboards are defined and what alerts. What al…”
Disclosure
Lopopolo: OpenAI uses an automated landing skill to delegate PR merges to Codex
“We invoke a dollar land skill and that coaches codecs to push the PR, wait for human and agent reviewers, wait for CI to be green, fix the flakes if there are any Merge upstream if the PR comes into conflict, wait for everything to pass, put it in the merge qu…”
Disclosure
Lopopolo: OpenAI Frontier codebase operates on approximately six core skills
“So like in our code base, we have, I think six skills. That's it. And if some part of the software development loop is not being covered, Our first attempt is to encode it in one of the existing setup skills, which means that we can change the agent behavior m…”
Disclosure
OpenAI Frontier runs daily agent loops over team logs to update repositories
“We're actually slurping these up for the entire team into blob storage and running agent loops over them every day to figure out where as a team can we do better? And how do we reflect that back into the repository? Yeah, though, everybody benefits from everyb…”
Disclosure
OpenAI Frontier grants coding agents full permission to file follow-up tickets
“Like, this thing is also able to cut its own tickets, because we give it full access. Yeah, yeah, yeah. You can make a ticket to have it cut tickets, you can put in the ticket that you expected to file its own follow-up work.”
Disclosure
Sun: Moonlake splits world modeling into multimodal reasoning and Reverie rendering
“Within our world modeling framework, we think there are two models that we train, right? Like there's the multimodal reasoning model that we just talked about that essentially handles Mainly the causality, the persistency, and logic, determinism, determinism o…”
Disclosure
Dreamer pays AI tool builders proportionally based on agent usage
“And we're actually sharing something for the first time on this podcast, which is tool builders on Dreamer get paid. So if you publish a tool to the platform and a lot of agents use it, you'll actually get paid in proportion to their usage.”
Disclosure
Rieseberg: Claude Cowork assembled existing prototypes rather than starting from scratch
“And what cowork actually became is like, we sort of picked the right pieces out of the many prototypes that we had. Right. And that's maybe also like, I think an important qualifier whenever people mention this like 10 day number, I do think it's important to …”
Disclosure
Anthropic: Claude Cowork executes inside a dedicated lightweight Linux VM
“So we currently run like a, we currently run like a lightweight VM and we put clock code into the VM and we do that for a number of reasons. Safety and security is a big one, but even if you ignore for a second safety and security and you're just like, okay, Y…”
Disclosure
Colvin: Monty will never support CPython ABI packages like NumPy
“There's no support for third party libraries just to be installed. There never will be directly as, and you'll never, we'll never be able to speak the C Python ABI and like install Pydantic or install NumPy or something.”
Disclosure
Colvin: Pydantic AI introduces declarative TOML serializable agents
“One of the things we're doing now in Pydantic AI is where we're about to introduce, I think there's a PR out for this. So I think this is public serializable agents. So basically you can define an agent entirely in a TOML file, everything from the model to the…”
Disclosure
Turbopuffer Bought Oregon Dark Fiber to Serve Notion Across Clouds
“It started getting really painful in like mid-twenty-twenty-four, because we were closing deals with Notion actually, that was running in AWS, and we're like, trust us, you really want us to run this in GCP?
And they were like, no, I don't know about that, lik…”
Disclosure
Shah: Supermemory infrastructure costs only two cents per million tokens
“For us, we have optimized infrastructure down to, it costs us two cents for a million tokens, because we have our own model, our own database, our own everything.”
Disclosure
Nelle: Full computer use shifted internal agent usage to driving new features
“Giving the model the tools to onboard itself and then use Full computer use end to end pixels in coordinates out and have sort of the cloud computer with different apps in it is the big unlock that we've seen internally in terms of usage of this going from, oh…”
Disclosure
Jeff Dean: Gemini was designed to ingest Waymo LIDAR and robotics telemetry
“I think one of the things about Gemini's multimodal aspects is we've always wanted it to be multimodal from the start. And so, you know, that sometimes to people means text and images and video sort of human-like and audio, audio, human-like modalities, but I …”