Opinion
Polu: The bulk of useful enterprise agent work can use APIs
“The bulk of the useful stuff that you can do within the company can be done through API. The data can be retrieved by API, the actions can be taken through API.”
Opinion
Polu: DeepMind IMO breakthrough relied on scaling RL and autoformalization
“I think the DeepMind team just did a good job of scaling. I think there's nothing too magical in their approach, even if it hasn't been published as a Dan Silver talk from seven days ago, where it goes a little bit into more details. It feels like there's noth…”
Opinion
Polu: Anthropic split was driven by disagreement over OpenAI's API commercialization
“What I understood of it is that there was a disagreement of the commercialization of that technology. I think the focal point of the disagreement was the fact that we started working on the API and wanted to make those models available through an API. Is that …”
Prediction Not checkable as stated
Polu: Post-hyper-growth tech companies may increasingly eliminate traditional SaaS
“So it's interesting that we might see kind of a bad time for SaaS in post-hyper-growth tech companies. So it's still a big market, but it's not that big, because if you're not a tech company, You don't have the capabilities to reduce desk cost. If you're a hig…”
Prediction Not checkable as stated
Polu: The next generation will see billion-dollar companies with 20 engineers
“All generations of company might be the first billion dollar companies with engineering teams of 20 people. That would be so exciting as well. That would be so great. You know, you don't have the management hurdle. You're just 20 focused people with a lot of a…”
Insight
Polu: OpenAI Managed Research Priorities Directly Through Compute Allocation
“In that space, there's a managing tool that is great, which is computer location. Basically, by managing the computer location, you can message the team of where you think the priority should go. And so it was really a question of you were free as a researcher…”
Assertion Not checkable as stated
Polu: OpenAI Believed in Transformer Scaling Pre-Kaplan Paper
“Before that, there really was a strong belief in, in scale. I think it was just the belief that the transformer was a generic enough architecture that you could learn anything, and that this was just a question of scaling.”
Assertion Partly supported
Polu: GPT-4 was ready internally at OpenAI months before September 2022
“I had seen GPT-IV internally at the time. It was September, 20, 22. So it was pre-chat GPT, but GPT-IV was ready since, I mean, I'd been ready for a few months internally.”
Opinion
Polu: Fully autonomous AI models 'get lost' and are not ready
“The AutoGPD approach, obviously, is extremely exciting, but we know that the agentic capability of models are not quite there yet. It just gets lost.”
Insight
Polu: Hierarchies of simple agents will unlock Auto-GPT level value
“Once you have those working really well, you can create meta agents that use the agents as actions, and all of a sudden you can kind of have a hierarchy of responsibility that will probably get you almost to the point of the auto GPT value.”
Opinion
Polu: GPT-4 Turbo performs better than GPT-4o on function calling
“I personally don't have proof, but I know many people, and I'm probably part of them, to think that GPT-IV Turbo is still better than GPT-IV on function calling.”
Assertion Supported
Polu: Claude Sonnet executes an unpublicized chain-of-thought step during function calling
“They kind of innovated in an interesting way, which was never quite publicized, but it's that they have that kind of chain of thoughts step whenever you use a Clouds model or Sonnet model with function calling. That chain of service step doesn't exist when you…”
Opinion
Polu: Airbyte's Notion connector output is not useful for AI models
“And the reality is that if you look at Notion, Airby does the job of taking Notion and putting it in a structured way, but that's a way that is not really usable to actually make it available to models in a useful way. Because you get all the blocks, details, …”
Insight
Polu: Vertical AI agents have easier GTM but limited enterprise upside
“Vertical solutions have a good market that is much easier because they're like, oh, I'm going to solve the lawyer stuff. But the potential within the company after that is limited. So there's really a nice tension there. We, we're true believers of the horizon…”
Assertion Not checkable as stated
Polu: Transformers failed at code fuzzing because they are too slow
“We're trying to apply transformers to code fuzzing. So code fuzzing, you have kind of an, sorry, an algorithm that goes really fast and tries to mutate the inputs of a library to find bugs. And we try to apply a transformer to that and do reinforcement learnin…”
Opinion
Polu: Sam Altman mastered technical ML details within two years at OpenAI
“One thing about Sam Altman, he really impressed me, because when I joined, he had joined not that long ago. And it felt like he was kind of a very high level CEO. And I was mind blown by how deep he was able to go into the subjects within a year or something, …”
Insight
Polu: LLM workflow development requires a dozen examples to prevent overfitting
“I had the strong belief from my research time that you cannot create an LLM-based workflow on just one example. Basically, if you just have one example, you overfit. So as you develop your interaction, your orchestration around the LM, you need a dozen example…”
Insight
Polu: LLM productization is currently only at the 'Pong' stage
“I think we're at the pong level of LLM productization, and we haven't invented the SIEV-III, we haven't invented Counter-Strike, we haven't invented Cyberpunk”
Insight
Polu: Models make mistakes when given high-level instructions and many tools
“If you provide a very high level Kind of an auto GPT-esque level in the instructions and provide 16 different tools to your model. Yes, we're seeing the models in that state making mistakes.”
Insight
Polu: Combining LLMs with formal math pairs creativity with proof verification
“Transformers are very creative, but yet they do mistakes. And formal math systems are the ability to verify a proof. And the tactics they can use to solve problems are very mechanical. So you miss the creativity. And so the idea was to try to explore both toge…”
Assertion Not checkable as stated
Polu: OpenAI's GPT-3 Was Internally Codenamed Project Nest
“Most of the compute was going to a product called Nest, which was basically GPT-free.”
Assertion Not checkable as stated
Polu: Ilya Sutskever Spent Surprising Amount of Time Communicating OpenAI Vision
“I think he was really focused on building the vision and communicating the vision within the company, which was extremely It's extremely useful. I was personally surprised that he spent so much time, you know, working on communicating that vision and getting t…”
Assertion Not checkable as stated
Polu: Dust averages 60% to 70% weekly active penetration in enterprise accounts
“The highest penetration we have is 88% daily active users within the entire employee of the company. The kind of average penetration and activation we have in our current enterprise customers is something like more like 60 to 70% weekly active.”