Nov 21, 2024 · 44m · no-priors
No Priors Ep. 91 | With Cohere Co-Founder and CEO Aidan Gomez
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Cohere Co-Founder and CEO Aidan Gomez discusses the evolution of Transformer architectures, Cohere's enterprise-focused strategy, and the economics of training frontier models. He offers deep insights into scaling laws, inference-time reasoning, and why foundation models remain far from commoditization.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 17.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
When Sarah notes frontier lab leaders claim nobody outside AGI labs should touch pre-training, Aidan bluntly dismisses the claim as empirically wrong.
Hardest push from the hosts ▶ 20:00 Sarah challenges enterprise pre-training feasibilitySarah directly challenges Aidan's framework by arguing that pre-training for enterprises is controversial and rejected by leading AI labs due to compute and data curation limits.
Biggest teaching moment ▶ 24:24 Aidan explains limitations of frozen open-source weightsAidan educates on why open-source models like Llama cannot simply replace proprietary base training, explaining that frozen, cooled-down models with zero gradients offer fewer levers than vertical training data integration.
The host holds their own ▶ 28:35 Sarah synthesizes the economic model shift of reasoning computeSarah demonstrates sharp industry insight by translating the technical concept of inference-time compute into a structural shift from upfront CapEx model training to consumption-based intelligence.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Founding Cohere and the Enterprise-First Mission | 4 | 2 | 1 | 1 | Sarah introduces Aidan and sets up his background from Google Brain to founding Cohere. Aidan amicably explains the founding rationale, targeting enterprise workloads rather than building a consumer ChatGPT competitor. | |
| Resolving Enterprise Pitfalls and Practical AI Use Cases | 4 | 4 | 1 | 2 | Sarah asks Aidan to make enterprise frustrations concrete. Aidan details common RAG pitfalls, model sensitivity to prompt formatting, and walks through practical vertical deployments like longitudinal healthcare record synthesis. | |
| Enterprise Adoption Strategy and Overcoming Adoption Barriers | 5 | 4 | 2 | 3 | Sarah questions the long-term equilibrium between specialized AI apps and in-house enterprise development, as well as whether enterprise AI is entering a trough of disillusionment. Aidan reframes the market using an adoption pyramid and argues the technology is too early for a disillusionment slump. | |
| Model Customization Tiers and Cohere's Cost Efficiency Strategy | 7 | 6 | 5 | 6 | Sarah pushes back against the premise of enterprises doing pre-training, citing frontier lab consensus, and challenges Cohere's supercomputer capital expense given open-source models like Llama. Aidan forcefully rejects the lab consensus as empirically wrong and explains the limitations of fine-tuning zero-gradient open-source weights. | |
| Scaling Laws, Inference-Time Compute, and Reasoning Paradigms | 6 | 5 | 1 | 2 | Aidan explains the plateau in generic vibe checks and the shift toward inference-time compute and reasoning models. Sarah demonstrates domain expertise by framing test-time compute as a transition from CapEx-driven capability gains to consumption-driven models. | |
| Data Limits, Scientific Frontiers, and AGI Realism | 6 | 5 | 4 | 4 | Sarah probes the limits of sequence-to-sequence scaling, AGI discrete milestones, and model representation bottlenecks. Aidan rejects sci-fi takeoff scenarios and checklist AGI definitions, while debunking model commoditization as temporary loss-leading price dumping. |