| 20% | “20%” | Beauchamp: Training LLMs to swear makes them 20% funnier | William Beauchamp | Jan 26, 2025 | — |
| 85% | “85%” | Beauchamp: Spotify would retain 85% of users with only five artists | William Beauchamp | Jan 26, 2025 | — |
| 5B | “five billion” | Zhang: Llama 405B sees very few enterprise users compared to 70B | Yining Zhang | Jan 19, 2025 | — |
| 5B | “five billion” | Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B | Yining Zhang | Jan 19, 2025 | Contradicted |
| 3× | “three times” | Zhang: SGLang achieved 3x throughput over vLLM in mid-2024 benchmarks | Yining Zhang | Jan 19, 2025 | Supported |
| 100% | “hundred percent” | McAteer: o1 is the first model to achieve one-shot codebase implementation | Dan McAteer | Jan 17, 2025 | — |
| 1K | “a thousand” | Bryk: Neural link prediction is strictly more powerful than Google's PageRank | Will Bryk | Jan 10, 2025 | — |
| 6M | “a five million” | Bryk: Exa purchased a $5 million Nvidia H200 GPU cluster | Will Bryk | Jan 10, 2025 | — |
| 200× | “200 X” | Bryk: 200x drop in LLM costs requires rethinking search from scratch | Will Bryk | Jan 10, 2025 | — |
| 50% | “50%” | Swyx: Gemini Flash accounts for 50% of OpenRouter requests | Shawn Wang | Jan 1, 2025 | Contradicted |
| $1K | “a thousand dollars” | Swyx: AI VCs and content creators should spend $1,000 monthly on AI tools | Shawn Wang | Jan 1, 2025 | — |
| 60M | “sixty million” | Swyx: VC appetite for GPU-rich early-stage startups is completely gone | Shawn Wang | Jan 1, 2025 | — |
| 2T | “two trillion” | Swyx: AI models hit a 2T parameter wall, won't reach 10T | Shawn Wang | Jan 1, 2025 | — |
| 10T | “10 trillion” | Swyx: AI models hit a 2T parameter wall, won't reach 10T | Shawn Wang | Jan 1, 2025 | — |
| 10× | “10 X” | Swyx: AI models hit a 2T parameter wall, won't reach 10T | Shawn Wang | Jan 1, 2025 | — |
| 20M | “twenty million” | Swyx: Suno grew from zero to $20M ARR running on Modal | Shawn Wang | Jan 1, 2025 | Partly supported |
| 20M | “twenty million” | Swyx: Claude wrapper Bolt.new reached $20M ARR | Shawn Wang | Jan 1, 2025 | Supported |
| 8M | “eight million” | Swyx: Claude wrapper Bolt.new reached $20M ARR | Shawn Wang | Jan 1, 2025 | Supported |
| 100M | “a hundred million” | Swyx: HeyGen has reached $100 million in ARR | Shawn Wang | Jan 1, 2025 | Supported |
| 13% | “13%” | Swyx: SWE-bench resolution rates surged from 13% to ~50% in 2024 | Shawn Wang | Jan 1, 2025 | Supported |
| 100M | “a hundred million” | Fanelli: 133 AI funding rounds exceeded $100M in 2024 | Alessio Fanelli | Jan 1, 2025 | Supported |
| 99% | “99%” | Neubig: Browsers, Terminals, And Code Editors Constitute The Core Agent Toolset | Graham Neubig | Dec 25, 2024 | — |
| 10× | “10 times” | Neubig Uses AI Coding Agents Five To Ten Times Daily | Graham Neubig | Dec 25, 2024 | — |
| 22.5% | “22.5%” | Neubig: Agent Workflow Memory Boosts WebArena Performance By 22.5 Percent | Graham Neubig | Dec 25, 2024 | Supported |
| 40% | “40%” | Neubig: Agents Solve 80 To 90 Percent Of Tasks With Feedback | Graham Neubig | Dec 25, 2024 | — |
| 90% | “90%” | Neubig: Agents Solve 80 To 90 Percent Of Tasks With Feedback | Graham Neubig | Dec 25, 2024 | — |
| 50B | “fifty billion” | Ben Allal: LLMs can be trained with entirely synthetic pipelines | Loubna Ben Allal | Dec 24, 2024 | Partly supported |
| 100% | “hundred percent” | Ben Allal: LLMs can be trained with entirely synthetic pipelines | Loubna Ben Allal | Dec 24, 2024 | Partly supported |
| 1.9T | “1.9 trillion” | Ben Allal: NVIDIA generated 1.9 trillion synthetic tokens for Nemotron-CC | Loubna Ben Allal | Dec 24, 2024 | Supported |
| 16T | “a 15 trillion” | Hugging Face Filtered 15T Token Dataset Down to 1.5T Educational Tokens | Loubna Ben Allal | Dec 24, 2024 | — |
| 15T | “15 trillion” | Hugging Face Filtered 15T Token Dataset Down to 1.5T Educational Tokens | Loubna Ben Allal | Dec 24, 2024 | — |
| 1.5T | “1.5 trillion” | Hugging Face Filtered 15T Token Dataset Down to 1.5T Educational Tokens | Loubna Ben Allal | Dec 24, 2024 | — |
| 1T | “one trillion” | Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA | Loubna Ben Allal | Dec 24, 2024 | Supported |
| 15T | “15 trillion” | Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA | Loubna Ben Allal | Dec 24, 2024 | Supported |
| 100M | “hundred million” | Meta research: Sub-1B language models benefit more from depth than width | Loubna Ben Allal | Dec 24, 2024 | — |
| 1T | “one trillion” | Ben Allal: Small models continue improving when trained on 11T tokens | Loubna Ben Allal | Dec 24, 2024 | — |
| 11T | “11 trillion” | Ben Allal: Small models continue improving when trained on 11T tokens | Loubna Ben Allal | Dec 24, 2024 | — |
| 2× | “two times” | Fu: Efficient AI Architectures Are Dead on Arrival Without Hardware Co-Design | Dan Fu | Dec 24, 2024 | — |
| 3M | “a two million” | Dan Fu: Nobody is actually submitting 2M token prompts into LLMs | Dan Fu | Dec 24, 2024 | — |
| 1M | “a million” | Cheah: Non-positional attention architectures remain stable beyond trained context | Eugene Cheah | Dec 24, 2024 | — |
| 1K | “a thousand” | Soldani: Frontier LLM pre-training requires at least 50,000 GPUs | Luca Soldani | Dec 23, 2024 | — |
| 6× | “six times” | Robinson: SAM 2's hierarchical encoder delivers 6x faster inference than ViT | Isaac Robinson | Dec 22, 2024 | Supported |
| 60% | “60%” | Robicheaux: Florence-2 Achieves 60% mAP on COCO | Peter Robicheaux | Dec 22, 2024 | Supported |
| 300M | “three hundred million” | Robicheaux: PaliGemma 1 Pre-Training Saturates at 300M Examples | Peter Robicheaux | Dec 22, 2024 | Supported |
| 2B | “two billion” | Robicheaux: 2B PaliGemma 2 Beats ChatGPT on MMVP with 47.3% | Peter Robicheaux | Dec 22, 2024 | Supported |
| 94% | “94%” | Robicheaux: 2B PaliGemma 2 Beats ChatGPT on MMVP with 47.3% | Peter Robicheaux | Dec 22, 2024 | Supported |
| 2B | “two billion” | Korupati: Moondream built a 0.5B model by pruning its 2B model | Vik Korupati | Dec 22, 2024 | — |
| 90% | “90%” | Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60% | Pranav Reddy | Dec 21, 2024 | Supported |
| 60% | “60%” | Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60% | Pranav Reddy | Dec 21, 2024 | Supported |
| 85% | “85%” | Reddy: Flagship OpenAI API costs fell 80% to 85% in roughly 18 months | Pranav Reddy | Dec 21, 2024 | Supported |
| $40B | “forty billion dollars” | Reddy: Mega model labs absorbed $30B-$40B of 2024 AI venture funding | Pranav Reddy | Dec 21, 2024 | Supported |
| 10% | “10%” | Mohan: GitHub might have under 10% full penetration in Fortune 500 | Varun Mohan | Dec 13, 2024 | — |
| 10% | “10%” | Mohan: Squeezing the final 10% on AI benchmarks encourages p-hacking | Varun Mohan | Dec 13, 2024 | — |
| 80% | “80%” | Mohan: Over 80% of software developers are on Windows | Varun Mohan | Dec 13, 2024 | Contradicted |
| 10M | “ten million” | Ramachandran: Codeium reached $10M enterprise ARR in under a year | Anshul Ramachandran | Dec 13, 2024 | — |
| 2.1K | “2.1 thousand” | Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs | Sarah Chieng | Dec 7, 2024 | Supported |
| 70× | “70 times” | Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs | Sarah Chieng | Dec 7, 2024 | Supported |
| 4T | “four trillion” | Cerebras WSE-3 features 900,000 cores, 44GB SRAM, and 4 trillion transistors | Sarah Chieng | Dec 7, 2024 | Supported |
| 90% | “90%” | Cerebras weight streaming prunes up to 90% of data without accuracy loss | Sarah Chieng | Dec 7, 2024 | Supported |
| 4M | “four million” | Fanelli: Bolt Generated $4M in Revenue in Its First Four Weeks | Alessio Fanelli | Dec 2, 2024 | — |
| 50% | “50%” | Friedman: GitHub Copilot user retention in enterprise is 38% to 50% | Itamar Friedman | Dec 2, 2024 | Contradicted |
| 1B | “one billion” | Friedman: StackBlitz must choose between Bolt and WebContainers or raise $1B | Itamar Friedman | Dec 2, 2024 | — |
| 1.5M | “half a million” | Simons: Bolt.new is adding $500,000 in ARR per day | Eric Simons | Dec 2, 2024 | — |
| 10× | “10 X” | Simons: Base AI models provide a 10x lift; agent scaffolding adds 3-4x | Eric Simons | Dec 2, 2024 | — |
| 4× | “four X” | Simons: Base AI models provide a 10x lift; agent scaffolding adds 3-4x | Eric Simons | Dec 2, 2024 | — |
| 30% | “30%” | Simons: Token add-ons drive 20% to 30% of Bolt's revenue | Eric Simons | Dec 2, 2024 | — |
| 50% | “50%” | Eugene Yan: Removing document chunking boosted pipeline downstream metrics by 50% | Eugene Yan | Nov 29, 2024 | — |
| 99% | “99%” | Schluntz: Reliability, not demo capability, is the bottleneck for robotics | Erik Schluntz | Nov 28, 2024 | — |
| 1K | “a thousand” | Fireworks AI's multi-LoRA system serves up to 1,000 adapters per base model | Lin Qiao | Nov 25, 2024 | Supported |
| 70% | “70%” | Crivello: 60% to 70% of user agent prompts are meaningless | Florent Crivello | Nov 15, 2024 | — |
| 1K | “a thousand” | Crivello: Lindy's calendar availability action requires roughly 1,000 lines of code | Florent Crivello | Nov 15, 2024 | — |
| 60% | “60%” | Crivello: WordPress is a commercial failure despite powering 60% of the web | Florent Crivello | Nov 15, 2024 | — |
| 10% | “10%” | Crivello: AI has 10% p(doom) and 90% chance of disease-free utopia | Florent Crivello | Nov 15, 2024 | — |
| 90% | “90%” | Crivello: AI has 10% p(doom) and 90% chance of disease-free utopia | Florent Crivello | Nov 15, 2024 | — |
| 88% | “88%” | Polu: Dust averages 60% to 70% weekly active penetration in enterprise accounts | Stanislas Polu | Nov 11, 2024 | — |
| 70% | “70%” | Polu: Dust averages 60% to 70% weekly active penetration in enterprise accounts | Stanislas Polu | Nov 11, 2024 | — |
| 2× | “two X” | Distilling sCM Requires Roughly Twice the Compute of Teacher Training | RJ Honicky | Nov 2, 2024 | Contradicted |
| 30% | “30%” | Chiang: Coding questions drive 20% to 30% of Chatbot Arena usage | Wei-Lin Chiang | Nov 1, 2024 | — |
| 1T | “one trillion” | He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain | Ethan He | Oct 29, 2024 | Supported |
| 5% | “five percent” | He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain | Ethan He | Oct 29, 2024 | Supported |
| 4% | “four percent” | He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain | Ethan He | Oct 29, 2024 | Supported |
| 15% | “15%” | Fanelli: Singapore accounted for 15% of NVIDIA's Q3 2024 revenue | Alessio Fanelli | Oct 19, 2024 | Supported |
| 3% | “three percent” | Hu: SWE-bench scores jumped from 3% on GPT-4 RAG to 43% | Jesse Hu | Oct 19, 2024 | Supported |
| 43% | “43%” | Hu: SWE-bench scores jumped from 3% on GPT-4 RAG to 43% | Jesse Hu | Oct 19, 2024 | Supported |
| 100% | “hundred percent” | Hu: AI Models Should Eventually Reach 100% on SWE-bench Verified | Jesse Hu | Oct 19, 2024 | Open |
| 12% | “12%” | SWE-bench Multimodal paper baseline scores 12 percent | Jesse Hu | Oct 19, 2024 | Supported |
| 17% | “17%” | Hu: OpenAI o1-preview achieves bronze medals in 17% of MLE-bench competitions | Jesse Hu | Oct 19, 2024 | Supported |
| 1B | “a billion” | Drew Houston: Google Maps did more for autonomy than actual self-driving | Drew Houston | Oct 18, 2024 | — |
| 90% | “90%” | Houston: Dropbox Went 90% Remote to Act as Distributed Work Lab | Drew Houston | Oct 18, 2024 | — |
| 8B | “eight billion” | Drew Houston brings an external GPU on planes to run Llama locally | Drew Houston | Oct 18, 2024 | — |
| 95% | “95%” | Drew Houston: Regular expressions solve 95% of email triage NLP tasks | Drew Houston | Oct 18, 2024 | — |
| 8B | “eight billion” | Houston: Production AI systems should use minimum viable fine-tuned models | Drew Houston | Oct 18, 2024 | — |
| 100× | “hundred X” | Drew Houston: Massive AI price-performance gains wouldn't happen without open source | Drew Houston | Oct 18, 2024 | — |
| 5× | “five X” | Drew Houston: AI hardware is currently in a 'rent, not buy' phase | Drew Houston | Oct 18, 2024 | — |
| 1M | “a million” | Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples | Vibhu Sapra | Oct 13, 2024 | Supported |
| 1M | “a million” | Pixmo Dataset Contains Approximately 1M Captions Across 700K Images | Vibhu Sapra | Oct 13, 2024 | Supported |
| 90% | “90%” | Goyal: Simple tool-calling prompts cover 80% to 90% of AI use cases | Ankur Goyal | Oct 11, 2024 | — |
| 75% | “75%” | Goyal: Over 75% of Braintrust Eval Users Use TypeScript SDK | Ankur Goyal | Oct 11, 2024 | — |
| 100% | “hundred percent” | Goyal: Braintrust saw nearly 100% OpenAI market share pre-Claude 3 | Ankur Goyal | Oct 11, 2024 | — |
| 5% | “five percent” | Goyal: Under 5% of Braintrust production customers use open source | Ankur Goyal | Oct 11, 2024 | — |
| 50% | “50%” | Goyal: Single-prompt manipulations make up about 50% of Braintrust AI workloads | Ankur Goyal | Oct 11, 2024 | — |
| 25% | “25%” | Goyal: AI workloads are roughly 25% simple agents and 25% advanced agents | Ankur Goyal | Oct 11, 2024 | — |
| 3M | “three million” | Huet: Over 3 million developers build on OpenAI | Romain Huet | Oct 4, 2024 | Supported |
| 2% | “two percent” | Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit | Shawn Wang | Oct 4, 2024 | Supported |
| 15× | “15 X” | Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit | Shawn Wang | Oct 4, 2024 | Supported |
| 200M | “two hundred million” | Weil: ChatGPT supports over 200 million weekly active users | Kevin Weil | Oct 4, 2024 | Supported |
| 20% | “20%” | Weil: OpenAI customer support team is 20% expected size thanks to AI | Kevin Weil | Oct 4, 2024 | — |
| 10M | “Ten million” | Altman: 10-million-token fast context windows are coming within months | Sam Altman | Oct 4, 2024 | Held up |
| 2% | “two percent” | Shunyu Yao says academic AI research overcomplicates methods on simplistic tasks | Shunyu Yao | Sep 27, 2024 | — |
| 90% | “90%” | Yao: Reliable tool design accounts for 90% of agent performance | Shunyu Yao | Sep 27, 2024 | — |
| 99% | “99%” | Yao: Enterprise customer support AI requires 99% reliability over simple tasks, not search | Shunyu Yao | Sep 27, 2024 | — |
| 100× | “hundred times” | Yao: Enterprise customer support AI requires 99% reliability over simple tasks, not search | Shunyu Yao | Sep 27, 2024 | — |
| 50% | “50%” | Karpathy: llm.c achieves nearly 50% MFU on single-node GPT-2 training | Andrej Karpathy | Sep 21, 2024 | Supported |
| 30% | “30%” | Karpathy: llm.c was 20% faster and used 30% less memory than PyTorch | Andrej Karpathy | Sep 21, 2024 | Supported |
| 20% | “20%” | Karpathy: llm.c was 20% faster and used 30% less memory than PyTorch | Andrej Karpathy | Sep 21, 2024 | Supported |
| 90% | “90%” | Schulhoff: Few-Shot Exemplar Order Can Shift Model Accuracy From 0% to 90% | Sander Schulhoff | Sep 20, 2024 | Supported |
| 1K | “a thousand” | Schulhoff: GPT-4 Fails to Output Reasoning on 1 in 100 to 1,000 Prompts | Sander Schulhoff | Sep 20, 2024 | — |
| 10× | “10 times” | Swix: Multi-sampling GPT-4o mini before GPT-4o judging yields net savings | Shawn Wang | Sep 20, 2024 | Supported |
| $1.5M | “a half million dollars” | Schulhoff: Hack-A-Prompt 2 Aims to Award $500,000 for Harmful AI Dataset | Sander Schulhoff | Sep 20, 2024 | — |
| 1M | “one million” | Jamil: Writing in the Margins avoids re-prefilling tokens, halving compute cost | Umar Jamil | Sep 19, 2024 | — |
| 2M | “two million” | Jamil: Writing in the Margins avoids re-prefilling tokens, halving compute cost | Umar Jamil | Sep 19, 2024 | — |
| 9× | “nine X” | Masking out prior KV cache tokens pushes autoregressive transformers out of distribution | Umar Jamil | Sep 19, 2024 | — |
| 100% | “100%” | Jamil: OpenAI and Cohere Overlap Prefill and Generation to Maximize GPU Utilization | Umar Jamil | Sep 19, 2024 | — |
| 100% | “hundred percent” | Function calling benchmarks like BFCL are largely saturated | Michelle Pokrass | Sep 17, 2024 | — |
| 95% | “95%” | Multi-step agentic apps fail at 95% reliability due to compounded errors | Michelle Pokrass | Sep 17, 2024 | — |
| 100% | “hundred percent” | Multi-step agentic apps fail at 95% reliability due to compounded errors | Michelle Pokrass | Sep 17, 2024 | — |
| 55% | “55%” | Swix: OpenAI Structured Outputs cut API costs by 55% vs Instructor | Shawn Wang | Sep 17, 2024 | — |
| 1K | “a thousand” | Fine-tuning requires only 100 to 1,000 high-quality examples | Michelle Pokrass | Sep 17, 2024 | — |
| 1M | “a million” | Carlini: 90% of scientific research is routine work that AI can automate | Nicholas Carlini | Aug 28, 2024 | — |
| 90% | “90%” | Carlini: 90% of scientific research is routine work that AI can automate | Nicholas Carlini | Aug 28, 2024 | — |
| 100% | “hundred percent” | Carlini: Anyone claiming 0% or 100% certainty on 5-year AI capabilities is probably wrong | Nicholas Carlini | Aug 28, 2024 | — |
| 2% | “two percent” | Prompting ChatGPT to repeat a word indefinitely leaks verbatim training data | Nicholas Carlini | Aug 28, 2024 | Supported |
| 66% | “66%” | Cosine's Genie achieved roughly 66% codebase retrieval accuracy across benchmark tasks | Alistair Pullen | Aug 22, 2024 | Supported |
| 43.8% | “43.8%” | Cosine's Genie achieved a state-of-the-art 43.8% on SWE-bench Verified | Alistair Pullen | Aug 22, 2024 | Supported |
| 80% | “80%” | Howard: 80% of top unique creators have unconventional or non-mainstream backgrounds | Jeremy Howard | Aug 17, 2024 | — |
| 99% | “99%” | Howard: Developers should distribute merged adapters rather than merged models | Jeremy Howard | Aug 17, 2024 | — |
| 49M | “forty nine million” | Nelson: Roboflow users labeled 49M images using SAM in one year | Joseph Nelson | Aug 7, 2024 | — |
| 5M | “five million” | Nelson: Roboflow users labeled 49M images using SAM in one year | Joseph Nelson | Aug 7, 2024 | — |
| 30M | “thirty million” | Ravi: SAM 2's largest model is 224M parameters, one-third of SAM 1 | Nikhila Ravi | Aug 7, 2024 | Supported |
| 24M | “twenty-four million” | Ravi: SAM 2's largest model is 224M parameters, one-third of SAM 1 | Nikhila Ravi | Aug 7, 2024 | Supported |
| 6× | “six times” | Ravi: SAM 2 runs roughly six times faster on video than SAM 1 | Nikhila Ravi | Aug 7, 2024 | Supported |
| 38M | “thirty-eight million” | Nelson: Smallest SAM 2 model has 38M parameters and runs at 45 FPS | Joseph Nelson | Aug 7, 2024 | Supported |
| 200M | “two hundred million” | Reddit makes over 200 million dollars in AI data licensing deals | Alessio Fanelli | Aug 2, 2024 | Supported |
| 1M | “a million” | E2B scaled from 10,000 to one million cloud containers in four months | Alessio Fanelli | Aug 2, 2024 | — |
| 30% | “30%” | Scialom: Llama 3's 128K tokenizer reduces token count by about 30% | Thomas Scialom | Jul 23, 2024 | Partly supported |
| 3% | “three percent” | Albrecht: Imbue's GPU cluster failure rate is well below industry 3% benchmark | Josh Albrecht | Jun 25, 2024 | — |
| 1K | “a thousand” | Albrecht: Imbue reproduced 500-1,000 examples per dataset to stop eval contamination | Josh Albrecht | Jun 25, 2024 | — |
| 10% | “10%” | Albrecht: Coding agents communicating uncertainty are far more useful than slightly more accurate ones | Josh Albrecht | Jun 25, 2024 | — |
| 100% | “hundred percent” | Albrecht: Coding agents communicating uncertainty are far more useful than slightly more accurate ones | Josh Albrecht | Jun 25, 2024 | — |
| 10× | “10 X” | Brady: Elicit routinely sees 10x p90 latency variation when prompting LLMs | James Brady | Jun 21, 2024 | — |
| 500M | “five hundred million” | Conover: Labor flow networks predict next-quarter S&P 500 market cap changes | Mike Conover | Jun 11, 2024 | Partly supported |
| 7M | “a six million” | Brightwave raises $6M Seed round led by Decibel Partners | Mike Conover | Jun 11, 2024 | — |
| 100% | “100%” | Conover: AI companies hire systems engineers, traditional software hires AI talent | Mike Conover | Jun 11, 2024 | — |
| 21B | “a twenty billion” | Conover: A $20B crossover hedge fund uses Brightwave for equity research | Mike Conover | Jun 11, 2024 | — |
| 1M | “a million” | Conover: Million-token context windows fail to extract deep insights from SEC filings | Mike Conover | Jun 11, 2024 | — |
| 35% | “35%” | Conover: One-year patient adherence rate for Ozempic is only 35% | Mike Conover | Jun 11, 2024 | Supported |
| 100% | “100%” | Conover: AI model developers are absolutely overfitting to public evaluation benchmarks | Mike Conover | Jun 11, 2024 | — |
| $40M | “forty million dollars” | Conover: Economic incentives to pre-train commodity foundation models from scratch are diminishing | Mike Conover | Jun 11, 2024 | — |
| 400B | “four hundred billion” | Conover: Economic incentives to pre-train commodity foundation models from scratch are diminishing | Mike Conover | Jun 11, 2024 | — |
| 40M | “forty million” | Conover: Economic incentives to pre-train commodity foundation models from scratch are diminishing | Mike Conover | Jun 11, 2024 | — |
| 1B | “a billion” | Huang: Context scaling requires positional interpolation rather than extrapolation | Mark Huang | May 31, 2024 | Supported |
| 1M | “a million” | Huang: PoSE breaks down on needle-in-a-haystack at 500k tokens | Mark Huang | May 31, 2024 | Open |
| 1B | “a billion” | Huang: Adding one billion tokens cannot teach trillion-token models new knowledge | Mark Huang | May 31, 2024 | — |
| 3M | “three million” | Liu: Fine-tuning on proprietary data at scale always beats off-the-shelf models | Jason Liu | Apr 24, 2024 | — |
| 1B | “a billion” | Liu: Fine-tuning on proprietary data at scale always beats off-the-shelf models | Jason Liu | Apr 24, 2024 | — |
| 1B | “a billion” | Liu: Instructor is not a billion-dollar startup opportunity | Jason Liu | Apr 24, 2024 | — |
| $100M | “a hundred million dollars” | Liu: LangChain and LlamaIndex can hit $100M revenue, but billions uncertain | Jason Liu | Apr 24, 2024 | — |
| 2M | “two million” | Stuhlmüller: Pure long-context LLMs are significantly harder to debug than RAG | Andreas Stuhlmüller | Apr 11, 2024 | — |
| 175B | “a hundred seventy-five billion” | Streaming throughput constraints prevent Suno from scaling to 175 billion parameters | Mikey Shulman | Mar 14, 2024 | — |
| 50× | “50 times” | Shulman: Gaming dwarfs music 50x because music is mostly passive consumption | Mikey Shulman | Mar 14, 2024 | — |
| 1% | “one percent” | Chintala: Less than 1% of open source AI usage yields feedback | Soumith Chintala | Mar 6, 2024 | — |
| 2M | “Two million” | Firshman: Replicate has 2M total users, not all developers | Ben Firshman | Feb 28, 2024 | — |
| $25M | “twenty five million dollars” | Firshman: Llama 2 costs $25M to train but $50 to fine-tune | Ben Firshman | Feb 28, 2024 | Partly supported |
| 1K | “a thousand” | Better.com Grew to 10,000 Employees Before Shrinking Back to 1,000 | Erik Bernhardsson | Feb 19, 2024 | Supported |
| 99% | “99%” | 99% of Data in Multi-Gigabyte Docker Images Is Never Read | Erik Bernhardsson | Feb 19, 2024 | — |
| 10× | “10 X” | Developer Productivity Will Continue Growing 10x Per Decade | Erik Bernhardsson | Feb 19, 2024 | — |
| 10,000× | “10,000 X” | Developer Productivity Will Continue Growing 10x Per Decade | Erik Bernhardsson | Feb 19, 2024 | — |
| 10× | “10 x” | AI Productivity Gains Will Ultimately Increase Demand for Software Engineers | Erik Bernhardsson | Feb 19, 2024 | — |
| 5M | “five million” | Prakash predicts up to 5 million AI GPUs will sell in 2024 | Vipul Ved Prakash | Feb 8, 2024 | Held up |
| 10× | “10 times” | Zhang: Next 10x AI inference gain requires multi-layer co-optimization | Ce Zhang | Feb 8, 2024 | — |
| 45% | “45%” | Prakash: 40% to 45% of Together AI's team is dedicated to research | Vipul Ved Prakash | Feb 8, 2024 | — |
| $1M | “a million dollars” | Hsu: Retool had 5-6 years of runway paying founders $30k-$40k | David Hsu | Feb 7, 2024 | — |
| 1M | “a million” | Hsu: Retool had 5-6 years of runway paying founders $30k-$40k | David Hsu | Feb 7, 2024 | — |
| 95% | “95%” | Hsu: 95% of Retool customers are developers | David Hsu | Feb 7, 2024 | — |
| 10% | “10%” | Hsu: AI coding assistants only speed up development by 10% to 20% | David Hsu | Feb 7, 2024 | — |
| 20% | “20%” | Hsu: AI coding assistants only speed up development by 10% to 20% | David Hsu | Feb 7, 2024 | — |
| 10× | “tenx” | Hsu: AI coding assistants only speed up development by 10% to 20% | David Hsu | Feb 7, 2024 | — |
| 70% | “70%” | Hsu: 60% to 70% of AI business utility comes from ChatGPT | David Hsu | Feb 7, 2024 | — |
| 20% | “20%” | Hsu: AI chat only boosts productivity 10-20% and won't replace employees | David Hsu | Feb 7, 2024 | — |
| 70B | “seventy billion” | Lambert: Scaling from 7B to 70B parameters fixes nuance and repetition | Nathan Lambert | Jan 11, 2024 | — |
| 100× | “hundred times” | Lambert: A 100x smaller language model filters output better than RLHF | Nathan Lambert | Jan 11, 2024 | — |
| 8M | “eight million” | Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data | Nathan Lambert | Jan 11, 2024 | — |
| 40% | “40%” | Lambert: Only 20% to 40% of Meta's Llama RLHF data is useful | Nathan Lambert | Jan 11, 2024 | — |
| 70% | “70%” | Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans | Nathan Lambert | Jan 11, 2024 | Supported |
| 80% | “80%” | Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans | Nathan Lambert | Jan 11, 2024 | Supported |
| 75% | “75%” | Lambert: RLHF reward models achieve only 65% to 75% validation agreement | Nathan Lambert | Jan 11, 2024 | Supported |
| 100% | “hundred percent” | Lambert: RLHF reward models achieve only 65% to 75% validation agreement | Nathan Lambert | Jan 11, 2024 | Supported |
| 70B | “seventy billion” | Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations | Nathan Lambert | Jan 11, 2024 | — |
| 22M | “twenty two million” | Ruiz: tldraw's 'Make It Real' demo garnered 22 million views in 30 days | Steve Ruiz | Jan 5, 2024 | — |
| 1B | “a billion” | Liu: Long-context recall depends directly on needle-in-haystack training loss | Beyang Liu | Dec 17, 2023 | — |