Shawn Lewis

7 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Jan 28, 2025 by Shawn Lewis · said 1 times in 1 episodes since 2025 · across every show →

On the record as a speaker too: Shawn Lewis's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Kyle Corbitt (1)

tap a year for its mentions
0011112025episodesmentions
0112025episodes it came up in
000.50.5112025episodesmentions per episode
2025 1 mention in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Shawn Lewis, oldest first

Jan 28, 2025 positive
Assertion Not checkable as stated
Phase Shift's Eval Studio user interface was entirely written using Cursor
“Everything in this in the phase shift UI, this, or this eval studio UI that I showed you was written by AI. So this entire UI was written by AI, but it was not written by. Oh my God. They shipped it was written by cursor.”
Shawn Lewis Jan 28, 2025 ▶ 26:22 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Jan 28, 2025
Insight
Automating AI trace inspection is impossible; manual review remains mandatory
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still i…”
Shawn Lewis Jan 28, 2025 ▶ 21:58 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Jan 28, 2025 neutral
Assertion Not checkable as stated
Running a 100-problem SWE-bench evaluation takes one to two hours
“So a Sweebench eval for me takes about an hour to two hours to run on like a subset of a hundred problems.”
Shawn Lewis Jan 28, 2025 ▶ 10:56 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Jan 28, 2025
Insight
Debug agent regressions by qualitatively clustering failures across execution traces
“So it's really, like, I'll flip through these traces and kind of, like for each one, I'll write down notes about, like, what I thought went wrong there, and I'll do that for, like, say, 20 or so, and then I kind of go, okay, what's the biggest problem that we …”
Shawn Lewis Jan 28, 2025 ▶ 22:50 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Jan 28, 2025 bullish
Assertion Supported
Shawn Lewis: o1 agent achieves 57% single-pass, 64% with parallel rollouts
“It solves, like, something like 57% of problems with a single Rollout and then using parallel rollouts and selecting the best one. With other techniques, we get something like 64%.”
Shawn Lewis Jan 28, 2025 ▶ 15:04 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Jan 28, 2025
Disclosure
Lewis: Ran approximately 1,000 evaluations while developing SWE-bench agent
“You can see in the course of this, I did something like a thousand evals.”
Shawn Lewis Jan 28, 2025 ▶ 20:02 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Oct 16, 2025 neutral
Assertion Not checkable as stated
Corbitt: Weights & Biases founders drove CoreWeave's OpenPipe acquisition
“So that was driven by actually mostly the weights and biases founding team. Lucas and Sean, particularly. So they, had recently been acquired by CoreWeave and CoreWeave was looking to continue growing up the stack. And so, yeah, they approached me and were lik…”
Kyle Corbitt Oct 16, 2025 ▶ 1:00:24 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.