Guy Gur-Ari

5 statements across 1 episodes · 5 bullish · 0 bearish · 1 people on the record · first statement Apr 2, 2025 by Guy Gur-Ari · across every show →

On the record as a speaker too: Guy Gur-Ari's record, appearances and statements → this page counts the times other people say the name.

Everything said about Guy Gur-Ari, oldest first

Apr 2, 2025 positive
Insight
Engineers should begin AI feature evaluation with 10-sample interactive notebooks
“So I would say the process that I like to follow is in the beginning when developing a feature, come up with a curated set of samples. Could be as small as 10, 10 samples that you run through, and then yeah, it's all notebooks basically. You run through the sa…”
Guy Gur-Ari Apr 2, 2025 ▶ 6:10 The #1 SWE-Bench Verified Agent
Apr 2, 2025 positive
Assertion Not checkable as stated
Sequential Thinking MCP outperformed Claude 3.7 native reasoning mode in evaluations
“We tried reasoning mode as well with the new three seven. And we didn't see that much of a bump in performance. We don't know if this is something that's code specific or not. I don't have an insight. We tried both and yeah, sequential thinking worked better.”
Guy Gur-Ari Apr 2, 2025 ▶ 3:57 The #1 SWE-Bench Verified Agent
Apr 2, 2025 positive
Opinion
Gur-Ari highlights the SWIRL research paper for reinforcement learning in coding
“There was an interesting paper called SWIRL, which I thought was a really nice paper on how to do RL for, specifically for coding. So that's by Wei and other authors.”
Guy Gur-Ari Apr 2, 2025 ▶ 27:06 The #1 SWE-Bench Verified Agent
Apr 2, 2025 bullish
Prediction Not checkable as stated
Running multiple parallel agents will unlock most value for software developers
“Being able to run multiple agents and not just one is going to be the way to unlock Honestly, most of the value out of these agents, and that's what we're working toward.”
Guy Gur-Ari Apr 2, 2025 ▶ 13:56 The #1 SWE-Bench Verified Agent
Apr 2, 2025 positive
Insight
Ground-truth evaluations without code execution provide significant mileage for AI
“So code execution, like checking the correctness of solutions and running tests automatically can help, although not, although you can get a lot of mileage out of evals that don't have code execution in them that just compare against ground truth.”
Guy Gur-Ari Apr 2, 2025 ▶ 6:48 The #1 SWE-Bench Verified Agent
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.