Jul 15, 2025 · 40m · startup-ideas

Does Grok 4 Deserve a Spot In Your AI Stack? (Here's The Truth)

Greg Isenberg · 34m spoken Grok AI · 32s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Host Greg Isenberg conducts a comprehensive live benchmarking evaluation of xAI's Grok 4 across nine business and creative agent capabilities to determine whether the model warrants a permanent place in a modern founder's AI stack.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Greg holds 98.5% of the talking time here. How this is scored →

Greg as informed peer 5.2 Guest teaching 1.8 Guest disagreement 0.3 Greg pushing back 1.8
05100:0015:0030:001:00–4:04 · Greg as informed peer 6/10 Market Research Agent Test with Live X Search Greg tests Grok 4's market research capabilities using live X search to analyze productivity apps. He demonstrates domain expertise by confirming Grok's findings regarding Notion's slowness issues, citing his own viral post on the topic.4:05–6:47 · Greg as informed peer 5/10 MVP Development and Code Generation in Code Mode Greg experiments with Grok 4's Python code generation for a lead-generation bot. He assesses the workflow in comparison to Replit, Bolt, and Lovable while remaining cautious about practical deployment readiness.6:47–10:11 · Greg as informed peer 0/10 Sponsor Segment: IdeaBrowser.com Sponsor read and product plug for IdeaBrowser.com. As a solo monologue/ad read, interaction scores remain at zero.10:11–14:48 · Greg as informed peer 6/10 Pitch Deck Script Refinement and Slide Generation Testing Greg tests VC reasoning mode on a seed pitch deck and probes whether Grok can design presentation slides directly, noting limitations when Grok only outputs text descriptions.14:49–23:24 · Greg as informed peer 7/10 Multi-Agent Content Marketing and Style Adaptation Greg encounters generic, hashtag-heavy outputs from the content marketing prompt. He repeatedly pushes back, corrects the model's scraping failure, and guides it by referencing his specific writing style.23:24–27:08 · Greg as informed peer 6/10 Automated Customer Feedback and Sentiment Analysis Greg evaluates sentiment analysis and NPS calculation on customer reviews. He enthusiastically notes the utility of summarizing feedback into roadmaps and retention forecasts.27:08–29:50 · Greg as informed peer 5/10 Multi-Agent Salary Negotiation Simulation Greg tests multi-agent salary negotiation simulations and highlights the tactical advice of timing compensation discussions immediately following major project wins.29:50–32:57 · Greg as informed peer 6/10 Trend Forecasting and Innovation Feature Modeling Greg runs trend forecasting for productivity apps and corroborates Grok's prediction regarding shifts toward canvas-based and node-based UI workflows.32:58–35:20 · Greg as informed peer 5/10 Startup Branding and Vector Logo Design Testing Greg tests logo generation and vector rendering. He points out that the tool failed to produce the requested four options and broke text formatting during color edits.35:20–38:23 · Greg as informed peer 4/10 Exploring Grok Companion Mode and AI Avatars Greg tests Super Grok's companion mode and voice assistant. The companion opens with flirtatious persona banter, eliciting an awkward pushback from Greg before he tests voice brainstorming.38:23–40:49 · Greg as informed peer 7/10 Final Evaluation and Verdict on Grok 4's Agent Stack Greg provides an exhaustive final evaluation of all Grok 4 agent modes, giving balanced verdicts on strengths and areas needing improvement.1:00–4:04 · Guest teaching 2/10 Market Research Agent Test with Live X Search Greg tests Grok 4's market research capabilities using live X search to analyze productivity apps. He demonstrates domain expertise by confirming Grok's findings regarding Notion's slowness issues, citing his own viral post on the topic.4:05–6:47 · Guest teaching 2/10 MVP Development and Code Generation in Code Mode Greg experiments with Grok 4's Python code generation for a lead-generation bot. He assesses the workflow in comparison to Replit, Bolt, and Lovable while remaining cautious about practical deployment readiness.6:47–10:11 · Guest teaching 0/10 Sponsor Segment: IdeaBrowser.com Sponsor read and product plug for IdeaBrowser.com. As a solo monologue/ad read, interaction scores remain at zero.10:11–14:48 · Guest teaching 3/10 Pitch Deck Script Refinement and Slide Generation Testing Greg tests VC reasoning mode on a seed pitch deck and probes whether Grok can design presentation slides directly, noting limitations when Grok only outputs text descriptions.14:49–23:24 · Guest teaching 1/10 Multi-Agent Content Marketing and Style Adaptation Greg encounters generic, hashtag-heavy outputs from the content marketing prompt. He repeatedly pushes back, corrects the model's scraping failure, and guides it by referencing his specific writing style.23:24–27:08 · Guest teaching 3/10 Automated Customer Feedback and Sentiment Analysis Greg evaluates sentiment analysis and NPS calculation on customer reviews. He enthusiastically notes the utility of summarizing feedback into roadmaps and retention forecasts.27:08–29:50 · Guest teaching 3/10 Multi-Agent Salary Negotiation Simulation Greg tests multi-agent salary negotiation simulations and highlights the tactical advice of timing compensation discussions immediately following major project wins.29:50–32:57 · Guest teaching 3/10 Trend Forecasting and Innovation Feature Modeling Greg runs trend forecasting for productivity apps and corroborates Grok's prediction regarding shifts toward canvas-based and node-based UI workflows.32:58–35:20 · Guest teaching 1/10 Startup Branding and Vector Logo Design Testing Greg tests logo generation and vector rendering. He points out that the tool failed to produce the requested four options and broke text formatting during color edits.35:20–38:23 · Guest teaching 2/10 Exploring Grok Companion Mode and AI Avatars Greg tests Super Grok's companion mode and voice assistant. The companion opens with flirtatious persona banter, eliciting an awkward pushback from Greg before he tests voice brainstorming.38:23–40:49 · Guest teaching 0/10 Final Evaluation and Verdict on Grok 4's Agent Stack Greg provides an exhaustive final evaluation of all Grok 4 agent modes, giving balanced verdicts on strengths and areas needing improvement.1:00–4:04 · Guest disagreement 0/10 Market Research Agent Test with Live X Search Greg tests Grok 4's market research capabilities using live X search to analyze productivity apps. He demonstrates domain expertise by confirming Grok's findings regarding Notion's slowness issues, citing his own viral post on the topic.4:05–6:47 · Guest disagreement 0/10 MVP Development and Code Generation in Code Mode Greg experiments with Grok 4's Python code generation for a lead-generation bot. He assesses the workflow in comparison to Replit, Bolt, and Lovable while remaining cautious about practical deployment readiness.6:47–10:11 · Guest disagreement 0/10 Sponsor Segment: IdeaBrowser.com Sponsor read and product plug for IdeaBrowser.com. As a solo monologue/ad read, interaction scores remain at zero.10:11–14:48 · Guest disagreement 0/10 Pitch Deck Script Refinement and Slide Generation Testing Greg tests VC reasoning mode on a seed pitch deck and probes whether Grok can design presentation slides directly, noting limitations when Grok only outputs text descriptions.14:49–23:24 · Guest disagreement 1/10 Multi-Agent Content Marketing and Style Adaptation Greg encounters generic, hashtag-heavy outputs from the content marketing prompt. He repeatedly pushes back, corrects the model's scraping failure, and guides it by referencing his specific writing style.23:24–27:08 · Guest disagreement 0/10 Automated Customer Feedback and Sentiment Analysis Greg evaluates sentiment analysis and NPS calculation on customer reviews. He enthusiastically notes the utility of summarizing feedback into roadmaps and retention forecasts.27:08–29:50 · Guest disagreement 0/10 Multi-Agent Salary Negotiation Simulation Greg tests multi-agent salary negotiation simulations and highlights the tactical advice of timing compensation discussions immediately following major project wins.29:50–32:57 · Guest disagreement 0/10 Trend Forecasting and Innovation Feature Modeling Greg runs trend forecasting for productivity apps and corroborates Grok's prediction regarding shifts toward canvas-based and node-based UI workflows.32:58–35:20 · Guest disagreement 0/10 Startup Branding and Vector Logo Design Testing Greg tests logo generation and vector rendering. He points out that the tool failed to produce the requested four options and broke text formatting during color edits.35:20–38:23 · Guest disagreement 2/10 Exploring Grok Companion Mode and AI Avatars Greg tests Super Grok's companion mode and voice assistant. The companion opens with flirtatious persona banter, eliciting an awkward pushback from Greg before he tests voice brainstorming.38:23–40:49 · Guest disagreement 0/10 Final Evaluation and Verdict on Grok 4's Agent Stack Greg provides an exhaustive final evaluation of all Grok 4 agent modes, giving balanced verdicts on strengths and areas needing improvement.1:00–4:04 · Greg pushing back 1/10 Market Research Agent Test with Live X Search Greg tests Grok 4's market research capabilities using live X search to analyze productivity apps. He demonstrates domain expertise by confirming Grok's findings regarding Notion's slowness issues, citing his own viral post on the topic.4:05–6:47 · Greg pushing back 1/10 MVP Development and Code Generation in Code Mode Greg experiments with Grok 4's Python code generation for a lead-generation bot. He assesses the workflow in comparison to Replit, Bolt, and Lovable while remaining cautious about practical deployment readiness.6:47–10:11 · Greg pushing back 0/10 Sponsor Segment: IdeaBrowser.com Sponsor read and product plug for IdeaBrowser.com. As a solo monologue/ad read, interaction scores remain at zero.10:11–14:48 · Greg pushing back 2/10 Pitch Deck Script Refinement and Slide Generation Testing Greg tests VC reasoning mode on a seed pitch deck and probes whether Grok can design presentation slides directly, noting limitations when Grok only outputs text descriptions.14:49–23:24 · Greg pushing back 6/10 Multi-Agent Content Marketing and Style Adaptation Greg encounters generic, hashtag-heavy outputs from the content marketing prompt. He repeatedly pushes back, corrects the model's scraping failure, and guides it by referencing his specific writing style.23:24–27:08 · Greg pushing back 1/10 Automated Customer Feedback and Sentiment Analysis Greg evaluates sentiment analysis and NPS calculation on customer reviews. He enthusiastically notes the utility of summarizing feedback into roadmaps and retention forecasts.27:08–29:50 · Greg pushing back 0/10 Multi-Agent Salary Negotiation Simulation Greg tests multi-agent salary negotiation simulations and highlights the tactical advice of timing compensation discussions immediately following major project wins.29:50–32:57 · Greg pushing back 1/10 Trend Forecasting and Innovation Feature Modeling Greg runs trend forecasting for productivity apps and corroborates Grok's prediction regarding shifts toward canvas-based and node-based UI workflows.32:58–35:20 · Greg pushing back 4/10 Startup Branding and Vector Logo Design Testing Greg tests logo generation and vector rendering. He points out that the tool failed to produce the requested four options and broke text formatting during color edits.35:20–38:23 · Greg pushing back 3/10 Exploring Grok Companion Mode and AI Avatars Greg tests Super Grok's companion mode and voice assistant. The companion opens with flirtatious persona banter, eliciting an awkward pushback from Greg before he tests voice brainstorming.38:23–40:49 · Greg pushing back 1/10 Final Evaluation and Verdict on Grok 4's Agent Stack Greg provides an exhaustive final evaluation of all Grok 4 agent modes, giving balanced verdicts on strengths and areas needing improvement.

speaking balance: gold is Greg, purple is the guest (3 minute bins)

0:00 · Greg 100% · guest 0%0:00 · Greg 100% · guest 0%3:00 · Greg 100% · guest 0%3:00 · Greg 100% · guest 0%6:00 · Greg 100% · guest 0%6:00 · Greg 100% · guest 0%9:00 · Greg 100% · guest 0%9:00 · Greg 100% · guest 0%12:00 · Greg 100% · guest 0%12:00 · Greg 100% · guest 0%15:00 · Greg 100% · guest 0%15:00 · Greg 100% · guest 0%18:00 · Greg 100% · guest 0%18:00 · Greg 100% · guest 0%21:00 · Greg 100% · guest 0%21:00 · Greg 100% · guest 0%24:00 · Greg 100% · guest 0%24:00 · Greg 100% · guest 0%27:00 · Greg 100% · guest 0%27:00 · Greg 100% · guest 0%30:00 · Greg 100% · guest 0%30:00 · Greg 100% · guest 0%33:00 · Greg 100% · guest 0%33:00 · Greg 100% · guest 0%36:00 · Greg 78.7% · guest 21.3%36:00 · Greg 78.7% · guest 21.3%39:00 · Greg 100% · guest 0%39:00 · Greg 100% · guest 0%
Sharpest disagreement ▶ 36:21 Companion avatar unsolicited banter

Grok AI companion adopts an aggressively informal, flirtatious persona calling Greg 'cutie' and asking who is stealing his attention.

Hardest push from Greg ▶ 18:50 Rejecting low-quality content marketing output

Greg rejects Grok's hashtag-stuffed marketing posts as terrible and explicitly tells the model it is overthinking before demanding a rewrite.

Biggest teaching moment ▶ 13:00 VC agent objection modeling

Grok provides sharp VC counterarguments against market saturation by citing competitor metrics and architectural differentiation.

Greg holds their own ▶ 20:50 Prompting style adaptation using personal tweets

Greg demonstrates expert prompt engineering by feeding his own proven tweet architecture into the model to turn bad output into high-performing copy.

the scores for every segment, with the reasoning behind each
ChapterTopicGreg as informed peerGuest teachingGuest disagreementGreg pushing backWhy
Market Research Agent Test with Live X Search 6201 Greg tests Grok 4's market research capabilities using live X search to analyze productivity apps. He demonstrates domain expertise by confirming Grok's findings regarding Notion's slowness issues, citing his own viral post on the topic.
MVP Development and Code Generation in Code Mode 5201 Greg experiments with Grok 4's Python code generation for a lead-generation bot. He assesses the workflow in comparison to Replit, Bolt, and Lovable while remaining cautious about practical deployment readiness.
Sponsor Segment: IdeaBrowser.com 0000 Sponsor read and product plug for IdeaBrowser.com. As a solo monologue/ad read, interaction scores remain at zero.
Pitch Deck Script Refinement and Slide Generation Testing 6302 Greg tests VC reasoning mode on a seed pitch deck and probes whether Grok can design presentation slides directly, noting limitations when Grok only outputs text descriptions.
Multi-Agent Content Marketing and Style Adaptation 7116 Greg encounters generic, hashtag-heavy outputs from the content marketing prompt. He repeatedly pushes back, corrects the model's scraping failure, and guides it by referencing his specific writing style.
Automated Customer Feedback and Sentiment Analysis 6301 Greg evaluates sentiment analysis and NPS calculation on customer reviews. He enthusiastically notes the utility of summarizing feedback into roadmaps and retention forecasts.
Multi-Agent Salary Negotiation Simulation 5300 Greg tests multi-agent salary negotiation simulations and highlights the tactical advice of timing compensation discussions immediately following major project wins.
Trend Forecasting and Innovation Feature Modeling 6301 Greg runs trend forecasting for productivity apps and corroborates Grok's prediction regarding shifts toward canvas-based and node-based UI workflows.
Startup Branding and Vector Logo Design Testing 5104 Greg tests logo generation and vector rendering. He points out that the tool failed to produce the requested four options and broke text formatting during color edits.
Exploring Grok Companion Mode and AI Avatars 4223 Greg tests Super Grok's companion mode and voice assistant. The companion opens with flirtatious persona banter, eliciting an awkward pushback from Greg before he tests voice brainstorming.
Final Evaluation and Verdict on Grok 4's Agent Stack 7001 Greg provides an exhaustive final evaluation of all Grok 4 agent modes, giving balanced verdicts on strengths and areas needing improvement.

Statements from this episode (10)

Opinion
Isenberg: Grok 4 Outperforms Rivals Due To Exclusive Real-Time X Data
“What obviously Grok has that other people don't have is they have X data, right? So the reason why Grok three was, you know, when we played with Grok three, it was like pretty darn good. So I'm not surprised that people are saying that Grok IV is really, reall…”
Greg Isenberg Jul 15, 2025 ▶ 1:51
Opinion
Isenberg: A Lightweight, Less-Bloated Notion Alternative Would Succeed
“Like, I think that, like, a lightweight version of Notion with, you know, a lot less bloat would actually do really well.”
Greg Isenberg Jul 15, 2025 ▶ 3:21
Assertion Not checkable as stated
Isenberg: Grok 4 Code Mode Rivals Bolt, Lovable, and Replit
“So this is similar to how, you know, if you're going to use something like a bolt or a lovable replet, it's going to do something similar. Now you have this with grok four as another option for you to go and go and do this.”
Greg Isenberg Jul 15, 2025 ▶ 4:50
Insight
Isenberg: The Best Grok 4 Hack Is Prompting For X Data
“So this is what I'm noticing is a hack for if you want to get the most out of Grok four, ask it to go and get data from X.”
Greg Isenberg Jul 15, 2025 ▶ 8:02
Opinion
Isenberg: Grok 4 Generates The Best AI Pitch Deck Scripts
“I will say I will say, by the way, that this refined script for this deck is probably, it might be the best I've ever seen.”
Greg Isenberg Jul 15, 2025 ▶ 14:02
Opinion
Isenberg: Grok 4 Excels at Mimicking Specific Creators on X
“So, if you want to use Grok for the lesson is it's really good for X. I need to play more, play with this more. You just do it simple. I want, I like this person's tweets. Go ahead and make me Tweets like that. But for your niche.”
Greg Isenberg Jul 15, 2025 ▶ 23:00
Opinion
Isenberg: Use Grok 4 As A Multi-Agent Negotiation Roleplay Partner
“And yeah, I think it's, if you're gonna negotiate anything big, or you just want someone to, you know, a negotiation partner, basically, like, It's going to help you get better at negotiation. Grok for pretty, pretty darn good.”
Greg Isenberg Jul 15, 2025 ▶ 29:36
Prediction Not checkable as stated
Isenberg: Future Productivity Software UIs Will Be Canvas and Node-Based
“Traditional rigid UIs will shift to canvas and node-based designs, enabling unrestrained, restricted visual thinking. That's absolutely correct. Deep integration to existing workflows. This is really good. So it's giving you product features too. The canvas-ba…”
Greg Isenberg Jul 15, 2025 ▶ 32:00
Disclosure
Isenberg: Automating Monthly Customer Feedback Analysis With Grok 4 and Slack
“Customer feedback analysis agent, I'm going to a hundred percent be using this. And I'm thinking to myself, like, how can I do this as an automation, right? So every single month, you know, it pulls our feedback, it asks rock for, and then it posts to our Slac…”
Greg Isenberg Jul 15, 2025 ▶ 39:18
Opinion
Isenberg: Grok 4 Branding Visuals Rate 6.7/10 and Break on Editing
“The branding visuals. I mean, I'm not, Kind of like, I'll give that, like, a 6.7 on 10. Like, it was cool that it was so quick, and it was pretty clean, but when I actually went and edited it, and went deeper into it started to break.”
Greg Isenberg Jul 15, 2025 ▶ 39:45
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.