Jun 15, 2023 · 1h 20m · catalyst
AI for climate: a real world test
⌖ your search result is the highlighted band (32:53–33:31). Playback starts there
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Host Shayle Kann and energy trading expert Seyed Madaeni put generative AI to the test by tasking non-programmer Duncan Campbell with building a wholesale battery storage dispatch optimization algorithm using ChatGPT. The experiment demonstrates how iterative prompt engineering and human domain oversight can compress weeks of specialized quantitative development into hours.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Shayle holds 28.9% of the talking time here. How this is scored →
speaking balance: gold is Shayle, purple is the guest (3 minute bins)
Seyed firmly pushes back against relying on basic regression forecasts by emphasizing that poor price forecasts produce garbage in, garbage out results in trading optimization.
Hardest push from Shayle ▶ 20:22 Shayle questions whether software engineering is turning into essay writingShayle directly challenges the conventional boundary of coding by asking Seyed whether software engineers will simply be relegated to writing prose prompts.
Biggest teaching moment ▶ 43:05 Seyed breaks down the 90 percent complexity reductionSeyed provides a masterclass on wholesale battery commercialization, explaining how excluding real-time market bidding and ancillary services removed 90 percent of the engineering complexity.
Shayle holds their own ▶ 32:55 Shayle details multivariate regression requirements for LMP forecastingShayle demonstrates deep quantitative understanding by outlining the exact input parameters and regression mechanics required to forecast day-ahead wholesale electricity prices.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Shayle as informed peer | Guest teaching | Guest disagreement | Shayle pushing back | Why |
|---|---|---|---|---|---|---|
| Mid-Roll Sponsorship: Bloom Energy and Engie | 0 | 0 | 0 | 0 | Sponsor advertisements followed by the host introducing the premise of the episode as an experimental test of large language models on battery dispatch optimization. | |
| Defining the CAISO Battery Bidding Challenge | 6 | 5 | 1 | 2 | Seyed defines the technical constraints of CAISO wholesale bidding while Shayle demonstrates familiarity with acceptance criteria, market parameters, and battery cycling limits. | |
| Traditional Software Workflows vs. Non-Coder Abilities | 5 | 4 | 1 | 2 | The panel discusses the software engineering timeline for building dispatch algorithms, with Duncan acknowledging his non-coding background and reliance on Excel. | |
| The First Attempt: The 'Buckshot' Prompting Failure | 5 | 4 | 1 | 2 | Duncan describes the failure of pasting the prompt whole-cloth into ChatGPT, concluding that LLMs require bottom-up structural guidance rather than top-down delegation. | |
| The Second Attempt: Step-by-Step Prompting Strategy | 5 | 5 | 1 | 3 | Shayle probes whether prompt engineering is turning coding into essay writing, while Seyed emphasizes software architecture and human quality assurance. | |
| Debugging Physics and Mathematical Constraints | 5 | 5 | 2 | 2 | Duncan explains how physical boundary conditions like simultaneous charge and discharge and daily cycling constraints required explicit human intervention to resolve bad math outputs. | |
| Scaling Across Multiple Days and Skipping Price Forecasting | 6 | 5 | 3 | 3 | Shayle articulates how multivariate regression models would need to be specified for LMP forecasting, while Seyed notes that flawed forecasts create garbage in, garbage out optimization. | |
| Formatting Data Outputs and the Human-AI Partnership | 4 | 3 | 0 | 1 | Duncan shares how wrestling with visualization libraries felt like working with a junior analyst, highlighting the collaborative feeling of prompting. | |
| Mid-Roll Sponsorship: Bloom Energy and Engie | 6 | 6 | 2 | 2 | Following the mid-roll ad break, Seyed clarifies that by omitting ancillary services and market bidding, Duncan reduced the challenge's overall complexity by roughly ninety percent. | |
| Code Review and Section-by-Section Grading | 5 | 7 | 1 | 2 | Seyed performs a code review, praising the clean modular structure and linear optimization formulation while noting technical gaps like missing integer programming decision variables. | |
| Quantifying Productivity Gains and Economic Value | 6 | 4 | 1 | 3 | The panel compares the four hours spent against traditional engineering cycles, with Seyed estimating that producing equivalent groundwork would take two engineers several weeks. | |
| Production-Grade Reliability vs. Analytical Pre-Construction Use | 5 | 4 | 1 | 2 | The discussion distinguishes between real-time production-grade software and exploratory pre-construction modeling, identifying the script as valuable for analytical workflows. | |
| Live Model Execution and Dispatch Behavior Analysis | 6 | 3 | 0 | 1 | Duncan executes the live notebook, showing the resulting dispatch schedule and price curve charts while the panel observes how the model optimized state of charge across multiple peaks. | |
| Macro Implications of Generative AI in the Energy Sector | 6 | 4 | 1 | 2 | The speakers reflect on broader industry implications, agreeing that generative AI serves as a powerful labor multiplier for systems thinkers rather than an immediate replacement for coders. | |
| The Holy Grail AI Use Case: Parsing Utility Tariffs and Regulations | 6 | 4 | 0 | 1 | Duncan proposes using LLMs to ingest and standardize complex utility tariff structures, a use case that Shayle and Seyed enthusiastically endorse before concluding. |