Sep 16, 2021 · 42m · we-live-to-build
Why Vague Data Searches Produce Vague Results — and Who Profits
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the We Live to Build podcast, host Sean Weisbrot interviews Nomad Data CEO Brad Schneider to explore how artificial intelligence and natural language processing revolutionize commercial data discovery. Their conversation delves into the alternative data landscape, the challenge of establishing ground truth amidst bias, and the critical need for safety guardrails in autonomous AI systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Sean holds 28.3% of the talking time here. How this is scored →
speaking balance: gold is Sean, purple is the guest (3 minute bins)
Brad directly rejects Sean's suggestion that an advanced AI could magically bypass system limits, noting that no matter how smart an AI is, it cannot defy the laws of physics.
Hardest push from Sean ▶ 40:16 Sean challenges Brad on autonomous self-rewriting capabilitiesSean refuses Brad's architectural safety defense and presses whether a sufficiently intelligent AI could discover its own flaws and rewrite its code regardless of permissions.
Biggest teaching moment ▶ 21:27 Brad educates Sean on ground truth limits and training biasBrad counters Sean's assumption that AI eliminates bias by explaining how machine learning models ingest human bias and fail without reliable ground truth verification.
Sean holds their own ▶ 22:44 Sean leverages personal China experience on data veracitySean demonstrates domain credibility by citing his decade living in China to assert that official governmental GDP metrics cannot be taken at face value.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Sean as informed peer | Guest teaching | Guest disagreement | Sean pushing back | Why |
|---|---|---|---|---|---|---|
| The Origins and Mission of Nomad Data | 3 | 5 | 1 | 1 | Sean asks standard introductory questions about Nomad Data's value proposition. Brad provides a clear breakdown of the fragmented alternative commercial data market and his entrepreneurial background. | |
| The Challenges of Navigating Commercial Data | 4 | 6 | 2 | 2 | Sean asks whether platforms like Facebook sell their data or keep it proprietary. Brad explains that tech giants hoard their data as a moat and describes why navigating thousands of external data vendors is like reading a massive diner menu without context. | |
| Blending Human Intelligence with AI for Data Discovery | 5 | 6 | 2 | 2 | Sean references search indexing within his own company to ask about search strategies across database columns. Brad reframes the problem, explaining that database columns describe what data is rather than the cross-disciplinary business use cases it can answer. | |
| The Limits of Human Recall and AI Matchmaking | 4 | 5 | 3 | 3 | Sean suggests machine learning is the only viable path to search commercial data, prompting Brad to nuance the claim by outlining filtering and tagging before explaining how human recall limits necessitate AI matchmaking. | |
| Sources of Alternative Commercial Data | 4 | 6 | 1 | 1 | Sean asks how vendors acquire data to sell commercially. Brad educates Sean on multiple data generation channels, including web scraping, business exhaust data, Freedom of Information Act filings, and municipal building permits. | |
| Digital Footprints, Aggregated Data, and Targeted Marketing | 4 | 6 | 2 | 2 | Sean asks about the efficacy of VPNs and privacy tools, sharing a targeted advertising anecdote. Brad explains that most commercial monetization relies on high-level data aggregation rather than deanonymized personal dossiers. | |
| User Data Monetization and Participant Selection Bias | 5 | 7 | 2 | 2 | Sean raises blockchain projects proposing to pay users for their data. Brad explains why user-paid panels suffer from severe selection bias, prompting Sean to connect this dynamic to psychological research flaws and historical scientific lobbying. | |
| The Ground Truth Challenge and Macro Data Verification | 6 | 7 | 3 | 3 | When Sean posits that AI can eliminate human bias, Brad corrects him by noting training data inherits human biases and introduces the concept of ground truth data. Sean draws on a decade of living in China to validate challenges in relying on official figures. | |
| Satellite Imagery and Unalterable Physical Data Sources | 5 | 6 | 2 | 3 | Sean asks if foreign governments can manipulate port and shipping figures. Brad explains how alternative data sources like transponder AIS beacons and satellite constellation imagery circumvent governmental manipulation. | |
| Big Tech Ecosystems and Data Privacy Restrictions | 4 | 6 | 2 | 2 | Sean asks if corporate data dominance will remain unchanged for decades. Brad shifts focus to privacy crackdowns by ecosystem owners like Apple, and demonstrates how high-frequency alternative data detected the COVID economic rebound before mainstream news. | |
| Data Quality Versus Volume and Ground Truth Limits | 5 | 7 | 3 | 2 | Sean asks if data volume will overwhelm AI architectures. Brad refutes this using compute economics, explaining that the real bottleneck is data quality and missing context rather than raw scale, citing genomics and protein folding. | |
| The Black Box Dilemma and Trusting Advanced AI | 5 | 6 | 4 | 4 | Sean questions how humans can trust machine learning models if they are opaque black boxes. Brad counters with a philosophical analogy, arguing humans routinely trust other people despite completely lacking visibility into their underlying neural firing. | |
| Humanoid Robots and AI Safety Guardrails | 5 | 6 | 5 | 5 | Sean raises concerns about autonomous systems taking over and AI rewriting its own code. Brad firmly grounds the debate, arguing architectural guardrails and physical impossibilities prevent runaway AI scenarios. | |
| The Future of Data Ecosystems and Episode Conclusion | 3 | 4 | 1 | 1 | Sean asks for closing thoughts and contact details. Brad summarizes the virtuous cycle between growing data availability and AI applications before Sean concludes the episode. |