Oct 16, 2014 · 25m · mad

Nick Sinai, White House Deputy CTO // Data Driven #30 // Oct 2014 (Hosted by FirstMark Capital)

Nick Sinai · 19m spoken Matt Turck · 2m spoken Ted Angelis · 53s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Former US Deputy CTO Nick Sinai discusses the White House's initiatives to unleash government data, build open-source public tech infrastructure, and foster developer ecosystems to drive economic growth and civic innovation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 10% of the talking time here. How this is scored →

Matt as informed peer 2.6 Guest teaching 3.2 Guest disagreement 0.8 Matt pushing back 0.8
05100:0010:0020:001:16–7:16 · Matt as informed peer 2/10 Distinguishing the Roles of US CTO and CIO Host Matt Turck opens with introductory questions asking Sinai to clarify the distinction between the US CTO and CIO roles. Sinai gives an extended, highly collaborative overview of White House tech policy and open data history without any dynamic friction.7:16–9:42 · Matt as informed peer 4/10 The DATA Act of 2014 and Modernizing Data.gov Turck demonstrates preparation by citing that data.gov currently hosts 156,000 datasets. Sinai gently corrects the host's reliance on dataset counts, schooling him on how bureaucratic incentives lead to chopped-up CSV files rather than meaningful data quality.9:42–12:23 · Matt as informed peer 2/10 Building Open Data Ecosystems Through Community Engagement Turck prompts Sinai to discuss community engagement tactics like Datapaloozas and hackathons. Sinai outlines the playbook pioneered by Todd Park in a friendly, conversational dialogue.12:23–16:26 · Matt as informed peer 4/10 Integrating Federal Open Data into Commercial and Consumer Platforms Turck asks why consumer apps like Yelp don't feature sanitation data and pushes back when Sinai mentions Yelp setting health grade standards, asking if Yelp is forcing everyone else to conform to its proprietary standards. Sinai clarifies how open standards versus scale work.16:26–22:09 · Matt as informed peer 1/10 Addressing Bureaucratic Friction and Accelerating Data Release Turck asks a general question about administrative friction and then opens the floor to audience Q&A. An audience member challenges Sinai on the lack of useful financial data on data.gov, which Sinai diplomatically addresses by highlighting new Chief Data Officer roles and the IRS GetTranscript tool.1:16–7:16 · Guest teaching 3/10 Distinguishing the Roles of US CTO and CIO Host Matt Turck opens with introductory questions asking Sinai to clarify the distinction between the US CTO and CIO roles. Sinai gives an extended, highly collaborative overview of White House tech policy and open data history without any dynamic friction.7:16–9:42 · Guest teaching 5/10 The DATA Act of 2014 and Modernizing Data.gov Turck demonstrates preparation by citing that data.gov currently hosts 156,000 datasets. Sinai gently corrects the host's reliance on dataset counts, schooling him on how bureaucratic incentives lead to chopped-up CSV files rather than meaningful data quality.9:42–12:23 · Guest teaching 2/10 Building Open Data Ecosystems Through Community Engagement Turck prompts Sinai to discuss community engagement tactics like Datapaloozas and hackathons. Sinai outlines the playbook pioneered by Todd Park in a friendly, conversational dialogue.12:23–16:26 · Guest teaching 3/10 Integrating Federal Open Data into Commercial and Consumer Platforms Turck asks why consumer apps like Yelp don't feature sanitation data and pushes back when Sinai mentions Yelp setting health grade standards, asking if Yelp is forcing everyone else to conform to its proprietary standards. Sinai clarifies how open standards versus scale work.16:26–22:09 · Guest teaching 3/10 Addressing Bureaucratic Friction and Accelerating Data Release Turck asks a general question about administrative friction and then opens the floor to audience Q&A. An audience member challenges Sinai on the lack of useful financial data on data.gov, which Sinai diplomatically addresses by highlighting new Chief Data Officer roles and the IRS GetTranscript tool.1:16–7:16 · Guest disagreement 0/10 Distinguishing the Roles of US CTO and CIO Host Matt Turck opens with introductory questions asking Sinai to clarify the distinction between the US CTO and CIO roles. Sinai gives an extended, highly collaborative overview of White House tech policy and open data history without any dynamic friction.7:16–9:42 · Guest disagreement 2/10 The DATA Act of 2014 and Modernizing Data.gov Turck demonstrates preparation by citing that data.gov currently hosts 156,000 datasets. Sinai gently corrects the host's reliance on dataset counts, schooling him on how bureaucratic incentives lead to chopped-up CSV files rather than meaningful data quality.9:42–12:23 · Guest disagreement 0/10 Building Open Data Ecosystems Through Community Engagement Turck prompts Sinai to discuss community engagement tactics like Datapaloozas and hackathons. Sinai outlines the playbook pioneered by Todd Park in a friendly, conversational dialogue.12:23–16:26 · Guest disagreement 1/10 Integrating Federal Open Data into Commercial and Consumer Platforms Turck asks why consumer apps like Yelp don't feature sanitation data and pushes back when Sinai mentions Yelp setting health grade standards, asking if Yelp is forcing everyone else to conform to its proprietary standards. Sinai clarifies how open standards versus scale work.16:26–22:09 · Guest disagreement 1/10 Addressing Bureaucratic Friction and Accelerating Data Release Turck asks a general question about administrative friction and then opens the floor to audience Q&A. An audience member challenges Sinai on the lack of useful financial data on data.gov, which Sinai diplomatically addresses by highlighting new Chief Data Officer roles and the IRS GetTranscript tool.1:16–7:16 · Matt pushing back 0/10 Distinguishing the Roles of US CTO and CIO Host Matt Turck opens with introductory questions asking Sinai to clarify the distinction between the US CTO and CIO roles. Sinai gives an extended, highly collaborative overview of White House tech policy and open data history without any dynamic friction.7:16–9:42 · Matt pushing back 1/10 The DATA Act of 2014 and Modernizing Data.gov Turck demonstrates preparation by citing that data.gov currently hosts 156,000 datasets. Sinai gently corrects the host's reliance on dataset counts, schooling him on how bureaucratic incentives lead to chopped-up CSV files rather than meaningful data quality.9:42–12:23 · Matt pushing back 0/10 Building Open Data Ecosystems Through Community Engagement Turck prompts Sinai to discuss community engagement tactics like Datapaloozas and hackathons. Sinai outlines the playbook pioneered by Todd Park in a friendly, conversational dialogue.12:23–16:26 · Matt pushing back 3/10 Integrating Federal Open Data into Commercial and Consumer Platforms Turck asks why consumer apps like Yelp don't feature sanitation data and pushes back when Sinai mentions Yelp setting health grade standards, asking if Yelp is forcing everyone else to conform to its proprietary standards. Sinai clarifies how open standards versus scale work.16:26–22:09 · Matt pushing back 0/10 Addressing Bureaucratic Friction and Accelerating Data Release Turck asks a general question about administrative friction and then opens the floor to audience Q&A. An audience member challenges Sinai on the lack of useful financial data on data.gov, which Sinai diplomatically addresses by highlighting new Chief Data Officer roles and the IRS GetTranscript tool.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 32.9% · guest 67.1%0:00 · Matt 32.9% · guest 67.1%3:00 · Matt 3.9% · guest 96.1%3:00 · Matt 3.9% · guest 96.1%6:00 · Matt 5% · guest 95%6:00 · Matt 5% · guest 95%9:00 · Matt 6.8% · guest 93.2%9:00 · Matt 6.8% · guest 93.2%12:00 · Matt 18.8% · guest 81.2%12:00 · Matt 18.8% · guest 81.2%15:00 · Matt 15.7% · guest 84.3%15:00 · Matt 15.7% · guest 84.3%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 0.6% · guest 99.4%21:00 · Matt 0.6% · guest 99.4%24:00 · Matt 3.6% · guest 96.4%24:00 · Matt 3.6% · guest 96.4%
Sharpest disagreement ▶ 8:37 Reframing dataset volume metrics

Sinai directly counters Turck's cited figure of 156,000 datasets, calling bureaucratic dataset chopping ridiculous and cautioning against using raw counts as a measure of success.

Hardest push from Matt ▶ 13:54 Challenging proprietary standard creation

Turck interrupts Sinai's explanation of Yelp to press him on whether Yelp is unilaterally dictating standards that all other market participants are forced to conform to.

Biggest teaching moment ▶ 8:37 Bureaucratic gaming of open data metrics

Sinai educates Turck on how quoting dataset numbers creates perverse incentives in federal agencies to split single data files into thousands of small CSVs.

Matt holds his own ▶ 8:31 Citing specific data.gov metrics

Turck shows active research and background knowledge by citing the precise current figure of 156,000 datasets hosted on data.gov.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Distinguishing the Roles of US CTO and CIO 2300 Host Matt Turck opens with introductory questions asking Sinai to clarify the distinction between the US CTO and CIO roles. Sinai gives an extended, highly collaborative overview of White House tech policy and open data history without any dynamic friction.
The DATA Act of 2014 and Modernizing Data.gov 4521 Turck demonstrates preparation by citing that data.gov currently hosts 156,000 datasets. Sinai gently corrects the host's reliance on dataset counts, schooling him on how bureaucratic incentives lead to chopped-up CSV files rather than meaningful data quality.
Building Open Data Ecosystems Through Community Engagement 2200 Turck prompts Sinai to discuss community engagement tactics like Datapaloozas and hackathons. Sinai outlines the playbook pioneered by Todd Park in a friendly, conversational dialogue.
Integrating Federal Open Data into Commercial and Consumer Platforms 4313 Turck asks why consumer apps like Yelp don't feature sanitation data and pushes back when Sinai mentions Yelp setting health grade standards, asking if Yelp is forcing everyone else to conform to its proprietary standards. Sinai clarifies how open standards versus scale work.
Addressing Bureaucratic Friction and Accelerating Data Release 1310 Turck asks a general question about administrative friction and then opens the floor to audience Q&A. An audience member challenges Sinai on the lack of useful financial data on data.gov, which Sinai diplomatically addresses by highlighting new Chief Data Officer roles and the IRS GetTranscript tool.

Statements from this episode (11)

Assertion Supported
Nick Sinai was a venture capitalist at Lehman Brothers during its collapse
“So I was in, in the business world for about 10 years, and I was a venture capitalist, ah, for about five years after business school, and I was actually on my honeymoon, and I opened up the newspaper, and it says Lehman Brothers goes bust. Ah, the challenge w…”
Nick Sinai Oct 16, 2014 ▶ 0:30
Assertion Supported
The U.S. federal government spends about $80 billion annually on IT
“We spend about eighty billion dollars in federal IT”
Nick Sinai Oct 16, 2014 ▶ 1:50
Assertion Supported
Barack Obama signed a 2013 executive order prioritizing open, machine-readable data
“In 2013, the president signed, ah an executive order making open and machine readable the new default for government information.”
Nick Sinai Oct 16, 2014 ▶ 3:50
Assertion Partly supported
The U.S. federal government spends $140 billion annually on research and development
“We spend a hundred and forty billion dollars a year in research and development.”
Nick Sinai Oct 16, 2014 ▶ 4:33
Assertion Supported
Data.gov is open-source on GitHub, built with CKAN and WordPress
“So we moved it to an open source project. So it's now CCAN and WordPress. It's on GitHub, so you can actually log an issue or put your, put some feedback. We're pushing code, I think every four weeks or so.”
Nick Sinai Oct 16, 2014 ▶ 8:06
Insight
Tracking dataset quantity creates perverse incentives for government open data
“If you're just counting and you tell the bureaucracy that, that the numbers are important, and even if you don't explicitly say that, but implicitly, if you have folks in the minute, senior officials that keep using higher and higher numbers, it encourages peo…”
Nick Sinai Oct 16, 2014 ▶ 8:42
Assertion Supported
Google and Bing source their drug knowledge panels from FDA and HHS
“If you use a major search platform Google, or Bing, or others, and I guess a major search platform and you put in a drug name you'll see in the knowledge box, or knowledge panel, or something like that, you'll see a bunch of information, and sourced will be th…”
Nick Sinai Oct 16, 2014 ▶ 13:03
Assertion Supported
NOAA collects 20 terabytes of weather data daily but releases only 10%
“They still have 20 terabytes of data that they collect and create a day around weather, and they're only making 10% of that available.”
Nick Sinai Oct 16, 2014 ▶ 16:06
Assertion Supported
Federal agencies like Commerce, Energy, and Transportation are appointing Chief Data Officers
“Department of Commerce is hiring for a Chief Data Officer. Department of Energy is hiring for Chief Data Officer. Department of Transportation just hired a Chief Data Officer. So you see a number of Cabinet-level agencies that are hiring them.”
Nick Sinai Oct 16, 2014 ▶ 19:30
Assertion Supported
The unmarketed IRS GetTranscript tool increased annual digital requests to 12 million
“The IRS launched this thing in 2014 with really no marketing called GetTranscript, and they went from three million, kind of, paper, ah, requests served a year. Now there are over twelve million, ah, ah, requests served a year”
Nick Sinai Oct 16, 2014 ▶ 21:07
Disclosure
The White House asks agencies to classify data into three access tiers
“One of the practical ways that we're asking agencies to deal with this is to really think about data in three, three classes. There's stuff that is public, there's stuff that is restricted public, and there's stuff that is non-public”
Nick Sinai Oct 16, 2014 ▶ 24:07
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.