Jan 15, 2015 · 21m · mad

Hanna Wallach, Microsoft Research // Data Driven #33 // Jan 2015 (Hosted by FirstMark Capital)

Hanna Wallach · 18m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this presentation at Data Driven NYC, researcher Hanna Wallach examines computational social science, emphasizing how interdisciplinary collaboration, question-driven research, and rigorous model evaluation are essential for addressing algorithmic bias and ethical challenges in big data.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1% of the talking time here. How this is scored →

Matt as informed peer 0.0 Guest teaching 0.0 Guest disagreement 1.2 Matt pushing back 0.0
05100:0010:0020:001:42–4:31 · Matt as informed peer 0/10 Defining Big Data and Social Granularity Hanna Wallach delivers a solo keynote presentation defining big data and highlighting why granular social data makes people uncomfortable. As this is a monologue presentation without host participation, host expertise and pushback are non-existent.4:31–7:00 · Matt as informed peer 0/10 Ethics, Bias, and Interdisciplinary Collaboration in CSS Wallach advocates for meaningful interdisciplinary collaboration between computer scientists and social scientists to tackle ethics, bias, and privacy. The segment is a solo lecture monologue with no host interaction.7:00–11:25 · Matt as informed peer 0/10 Computational Challenges in Analyzing Heterogeneous Data Wallach gently critiques data-first and method-first research approaches in computer science, urging research to start with social questions rather than convenience datasets. Host scores remain zero due to the monologue format.11:25–16:16 · Matt as informed peer 0/10 Machine Learning Models, Error Analysis, and Uncertainty Wallach explains machine learning model error analysis, emphasizing the necessity of representing uncertainty when modeling social data. The host does not interrupt or participate during this presentation section.16:16–19:21 · Matt as informed peer 0/10 Drawing Fair Findings and Improving Scientific Communication Wallach discusses implicit bias, relying on methodology over human intuition, and public understanding of data science to conclude her talk. The segment contains no host dialogue.1:42–4:31 · Guest teaching 0/10 Defining Big Data and Social Granularity Hanna Wallach delivers a solo keynote presentation defining big data and highlighting why granular social data makes people uncomfortable. As this is a monologue presentation without host participation, host expertise and pushback are non-existent.4:31–7:00 · Guest teaching 0/10 Ethics, Bias, and Interdisciplinary Collaboration in CSS Wallach advocates for meaningful interdisciplinary collaboration between computer scientists and social scientists to tackle ethics, bias, and privacy. The segment is a solo lecture monologue with no host interaction.7:00–11:25 · Guest teaching 0/10 Computational Challenges in Analyzing Heterogeneous Data Wallach gently critiques data-first and method-first research approaches in computer science, urging research to start with social questions rather than convenience datasets. Host scores remain zero due to the monologue format.11:25–16:16 · Guest teaching 0/10 Machine Learning Models, Error Analysis, and Uncertainty Wallach explains machine learning model error analysis, emphasizing the necessity of representing uncertainty when modeling social data. The host does not interrupt or participate during this presentation section.16:16–19:21 · Guest teaching 0/10 Drawing Fair Findings and Improving Scientific Communication Wallach discusses implicit bias, relying on methodology over human intuition, and public understanding of data science to conclude her talk. The segment contains no host dialogue.1:42–4:31 · Guest disagreement 1/10 Defining Big Data and Social Granularity Hanna Wallach delivers a solo keynote presentation defining big data and highlighting why granular social data makes people uncomfortable. As this is a monologue presentation without host participation, host expertise and pushback are non-existent.4:31–7:00 · Guest disagreement 1/10 Ethics, Bias, and Interdisciplinary Collaboration in CSS Wallach advocates for meaningful interdisciplinary collaboration between computer scientists and social scientists to tackle ethics, bias, and privacy. The segment is a solo lecture monologue with no host interaction.7:00–11:25 · Guest disagreement 2/10 Computational Challenges in Analyzing Heterogeneous Data Wallach gently critiques data-first and method-first research approaches in computer science, urging research to start with social questions rather than convenience datasets. Host scores remain zero due to the monologue format.11:25–16:16 · Guest disagreement 1/10 Machine Learning Models, Error Analysis, and Uncertainty Wallach explains machine learning model error analysis, emphasizing the necessity of representing uncertainty when modeling social data. The host does not interrupt or participate during this presentation section.16:16–19:21 · Guest disagreement 1/10 Drawing Fair Findings and Improving Scientific Communication Wallach discusses implicit bias, relying on methodology over human intuition, and public understanding of data science to conclude her talk. The segment contains no host dialogue.1:42–4:31 · Matt pushing back 0/10 Defining Big Data and Social Granularity Hanna Wallach delivers a solo keynote presentation defining big data and highlighting why granular social data makes people uncomfortable. As this is a monologue presentation without host participation, host expertise and pushback are non-existent.4:31–7:00 · Matt pushing back 0/10 Ethics, Bias, and Interdisciplinary Collaboration in CSS Wallach advocates for meaningful interdisciplinary collaboration between computer scientists and social scientists to tackle ethics, bias, and privacy. The segment is a solo lecture monologue with no host interaction.7:00–11:25 · Matt pushing back 0/10 Computational Challenges in Analyzing Heterogeneous Data Wallach gently critiques data-first and method-first research approaches in computer science, urging research to start with social questions rather than convenience datasets. Host scores remain zero due to the monologue format.11:25–16:16 · Matt pushing back 0/10 Machine Learning Models, Error Analysis, and Uncertainty Wallach explains machine learning model error analysis, emphasizing the necessity of representing uncertainty when modeling social data. The host does not interrupt or participate during this presentation section.16:16–19:21 · Matt pushing back 0/10 Drawing Fair Findings and Improving Scientific Communication Wallach discusses implicit bias, relying on methodology over human intuition, and public understanding of data science to conclude her talk. The segment contains no host dialogue.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 6.4% · guest 93.6%18:00 · Matt 6.4% · guest 93.6%21:00 · Matt 6.1% · guest 93.9%21:00 · Matt 6.1% · guest 93.9%
Sharpest disagreement ▶ 8:50 Critique of data-first and tool-first approaches

Wallach offers her strongest intellectual pushback against mainstream computer science norms, criticizing researchers who create tools looking for nails or pick convenience datasets without framing meaningful social questions.

Hardest push from Matt ▶ 19:27 Host time management wrap-up

Because the presentation was a uninterrupted monologue, the only host intervention was Matt Turck stepping in at the end to limit Q&A to one quick question due to time constraints.

Biggest teaching moment ▶ 16:50 Refuting human intuition in social analysis

Wallach explicitly refutes the idea that computer scientists can rely on intuition about human behavior, educating the audience on unconscious implicit bias and the need for rigorous social science methods.

Matt holds his own ▶ 19:27 Host post-talk feedback

Matt Turck reassumes control of the session, contextualizing the presentation as a unique, highly substantial talk compared to typical data science pitches.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Big Data and Social Granularity 0010 Hanna Wallach delivers a solo keynote presentation defining big data and highlighting why granular social data makes people uncomfortable. As this is a monologue presentation without host participation, host expertise and pushback are non-existent.
Ethics, Bias, and Interdisciplinary Collaboration in CSS 0010 Wallach advocates for meaningful interdisciplinary collaboration between computer scientists and social scientists to tackle ethics, bias, and privacy. The segment is a solo lecture monologue with no host interaction.
Computational Challenges in Analyzing Heterogeneous Data 0020 Wallach gently critiques data-first and method-first research approaches in computer science, urging research to start with social questions rather than convenience datasets. Host scores remain zero due to the monologue format.
Machine Learning Models, Error Analysis, and Uncertainty 0010 Wallach explains machine learning model error analysis, emphasizing the necessity of representing uncertainty when modeling social data. The host does not interrupt or participate during this presentation section.
Drawing Fair Findings and Improving Scientific Communication 0010 Wallach discusses implicit bias, relying on methodology over human intuition, and public understanding of data science to conclude her talk. The segment contains no host dialogue.

Statements from this episode (13)

Insight
Wallach: Big data sets differ from physics data by documenting human behavior
“Unlike the data sets arising in physics, the data sets that typically fall under the big data umbrella are about people, Their attributes, their actions, and their interactions. That is to say these are social data sets that document people's lives, behaviors …”
Hanna Wallach Jan 15, 2015 ▶ 3:44
Insight
Wallach: Big data's defining feature is granularity, not size
“The issue is granularity. In other words, not only do these data sets document social phenomena, they do so at the granularity of individual people and their second-to-second activities.”
Hanna Wallach Jan 15, 2015 ▶ 4:18
Assertion Contradicted
Wallach: Research shows diverse teams best facilitate rapid breakthrough innovation
“In fact, there's even substantial research in social psychology and sociology indicating that rapid breakthrough innovations are best facilitated by bringing together people with really diverse backgrounds.”
Hanna Wallach Jan 15, 2015 ▶ 6:13
Insight
Wallach: Tech and government players must hire social scientists to address bias
“So as a result, if technology companies and government organizations, the biggest players in the big data game, Are going to take issues like bias and fairness and inclusion seriously. They need to hire social scientists, the people with the best training and …”
Hanna Wallach Jan 15, 2015 ▶ 6:27
Insight
Wallach: Addressing algorithmic bias requires focusing on big data's granular nature
“When it comes to addressing issues like bias, fairness, and inclusion, perhaps we should instead be focusing our attention on the granular nature of big data.”
Hanna Wallach Jan 15, 2015 ▶ 7:20
Prediction Not checkable as stated
Wallach: The big data game will be won by asking interesting questions
“Ultimately, the big data game will be won by those who know how to ask and answer interesting questions.”
Hanna Wallach Jan 15, 2015 ▶ 9:23
Insight
Wallach: Data-first research approaches amplify issues with algorithmic bias and fairness
“Although this is kind of conducive to fast-paced work, these data-first or method-first approaches can actually amplify issues related to bias, fairness, and inclusion of minorities.”
Hanna Wallach Jan 15, 2015 ▶ 10:25
Insight
Wallach: Convenience datasets bias analytical models toward demographic majorities
“But the problem with these convenience data sets is that they typically reflect only a particular segment of society. For example, young people with smartphones. And so, as a result, many of the methods developed to analyze these data sets end up prioritizing …”
Hanna Wallach Jan 15, 2015 ▶ 11:04
Assertion Supported
Wallach: ML fairness research overwhelmingly focuses on predictive over exploratory models
“So much of the existing work on fairness and transparency in machine learning, not that there's very much of it, focuses on predictive models rather than models for exploratory or explanatory analyses”
Hanna Wallach Jan 15, 2015 ▶ 12:22
Insight
Wallach: Bias in exploratory models distorts subsequent predictive models
“As a result, the models used to perform these previous analyses and any kind of bias or unfairness in them will necessarily influence the resultant findings, and hence the representations that we then choose to use in our predictive models.”
Hanna Wallach Jan 15, 2015 ▶ 13:12
Insight
Wallach: Predictive social data models must maintain and report uncertainty
“When building and using predictive models for social data, whether resultant decisions can actually affect real-world people, representing, maintaining, and reporting uncertainty, along with any subsequent decisions, should be standard practice.”
Hanna Wallach Jan 15, 2015 ▶ 15:45
Prediction Not checkable as stated
Hanna Wallach: Most machine learning methods will be deployed without creator involvement
“It's really likely that most machine learning methods or data science methods will at some point be used by some people, end users, without the involvement of those people who created them.”
Hanna Wallach Jan 15, 2015 ▶ 18:01
Insight
Wallach: Investigating correct model predictions helps contextualize how models treat certainty
“I don't know if there are necessarily any sort of general purpose, like this is going to fix everything kind of solutions, but I would say yes, digging into why your model is making certain predictions, even when those predictions are correct, can kind of help…”
Hanna Wallach Jan 15, 2015 ▶ 21:06
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.