Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered produced feed
D 5 · C 5 · P 5 · Cm 5 5.00
Q And you're talking about the Good Judgment Project. Can you maybe introduce us to that a little?
A Sure. Well, the Good Judgment Project is a research program that my wife, Barbara Mellors, and I Uh, started, uh, several years ago. Uh, it was supported by a branch, research and development branch of the U.S. intelligence community, known as IARPA, Intelligence Advanced Research Projects Activity, which models itself after DARPA in the Defense Department, and their mandate is to support research that has the potential to revolutionize intelligence analysis. So working from that mandate, they decided in 2010 to support, uh, a series of forecasting tournaments in which Major universities would, um, uh, compete. Researchers at major universities would compete to, uh, generate accurate probability estimates of possible futures of national security, uh, relevance. And, uh, we were one of the five teams selected for the competition in 2010. The tournaments ran from 2011 to 2015. They ended in June, uh, of this year. And, uh, the Good Judgment Project, uh, I am proud to say, was the winner of those forecasting tournaments. Uh, and I can explain more about what winning a forecasting tournament means later if you want.
AI assessment note: “the Good Judgment Project is a research program that my wife, Barbara Mellors, and I”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q and your ability to predict is kind of, uh, keeping score. And do you think it takes a certain type of person to want to keep score? I mean, most of us are happy to kind of weasel out of or use, uh, uncertain wording or jargon when we're going about making decisions so that even if we're wrong, we can kind of say, well, that's not what I meant.
A Absolutely. It does take a particular type of person, and, uh, there are many factors that come into play. I, I think it certainly helps to be open-minded, um, but there are other things that come into play that are a little more, say, sociological. I've been doing forecasting tournaments for over 30 years now, um, and I started when I was about 30 in 1984. I'm 61 years old now. So I'm, if I were an intelligence analyst, a sixty-one-year-old intelligence analyst, I would be a very senior analyst. Um, and let's just say for sake of argument that, that I, uh, I am a senior analyst in, in, in the US, uh, intelligence community. I'm on the National Intelligence Council, say, just for sake of argument, and I, I'm the go-to guy on China. So when Xi Jinping comes into town, people say to me, you know, what's, what's going on. I have inputs into the presidential daily briefing and help with national intelligence estimates. And I'm at the top of the status pecking order within the IC on China. And someone comes along like IARPA is this, this upstart research and development, uh, branch of the, for the office of director of national intelligence. And they say, Hey, you know what we're going to do? We want to run forecasting tournaments now. And, um, Everyone's going to compete on a level playing field, ah, and 25 year old China analysts are going to compete against, ah, 61 year old analy…
AI assessment note: “Absolutely. It does take a particular type of person”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q And what did that consist of, this 50 minute training exercise?
A Some basic ideas about, um, uh, heuristics and biases and how to check biases. Uh, for example, one of the classic Kahneman arguments is that, um, people don't give enough weight to statistical or base rate information, uh, in assessing the probabilities of events. Um, they, um, they're, they're, they're too quick to take the inside view. So if you're attending a wedding and you see the happy couple and you, you're impressed by how much in love they are and the enthusiasm of the moment, and someone asks you how likely are they to get divorced? Uh, you're not likely to consult national divorce statistics for that SES subgroup. Uh, you're likely to say, hmm, they look really happy and compatible. I'm going to touch a very high probability to they're not getting divorced. Um, and the net result of making predictions in that way is that you're not going to be somewhat inaccurate, but less accurate than you would have been if you had at least started your estimation process by saying, what are the base rates of divorce? And now I'm going to adjust that based on whatever idiosyncratic factors are present in this particular relationship.
AI assessment note: “Some basic ideas about, um, uh, heuristics and biases and how to check biases.”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q And what was so interesting about the way that Fermi approached it?
A He really believed in flushing out your ignorance and decomposing, decomposing the problem into as many tractable components as possible. So you would start by, you know, how, how many, how many stars are there in the Milky Way? Roughly about a hundred billion Uh, you'd say, well, how many, um, of these stars have planets orbiting around them? You might look at the most recent data from Kepler, which has done some reconnaissance in our local area, but about 60 light years around, and, um, you say, well, you know, it looks like a fair number, a pretty high percentage of stars do seem to have planets going around them. Uh, let's say it could be as much as half, um, or maybe slightly less. I, I, I, but you, but you, I don't really know the answer to that question, but you make, you make, you make initial guesses. You flush out your ignorance, and then other people can come back, and they can see that Tetlock said about half, and they say, oh, Tetlock doesn't understand what Kepler's doing. It should have been 70%. No, it should have been 30%. Uh, but what we, it's not that Tetlock is getting it right. It's that we're flushing out Tetlock's zone of ignorance, and we're making it clear, and, uh, it's all open and transparent. Uh, and then we would, you know, in that process of inquiry, we would continue how many planets are in a habitable zone, and you can direct some further guesst…
AI assessment note: “flushing out your ignorance and decomposing, decomposing the problem into as many tractable components”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q And so, were you using, ah, a representative subset of the Good Judgment Project, or were you using superforecasters from the project, or how were you competing in that?
A Well, different universities and different teams of researchers took different approaches to generating accurate probability estimates. We, ah, recruited thousands of forecasters, and Ah, we explored a number of different techniques for eliciting the best possible probability estimates from those forecasters. We are continually running experiments, and one of the experiments we conducted was, ah, to identify top performers in each year, ah, the top two percent of performers each year, ah, cream them off into, ah, teams, elite teams with super teams of super forecasters, And, uh, give them as much support as we could, intellectual, uh, support as we could, uh, for their task. And see how, see, see what, what would happen. And they really went to town. They, they did a phenomenally good job. They, they, they blew the ceiling off all of the performance expectations that are up ahead, ah, for what was possible. And frankly, they, they certainly exceeded my expectations as well.
AI assessment note: “one of the experiments we conducted was, ah, to identify top performers”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q What would you say to people inside an organization? How can they use your research to make better decisions inside their company?
A Well, I think it's, it's something you want to consider seriously, uh, that when, when people make forecasts inside organizations, most organizations today, accuracy is only one of the goals that they're pursuing. Uh, they're also interested in making forecasts that are going to be difficult to falsify. I said they can't be embarrassed. So a lot of the forecasting inside organizations doesn't involve numbers. It involves a lot of vague verbiage. Um, they're also interested in making forecasts that don't annoy other people in the organization. They don't want to tip the political apple card over. Uh, so They're compromising accuracy in a whole host of ways that, um, help promote their careers inside the organization, help to maintain political stability in the organization, uh, but that aren't all that centrally focused on accuracy. Forecasting tournaments are really weird, uh, because they focus 100% on accuracy. That's all that matters. So I guess the thing you'd want to consider as an executive would be, do I want to reserve part of my organization's, uh, Analytical processing capacity for a pure accuracy game. So I want to incentivize some small group of the people in my organization to play pure accuracy games in forecasting tournaments, and those probability estimates would then, uh, filter up to senior executives to guide decision making. I think it's a, it's really an in…
AI assessment note: “incentivize some small group of the people in my organization to play pure accuracy games”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q You mentioned open-mindedness at the beginning. How do we go about fostering open-mindedness? Are there ways that we can improve that in ourselves or other people?
A Well, we try to, that's another thing we do try to emphasize in the training. Um, exerting people simply to be open-minded Is, um, most people don't think they're close-minded. Um, most people think, most, most people think they're quite reasonable. And, uh, simply exhorting people to be open-minded, people struggle and say, well, yeah, I already am. I think you want to start in a more specific way. So you want to start with very specific problems in which you assess whether people change their minds in an appropriate way. So there are some normative models like Bayes theorem that tell you how much you should change your mind in response to evidence that has certain diagnostic value, and you can create simulated problems. It might be medical diagnosis problems. It might be economic problems. It might be military problems, but you can create simulated problems with simulated data, and you can see whether people learn through practice to update their beliefs the way they should. Uh, now there's always a question whether they're gonna, those lessons are going to stick, and, um, you know, we, we found that they do stick a little bit because they can produce 10% improvement throughout the year, but it's, it's one of the great challenges. I don't think we've solved the problem of how to make people more open-minded. I, I think we can make people, uh, better belief updaters on problem…
AI assessment note: “you want to start with very specific problems in which you assess whether people change”
Answered produced feed
D 4 · C 5 · P 4 · Cm 4 4.30
Q Do you think the super forecasters were better at learning from the other super forecasters than the, say, average forecaster? Like, if somebody had a better approach, would they copy it? Would they just drop their own internal approach and
A I think they listen to each other quite carefully in the super forecaster teams. Um, and even when they disagree with each other, they, they, they, they, they disagree diplomatically, but they can disagree quite forcefully, uh, about what lessons they should draw from particular forecasting failures, uh, or even forecasting successes. I mean, it's, it's fairly common for regular forecasters even to say, well, what, what did we do wrong with the forecasting failure? Um, And supers do that too. Um, but they also second guess their successes. Um, they say, well, were we lucky just where we, we, we, we, we really nailed this question, but were we, um, lucky? Uh, could it, could it have gone otherwise? Were we almost wrong? Um, That's an unusual question for people to ask themselves. Um, people don't normally look a gift, gift horse in the mouth. And when they're right, they, they, they want to take credit for it. Uh, and super forecaster skepticism even extends to their, um, their forecasting successes.
AI assessment note: “I think they listen to each other quite carefully in the super forecaster teams.”
Partly produced feed
D 3 · C 5 · P 4 · Cm 4 4.00
Q So some of us are good and some of us are bad and some of us seem like way off the chart at making predictions. Why are some people so good?
A That is indeed the 64,000 dollar question. Why are some people so good? So the skeptics argue that, that if you toss enough coins enough times, some of them are bound to come up heads. So the super forecasters are just super lucky. So let's treat that as kind of the default skeptical hypothesis. There's nothing special about super forecasters. Uh, if we ran a tournament in which the task was a to predict, uh, whether a fair coin would land heads or tails, some, some, some people would do better than others just by chance in a given year. Uh, we could anoint those people as super coin toss predictors, and we could say, well, how are they going to do the next year? And what we would find is perfect regression toward the mean. The best prediction is that the super coin toss predictors in year one will be, uh, essentially around the average in year two. And the worst predictors will progress upward toward the mean, of course. Uh, so that's what pure, what's what a pure chance environment would look like. Uh, well, what we find in the ARPA tournament is, is that there certainly is an element of chance in predicting geopolitical and geoeconomic outcomes, uh, but the skill luck ratio seems to be about 70 30. Uh, so you're not observing a great deal of regression toward the mean among super forecasters, but there, there inevitably is some regression toward the mean among the top perfor…
AI assessment note: “the skill luck ratio seems to be about 70 30”
Partly produced feed
D 3 · C 5 · P 4 · Cm 4 4.00
Q So if you were an organization, you wanted to set up a team environment, like a forecasting team within a large company, say IBM, how would you go about doing that with your knowledge?
A That's a great question. And I'm a little bit wary about saying that, uh, organizations should try to construct super teams the way the Good Judgment Project did. Um, because, uh, team construction has a lot of implications for other parts of the organization. That can be tricky. I mean, imagine that if you, if you just did what we did in the ARFA tournament to win it, and you just identified the very best people and brought them together and nurtured them and helped them and pushed them hard. That would be a very elitist and somewhat divisive thing to do in many organizations. Um, and it could cause a lot of political friction. Now we didn't care a lot about that because we were in a forecasting tournament. We didn't really have an organization in the traditional sense of the term. We wanted performance engine.
AI assessment note: “I'm a little bit wary about saying that, uh, organizations should try to construct super teams”
Redirected produced feed
D 2 · C 4 · P 4 · Cm 3 3.25
Q And do you think, what transfers from your research into the decision making process in a corporation, not necessarily about forecasting, but about how we go about organizing, unpacking, synthesizing, multiple views, how does that transfer, do you think, into a learnable skill that people can have inside of an organization?
A There are many ways that could happen. We put a lot of emphasis in the Good Judgment Project on synthesizing diverse views into aggregate forecasts, and I think one of our major performance engines was the statistical or aggregation algorithms that our statisticians developed for doing that. When IARPA started this whole exercise, they thought it would be really hard to do better than 20 or 30 or 40% better than the unweighted average of the group of control group forecasters. And our super forecasters exceeded that performance benchmark quite substantially each year of the tournament. Um, they did so well that IARP essentially suspended the tournament after two years, and, um, our, our, our, and they, we were able to absorb the other teams into our team in substantial ways and compete against the intelligence community and against the prediction market baselines, um, instead of the other universities. Now, how did all that Come to pass. I think the aggregation algorithm developed, if I had to credit two big things as responsible for the victory of the Good Judgment Project, one of them would be the super forecasters, and the other would be, they call them super algorithms, the great algorithms that our statisticians develop. Now, when I, when I describe these algorithms, at some level, you're not going to be too surprised at first, but there is one aspect of them that does sur…
AI assessment note: “synthesizing diverse views into aggregate forecasts, and I think one of our major performance”