OpenAI researcher Szymon Sidor explains how their reinforcement learning model spontaneously developed psychological baiting tactics against human Dota 2 players.
Opinion
Sidor: Rigorous baseline optimization will advance AI more than complex architectures
“And you know, it's not the kind of sexy research that people want to see, where you have like some hierarchy of big RNN, but it actually, this kind of research, I think at this point will advance field the most.”
Opinion
Sidor: Strong Engineering Is Much More Valuable at OpenAI than Writing Exotic Models
“So, so essentially becoming a good engineer for our team is much more valuable than, for example, people spending you know, months upon months implementing exotic models in TensorFlow.”
Insight
Sidor: Machine learning math is easier to learn than good engineering
“Getting good basics in linear algebra and in basic statistics, that's, especially when Doing experiments, it's easy to make like elementary statistics mistakes and linear algebra is just kind of most of what you need to know to like basic optimization as well …”
Opinion
Sidor: AI research neglects understanding existing methods and limits
“And what people do is people kind of try to invent problems and the, such as solving some complicated games of character structure, and they try to Add kind of extra features to their models to combat those problems, but I think there's very little research ha…”
Assertion Partly supported
Sidor: OpenAI's RL bot beat its hand-coded bot after one to two weeks
“So I leave, there is nothing, I come back, there is this reinforcement learning bot, and actually, it's beating our scripted bot after, like a week worth of engineering effort. Possibly it was two weeks, but it was something very miniature compared to the deve…”
Assertion Not publicly verifiable
Sidor: Pro Dota player reached 30% win rate against OpenAI bot
“I think it might be actually 30. And that player played hundreds of games with the bot.”