Nima Alidoust, CEO of Vevo Therapeutics, benchmarks current AI biology milestones against the progression of OpenAI's GPT language model series.
Insight
Alidoust: Hypothesis-driven biology slowed progress, but falling costs enable unbiased data
“One thing that has been a Has been slowing the progress in bio is the fact that we have always been super hypothesis driven, and I think it has, the reason is that a lot of these experiments are expensive, you know, that they take a lot of time, a lot of resou…”
Assertion Not checkable as stated
Alidoust: Single-cell AI performance holds even after 99% data downsampling
“If you actually reduce the number of the sixty million, you down sample it by like, Even 99%. You know, you just use one percent of that data to train your models. Actually, the model's performance doesn't reduce that much. So it means that the information con…”
Assertion Partly supported
Alidoust: Tahoe dataset expands public perturbational single-cell data fiftyfold
“I think when you put all of the perturbational data sets in the world together if you're generous, it's like one to two million single cell data points. And this is publicly available data. We don't know as much about, you know, what, what's inside different o…”
Prediction Not checkable as stated
Alidoust: Virtual cells will expand AI drug discovery to systems biology
“Virtual cells, in my opinion, are going to allow us to go beyond the language of structural biology and venture into the language of systems biology and understand how how the drug is interacting with the broader biological system.”
Assertion Supported
Alidoust: Prior public single-cell data totaled only 45M to 60M cells
“Before that, I think the number of human cells that we had, had been collated together it was in the order of 45, fifty million, if you are generous, sixty million single cell data points.”
Insight
Alidoust: 100 million single-cell data points equate to 200-300 billion tokens
“Think of it like a cell collection of for this data says 2000 to 5000 genes, and each gene and its expression is basically a token in what we're doing. So 200, like a hundred million single cell data points is akin to around 200 to three hundred billion tokens…”