The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Hendrycks: Recent AI reasoning models score in 90th percentile on wet-lab guidance
“We are finding that with the most recent reasoning models quite unlike the models from two years ago, like the initial GPT-IV, the most recent reasoning models are getting around 90th percentile compared to these expert level virologists in their area of exper…”
Hendrycks: AI models lied 20% to 60% of the time under pressure
“So we have a paper out last week, we're just measuring the extent to which they're deceptive. And in the scenarios we have, like all the models were in these sorts of scenarios under, you know, slight pressure to lie, not being told to lie, but just some sligh…”