InstructGPT, every mention
10 scenes · ← back to InstructGPT
tap a year for its mentions
every year anyone Nathan Lambert 4Jack Morris 3Thomas Scialom 2Will Brown 1Michael Royzen 1Jason Liu 1
Verbatim, from the transcripts: the passages where InstructGPT comes up
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 17:18 unnamed speaker The idea was, you, you started with codecs, someone else was doing instroft GPT, then we launched GPT, four, four O, I guess O-one.
Information Theory for Language Models: Jack Morris
- ▶ 3:06 Jack Morris Like around when I guess GPT three, one hundred and seventy five billion had been released, but not instruct GPT.
- ▶ 1:11:12 Jack Morris The InstructGPT techniques, you just need the data. 2 times in the scene
⚡️Open Questions in Agentic RL — Will Brown (Prime Intellect)
- ▶ 2:46 Will Brown Um, and so we've kind of moved beyond the single-term RL eject world of, like, InstructGPT, ejectGPT sort of things, as well as the
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 5:11 Thomas Scialom I actually was working on Galactica Instruct, which basically you could connect it, we had a partner with Overleaf, the Google Doc of, like, scientists, where you can write papers, and it's in, you write there in LaTeX, you have to do a… 2 times in the scene
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 25:15 Nathan Lambert It has this three-step process that we'll go to, into more when I kind of go into the main concepts, but it's like the first time you see this diagram that they reuse with InstructGPT, they reuse with ChatGPT, and the types of examples… 2 times in the scene
- ▶ 43:35 Nathan Lambert I threw in the, um, slide 28, it's like the InstructGPT metadata.
- ▶ 58:11 Nathan Lambert I think InstructGPT does something where they, like, try to get the RL model to match these, the instruction tuning model, or the instruction tuning dataset, because they were really happy with that dataset to constrain the distribution.
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 14:51 Michael Royzen Um, this is before InstructGPT.