Is RL + LLMs enough for AGI? β Sholto Douglas & Trenton Bricken
Dwarkesh Patel
View ChannelAbout
Deeply researched interviews
Latest Posts
Video Description
New episode with my good friends Sholto Douglas & Trenton Bricken. Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic. We talk through whatβs changed in the last year of AI research; the new RL regime and how far it can scale; how to trace a modelβs thoughts; and how countries, workers, and students should prepare for AGI. See you next year for v3. Enjoy! πππππππ πππππ * Transcript: https://www.dwarkesh.com/p/sholto-trenton-2 * Apple Podcasts: https://podcasts.apple.com/us/podcast/dwarkesh-podcast/id1516093381 * Spotify: https://open.spotify.com/episode/3H46XEWBlUeTY1c1mHolqh?si=b645971b1af546fa * Last year's episode: https://www.youtube.com/watch?v=UTuuTTnjxMQ ππππππππ * WorkOS ensures that AI companies like OpenAI and Anthropic don't have to spend engineering time building enterprise features like access controls or SSO. Itβs not that they don't need these features; it's just that WorkOS gives them battle-tested APIs that they can use for auth, provisioning, and more. Start building today at https://workos.com. * Scale is building the infrastructure for safer, smarter AI. Scaleβs Data Foundry gives major AI labs access to high-quality data to fuel post-training, while their public leaderboards help assess model capabilities. They also just released Scale Evaluation, a new tool that diagnoses model limitations. If youβre an AI researcher or engineer, learn how Scale can help you push the frontier at https://scale.com/dwarkesh. * Lighthouse is THE fastest immigration solution for the technology industry. They specialize in expert visas like the O-1A and EB-1A, and theyβve already helped companies like Cursor, Notion, and Replit navigate U.S. immigration. Explore which visa is right for you at https://lighthousehq.com/ref/Dwarkesh. To sponsor a future episode, visit https://dwarkesh.com/advertise. ππππππππππ 00:00:00 β How far can RL scale? 00:16:27 β Is continual learning a key bottleneck? 00:31:59 β Model self-awareness 00:50:32 β Taste and slop 01:00:51 β How soon to fully autonomous agents? 01:15:17 β Neuralese 01:18:55 β Inference compute will bottleneck AGI 01:23:01 β DeepSeek algorithmic improvements 01:37:42 β Why are LLMs βbaby AGIβ but not AlphaZero? 01:45:38 β Mech interp 01:56:15 β How countries should prepare for AGI 02:10:26 β Automating white collar work 02:15:35 β Advice for students
AI Enthusiast's Tech Upgrade
AI-recommended products based on this video




