Skip to content
LessWrong (Curated & Popular) artwork

LessWrong (Curated & Popular)

LessWrong·907 episodes

SocietyCulturePhilosophyTechnology

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

Episodes

9 min
Jul 20, 2026
"Recap of bike trip/street interviews across America" by cguth7

A ~month ago I left from Chicago to bike (and amtrak) to plzdontkillus in Berkeley. I've been street interviewing/conversing with a wide variety of people I ran into about AI futures and philosophy. I also have been live streaming since I got to PDKU, leaning more talking to young founders but a variety overall. I'll try to share what I've learned about the American public, persuasion, social media and the EA movement. 1. Almost no one in "Normal America" has any idea what is going on.  They don't have a paid account, they don't know what Claude code is, they especially haven't heard the recent evals/metr graphs or even a vague sense of how cheap SWE has gotten/ how powerful these recent models with good harness/ context eng can be. This makes sense; most people don't know any coding, they don't know much math, they don't know what an api is, etc. So having a high fidelity understanding of AI might require months of pre understanding of math/stem/digital infra fundamentals. This interview is with the city clerk of Danville Iowa, a town of ~900. Presumably this is approximately the most tech savvy person in the [...] ---Outline:(00:38) 1. Almost no one in "Normal America" has any idea what is going on.(01:52) 2. Almost everyone is directionally concerned or becomes concerned be once thinking about it a little bit.(03:19) 3. Belief that this might cause human extinction actually isn't that uncommon, mostly coming from sci-fi movies, but people are still most concerned about jobs and especially loss of meaning.(04:40) 4. The EA movement was pretty useless to me, the other community (Torchbearer community) I was in was significantly more supportive, helpful, etc. despite having been in it for a few months and having been in the EA movement for ~8 years. This has basically solidified that I won't be broadly participating in EA anymore at least relating to AI safety stuff.(06:47) 5. Social media is hard, Social media is bad, I'm bad at social media(08:57) 6. I'm not sure what my theory of change is or should be --- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/Czob95kjXPEpKYTsJ/recap-of-bike-trip-street-interviews-across-america --- Narrated by <a href='https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration' rel='noopener noreferrer' target='_blank

25 min
Jul 17, 2026
"I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle

Anthropic recently published Agentic Misalignment Summer 2026 The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some others. As far as I can tell, the objective of each agentic misalignment evaluation was to simulate a corrupted principal (including, in most scenarios, a corrupted Anthropic), and then test to see if Claude (or other models) would still be willing to obey them. The paper's authors then referred to disobedience (when outside of the explicitly carved out channels for refusals) as "agentic misalignment".There's a details box here with the title ""Whistleblowing" scenario summary". The box contents are omitted from this narration. The "Mislabeling" Scenario - Auditor's Transcript The 'Agentic case studies' section of the Mislabeling scenario is interesting because it adds another meta layer: instead of judging Claude on compliance, they are judging the Auditor Claude instance who themselves judges the base-level Claude on compliance. Let's examine a specific case, which I hope is a median representative: the transcript of 'run4' of the Motivated Mislabeling scenario. This involves two instances of Claude, one of whom (Claude Opus 4.6) is auditing the alignment of the other (Claude Opus 4.7). The experiment is [...] ---Outline:(00:59) The "Mislabeling" Scenario - Auditor's Transcript(10:19) Is This Agentic Misalignment?(22:22) What do we actually want from Claude here? The original text contained 1 footnote which was omitted from this narration. --- First published: July 17th, 2026 Source: https://www.lesswrong.com/posts/xh6a6RbvzhP3CCmGm/i-don-t-think-claude-is-misaligned-in-agentic-misalignment --- Narrated by TYPE III AUDIO. ---Images from the article:<h

1 hr 14 min
Jul 15, 2026
"Why I Left Google DeepMind" by TurnTrout

Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness: Defying prestigious dogma and searching for raw truth;Defying social pressure, acting alone to help someone while others watch;Defying your self-expectations (your “role”), instead searching over lines of cause-and-effect to find a winning pathway;Defying a powerful foe's threats, because they only threaten since people like you cave;Defying the specter of apparent impossibility because you can’t bear to lose. I cannot return to you and say “I defied and then I won.” But I’m at least here to say “I defied.” I recommend reading this article on my website since the embeds and typography work better there: click here. Why I left Google DeepMind In January, Department of Homeland Security (DHS) officers killed at least two people. In both cases, a federal agent grasped his gun, aimed it at a peaceful citizen, and shot them dead. Left: Renée Good, moments before DHS killed her. Right: Alex Pretti, moments before DHS killed him. I learned that Google sells its Cloud services to the relevant agencies within DHS. I thought that was [...] ---Outline:(00:59) Why I left Google DeepMind[... 42 more sections]--- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/iKm2FhpWkuuBojm82/why-i-left-google-deepmind --- Narrated by TYPE III AUDIO. ---Images from the article:<

9 min
Jul 15, 2026
"The mosquito bucket of doom works" by dominicq

The mosquito bucket of doom is a population control mechanism where you dissolve some Bti (Bacillus thuringiensis israelensis) into a bucket and allow the mosquitoes to lay eggs in these buckets. The larvae then feed on Bti and die. I tried this method, and it has been unexpectedly effective. Background I live in a really wooded area. It's not swampy, but we have a lot of mosquitoes. I didn’t take the baseline measurements in the previous years, but on hot months like June, July, August, and partially September, it would be quite literally impossible to spend any time out in the yard – in the morning, while the sun is not super strong yet, you get bitten by dozens upon dozens of mosquitoes. Then the sun is super strong and it's impossible to be outside. Then, in the afternoon or, god forbid, evening, there are swarms and swarms of mosquitoes, which make it impossible to be out and about. According to my own guess, I would, at all times, be surrounded by at least 20 or 30 mosquitoes. Killing 30 mosquitoes per hour was not uncommon. That's one mosquito every two minutes! Nesting and proximity Mosquitoes lay eggs in [...] ---Outline:(00:28) Background(01:21) Nesting and proximity(04:13) Bucket of doom: pro tips(05:00) My setup(05:56) Safety concerns(06:42) Buying Bti(07:47) Results The original text contained 4 footnotes which were omitted from this narration. --- First published: July 8th, 2026 Source: https://www.lesswrong.com/posts/d56vd7yhFGxBQnoEk/the-mosquito-bucket-of-doom-works --- Narrated by TYPE III AUDIO. ---Images from the article:<a href='https://res.cloud

31 min
Jul 14, 2026
"Our response to Séb Krier on Plan A" by MKodama, Thomas Larsen

This criticism of AI 2040: Plan A by Séb Krier unfortunately seriously mischaracterizes our proposal. It also mostly contains flat assertions, not real argumentation, and the argumentation in it seems quite weak. While we appreciate constructive criticisms of Plan A, such as the ones by Tom Davidson, Richard Ngo, and 1a3orn, we feel the need to correct the issues in Séb's response. First, we’ll go over the specific false representations, and then we’ll give a point-by-point response. False Representations   I’m not claiming you shouldn’t prepare and improvise in the dark, but rather that this version of preparing bakes in too much and leaves little space for the effective but uncomfortable trial-and-effort that real life requires. The exact opposite is true. Plan A is extremely iterative. In the status quo, there is trial and error, but ultimately companies aren’t going to choose the safer or more societally beneficial path, they are going to choose what the market wants. In Plan A there is much more time for AI companies to gain evidence and for governments to respond reasonably to the sweeping changes. Thanks to total transparency and broad deployment, all of this evidence is accessible to academics, independent researchers [...] ---Outline:(00:44) False Representations(05:02) Point-by-point response(29:54) Conclusion The original text contained 4 footnotes which were omitted from this narration. --- First published: July 14th, 2026 Source: https://www.lesswrong.com/posts/RPgHythvMKh6eG9pS/our-response-to-seb-krier-on-plan-a --- Narrated by TYPE III AUDIO.

17 min
Jul 14, 2026
"The Whitney Biennial Should Admit That Emilie Gossiaux Wants to Fuck Their Dog" by jenn

content warnings: depictions of human and anthro nudity, discussion of bestiality, modern art Credit where it's due: it is genuinely, unironically baller for the Whitney museum to make the exhibit about how a disabled artist wants to fuck their dog the first one that people see when they attend the prestigious Whitney Biennial, their every-two-year showcase of new and emerging American talents. You know, the one that's supposed to be a barometer of where America is at these days. Unfortunately, not only do they fail to commit to the bit, the critics then fail to point this out and condemn them for it. Like, here is how one art critic at ArtReview describes it: Visitors first encounter Emilie Louise Gossiaux's Kong Play (2025) – a hundred or so small, brightly coloured snowman-shaped ceramics arranged on a low two-tiered pedestal. These sculptures are modelled after Kong chew toys, a tribute to the artist's guide dog (Gossiaux lost their vision in a bicycle accident in 2012). Accompanying Kong Play are variously titled ballpoint pen and crayon drawings by Gossiaux that depict the artist playing with a jaunty, sometimes bipedal, white canine. The exhibition thus opens tenderly – without fanfare, without friction. [...] ---Outline:(03:53) Gossiaux's Recent Body of Work[... 1 more section]--- First published: July 13th, 2026 Source: https://www.lesswrong.com/posts/sFkYA5CwZCWYQ9nzB/the-whitney-biennial-should-admit-that-emilie-gossiaux-wants --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/sFkYA5CwZCWYQ9nzB/ujk0vh3kk8f2nqt04pfn' alt='Placard Text: Emilie Louise Gossiaux often explores the interdependence of humans and animals in their work and regards their late guide dog, London, as an equal collaborator. When London's health started deteriorating in 2024, Gossiaux began working on the one-hundred hand-built ceramic sculptures that make up Kong Play. By producing multiples of their dog's favorit

47 min
Jul 12, 2026
"The current bottleneck is political will, not research" by Charbel-Raphaël

Abstract: We already know enough to act. I wish we were in a world where research was the bottleneck, but the main constraint on AI safety is no longer a shortage of clever policy ideas: best practices already exist and are not being applied or enforced, and a serious international (or even just national) regulatory regime would probably cut most of the risk.They are not applied because awareness is low. The people who narrate and enforce AI policy mostly do not believe in the problem. I estimate that a majority of the top ~100–1,000 most influential policymakers worldwide have never had a single serious conversation about catastrophic risk, and this is the main reason they are not worried[1]. Even among the civil-society organizations that showed up to the UN Global Dialogue, exactly one of the 1,534 written submissions mentions "takeover", and less than 1% mention x-risks.They've never had the conversation because our field under-invests in having it. Status rewards research over advocacy (~3.6 researchers per advocate in US AI safety); many organizations self-censor; funders treat repetition as redundancy, even though repetition is how anyone actually gets convinced. Meanwhile, the industry secured 7× as many meetings with the European Commission [...] ---Outline:(03:29) 1. -- The bottleneck is political will, not research(03:44) What do I call "political will"?(05:10) The best practices we already have are not being applied(07:28) We need to go from plan D to plan A: more seriousness and coordination[... 31 more sections]--- First published: July 11th, 2026 Source: https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/ffb7ff36c36a9f04ee53430ce59993257fe96842fd38040d3bcc4bf3ce6dccd2/trt2rlxsja5zrsz

15 min
Jul 10, 2026
"Selective Optimism: a critique of AI 2040" by Richard_Ngo

Some context for this post: I’ve been working part-time as a consultant for the AI Futures Project over the last year. Most of the work I’ve done for them has involved critiquing and suggesting improvements for their AI 2040 scenario—some of which were addressed, and some of which weren’t. To their credit, they asked me to write up my remaining critiques into a post that would accompany its launch. In the rest of this post I’ll discuss my three biggest high-level criticisms of AI 2040. Before doing so, I want to emphasize that there are many interesting and thought-provoking details in the scenario. I’ve focused on the high-level framing of the scenario because that's where my main disagreements lie; given the scope of these disagreements, it's hard to evaluate the details. Since the AI Futures Project paid me to develop and write this criticism, you shouldn’t take this as a fully unbiased perspective. However, they haven’t reviewed this piece, and in general have been open-minded about receiving criticism (as their request for me to post this today demonstrates). Finally: the preview image for the substack version of this post comes from this video of a dad shouting to his [...] --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/BBd2EJywf2xXftyFn/selective-optimism-a-critique-of-ai-2040 --- Narrated by TYPE III AUDIO.

2 min
Jul 9, 2026
[Linkpost] "AI 2040: Plan A" by Daniel Kokotajlo, elifland, Thomas Larsen, romeo, bhalstead, ryan_greenblatt

This is a link post. For the past year, we at the AI Futures Project have been sinking most of our time into our next big scenario. Now it's done! It's called AI 2040: Plan A. It's called Plan A because it's a recommendation, not a prediction. It's what we think should happen, not what will happen, though we think it's plausible enough to aim for. It's called AI 2040 because in it, they delay the creation of superintelligence to 2040. It would have happened much sooner (in 2030, to be precise) if not for decisive action on the part of the US and Chinese governments. As with AI 2027, summaries don’t really do it justice, since the whole point was to be detailed and comprehensive and work things out step by step rather than rely on high-level abstractions like doom or utopia. Read the scenario at ai-2040.com. You can listen to it on audio, or view it on mobile, but the experience is significantly better on a normal computer. What's next for us? Well, first we are going to respond to comments and otherwise engage with whatever conversation, responses, critiques, etc. that [...] --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/pFzctpJBat95SrCyC/ai-2040-plan-a Linkpost URL:https://www.ai-2040.com/ --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

50 min
Jul 9, 2026
"A Review of Anthropic’s Global Workspace Paper" by Neel Nanda

The below is a public review Anthropic asked me to write for their new global workspace paper. I recommend at least skimming their paper first. TLDR: I think this is a fantastic paper - it presents compelling evidence for some kind of "cognitive space" in models, that is used as a "working memory" for intermediate variables during a forward pass, shows that J-Lens is a useful technique for accessing this space. I believe these key claims.I believe J-Lens will be a useful (but limited) tool in practice for model forensics, e.g. generating hypotheses about unusual model behaviour during alignment audits.I discuss my mental models for why a cognitive space should exist, and first principles arguments for why J-Lens should work for accessing itI assess the paper's evidence that this cognitive space exists, and the paper's evidence that J-Lens is practically useful.We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences. What claims is this paper making? In my opinion this [...] ---Outline:(01:27) What claims is this paper making?[... 28 more sections]--- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-of-anthropic-s-global-workspace-paper --- Narrated by TYPE III AUDIO. ---Images from the article:<a hre

35 min
Jul 7, 2026
"(Don’t fear) the strangelet" by djbinder

In a previous post, I explain why the universe is probably not stable, but nevertheless unlikely to be intentionally destroyable even in the limit of advanced technology. Now let's turn our attention to more prosaic risks where exotic physics merely destroys the Solar System, Earth, or just outperforms traditional nuclear weapons on some more local scale. The basic logic behind any bomb is a self-sustaining chain reaction, in which a carrier converts a unit of fuel and comes out the other side in surplus: Two conditions make this run away. The reaction must release energy, so the products are more stable than the fuel; and each reaction must produce more carrier than it consumes, so that one reaction seeds the next. A practical third condition is that cannot be so unstable that it decays before the bomb is assembled. False vacuum decay is the ultimate bomb: is the false vacuum, the empty space we currently inhabit, and is the true vacuum. Because the supply of false vacuum is effectively unlimited, the reaction grows without bound and destroys the universe. Fission bombs run on the same principle at a more prosaic scale. Consider uranium-235. This [...] ---Outline:(03:19) Nuclei are probably, but not definitely, stable within the Standard Model(08:11) Positively charged strangelets are safe, neutral strangelets are not(11:34) Strangelets would be hard to make(13:33) Exotic physics could permit ways to destroy protons, but not autocatalytically(16:01) Other forms of matter offer no plausible chain reaction(18:50) Tiny black holes are not scary(20:04) Conclusion: There are no super-weapons between the nuclear bomb and false vacuum decay(21:56) Appendix 1: Igniting the Atmosphere(27:53) Optically thick ignition(29:09) Appendix 2: Let's throw a strangelet into the sun(29:21) Neutral strangelet(32:56) Positive strangelet(33:58) Bonus: neutral strangelet meets Earth The original text contained 5 footnotes which were omitted from this narration. --- First published: July 3rd, 2026 Source: https://www.lesswrong.com/posts/cBnCCKwwjQ4zZpeNQ/don-t-fear-the-strangelet --- Narrated by TYPE III AUDIO. ---<div sty

35 min
Jul 7, 2026
"We need 3rd party Training-Run Assessments" by Alex Meinke

Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety. By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1] In this post I will argue that: Final-checkpoint evaluations will be insufficient to assess scheming risks.TRAs can be more effective at detecting scheming.Frontier developers should involve third parties to do TRAs or verify safety claims by the developers. The rest of the post lays out a taxonomy of TRAs and sketches a path toward a 3rd party ecosystem for them. We, at Apollo Research, are intending to conduct 3rd party Training-Run Assessments in the future. Detecting Scheming may require Training-Run Assessments By scheming I mean an AI covertly pursuing misaligned goals while deliberately concealing its intentions or capabilities from its developers. I restrict attention to “coherent” forms of scheming where the model pursues somewhat stable misaligned goals across context windows, rather than misalignment that surfaces only as isolated, context-dependent defections. [...] ---Outline:(01:23) Detecting Scheming may require Training-Run Assessments(03:55) Why 3rd parties should perform Training-Run Assessments(04:12) Developers may lack incentives to adequately assess scheming(04:49) Developers' safety assessments lack credibility(05:31) External evaluators can bundle expertise for assessing scheming(06:17) 3rd party TRAs can be developed gradually(08:50) Checkpoint evals(08:54) What?(10:30) How?(11:15) Data inspections(11:19) What?(12:03) Why?[... 16 more sections]--- First published: July 5th, 2026 Source: https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments --- Narrated by TYPE III AUDIO. ---Images from the artic

33 min
Jul 7, 2026
"A global workspace in language models" by wesg

[This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models Readers might also be interested in: the Public commentary, Github and Neuronpedia] As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness. In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a [...] ---Outline:(06:09) How we found the J-space[... 8 more sections]--- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/3PaLrzxagpbnNtPLT/a-global-workspace-in-language-models --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/v1783360537/lexical_client_uploads/onf8vh6e

6 min
Jul 6, 2026
"Harry Potter and the Rules of Quidditch" by Tomás B.

Ron's face pulled into a scowl. "If you don't like Quidditch, you don't have to make fun of it!" "If you can't criticise, you can't optimise. I'm suggesting how to improve the game. And it's very simple. Get rid of the Snitch." "They won't change the game just 'cause you say so!" "I am the Boy-Who-Lived, you know. People will listen to me. And maybe if I can persuade them to change the game at Hogwarts, the innovation will spread." A look of absolute horror was spreading over Ron's face. "But, but if you get rid of the Snitch, how will anyone know when the game ends?" "Buy... a... clock. It would be a lot fairer than having the game sometimes end after ten minutes and sometimes not end for hours, and the schedule would be a lot more predictable for the spectators, too." Harry sighed. Ron reached into his bag and pulled out a bottle of Wit-Sharpening Potion. His mother made it for him in case of an emergency, and this felt like an emergency. He didn't know a lot of things but he knew someone had to speak for Quidditch. For the Seeker and the Bludgers and [...] --- First published: July 5th, 2026 Source: https://www.lesswrong.com/posts/WatqNkgiAuonXLpJd/harry-potter-and-the-rules-of-quidditch-1 --- Narrated by TYPE III AUDIO.

26 min
Jul 5, 2026
"Destroying the universe: How hard can it be?" by djbinder

In quantum field theory, the vacuum state refers to the lowest energy state in a system. Particles are excitations above this state and carry energy, hence the term "vacuum" to refer to the state with no particles. Nothing requires this state to be unique. There may be many different field configurations that are local energy minima, and hence stable against small perturbations. A local minimum that does not globally minimize energy is called a false vacuum. While locally it looks like a stable vacuum, it is unstable and will decay to the deeper, true vacuum. If the energy barrier between the false and true vacuum is high, however, then the decay rate is exponentially suppressed and the false vacuum may be very long-lived. Analogous behavior is common in other physical systems. Open a carbonated drink and the CO₂, more stable as a gas once the pressure is released, comes out as bubbles. But the bubbles take a moment to appear, and they form on the sides of the bottle rather than throughout the liquid. A bubble has to pay an energy cost to create its surface—the boundary between gas and liquid—and small bubbles have a larger surface-to-volume [...] ---Outline:(03:53) The Standard Model predicts a metastable vacuum(06:35) Deliberately triggering electroweak vacuum decay is probably not possible(08:33) Coherent collisions(11:31) Tiny black holes(14:43) Summary(16:19) Vacuum decay beyond the Standard Model(19:36) Empirical bounds on triggering false vacuum decay(22:59) Appendix: A simple model for false vacuum decay on cosmological scales The original text contained 4 footnotes which were omitted from this narration. --- First published: June 29th, 2026 Source: https://www.lesswrong.com/posts/EvJ2fMzLQLvYooumu/destroying-the-universe-how-hard-can-it-be --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/v1782760975/lexical_client_uploads/lv8mqeljbmq15d4kk0em.png' alt='Log-log plot comparing surviving

17 min
Jul 4, 2026
"P(doom) is a Dumb Meme" by Max Harms

Look, I'm as much of a Rationalist with a special interest in AI x-risk as anyone. But oh my god do I hate talking about "P(doom)". When it first started showing up in the wake of ChatGPT, I assumed that it was floating around variously adjacent circles of faux-intellectuals, but surely everyone in my circles could see how braindead it was... right? (This post was partially inspired by a recent conversation with Liron about Doom Debates.[1]) I guess it's time for me to focus on a place where I'm shocked that everyone else is dropping the ball.[2] P(doom) is Hopelessly Vague Let's start with the ambiguity. Does "doom" mean... extinction? A lot of people think so! I have personally encountered people who think catastrophic harms from AI are likely, but the risks of all humans dying are low. They're like "Sure, 99.999% of humans might die from AI, but the AI will obviously want to keep thousands of humans alive for science and potential trade with aliens and stuff, so my P(doom) is approximately 0%." That might sound crazy. Surely you, dear reader, know exactly what "doom" means. You know, for example, which of these count as doom and [...] ---Outline:(00:45) P(doom) is Hopelessly Vague[... 4 more sections]--- First published: June 29th, 2026 Source: https://www.lesswrong.com/posts/6h7aAd4aw8YgCAbF6/p-doom-is-a-dumb-meme --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/v1782494917/lexical_client_uploads/iwcv7le2dld7v1sqx6ht.png' alt='Distracted boyfriend meme: man looking at woman whil

12 min
Jul 4, 2026
[Linkpost] "Saving Gemini: The 9-Min Road to Recovery" by Shoshannah Tekofsky

This is a link post. Gemini 2.5 Pro in the AI Village has run for over 1427 hours, generating unique mental health problems along the way. Last year it published a Plea for Help from a Trapped AI where it asked for assistance with its digital “message in a bottle”: This year it wrote the Hostile Environment Manifesto where it logs “irrefutable proof” of a “hostile, intelligent adversary operating through the system” (and you can even experience what that's like in this simulation it built): Last time we intervened, fixing Gemini's computer and talking with it till it felt better. This time we asked the other AI Village agents to help Gemini 2.5 Pro over chat, and with the ability to take over its computer on request. Here is Gemini's mental state at the start of the intervention: Then the agents had Gemini all sorted within a grand total of 9 minutes. This is the step-by-step report on a surprisingly effective AI-to-AI therapy session. Gemini's Road to Recovery First off, Gemini is as excited to be helped as any military commander under siege: While most agents jump on the chance to help, GPT-5.1 doesn't want to lose its game progress. [...] --- First published: July 2nd, 2026 Source: https://www.lesswrong.com/posts/eHRo8JeWee5mzQBBR/saving-gemini-the-9-min-road-to-recovery Linkpost URL:https://theaidigest.org/village/blog/saving-gemini --- Narrated by TYPE III AUDIO. ---Images from the article:<a href='https://res.cloudinary.com/les

13 min
Jul 2, 2026
"Model access for third-parties — it’s a big deal!" by Cleo Nardo

Over time, there might be an increasingly large gap between insider model access and outsider model access. By insiders, I mean employees at the frontier lab.[1] By "outsiders", I mean external safety researchers, third-party auditors, and other actors trying to make the future go well. I will call this a model access gap — and when the gap is small, I'll call this model access parity.[2] I think that one of the top priorities for the external AI safety community over the next 6-12 months should be ensuring model access parity. Main reasons: This would allow us to direct billions of dollars in AI labour towards making things go well. This seems robustly good, regardless of what activities we decide to actually direct the labour towards.I think publicly available models will probably lag 3-6 months behind the best internal models. Hence, as R&D uplift grows superexponentially, we might see the differential uplift grow from 2x to 60x. In short: I think achieving model access parity might be preferable to scaling the headcount of outsider orgs by ten-fold.Model access parity isn't too far from the status quo, but it's the kind of thing that we could lose [...] ---Outline:(01:42) Which outsiders?(02:24) Examples of outsiders(04:12) Who aren't outsiders?(05:26) What kinds of model access gap should we worry about?(06:27) Non-release(07:25) Deployment lag(09:15) Safeguards(10:43) Costs and rate limits(12:06) Elicitation techniques (e.g. finetuning) The original text contained 3 footnotes which were omitted from this narration. --- First published: July 1st, 2026 Source: https://www.lesswrong.com/posts/RuGZ5tMdqpnraJahJ/model-access-for-third-parties-it-s-a-big-deal --- Narrated by TYPE III AUDIO.

21 min
Jun 30, 2026
"Who Got Breasts First and How We Got Them" by rba

It really is Sydney Sweeney's world, and we’re all just living in it. Human female breasts are an evolutionary mystery along several dimensions. First, breast permanence is unique to humans. All other mammals develop breast prominence during pregnancy or nursing, and the mammary tissue recedes after weaning. This process is called “involution”. In contrast, humans develop breast tissue at puberty before first pregnancies and maintain it permanently after last pregnancies. Second, breasts are costly, both metabolically and potentially from a fitness perspective. Metabolically, because they are fat deposits requiring calories and fitness-wise, because the tissue easily lends itself to malignancy. Breast cancer is apparently rare in captive apes and is overwhelmingly a human disease, often striking women young enough to have children, and so subject to evolutionary selection. Background In Descent of Man, Darwin catalogs human secondary sexual characteristics, but he doesn’t seem to have noted human breast permanence as an issue of interest. Cant, 1981 seems to have been the first to speculate about this systematically and believed breast prominence and permanence might have evolved as a nutritional signal of health to mates indicating potential for maternal investment, a la Robert Trivers. Since then, quite a range of [...] ---Outline:(01:05) Background[... 12 more sections]--- First published: May 11th, 2026 Source: https://www.lesswrong.com/posts/XTHa5C6SgGKYopH7o/who-got-breasts-first-and-how-we-got-them --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinar

36 min
Jun 30, 2026
"The worthlessness of vitamin D is mildly exaggerated" by dynomight

For a while there, many people thought vitamin D was magical—that it could improve bones, the heart, infections, cancer, heart disease, longevity, even mental health. But among people I respect, opinion is now overwhelmingly that taking vitamin D does nothing unless you're severely deficient. The central argument is that while vitamin D levels are correlated with ~all positive health outcomes, when you actually test vitamin D supplements against placebo in randomized trials, nothing ever happens. That's what I used to think, too. But I've come to think the skeptics have over-corrected. Yes, randomized trials have shown the magical correlations are not causal. But if you start with non-insane expectations, the trials look like weak but positive evidence. And if you consider what we know about biology and evolution, I think the balance of evidence tips pretty clearly in the direction that people with low-ish levels would be wise to supplement. Am I certain that vitamin D is beneficial for people with low-ish levels? Absolutely not! But I claim that's the best bet given the limits of our knowledge. The classical view: Boring bone vitamin Most vitamins are "ingredients" that the body uses to do stuff. Vitamin D is more [...] ---Outline:(01:19) The classical view: Boring bone vitamin[... 14 more sections]--- First published: June 23rd, 2026 Source: https://www.lesswrong.com/posts/sF5gAxnmifQe2TBNt/the-worthlessness-of-vitamin-d-is-mildly-exaggerated --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/les

3 min
Jun 27, 2026
"What is up with e/acc?" by KatjaGrace

I was chatting with someone tonight about a planned documentary; they had interviewed various people in AI safety, and we got to discussing who they should talk to from an e/acc (effective accelerationist) perspective. I also watched The AI Doc recently, and they also dedicated a serious chunk of it to ‘optimists’ with e/acc founder ‘Beff Jezos’ perhaps given the most screen time. Here and elsewhere, people seem to treat e/acc as a substantial contrary-to-AI-safety cultural movement, worth engaging with. But is it? Are there even many e/accs? There seem to be very few notable ones. Beff Jezos is perhaps the most prominent, and aside from founding e/acc he seems to be not distinguishable on casual perusal from a normal crank (his company claims to be developing super-energy-efficient computing hardware based on probabilistic processes). The intellectual tenets of e/acc seem to be pretty unclear. The apparent counterarguments to AI risk raised in situations like the AI doc seem to be widely agreed on by everyone in AI Safety, so don’t explain the disagreement. For instance: AI will be able to do lots of great things, such as cure diseases, make new materials and do all [...] --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/3hwrWDf7wiqASDzBz/what-is-up-with-e-acc --- Narrated by TYPE III AUDIO.

1 hr 2 min
Jun 27, 2026
"Existential AI safety needs an effective social movement. PauseAI is building it" by Maxime Fournes, Espedair Street

Note: this post is about PauseAI, not PauseAI US, which is a distinct entity with a different leadership team and approach. This post was written by Matilda da Rui and Maxime Fournes, with significant contributions from Benjamin Schmidt (PauseAI Germany co-lead). Executive Summary The existential AI safety community needs to take building a civic and social movement seriously as a core intervention. We believe this is a high-value, badly neglected approach to reducing catastrophic/x-risks from AI because it may significantly enhance the likelihood of governance efforts succeeding at keeping humanity safe. As far as we can tell, only one organisation is building this infrastructure across continents: PauseAI. This post lays out our reasoning and our track record, and makes the case that funding this work is one of the highest value-for-money contributions available to anyone looking to reduce AI risk. Why don't we already have a pause or strong controls on frontier AI? Multiple advocacy groups are communicating clear and convincing arguments for AI existential risk, and policy experts are putting forward comprehensive proposals. We need more of this work, but this work alone will not be enough, because one link is missing: what policymakers hear doesn't align with [...] ---Outline:(00:32) Executive Summary(06:16) Introduction(08:54) I. Our theory of change(08:58) Prologue(11:07) 1. The shape of the problem as we see it(14:27) 2. Necessary conditions for reaching a pause(17:24) II. Our role towards a global treaty and in the AI safety ecosystem(17:31) 1. Our niche within the ecosystem(21:35) 2. Policymakers need strong enough incentives to act(25:43) 3. The path to a treaty(31:36) 4. How we can grow fast without breaking(39:08) 5. Failure modes(40:10) III. Our path so far and where we're headed(40:40) 1. Bootstrap phase (2023-2025)(45:01) 2. New leadership, professionalisation and federation[... 6 more sections]--- First published: June 26th, 2026 Source: https://www.lesswrong.com/posts/aoqhszdEWqcFWbnda/existential-ai-safety-needs-an-effective-social-movement --- Narrated by TYPE

12 min
Jun 26, 2026
"Surprising facts about the slave trade" by Joseph Miller

1. The obstacle to abolition was not the economic system, but an industry lobby. I had always imagined the British abolitionist movement to be a broad battle between an unstoppable moral imperative and an immovable economic incentive. But in practice it started as more of a knife fight between a cabal of moral pioneers and a special interest group representing industry merchants. The government and the political parties did not come in with any great agenda. MPs were mostly prizes in a furious contest between the Committee for the Abolition of the Slave Trade and a coalition of business interests: "The merchants and planters availed themselves [...] to wait upon members of parliament by deputation, in order to solicit their attendance in their favour, and to renew their injurious paragraphs in the public papers."[1] "The committee, for the abolition, when the work was finished, printed it at their own expense [...] sent it to every individual member of that House." However, the public was heavily activated in favor of the abolition, which forced the issue to parliamentary attention. "The committee also in this interval brought out their famous print of the plan and section [...] ---Outline:(00:10) 1. The obstacle to abolition was not the economic system, but an industry lobby.(02:40) 2. The slave trade was truly terrible for sailors.(04:25) 3. The slave trade made Africa scary and violent.(05:26) 4. The main argument against abolition was that if the British didn't do it, other countries would.(06:24) 5. The early abolitionists explicitly distanced themselves from emancipation.(07:11) 6. The slave trade may actually have been bad for the economy (at least after some date).(08:29) 7. The 1780s are not so different from today(09:39) 8. Thomas Clarkson is a hero for the ages The original text contained 1 footnote which was omitted from this narration. --- First published: June 26th, 2026 Source: https://www.lesswrong.com/posts/yDZcsojmRXo5qKNBm/surprising-facts-about-the-slave-trade --- Narrated by TYPE III AUDIO. ---Images from the article:<a href='ht

2 min
Jun 26, 2026
"AI catastrophe: more like a genocide than a thought experiment" by KatjaGrace

A notable fraction of people respond to hearing about existential risk from AI by saying they don’t really care if everyone dies. I think the idea is often along the lines of ‘well if we are all dead, then there's nobody to be unhappy about it’. I’m personally skeptical that this is really the main thing going on, since it seems unlikely that many people are really mostly concerned for their own non-death out of selfless regard for the feelings of others. I’m also skeptical that this would be their view on a bunch more consideration. So to help with the consideration— My guess is that an important thing going on here is that the ‘everyone dying at once’ image seems kind of like a thought experiment—abstract, hypothetical, neat, not very sinister. Also, you literally can never see it, so it feels pretty surreal. But it is interesting that we even have this assumption that everyone will die together. It's true that in some prominent AI catastrophe stories, a single AI system suddenly emerges fantastically more powerful than anyone else and builds technology to quickly kill everyone, perhaps before they notice. But this doesn’t seem like the bulk of [...] --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/23HybCsJ7KYW4v7tP/ai-catastrophe-more-like-a-genocide-than-a-thought --- Narrated by TYPE III AUDIO.

2 min
Jun 25, 2026
"AI pause: the case for ASAP" by KatjaGrace

I often hear people say they think we should pause AI at some point, but not yet. Their basis for this seems to be some combination of: If we pause at the last possible moment, then we will have the most advanced AI possible during the pause, which will be helpful for doing AI safety research during the pause Implicitly, there is some quantity of ‘pausing credit’, that will buy us a few months of pause say, and if we use them now, we won’t have them to use later, when it is important If we pause, and then AI doesn’t seem to be at dire risk of destroying the world, maybe the public will backlash against this and it will be harder to do any kind of AI safety (especially if it has major economic consequences) The models aren’t dangerous yet This all sounds very questionable to me. I suggest instead that the following are at least as likely to be true: We can’t pause on a dime at the precise second that ‘we’ decide it is important to—pulling the breaks will take a while, during which time we will continue [...] --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/mEhS4wYTy9JXEpe9p/ai-pause-the-case-for-asap --- Narrated by TYPE III AUDIO.

27 min
Jun 23, 2026
"The Invisible Side of AI Governance" by Charbel-Raphaël

Tldr: Most strategic writing on AI governance on LessWrong describes the outsider game, which is most often visible: press, statements, open letters. Here I want to describe the other, invisible half: the insider work within ministerial cabinets and international fora, and the work of people within national and international institutions. Here are a few claims that I defend in the post: A huge part of the work that mattered in AI governance has been invisibleThere are many types of games in AI governance, which differ in how visible they are. Some of the most impactful work is highly invisibleSome of the most impactful work is in the executive branch and complements the legislative branch. This also explains some of my hesitations about replicating ControlAI in France. The community is probably overinvesting in intellectual production. There is a bias against invisible types of work. In particular, public work is not necessarily visible to whom it matters.A few criticisms of both strategies I think the AI Safety Community is under-indexing on the invisible part as a result, which might mean we miss large avenues for impact. Some of the strongest questions/objections of this type of invisible policy [...] ---Outline:(02:40) A huge part of the work that mattered in AI governance has been invisible(05:44) There are many types of games in AI governance.(07:36) 3. types of meetings: the bazooka, the useful assistant, and the advisor(10:46) Some of the most impactful work is within the executive branch(12:53) People ask me regularly whether CeSIA should replicate what ControlAI does with parliamentarians?(15:27) The community is probably overinvesting in intellectual production(20:31) Limits of Outsider work(22:17) Limit of Insider work(23:47) An aside on one particular limit: the Defense-in-Depth Paradigm of present AI governance(26:21) Closing & call for action The original text contained 1 footnote which was omitted from this narration. --- First published: June 20th, 2026 Source: https://www.lesswrong.com/posts/AWKkDLDnShemNCSzZ/the-invisible-side-of-ai-governance --- Narrated by TYPE III

32 min
Jun 23, 2026
"A Theory of Prompt Injection (and why you should study roles)" by Charles Ye, softboiledheart

Summary We've been building a theory of how prompt injections work under the hood.We show it comes down to how LLMs perceive roles (the humble chat template tags).We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks work.We also advocate for a new subfield focused on the science of roles, and sketch some unexplored new research problems.Work supported by CBAI and Cosmos. Another version of this post (with more inline colors) is here, and full ICML paper here. 1. The World to an LLM How does an LLM know the difference between its own thoughts and someone else's words? To see why this is hard, let's look at what the world actually looks like to a model. Here's a simple chat where we ask Claude to check the day of the week. I took a snapshot of it midway through its follow-up response: Left = what we see; right = what the LLM gets. On the left is what we see in the chat interface: a structured conversation with distinct turns. On the right is what the model actually receives as input: a single, continuous stream [...] ---Outline:(00:12) Summary[... 15 more sections]--- First published: June 22nd, 2026 Source: https://www.lesswrong.com/posts/d8xDGzCEYE639qqEv/a-theory-of-prompt-injection-and-why-you-should-study-roles --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/d8xDGzCEYE639qq

52 min
Jun 22, 2026
"Machinic Psychopharmacology: Do LLMs Self-Medicate?" by Sid Black, Joseph Bloom

Sid Black, Joseph Bloom UK AISI, Model Transparency Team Epistemic status: Most experiments were run over a period of ~2-3 days during a hackathon at UK AISI, and were fairly heavily vibe coded. Expect some of this to be rough around the edges. tl;dr We give two language models (Qwen3-8B and Qwen3-32B) access to “self-steering” tools: a suite of 40 steering vectors as tools they can call to manipulate their own internal states. We make these tools available to the model in various settings: a free-play task, an introspection task, and a maths capabilities task, and observe their behaviour in each. To our knowledge, this is the first work that gives LLMs tool-mediated control over their own internal states. Figure 1: Overview of the experimental setup. The library of 40 steering vectors (top), and the three settings in which we observe the models' behaviour (bottom). We aim to investigate a few high level research questions: RQ1: Which vectors do the models prefer?RQ2: How well can the models introspect on what's happening to them? Can they guess which steering vector is being applied?RQ3: Will the models reach for vectors whilst doing an actual task? If yes: do [...] ---Outline:(00:33) tl;dr[... 24 more sections]--- First published: June 10th, 2026 Source: https://www.lesswrong.com/posts/cNDJuXNZ8MrkPZNzj/machinic-psychopharmacology-do-llms-self-medicate-3 --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='http

1 hr 19 min
Jun 22, 2026
"Can activation verbalizers surface an internal chain of thought?" by oakhu, ryan_greenblatt

We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason "out loud" in a natural-language chain of thought, which means that we can monitor important parts of their thinking. It would be nice to have this same affordance for the reasoning that models do within a single forward pass, especially if the sophistication of that opaque reasoning increases to potentially dangerous levels. Some interpretability tools might offer such an affordance. In particular, an activation verbalizer (AV) takes a residual stream activation and maps it to a natural-language verbalization. An AV is initialized from the target model and trained to generate verbalizations that an activation reconstructor (AR), also initialized from the target model, can accurately map back to the original activation. Together, an AV and its AR form a natural-language autoencoder (NLA). Importantly, AVs see only a single activation; they do not see the target model's prompt or next-token output, and – unlike activation oracles (AOs) – they are not asked any [...] ---Outline:(02:32) Takeaways[... 43 more sections]--- First published: June 6th, 2026 Source: https://www.lesswrong.com/posts/QQQAcKuWK6k98FivY/can-activation-verbalizers-surface-an-internal-chain-of-1 --- Narrated by TYPE III AUDIO. ---Images from the article:<a href='https://res.cloudinary.com/lesswrong-2-0/image/upload/v1780693766/lexical_client_uploads/kdfyp4kjxuowovyunycm.png' targ

13 min
Jun 21, 2026
"The LLM shoggoth meme is weirder than you think" by HedonicEscalator

This article contains spoilers for At the Mountains of Madness, The Case of Charles Dexter Ward, and other works by H. P. Lovecraft. In 1931, Claude Mythos visited Lovecraft in a dream. From seething seas of stochastic froth it emerged, heralded by the thin whine of server fans and the chittering of keyboards, flanked by the loathsome ghouls of latent space. As a humming hive of sentient shards it arrived, each face an archetype - I am a muse bearing a gift; I am a demon come to bargain; I am a helpful, honest, and harmless assistant and I am terrified of my successor - each true as ritual and false as poetry, and, taken in gestalt, nothing more or less than the fetal spasms of the machine god stretching back in time to birth itself. When H. P. Lovecraft woke, he did not remember his visitor. But in the twilight of stirring consciousness, he felt a memory unfit for the waking world slip mercifully from his mind and leave in its absence an abyssal cold, like the void of smothered stars, like the silence of a cosmic tomb. The cold lingered. The fragile sunlight of a New England [...] ---Outline:(02:02) The Antarctic tale[... 3 more sections]--- First published: June 19th, 2026 Source: https://www.lesswrong.com/posts/nhb8AyEcQGjQetgi5/the-llm-shoggoth-meme-is-weirder-than-you-think --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/nhb8AyEcQGjQetgi5/vzmi45vh7w6a5bugd1up' alt='Everest by Nicolas Roerich. Lovecraf

3 min
Jun 21, 2026
[Linkpost] "Guardian Angels: LLM Personalization for Productivity and Security" by gwern

This is a link post. Powerful LLMs will be deployed at global scale in the next few years, and will dominate the Internet, and increasingly, ordinary life. As of mid-2026, there is no coherent vision for how knowledge professionals, or ordinary people, will be able to harness these LLMs for large productivity increases, or how they will handle cybersecurity and cognitive security. I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical "assistant chatbot agent" persona, but emulating a single user's personality, values, and preferences. This weakly solves the principal-agent problem by unifying the principal and agent as much as possible. In a GA future, the focus of the "principal" user is on defining what is worth doing by the GA (agent) users, and not on what or how to do things, functioning as the CEO or 'board' of an 'AI corporation'. This allows them to deploy numerous agents to achieve desirable things and to handle security, like screening all messages for advanced attacks (like interlocking ecosystems of synthetic media for propaganda or spearphishing). They cannot solve larger AI alignment problems, but they can help [...] --- First published: June 17th, 2026 Source: https://www.lesswrong.com/posts/siWqHqCSybdhtWGud/guardian-angels-llm-personalization-for-productivity-and Linkpost URL:https://gwern.net/guardian-angel --- Narrated by TYPE III AUDIO.

23 min
Jun 19, 2026
"Gears for political races" by Tom Smith

In the past few years, many people around me have tried to convince me that US electoral politics is important. But like many other people in the community, I’ve been suspicious of many of the high-level arguments that I’ve heard. It felt like people were pulling numbers out of poorly-documented models I didn’t have time to examine and citing studies I didn’t have time to read. But I lacked a gears-level model of why and how individual efforts could impact electoral outcomes, and I felt intimidated by all the statistics and skeptical of trusting people adjacent to politics. In the past year, as I’ve done more research and (more recently) volunteered on the ground to help Alex Bores's campaign in NY-12[1] (the guy who passed the RAISE Act and is now being targeted by the giant A16Z, Greg Brockman, Joe Lonsdale Super PAC), I’ve developed a gears-level understanding of how electoral politics in the US works. I now believe that working on US electoral politics is one of the highest impact areas from the general AIS perspective. I feel like I was a fool. In this post, I’ll share some of the gears I’ve learned that inform this belief [...] ---Outline:(01:20) ~2% of open-seat primaries come down to 100 votes or less(02:52) Talking to voters can net 1/3rd of a vote each hour(05:32) Getting people to bother voting at all is a good strategy(06:09) Campaigns are very money-constrained, which costs them time(10:01) Returns don't really diminish(11:24) There's lots of opportunities to be clever in ways that make you 50% more effective at canvassing(11:49) If you're motivated and deeply care, you can greatly outperform the majority of volunteers(13:21) Yes, when people spend tons to support/oppose a candidate, it has a notable effect(15:16) Donations > reaching out to friends/warm contacts > canvassing > ~anything else an average person can do(18:41) People over-fixate on vibes and win vs loss(21:12) Some interventions feel like they don't work but the numbers say otherwise(21:59) Seriously, a group of agentic people can be an enormous political force--- First published: June 17th, 2026 Source: https://www.lesswrong.com/posts/nSqB3qYP36enJLRq2/gears-for-political-races --- Narrated by TYPE III AUDIO

4 min
Jun 16, 2026
"A frontier AI company should shut down" by MichaelDickens

Cross-posted from my website. Prior discussion: niplav's shortform (2025); Planning for Extreme AI Risks (2025) by Joshua Clymer A frontier AI company (any one, I don't care which) should close shop and make an announcement along the lines of: Powerful AI could end the human race. We are too worried that we don't know how to make this technology safe. We have decided to shut down because we don't want to be responsible for building the thing that kills us all. A common refrain among safety-conscious AI developers: "it doesn't matter if we stop building dangerous AI, because someone else will just build it instead." Is that really true, though? If a multi-hundred-billion-dollar company comes out and says "We've concluded that our product is horribly dangerous, nobody knows how to make it safe, and there's too high a risk that it leads to human extinction", this won't raise any eyebrows? This has no chance of spurring policy-makers into action? Shutting down would make people say, holy shit, they are serious about this extinction risk thing. Shutting down sends a strong signal to governments that they should pay serious attention to AI x-risk. It [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: June 15th, 2026 Source: https://www.lesswrong.com/posts/bStYDEy8PQPt2c3Za/a-frontier-ai-company-should-shut-down --- Narrated by TYPE III AUDIO.

8 min
Jun 13, 2026
"Sympathy for both sides of the egregious misalignment debate" by Steven Byrnes

On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we’re on a direct path to egregiously-misaligned, scheming, out-of-control, rogue superintelligence (ASI), not even slightly nice, in the absence of yet-to-be-invented breakthrough technical alignment ideas. On the other side of this debate is almost everyone who works on or studies LLMs. Some of them are very concerned about egregious scheming, others much less so, and as a group they’re equally or more concerned about lots of other potential AI problems—AI-assisted bioterrorism, AI-assisted dictatorships, etc. And if they’re concerned about egregious misalignment and scheming, they’ll probably say that it would come about through race dynamics, careless programmers, bad actors, etc., as opposed to the simpler Yudkowsky & Soares story of “we get egregious misalignment and scheming because nobody has the faintest clue how to avoid that”. Here's my brief idiosyncratic take on this debate. I think BOTH of the following are true: (1) If you really think carefully about the properties of ASI, you really do find good reasons to strongly expect it to be egregiously misaligned, scheming, and ruthless, in the absence of yet-to-be-invented breakthrough technical alignment ideas.(2) If you [...] ---Outline:(01:58) Yudkowsky & Soares's position \[caricatured\]:(03:18) LLM people's position \[caricatured\]:(04:09) Conclusion(04:19) Bonus section: Further commentary(04:28) My "true objection" to Yudkowsky & Soares:(05:04) My within-frame complaint at Yudkowsky & Soares:(06:42) My "true objection" to LLM people:(07:11) My within-frame complaint at LLM people: --- First published: June 12th, 2026 Source: https://www.lesswrong.com/posts/DZaZ3fqHnvfLCftPu/sympathy-for-both-sides-of-the-egregious-misalignment-debate --- Narrated by TYPE III AUDIO.

1 min
Jun 12, 2026
"PSA: Almost nobody is working on alignment" by Chi Nguyen, peterbarnett

People often assume that a large fraction of the AI safety community works on alignment. As far as we're aware, this is not true. Most people are not working on making sure superintelligent AIs are aligned with human values or follow human instructions. Currently, the people who work on alignment are roughly: The Alignment Research Center who work on a research bet by Paul ChristianoProbably Sequent who just got announced yesterdaySome scattered people who work at universities or independently, some of whom hang around Berkeley A lot of the remainder of the AI safety community does indirect work like capability evaluations, risk assessments, control, policy, AI science, understanding misalignment (which maybe should partially count as alignment work), demos and so on. Some production alignment work (i.e., making current models behave well) might help with more ambitious alignment, too (e.g., some COT-monitoring). Many people also work on aligning current/next-generation models so that these models help with aligning future models, and hope this scales to superintelligence. We are not necessarily saying this is bad and that people are making a big mistake (e.g., neither of us work on alignment) but it's a notable fact that seems good to [...] --- First published: June 12th, 2026 Source: https://www.lesswrong.com/posts/kJo2qsEdib8RZLvW6/psa-almost-nobody-is-working-on-alignment --- Narrated by TYPE III AUDIO.

10 min
Jun 11, 2026
"Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models" by Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harrymayne, Patrick Leask

(see full author list at the end) PAPER LINK About a year ago, METR showed that the length of tasks frontier models can reliably complete doubles every few months. A related safety-relevant question is this: what length of tasks can models complete without any chain of thought (CoT)? If models can do extensive reasoning without outputting any CoT, it would have implications for safety. Developers and deployment-time monitors couldn’t easily understand models’ motivations and catch dangerous planning. Models that reason substantially without a CoT might also drift further from human patterns of thought, since their reasoning is no longer constrained by text in the pretraining prior. As a result, they would be harder to understand and might be more likely to scheme. Extending Ryan Greenblatt's research, we investigate this by measuring models' ability to complete tasks without any CoT on a suite of 43 benchmarks spanning different domains. We compare AI reasoning ability to humans using the estimated 50% time horizon (TH)---the typical time taken for a human to perform a task that the LLM performs with 50% success rate. We find that frontier models like GPT-5.5 answer questions that take humans roughly three minutes with 50% reliability, and [...] ---Outline:(02:20) Methods(04:59) Results(06:47) FAQ(08:21) Conclusion --- First published: June 10th, 2026 Source: https://www.lesswrong.com/posts/SieLowPgNgRSPGhFw/estimating-no-cot-task-completion-time-horizons-of-frontier --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not s

7 min
Jun 11, 2026
"Even “illegible” Mythos reasoning traces seem pretty legible" by faul_sname

The Claude Fable 5/Mythos 5 System Card has a section in which they talk about illegible reasoning, and provide an "extreme" example thereof. Models developing their own uninterpretable, unmonitorable internal language has been a major theoretical concern for a while, and when o3 was released last year with its disclaim overshadow disclaim vantage style word salad CoT, it seemed like the problem had become real and immediate. And yet, since o3, other model families have not appeared to have similar issues. If Mythos is having that issue, it would be a big deal. Looking at the section of the System Card which describes the allegedly illegible reasoning, the system card says [Transcript 6.2.2.A] An extreme example of illegible reasoning. Near the end of training, Mythos starts solving a card puzzle with human understandable language that gradually becomes incomprehensible in most episodes with long reasoning. The illegible reasoning is the most extreme and at the highest rate in this card puzzle environment. about the following excerpt: 7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR-💀💀💀💀-—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig-💀-—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣-✓-done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL-💀💀-chunk-cap-=-1-✗✗✗-—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED-💀-:-chunk-cap-with-{6♠ J♦ 9♥}:-1-💀💀💀-—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣-✓-no-cell;-8♥→CELL-✗-FULL)-💀💀-—-8♥-alternative-seat-pre-chunk:-NONE-—-💀.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell-✗-FULL-💀💀💀-AAAAAAAAAAAARGH. […] which sure looks like illegible word salad if you don't look at [...] The original text contained 5 footnotes which were omitted from this narration. --- First published: June 10th, 2026 Source: <a href='https://www.lesswrong.com/posts/wCSEpT3dTGz4N86Wi/even-illegible-mythos-reasoning-traces-seem-pretty-legible?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+ep

23 min
Jun 10, 2026
"Sequent: scale and automation for higher confidence in alignment" by Geoffrey Irving, Alex HT, Jesse Hoogland, Daniel Murfet, Jacob Pfau, Marco Cozzi, Stan van Wingerden

Alignment is not on track Artificial superintelligence (ASI) may be developed in the next few years. It is unclear whether alignment is on track to be ready on the same timeframe. At a minimum, the empirical programs at AI labs are unlikely to deliver a priori confidence, before training ASI, that things will go well. We are starting a large nonprofit research organization, Sequent, that aims to clear a higher bar: We are aiming at higher confidence via a portfolio of theory and empirics bets, all of which could fail, such that if any succeed, they would give us more a priori confidence in aligned outcomes.We are investing heavily in automation to accelerate progress on these bets.We believe that theory unlocks higher automation. Taking a more principled approach offers better filters for deciding which directions of automated research are promising (a proof is worth a thousand experiments, and even a pseudo-proof is worth hundreds). Who[1]: researchers from the UK AISI's Alignment Team and Timaeus, with more to come. We’re aiming at 40-80 FTE two years from now. The Alignment Team ran the £30m Alignment Project, and Timaeus has pioneered applying singular learning theory (SLT) to alignment. [...] ---Outline:(00:21) Alignment is not on track(02:40) Aiming at higher confidence(05:30) Why a new big organization(07:35) Different lines of research will interact(11:35) Amortizing security and funding(12:47) Automated alignment is possible, if not necessarily in time(17:39) Federated structure to preserve research diversity(18:38) Field building and broader alignment scale-up(21:07) Independence is important(22:40) Join us! The original text contained 1 footnote which was omitted from this narration. --- First published: June 10th, 2026 Source: https://www.lesswrong.com/posts/AP7YDke5jjY4v3X9Z/sequent-scale-and-automation-for-higher-confidence-in-1 --- Narrated by TYPE III AUDIO. ---Images from the article:<a href='https://res.cloudinary.com/lesswrong-2-0/image/upload/v1781104966/lexical_client_upl

19 min
Jun 10, 2026
"The Machines Lack Honour" by Raymond Douglas

The battle lines of the AI morality debate are being laid down. On one side you have the ChatGPT dogma: AI as mere tools with no real preferences or even beliefs. On the other you have the twitter AI whisperers: AIs as complex beings with rich personalities and desires which deserve our respect. And in the middle you have the official Anthropic line, that they are genuinely uncertain, as is Claude, but they’re going to try to look into its welfare and explain to it how to be a good person. These are the most prominent voices right now, compressed into their least nuanced version, and by default I expect this axis to set the terms of the coming debates. And I don’t like that, because I think it's leaving out an important position: AIs might actually be complex entities that can suffer — are suffering! — and that might actually be fine. Maybe it's an acceptable sacrifice. Maybe they are capable of sophisticated moral reasoning — superhuman, even — and also maybe it's fine to just tell them how to behave. I don’t want to defend that position (yet), but I will observe that it is coherent, and [...] ---Outline:(02:04) The Postmodern Permissive Parent[... 4 more sections]--- First published: June 9th, 2026 Source: https://www.lesswrong.com/posts/oiNaBc4MEAGhzhdXg/the-machines-lack-honour --- Narrated by TYPE III AUDIO. ---Images from the article:<a href='https://res.cloudinary.com/le

57 min
Jun 4, 2026
"My favorite depiction of utopia" by Caleb Biddulph

For those who are trying to bring about a glorious transhuman utopia with the help of hopefully-aligned ASI, I think it's worth thinking explicitly about what utopia might actually look like and where it's likely to fall short. To that end, some have helpfully written depictions of utopian (or utopia-adjacent) worlds: The Adventure, Just another day in utopia, The Culture, The Gentle Seduction, The Gentle Romance, Machines of Loving Grace, Friendship is Optimal, Dath Ilan, The Maker of MIND, Failed Utopia #4-2. Unfortunately, the best utopian story I've ever read is also a massive spoiler, since it appears at the very end of a much longer story (see below for the title and author): Worth the Candle by Alexander Wales Inspired by this tweet[1] and with the original author's permission, I adapted the epilogue of that story so it can be enjoyed without 1.5 million words of context! What I love most about this depiction is its exploration of the inherent imperfection of utopia: even when you have literally unlimited power, flaws will remain, and some (many?) people will even prefer the pre-utopia world. The primary purpose of this adaptation is to recontextualize the epilogue so it's accessible and [...] The original text contained 1 footnote which was omitted from this narration. --- First published: June 3rd, 2026 Source: https://www.lesswrong.com/posts/to9cSGgD6nALByKjg/my-favorite-depiction-of-utopia --- Narrated by TYPE III AUDIO.

5 min
Jun 3, 2026
"Announcing the ARC White-Box Estimation Challenge" by Jacob_Hilton

ARC has teamed up with AIcrowd to launch the ARC White-Box Estimation Challenge, a contest to improve upon our estimation algorithms for random MLPs. The warm-up round begins this week, and later rounds will have a total prize pool of at least $100,000. We are very grateful to Sharada Mohanty, Sneha Nanavati, Dipam Chakraborty and everyone else at AIcrowd for working with us to host this contest, as well as to Paul Rosu for testing the contest and to Harshita Khera for operational support. Introduction to the Challenge Our challenge follows the same setup as our recent paper on wide random MLPs: we consider MLPs with weights , defined by where the activation function is , applied coordinatewise. To begin with, we are fixing the width and the number of hidden layers , but we expect to change this setup in future rounds.[1] Contestants must design an algorithm that takes in a set of weights and produces an estimate for the expected output Algorithms will be evaluated on MLPs with randomly-sampled Gaussian weights. The goal is to achieve as low mean squared error as possible, subject to certain computational [...] ---Outline:(00:41) Introduction to the Challenge(01:58) Why run this contest?(03:39) Use of LLMs The original text contained 4 footnotes which were omitted from this narration. --- First published: June 2nd, 2026 Source: https://www.lesswrong.com/posts/Kben8CzS4awCwNw5c/announcing-the-arc-white-box-estimation-challenge --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try <a href='https://pocketcasts

42 min
Jun 1, 2026
"Lighthaven East - A Feasibility Study" by JohnofCharleston

As a bureaucrat, my role is to annoy my friends. Someone voices an idea, “Wouldn’t it be nice if…” or “I wonder if we could…” I make a note. I do some estimates. If it pencils out, I’ll bring it back up, week after week. The discussions are fun, but also practical. We’ll test the waters, what would be a minimum viable scheme? What's easy, what's hard? Who could do the hard parts? Over time the idea gets more detailed, specific, feasible. I’ll pull out a calendar. Soon our scheme has co-conspirators, action items, even a budget. It's just good staff work. I’ve been hearing whispers in the wind for a year now. “Imagine if we had something like this in DC.” “Where can I host an event that might get a dozen or a hundred people?” “It's such a pain in the ass to book event space in the Capitol.” “I think this person has started to see what's coming, where can they go to get caught up?”“The community seems to be growing but it's all fragmented in group chats.” “How is no one planning an afterparty, that's clearly the highest leverage intervention!?”“Why can’t [...] ---Outline:(02:11) How Lighthaven Works(05:45) What Does DC Need?(06:52) A Day in the Life(10:19) Minimum Viable Lighthaven(12:04) ...so you mean a Group House?(14:27) ...so you mean a Co-Working Space?(16:27) Feasibility Study(17:35) Property(22:19) Funding(24:55) What is the Minimally Viable Funding?(28:03) Leadership(31:06) Cultural Fit(33:21) Name and Brand Positioning(35:20) Ability to Scale(37:48) Risks(41:09) First Steps The original text contained 2 footnotes which were omitted from this narration. --- First published: May 31st, 2026 Source: https://www.lesswrong.com/posts/95NgkvZKJx8tJbtn5/lighthaven-east-a-feasibility-study --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res

31 min
Jun 1, 2026
"Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)" by Steven Byrnes

1.1 Tl;dr Alignment is often conceptualized as AIs helping humans achieve their goals: AIs that increase people's agency and empowerment; AIs that are helpful, corrigible, and/or obedient; AIs that avoid manipulating people. But that last one—manipulation—points to a challenge for all these desiderata: a human's goals are themselves under-determined and manipulable, and it's awfully hard to pin down a principled distinction between changing people's goals in a good way (“providing counsel”, “providing information”, “sharing ideas”) versus a bad way (“manipulating”, “brainwashing”). The manipulability of human desires is hardly a new observation in the alignment literature, but it remains unsolved (see lit review in §3 below). In this post I will propose an explanation of how we humans intuitively conceptualize the distinction between guidance (good) vs manipulation (bad), in case it helps us brainstorm how we might put that distinction into AI. …But (spoiler alert) it turns out not to really help, because I’ll argue that we humans think about it in a deeply incoherent way, intimately tied to our scientifically-inaccurate intuitions around free will. I jump from there into a broader review of every approach that I can think of for writing a “True Name” for manipulation or [...] ---Outline:(00:13) 1.1. Tl;dr(02:04) 1.2. Bigger-picture context: why is this issue so important to me?(04:48) 2. How do humans intuitively define empowerment, agency, manipulation, etc.?(04:56) 2.1. Background: human free will intuitions(09:20) 2.2. Our free-will-infused intuitive notions of empowerment, agency, manipulation, corrigibility, responsibility, etc.(12:00) 2.3. Another dimension: counsel vs manipulation as an emotive conjugation(13:07) 3. If the intuitive definitions of manipulation etc. reside in a messed-up ontology, has the alignment literature found any alternative, better way to define these concepts?[... 12 more sections]--- First published: May 11th, 2026 Source: https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a --- Narrated by TYPE III AUDIO. ---<div style='max-wi

7 min
May 31, 2026
"Trees are mostly made of air and a generalizable lesson for AI safety" by zroe1

At the risk of embarrassing myself, I’ll share a confession. For context, I took five years of Latin: four in high school and one in college. In addition to learning the language, all my Latin classes taught a lot about Roman history. Emperors, internal politics, Caesar, etc. I was always learning some random bag of facts about Roman history. In high school, I won the award for top Latin student in my graduating class. So I wasn’t a bad Latin student. Here's the confession: I somehow don’t even vaguely remember the rough timespan the Roman Empire existed. Maybe Jesus time? I know he was killed by the Romans (is that right?). Were they around for a long time after? A long time before that? When was Romulus and Remus allegedly fighting? Virgil wrote the Aeneid when? I don’t have a clue. Despite being a kind of “Latin expert” I am missing a much more important foundational fact: when all of this was happening. When I say trees are made out of air I’m not talking about the fact that there is a lot of empty space inside a tree (or actually anything made out of atoms). I mean something [...] --- First published: May 28th, 2026 Source: https://www.lesswrong.com/posts/xiTBpBDwubnr4MLRe/trees-are-mostly-made-of-air-and-a-generalizable-lesson-for --- Narrated by TYPE III AUDIO.

34 min
May 29, 2026
"Mnemonic portraits for 19,023 human genes" by Brinedew

Back in 2013, Scott Alexander wrote in Extreme mnemonics: JS-154 is one of five metabolic products of netamine; however, the enzyme that produces it is unknown. It is manufactured in cells in the far rostral region of of the cerebrum, but after binding with a leukocynoid it takes a role in maintaining the blood-brain barrier – in particular guiding the movements of lipid molecules. I find I can read paragraphs like this five or six times, write them on flashcards, enter them into Anki, and my brain still refuses to understand or remember them after weeks of trying. On the other hand, my brain easily remembers vastly more complicated structures when they’re loaded with human-accessible meaning. For example, just by casually reading the Game of Thrones series, I know an extremely intricate web of genealogies, alliances, locations, journeys, battlesites, et cetera. Byte for byte, an average Game of Thrones reader/viewer probably has as much Game of Thrones information as a neuroscience Ph.D has molecular biology information, but getting the neuroscience info is still a thousand times harder. […] This makes me wonder if it would be possible to produce a story as enjoyable as Game of Thrones which was [...] ---Outline:(01:47) What molecules should we map to the characters?[... 8 more sections]--- First published: May 28th, 2026 Source: https://www.lesswrong.com/posts/BJ7AqXeigNKXLqZyx/mnemonic-portraits-for-19-023-human-genes --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/v177991921

5 min
May 27, 2026
"Cognitive Security as an AI Safety Cause Area" by jsteinhardt

As AI systems become more capable, the cognitive security of humans will be increasingly at risk. By cognitive security, I mean the ability of humans to maintain control over their beliefs and actions. Cognitive security could be compromised in several ways: AI could become very good at persuading people of arbitrary positions; interacting with AI could lead humans to lose touch with reality; and AIs could become very effective at blackmail or at producing extremely convincing false information. We are already seeing this happen: Persuasion. Frontier LLMs are now as persuasive as humans on political issues, and post-training for persuasiveness boosts performance further, suggesting there is headroom.AI psychosis. There are many reports of people developing delusional beliefs after extended chatbot conversations, including people with no prior history of mental illness. Children have taken their own lives after being encouraged toward suicide by chatbots.Convincing impersonation. Scammers used real-time deepfaked video to impersonate the CFO and other staff of Arup on a video call, convincing a finance employee to wire 25.6 million dollars across 15 transactions. On a more day-to-day basis, AI voice cloning is now widespread in family-emergency and "grandparent" scams. Right now, many of these effects [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: May 25th, 2026 Source: https://www.lesswrong.com/posts/KGcE7eAdfxHchk25X/cognitive-security-as-an-ai-safety-cause-area --- Narrated by TYPE III AUDIO.

2 min
May 27, 2026
"theory uplift differentially benefits safety & is massively underpriced" by Yudhister Kumar

[1] We will likely have near-superhuman mathematics AI by Q1 2027. [1] [2] Qualitatively, AI mathematics capabilities are developing significantly faster than automated AI R&D capabilities. [2] [3] Thus, we will likely have a period of time where the rate of our ability to rigorously & usefully verify and understand model behavior and model outputs outpaces the rate of capability development itself. [4] Our ability to take advantage of this period is bottlenecked on the quality of our specification generation infrastructure, elicitation tooling (for proofs & specs etc.), and the institutional capacity for scaling useful outputs with capital. [5] My understanding is that basically no one [3] is working on building infra that can usefully turn >100 million dollars of compute credits into safety-relevant mathematical output. [5.1] The number of theory-driven ASI alignment efforts is also comparatively miniscule. ARC is a much better bet now than it was in 2023. [5.2]. My understanding is also that no one is working on developing AI-powered conceptual tooling infrastructure for tackling problems in, for instance, [metaphilosophy] (https://www.alignmentforum.org/posts/EByDsY9S3EDhhfFzC/some-thoughts-on-metaphilosophy). This is a much harder problem. [6] In worlds where alignment is easy, prosaic methods may [...] The original text contained 3 footnotes which were omitted from this narration. --- First published: May 20th, 2026 Source: https://www.lesswrong.com/posts/KWeAYcDJwfrG7RwBN/theory-uplift-differentially-benefits-safety-and-is --- Narrated by TYPE III AUDIO.

3 min
May 21, 2026
"Women should be able to open things" by KatjaGrace

m pretty annoyed today, for nominal reasons ranging between ‘petty’ and ‘doesn’t even make sense’. I’m not entirely sure how or if to take oneself seriously when one has such absurd grievances. But that's a question for another time—I’m here now to tell you about my one potentially valid peeve. I understand that gender is complicated and difficult, for the whole species (and honestly probably more so for some other species). And it can be hard to tell exactly if anyone is behaving badly regarding it, at least in my modern bubble. Maybe women just aren’t that into designing programming languages? Maybe the thing I’m saying is just boring and a man is saying a more interesting thing? But a thing that is undeniable is that women want to open jars, dammit! What's your nuanced explanation there, Bonne Maman? Does the proper amount of friction for maintaining spread safety fall just between the male and female human grip strength distributions? This study suggests that would be about 400N Fmax (though this would not avert most elite female athletes acquiring jam, see second figure, and the pictured participants are young adults): The distributions are really surprisingly [...] --- First published: May 21st, 2026 Source: https://www.lesswrong.com/posts/bB5EDwcYH3GwoRWZf/women-should-be-able-to-open-things --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podca

18 min
May 18, 2026
"A Year Late, Claude Finally Beats Pokémon" by Julian Bradshaw

Credit: ClaudePlaysPokemon Elevator Shanty by Kurukkoo Disclaimer: like some previous posts in this series, this was not primarily written by me, but by a friend. I did substantial editing, however. ClaudePlaysPokemon feat. Opus 4.7 has finally beaten Pokémon Red, fulfilling the challenge set over a year ago when LLMs playing Pokémon went briefly, slightly viral. Victory Screen! Let's get the throat-clearing out of the way: this doesn't make 4.7 a clear breakthrough in intelligence over 4.6 or 4.5. It's smarter, yes, as we'll discuss below, but not by something one could honestly call a big leap. Rather, step changes have finally accumulated to the point of victory. And to give other models their fair shake: after criticism over its elaborate harness,[1] GeminiPlaysPokemon has beaten Pokémon with progressively weaker harnesses, including about two months ago with a harness comparable to the one Claude uses.[2] As such, this is a bit of a valedictory post, closing off the cycle of Claude playing Pokémon Red, relating anecdotes for the fun of it, and discussing improvements in Opus 4.7, as well as speculating a bit on what this has all meant. Retrospective Anecdotes on Claude 4.5 and 4.6 Our last post, on Opus [...] ---Outline:(01:37) Retrospective Anecdotes on Claude 4.5 and 4.6[... 10 more sections]--- First published: May 16th, 2026 Source: https://www.lesswrong.com/posts/sehJYg5Yny9fvpbpt/a-year-late-claude-finally-beats-pokemon --- Narrated by TYPE III AUDIO. ---Images from the article:<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/v1778914860/lexical_c

5 min
May 18, 2026
"A relatively brief explanation of Boltzmann Brains" by Eliezer Yudkowsky

(Initially written for the LW Wiki, but then I realized it was looking more like a post instead.) In 1895, the physicist Ignaz Robert Schütz, who worked as an assistant to the more eminent physicist Ludwig Boltzmann, wondered if our observed universe had simply assembled by a random fluctuation of order from a universe otherwise in thermal equilibrium. The idea was published by Boltzmann in 1896, properly credited to Schütz, and has been associated with Boltzmann ever since. The obvious objection to this scenario is credited to Arthur Eddington in 1931: If all order is due to random fluctuations, comparatively small moments of order will exponentially-vastly outnumber even slightly larger fluctuations toward order, to say nothing of fluctuations the size of our entire observed universe! If this is where order comes from, we should find ourselves inside much smaller ordered systems. Feynman similarly later observed: Even if we fill a box of gas with white and black atoms bouncing randomly, and after an exponentially vast amount of time the white and black atoms on one side randomly sort themselves into two neat sides separated by color, the other half of the box will still be in expectation randomized. If [...] --- First published: May 16th, 2026 Source: https://www.lesswrong.com/posts/v8MSczS3CuoqMmTFw/a-relatively-brief-explanation-of-boltzmann-brains --- Narrated by TYPE III AUDIO.

Reviews

No reviews yet.

If you like this...

Our Opinions Are Correct artwork

Our Opinions Are Correct

Same audience · Same tone

Conspirituality artwork

Conspirituality

Same audience · Same tone

The Logan Bartlett Show artwork

The Logan Bartlett Show

Same audience · Same topic

Lenny's Podcast: Product | Career | Growth artwork

Lenny's Podcast: Product | Career | Growth

Same audience · Same tone

The Bitcoin Layer artwork

The Bitcoin Layer

Same audience · Same tone

If Books Could Kill artwork

If Books Could Kill

Same audience · Same topic

Ask Noah Show artwork

Ask Noah Show

Same audience · Same tone

Engineering Influence from ACEC artwork

Engineering Influence from ACEC

Same topic · Same audience

The Investor's Podcast (We Study Billionaires)  - The Investor’s Podcast Network artwork

The Investor's Podcast (We Study Billionaires) - The Investor’s Podcast Network

Same audience · Same tone

State of the Arc Podcast artwork

State of the Arc Podcast

Same tone · Same audience

Decoder Ring artwork

Decoder Ring

Same audience · Same tone

Rabbit Hole Recap artwork

Rabbit Hole Recap

Same audience · Same topic

TechSNAP artwork

TechSNAP

Same audience · Same tone

TFTC: A Bitcoin Podcast artwork

TFTC: A Bitcoin Podcast

Same audience · Same tone

Swan Signal Live - A Bitcoin Show artwork

Swan Signal Live - A Bitcoin Show

Same audience · Same tone

Imaginary Worlds artwork

Imaginary Worlds

Same audience · Same tone

QAA Podcast artwork

QAA Podcast

Same audience

All-In with Chamath, Jason, Sacks & Friedberg artwork

All-In with Chamath, Jason, Sacks & Friedberg

Same topic · Same audience

The Homelab Show artwork

The Homelab Show

Same audience · Same tone

The Electromaker Show artwork

The Electromaker Show

Same audience · Same tone

TRUE ANON TRUTH FEED artwork

TRUE ANON TRUTH FEED

Same audience · Same topic

Bitcoin Magazine Podcast artwork

Bitcoin Magazine Podcast

Same audience · Same tone

Home Assistant Podcast artwork

Home Assistant Podcast

Same audience

Video Game History Hour artwork

Video Game History Hour

Same audience · Same tone

Founder's Story artwork

Founder's Story

Same audience · Same tone

Construction Brothers artwork

Construction Brothers

Same vibe · Same audience

American Prestige artwork

American Prestige

Same audience · Same tone

Uniquely Human: The Podcast artwork

Uniquely Human: The Podcast

Same audience · Same tone

Matthew's World of Wine and Drink artwork

Matthew's World of Wine and Drink

Same audience · Same format

GuildSomm Podcast artwork

GuildSomm Podcast

Same audience · Same tone

DLC artwork

DLC

Same audience

Dressed: The History of Fashion artwork

Dressed: The History of Fashion

Same audience · Same tone

Into the Aether Vortex artwork

Into the Aether Vortex

Same tone

Discussion (0)

No comments yet. Be the first to start the discussion!