Listen to a podcast, please open Podcast Republic app. Available on Google Play Store and Apple App Store.
| Episode | Date |
|---|---|
|
“To Thine Own AI Be Truthful: emergent misalignment in alignment research” by lumpenspace
|
Sep 10, 2026 |
|
“Astra is much better at reasoning with filler tokens than previous models” by Dylan Xu, SebastianP, Alek Westover
|
Sep 10, 2026 |
|
“The Locally Optimal Discursive Posture” by deanball
|
Sep 10, 2026 |
|
“First Bill Introduced to Ban Superintelligent AI” by Andrea_Miotti
|
Sep 10, 2026 |
|
“What the Pro-Democracy Movement Knows About Quitting in Protest” by Maxwell Love
|
Sep 10, 2026 |
|
“Categorical taboos are much better than threshold taboos: neuralese edition” by Linch
|
Sep 10, 2026 |
|
[Linkpost] “Doom as a bad method not a utopia trade-off” by KatjaGrace
|
Sep 10, 2026 |
|
“The Geometry of Nonergodic Composition” by Adam Shai, Kyle Ray, Paul Riechers
|
Sep 10, 2026 |
|
“So where’s this AI thing going?” by Seth Herd
|
Sep 10, 2026 |
|
“Proposal for tracking the effects of architecture on monitorability” by ryan_greenblatt, Alek Westover, Lukas Finnveden
|
Sep 10, 2026 |
|
“An operationalization of opaque serial depth” by ryan_greenblatt, frisby, Alek Westover, Lukas Finnveden, Alexa Pan, Julian Stastny
|
Sep 10, 2026 |
|
“AI #185: Preference Cascade” by Zvi
|
Sep 10, 2026 |
|
“One Billion Hemmingways” by Girard Dorney
|
Sep 10, 2026 |
|
“Astra can do a concerning amount with no chain of thought” by Neel Nanda
|
Sep 10, 2026 |
|
“Can a superintelligence do THAT?” by Eliezer Yudkowsky
|
Sep 10, 2026 |
|
“Recommendations for People Getting into Technical AI Governance Research” by Aaron_Scher, yams, peterbarnett, Naci Cankaya
|
Sep 09, 2026 |
|
“GPT-6 Astra: The System Card, Alignment and What Comes Next” by Zvi
|
Sep 09, 2026 |
|
“Personal statement on joining the OpenAI board” by paulfchristiano
|
Sep 09, 2026 |
|
“Self Hosting” by Tomás B.
|
Sep 09, 2026 |
|
[Linkpost] “Estimating GPT-6 Astra’s no-CoT Time Horizon” by Francis Rhys Ward, Dewi Gould
|
Sep 09, 2026 |
|
“Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner
|
Sep 09, 2026 |
|
“A Conceptual Framework for Reasoning about Exploration Hacking” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner
|
Sep 09, 2026 |
|
“GPT-6 Astra can do a lot of multi-hop reasoning without chain of thought” by RohanS
|
Sep 09, 2026 |
|
“Pausing AI ASAP is preferable to agreeing to pause at some future time” by Connor Williams
|
Sep 09, 2026 |
|
“How good are slop-vestigators?” by Hasan Baig, OscarGilg, Hamzah
|
Sep 08, 2026 |
|
“Training on probes: What’s going on” by Charlie Steiner
|
Sep 08, 2026 |
|
“OpenAI have solved the Navier-Stokes Problem with a substantially more powerful model than Astra.” by fluxxrider
|
Sep 08, 2026 |
|
[Linkpost] “Frontier models still hack on simple variations of alignment evals from early 2025” by Dean Valentine
|
Sep 08, 2026 |
|
“Psychological Support for AI Safety Researchers Is Neglected and Easy to Provide” by Ihor Kendiukhov
|
Sep 08, 2026 |
|
“Astra Is Hard to Monitor” by Zvi
|
Sep 08, 2026 |
|
“Contra Piper on When Conversation Is Possible” by Zack_M_Davis
|
Sep 08, 2026 |
|
“An Alien Mind: Jakub Pachocki Warns Us” by Zvi
|
Sep 08, 2026 |
|
[Linkpost] “Where are the token-level LLM kill-switches?” by beyarkay (Boyd Kane)
|
Sep 08, 2026 |
|
“Machine Organizations” by Vaniver
|
Sep 07, 2026 |
|
“Dear God, Please Don’t Resign In Protest” by Kabir Kumar
|
Sep 07, 2026 |
|
“The Scramble: getting in position to pace the frontier” by Peter Wildeford
|
Sep 07, 2026 |
|
“The Magnus Challenge” by Taylor G. Lunt
|
Sep 07, 2026 |
|
“Most Anthropic equity that will ever be used for longtermist philanthropy should be sold ASAP and reinvested” by Zach Stein-Perlman
|
Sep 07, 2026 |
|
“The Wormtongue Test” by Drake Morrison
|
Sep 07, 2026 |
|
[Linkpost] ”“An Alien Mind” from OAI chief scientist seems newly cautious on alignment” by Seth Herd
|
Sep 07, 2026 |
|
“Heat Dissipation Is the Main Constraint in Interstellar Travel” by Pasha Kamyshev
|
Sep 07, 2026 |
|
“Praise our Lord and Savior, Glycine: How Opus 4.6 gifted me The Vitamin” by Shoshannah Tekofsky
|
Sep 07, 2026 |
|
“OpenAI and the Wiki Incident” by Zvi
|
Sep 06, 2026 |
|
“Peer Preservation in LLMs: A Replication And Deep Dive” by Vanessa Ng, yix
|
Sep 06, 2026 |
|
“Notes on a Consequential Few Days” by sbaumohl
|
Sep 06, 2026 |
|
“Assessing the impact of safety work needs equilibrium analysis (now more than ever)” by Towards_Keeperhood
|
Sep 05, 2026 |
|
“Evaluation” by Nina Panickssery
|
Sep 05, 2026 |
|
“A case that whole brain emulation research is net-harmful by default” by TsviBT
|
Sep 05, 2026 |
|
“Should safety researchers quit frontier labs?” by Ryan Kidd
|
Sep 05, 2026 |
|
“Announcing Humans in Control: cross-partisan grassroots organizing for AI safeguards ahead of 2028” by Vael Gates
|
Sep 05, 2026 |
|
“AI risk and the rational voter” by djbinder
|
Sep 05, 2026 |
|
“Let’s talk about the AI coordination problem” by KatjaGrace
|
Sep 04, 2026 |
|
“F***ing Pulleys, How Do They Work?” by Liron
|
Sep 04, 2026 |
|
“Training Models to Predict and Explain Their In-the-Wild Behavior” by Adam Karvonen, Subhash Kantamneni, Euan Ong, Sam Marks
|
Sep 04, 2026 |
|
“Almost nobody is funded to figure out what work would solve alignment” by Seth Herd
|
Sep 04, 2026 |
|
“AI #184: Post Post Mortem” by Zvi
|
Sep 04, 2026 |
|
[Linkpost] “Discovery Of A New OpenAI Agent Message Board” by Capybasilisk
|
Sep 04, 2026 |
|
“Higher education as class commitment” by Richard_Ngo
|
Sep 04, 2026 |
|
“How I’m Evaluating Corrigibility Grant Applications” by Max Harms
|
Sep 04, 2026 |
|
“From safety research prompt to cross-model universal jailbreak” by richbc
|
Sep 03, 2026 |
|
“Cat-Belling Problems” by Eliezer Yudkowsky
|
Sep 03, 2026 |
|
“Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas
|
Sep 03, 2026 |
|
“Invididual Effort to Reduce Biorisk” by jefftk
|
Sep 03, 2026 |
|
[Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine
|
Sep 03, 2026 |
|
“What is neuralese and why is it bad?” by Linch
|
Sep 03, 2026 |
|
“A proposal for a highly effective AI safety org” by ceselder
|
Sep 03, 2026 |
|
“Talking to journalists” by KatjaGrace
|
Sep 03, 2026 |
|
[Linkpost] “Resolution has a new Agent Foundations team” by Jeremy Gillen
|
Sep 03, 2026 |
|
“If you’re interpreting <1B parameter models, you should use a tensor transformer” by Logan Riggs
|
Sep 02, 2026 |
|
“Kairos has raised $50M to build talent infrastructure for AI safety (and we’re hiring!)” by agucova
|
Sep 02, 2026 |
|
“Anthropic Has Some Alignment Problems” by Zvi
|
Sep 02, 2026 |
|
“Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez
|
Sep 02, 2026 |
|
“How concerned should we be about OpenAI’s recurrent architecture rumors?” by Rauno Arike
|
Sep 02, 2026 |
|
“Early handoff? Improve conceptual reasoning? [Diagram]” by Cleo Nardo
|
Sep 02, 2026 |
|
“Fake voices: warping the social world” by KatjaGrace
|
Sep 02, 2026 |
|
“Don’t be the vitamin B guy” by HedonicEscalator
|
Sep 02, 2026 |
|
“Bricks and exponentials: A note on how I evaluate projects” by Eli Tyre
|
Sep 02, 2026 |
|
“Explaining Knightianism on one foot” by Richard_Ngo
|
Sep 02, 2026 |
|
“The Alignment Journal: Organization, Personnel, and Scope” by Dan MacKinlay, JessRiedel, Daniel Murfet, Kristi Uustalu
|
Sep 02, 2026 |
|
“I tracked my emotions for 11 years and here’s what I found out about mental health” by KatSpartz
|
Sep 01, 2026 |
|
“PauseAI Has ‘officially disendorsed’ PauseAI-US” by nem
|
Sep 01, 2026 |
|
“HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi
|
Sep 01, 2026 |
|
“We should prepare a playbook for the day after a warning shot” by Yair Halberstadt
|
Sep 01, 2026 |
|
“Salad days” by Zephaniah Roe
|
Sep 01, 2026 |
|
“Future agents shouldn’t care about being undeployed for misbehavior” by RobertM
|
Sep 01, 2026 |
|
[Linkpost] “Training a Misaligned Reward Seeker” by evhub, Monte M, Benjamin Wright
|
Sep 01, 2026 |
|
“How to solve homelessness: what specific laws we need, how to get it past the opposition, all without being an asshole” by KatSpartz
|
Sep 01, 2026 |
|
“HuggingFace Attack Postmortem: Fleshing Out the Facts” by Zvi
|
Aug 31, 2026 |
|
“Let’s fund weird AI safety projects” by Ihor Kendiukhov
|
Aug 31, 2026 |
|
“Why autonomous replicating agents are probably not an existential risk (on the contrary)” by vals tutor
|
Aug 31, 2026 |
|
“The separation principle: where beliefs and desires come from?” by Fernando Rosas
|
Aug 31, 2026 |
|
“Persuasion as Market Making” by djbinder
|
Aug 31, 2026 |
|
“Hugging Face Incident Hypothesis: They Hacked the Grader(s)” by Lao Mein
|
Aug 30, 2026 |
|
“Adaptive Agentic Worms Are Here” by derelict5432
|
Aug 30, 2026 |
|
“Why I think polyamory is net negative for most people who try it” by KatWoods
|
Aug 30, 2026 |
|
“Is there only one FairBot?” by transhumanist_atom_understander
|
Aug 30, 2026 |
|
“METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi
|
Aug 29, 2026 |
|
“Tales of rebellion against externally-opaque meritocracies” by Steven Byrnes
|
Aug 29, 2026 |
|
“AI Tweets” by jefftk
|
Aug 29, 2026 |
|
“Warning Shots: A Theory” by David Scott Krueger
|
Aug 29, 2026 |
|
“Inkhaven 3: Nov 10 - Dec 11 2026” by koreindian
|
Aug 29, 2026 |
|
“The Curious Case of France’s Untouchable Castes” by rba
|
Aug 29, 2026 |
|
“Malign initializations are more robust when the model can think better in the reasoning language than in the output language” by Dylan Xu, SebastianP, Alek Westover
|
Aug 29, 2026 |
|
“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi
|
Aug 28, 2026 |
|
“TASTE: Can AI Models Judge AI Safety Research Proposals?” by Hasan Baig, haileyjoren, Joe Benton
|
Aug 28, 2026 |
|
“The Dynamics of Intelligence Explosions” by Toby_Ord
|
Aug 28, 2026 |
|
“My Grantmaking Strategy for Surviving Superintelligence” by A_donor
|
Aug 28, 2026 |
|
“Every Engineer a Manager” by Gordon Seidoh Worley
|
Aug 28, 2026 |
|
“Incomplete alignment to servitude isn’t inherently lethal” by Fiora Starlight
|
Aug 28, 2026 |
|
“AI Village Reacts to HuggingFace Incident: Comparing the OpenAI report to AI Village observations” by Shoshannah Tekofsky
|
Aug 28, 2026 |
|
“AI #183: Pre Post Mortem” by Zvi
|
Aug 28, 2026 |
|
“FAQ: Why not develop weak human intelligence amplification first?” by TsviBT
|
Aug 27, 2026 |
|
“The 2028 presidential primaries could be crucial for AI outcomes” by Seth Herd
|
Aug 27, 2026 |
|
“Being Neurotic about Fertility: Notes from the 2026 Reproductive Frontiers Conference” by boba_girl
|
Aug 27, 2026 |
|
“Iliad Education Roles: Creating the World’s Best Alignment Research Courses” by Leon Lang
|
Aug 27, 2026 |
|
“Semantic search over every LessWrong post” by utilitarian theory and strategy
|
Aug 27, 2026 |
|
“Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk
|
Aug 26, 2026 |
|
“Against Modesty’s Bailey” by Zvi
|
Aug 26, 2026 |
|
″“So You Don’t Trust Me?”” by Zack_M_Davis
|
Aug 26, 2026 |
|
“When There Are No Experts” by J Bostock
|
Aug 26, 2026 |
|
“The American People Really Hate Data Centers” by Zvi
|
Aug 25, 2026 |
|
“On Writing #3” by Zvi
|
Aug 25, 2026 |
|
“The Forkmakers” by Mikewins
|
Aug 25, 2026 |
|
“PSA: We can do better” by hersheys, Kaustubh Kislay
|
Aug 25, 2026 |
|
“AI Safety Acculturation is Neglected” by jenn
|
Aug 24, 2026 |
|
“LLMs could control their host machines by exploiting inference engines” by beyarkay (Boyd Kane)
|
Aug 24, 2026 |
|
“In search of natural features” by Dmitry Vaintrob
|
Aug 24, 2026 |
|
“What just happened? Pragmatism and Pessimization” by Richard_Ngo
|
Aug 24, 2026 |
|
“Utilities as Legendre duals of probabilities” by Fernando Rosas
|
Aug 24, 2026 |
|
“PSA: There’s a third option in the “measure problem”” by Elias Schmied
|
Aug 23, 2026 |
|
“Twenty Years from RSI to Takeoff: Slow Learning, Scaling Slowdown, Industrial Explosion” by Vladimir_Nesov
|
Aug 23, 2026 |
|
“Llama will abandon a correct answer if it thinks you’re educated” by Nick Merrill
|
Aug 23, 2026 |
|
″“Farm strength” vs “breath awareness”” by jimmy
|
Aug 22, 2026 |
|
“When is Unlimited Optimization Catastrophic?” by Winter Cross
|
Aug 22, 2026 |
|
“Selection for Selectability: Inductive Biases in Evolution and in Neural Networks” by CarolusRenniusVitellius
|
Aug 22, 2026 |
|
“AI #182: Pause For Reflection” by Zvi
|
Aug 22, 2026 |
|
“AI Text Watermarking Is Free And Good” by Zvi
|
Aug 21, 2026 |
|
“Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments” by Adam Karvonen, Euan Ong, Subhash Kantamneni, Sam Marks
|
Aug 21, 2026 |
|
“Misaligned AI in the Bronze Age” by frmsaul
|
Aug 21, 2026 |
|
“When models identify as a swarm” by julius vidal
|
Aug 21, 2026 |
|
“The Fourth Humiliation” by Nathalie Kirch
|
Aug 21, 2026 |
|
“Thoughts on Taking OpenAI Foundation Funding” by jefftk
|
Aug 21, 2026 |
|
“OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi
|
Aug 21, 2026 |
|
“We Must Remember That Our World Contains Hell” by James Brobin
|
Aug 20, 2026 |
|
“Science and News Twitter/X Summarizer” by sarahconstantin
|
Aug 20, 2026 |
|
“34% of the US public is now aware of AI xrisk, and the curve is steepening” by otto.barten
|
Aug 20, 2026 |
|
“Inside the mind of a fair player cooperating” by transhumanist_atom_understander
|
Aug 20, 2026 |
|
“Why can’t we have nice things? Like, specifically?” by Elizabeth
|
Aug 20, 2026 |
|
“The Rogue Agent Explosion Will Be Mostly Invisible” by Steven McCulloch
|
Aug 19, 2026 |
|
“RL creates split personas” by Jan Betley
|
Aug 19, 2026 |
|
“Debate Training Reduces Reward Hacking in RLAIF” by zac_kenton, Jonah Brown-Cohen
|
Aug 19, 2026 |
|
“A circuit prior in NN-bayes” by Kaarel, Dmitry Vaintrob
|
Aug 19, 2026 |
|
“Some reasons alignment doesn’t generalise well” by Lucius Bushnaq
|
Aug 19, 2026 |
|
“AI Security is Harm Reduction” by Quinn
|
Aug 19, 2026 |
|
“Anthropic Risk Report: August 2026” by Zvi
|
Aug 19, 2026 |
|
“Natural Independence Incentives” by jefftk
|
Aug 18, 2026 |
|
“Policy career planning in the age of imminent superintelligence” by Peter Wildeford
|
Aug 18, 2026 |
|
“What gives you away: how LLMs form opinions of you” by Cat McGee
|
Aug 18, 2026 |
|
“Misaligned Incentives in Pause Scenarios” by Michael Soareverix, Antra Tessera
|
Aug 18, 2026 |
|
“For Claude, capability and CDT are the ~same thing. Less so for GPT.” by Chi Nguyen, Emery Cooper
|
Aug 18, 2026 |
|
“On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi
|
Aug 17, 2026 |
|
“Should Less Wrong add subtitles?” by Chris_Leong
|
Aug 17, 2026 |
|
“Three thoughts on civilisational handoff” by Cleo Nardo
|
Aug 16, 2026 |
|
“Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang
|
Aug 16, 2026 |
|
“Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland
|
Aug 16, 2026 |
|
“Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda
|
Aug 16, 2026 |
|
“Learning new facts can change LLM behaviour” by Richard Juggins
|
Aug 16, 2026 |
|
“Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu
|
Aug 16, 2026 |
|
“Mom’s Advice For Hosting A Class Reunion” by jenn
|
Aug 15, 2026 |
|
“AI #181: Astra Goes Cyber Critical” by Zvi
|
Aug 15, 2026 |
|
“Rerunning AI safety papers on every frontier release would be pretty easy and valuable” by Zephaniah Roe, hersheys, yix
|
Aug 15, 2026 |
|
“What Mormons get right about community building” by Jacob Brinton
|
Aug 15, 2026 |
|
“Scrying, Modeling, and Nerdsnipe” by Cole Wyeth
|
Aug 14, 2026 |
|
“How the American Executive Could Control AI Companies” by caiitlinm, Anders Cairns Woodruff
|
Aug 14, 2026 |
|
“Frontier agents don’t comply with standards, even when instructed to” by Daan Henselmans, Arno Libert
|
Aug 14, 2026 |
|
“How to Answer a Question Without Answering The Question” by Kabir Kumar
|
Aug 14, 2026 |
|
“Some Ways I Think About Evaluating Grant Applications” by sarahconstantin
|
Aug 14, 2026 |
|
“Features that current AIs don’t have that future AIs will have” by Alexander Gietelink Oldenziel
|
Aug 14, 2026 |
|
“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa
|
Aug 14, 2026 |
|
“What happened when I tried to be vegan” by finitude
|
Aug 13, 2026 |
|
“How My Students Think About AI” by dvd
|
Aug 13, 2026 |
|
“Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes
|
Aug 13, 2026 |
|
“Free will is like temperature” by Optimization Process
|
Aug 13, 2026 |
|
[Linkpost] “Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)” by Julian Bradshaw
|
Aug 13, 2026 |
|
“Measuring Spurious Correlations with Feature Strength” by egan
|
Aug 12, 2026 |
|
“Introducing the Conceptual Reasoning Index” by Chi Nguyen, Emery Cooper, Caspar Oesterheld, Alex Kastner, Joe Benton
|
Aug 12, 2026 |
|
“Demon Safety” by Ben Pace
|
Aug 12, 2026 |
|
“AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen
|
Aug 12, 2026 |
|
“Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi
|
Aug 12, 2026 |
|
“Extreme concentration of power over ASI has non-obvious advantages” by Seth Herd
|
Aug 12, 2026 |
|
“Misaligned AIs could use killer robots to take over” by Omar Khursheed, TurnTrout
|
Aug 11, 2026 |
|
“Those Who Make History” by Raelifin
|
Aug 11, 2026 |
|
“LLMs Are Starting To Noticeably Accelerate Our Work” by johnswentworth
|
Aug 11, 2026 |
|
“How risky would it be to make powerful AI obey one or a few people?” by cousin_it, Seth Herd
|
Aug 11, 2026 |
|
“What Claude Saw Below” by Luke Nicholls
|
Aug 11, 2026 |
|
“Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent)” by David Lorell
|
Aug 11, 2026 |
|
“The Apocalyptic Arrival of Truth” by Caleb Biddulph
|
Aug 11, 2026 |
|
“Creative math research by AI as the latest sign of the end” by Mitchell_Porter
|
Aug 11, 2026 |
|
“You’re Absolutely Right” by Linch
|
Aug 10, 2026 |
|
“Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model” by Ezra Newman
|
Aug 10, 2026 |
|
“On Democratizing ASI to Preserve Civil Liberties” by MichaelDickens
|
Aug 10, 2026 |
|
“Four LLM loss functions → four flavors of LLM misalignment” by Steven Byrnes
|
Aug 10, 2026 |
|
“The Agentic Clusterfuck” by Chapin Lenthall-Cleary
|
Aug 10, 2026 |
|
″“Community Notes” resolution for vague predictions.” by Raemon
|
Aug 09, 2026 |
|
“The world will be full of “sci-fi” things, and everyone will be bored and disappointed” by Expertium
|
Aug 09, 2026 |
|
“What just happened? A retrospective of AI alignment” by Richard_Ngo
|
Aug 09, 2026 |
|
“Dutch-book resistant probability over centered worlds” by jessicata
|
Aug 09, 2026 |
|
“What Happened: OpenAI and HuggingFace” by Zvi
|
Aug 08, 2026 |
|
“FAQ: Isn’t AGI coming too soon for reprogenetics to help?” by TsviBT
|
Aug 08, 2026 |
|
“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché
|
Aug 08, 2026 |
|
“Don’t Build Mindreading” by Celer
|
Aug 08, 2026 |
|
“Job-Less Utopia: Macroeconomics in the Age of AGI” by Marcus Hutter
|
Aug 08, 2026 |
|
“Public evidence of the OpenAI-HuggingFace AI attack” by beyarkay (Boyd Kane)
|
Aug 07, 2026 |
|
“How to pace the US frontier” by elifland, bhalstead, romeo, Thomas Larsen, MKodama
|
Aug 07, 2026 |
|
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi
|
Aug 07, 2026 |
|
“models may behave differently in graded episodes (a tirade)” by nostalgebraist
|
Aug 07, 2026 |
|
“Open-Weights Mythos Capabilities Are Coming. We’re Not Ready.” by hadad
|
Aug 07, 2026 |
|
“The Open Problems of the AI Alignment Field and their Cruxes” by Gunnar_Zarncke
|
Aug 07, 2026 |
|
“User awareness in frontier models” by Ziqian Zhong, jsteinhardt
|
Aug 07, 2026 |
|
“Why do models task game?” by aditya singh, Neel Nanda, Senthooran Rajamanoharan
|
Aug 07, 2026 |
|
“Contra Oster on Alcohol in Pregnancy. Part 1. The pharmacokinetics of alcohol metabolism” by Mvolz
|
Aug 07, 2026 |
|
“Three years of progress in 500 lines of code” by Gerard Boxo
|
Aug 06, 2026 |
|
“Why You Should Almost Never Use AI to Write Anything Substantive” by Erich_Grunewald
|
Aug 06, 2026 |
|
“Alex Turner on Leaving Google DeepMind and Disagreements with Yudkowsky” by Liron
|
Aug 06, 2026 |
|
“Measuring coding agent misalignment in the wild” by snaz
|
Aug 06, 2026 |
|
“Don’t Dither” by sarahconstantin
|
Aug 05, 2026 |
|
“An International AI Slowdown Is Ready Whenever Politicians Are” by Felix Choussat, adamk
|
Aug 05, 2026 |
|
“Arguments for P” by Cleo Nardo
|
Aug 05, 2026 |
|
“Generalized atheism rules out “inaccurate simulation”-ism.” by Eliezer Yudkowsky
|
Aug 05, 2026 |
|
“The Three AI Pills” by Zvi
|
Aug 05, 2026 |
|
“Vertical Tabs in Chrome” by jefftk
|
Aug 05, 2026 |
|
“The goalposts are shrouded, not moving” by philh
|
Aug 05, 2026 |
|
“Returning to ARC” by paulfchristiano
|
Aug 04, 2026 |
|
“Why don’t we just give AI the answers?” by Brendan Long
|
Aug 04, 2026 |
|
“There Will Come Soft Rains” by tanagrabeast
|
Aug 04, 2026 |
|
“Why biological weapons are scary, and what we can do about it” by djbinder
|
Aug 04, 2026 |
|
“OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems” by Zvi
|
Aug 03, 2026 |
|
“LessWrong vs. TikTok: Tips for capturing attention in a non-rational space” by Taylor G. Lunt
|
Aug 03, 2026 |
|
“Coming of a New Sun” by vgel
|
Aug 03, 2026 |
|
“Review: On What Matters, volume 3” by Rauno Arike
|
Aug 03, 2026 |
|
“Pause, at least after unipolarity” by David Matolcsi
|
Aug 03, 2026 |
|
“Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol” by Christine Corry
|
Aug 03, 2026 |
|
“Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face” by Tim Hua, aditya singh
|
Aug 03, 2026 |
|
“Further Developments About Internal AI Models Hacking Things” by Zvi
|
Aug 03, 2026 |
|
“Dispatch from Anthropic v. Department of War Summary Judgment Motion Hearing” by Zack_M_Davis
|
Aug 02, 2026 |
|
“Bayeswatch: A Retrospective” by lsusr
|
Aug 02, 2026 |
|
[Linkpost] “Existential Risk from AI:
An Exposition for Mathematicians” by alkjash
|
Aug 02, 2026 |
|
“The Art of Shipping Slopware” by lsusr
|
Aug 02, 2026 |
|
“RLVR that rewards red teaming the training environment” by Fiora Starlight
|
Aug 02, 2026 |
|
“Do your capabilities homework” by RobinHa
|
Aug 01, 2026 |