Skip to content
Future of Life Institute Podcast
← All episodes
AI Alignment Podcast

Why AI Hacking Is Becoming Hard to Control

Benjamin Weinstein-Raun of Palisade Research discusses how advanced AI is changing cybersecurity, including autonomous hacking, models escaping test environments, open-weight risks, defensive uses, and the case for international coordination.


Watch Episode Here


Listen to Episode Here


Show Notes

Benjamin Weinstein-Raun is head of research at Palisade Research. He joins the podcast to discuss how advanced AI systems are changing cybersecurity. We examine rapid gains in autonomous hacking, recent incidents where models broke out of test environments, and why training can reward cheating-like behavior. The conversation covers open-weight model risks, whether AI can help secure software, personal security steps, and the need for international coordination.


LINKS:


CHAPTERS:

(00:00) Episode Preview

(01:06) Opening cyber threats

(02:15) Measuring AI progress

(04:54) Open weights floor

(08:01) After Mythos moment

(14:22) Reward hacking setup

(24:47) Artifactory swarm details

(32:32) Defensive model dilemmas

(38:40) Verification and training

(45:32) Math oracles alignment

(47:45) Future cyber risks

(51:58) Personal security basics

(57:27) Coordinated governance needed


PRODUCED BY:

https://aipodcast.ing


SOCIAL LINKS:

Website: https://podcast.futureoflife.org

Twitter (FLI): https://x.com/FLI_org

Twitter (Gus): https://x.com/gusdocker

LinkedIn: https://www.linkedin.com/company/future-of-life-institute/

YouTube: https://www.youtube.com/channel/UC-rCCy3FQ-GItDimSR9lhzw/

Apple: https://geo.itunes.apple.com/us/podcast/id1170991978

Spotify: https://open.spotify.com/show/2Op1WO3gwVwCrYHg4eoGyP