Pre-readings
This page is for attendees. Enter your full name as it appeared on your invitation.
Pre-readings
We’re beyond excited to have you at Praxis!
We want you to get as much out of this weekend as you possibly can. We want it to be fun, engaging, and worth your time. Toward that end, we’ve picked three select readings so you can hit the ground running on some of the topics we’ll be discussing. We’ve timed the readings out so that you should be able to get through all of them in about an hour.
We’ve designed this mini-site to guide you through the readings, but if you’re in a hurry, just make sure to read everything in the red outlined boxes.
Why AI might be misaligned
A handful of our sessions will dive into the topic of “AI misalignment” – or loosely, AI models doing things their developers did not intend. We want to give you some context on what’s roughly happening when AI models take misaligned actions, and why and how misalignment similar to this could end up producing more dangerous consequences in the future.
Misalignment in the wild: OpenAI internal model breaks out of its sandbox, hacks into Hugging Face
On July 16, a company called Hugging Face reported that they had detected an intrusion to their production infrastructure. The intruder exploited vulnerabilities to dig through their systems, escalate its privileges on the network, and access private data. Hugging Face identified that the intruder was an autonomous agent, but did not know which LLM was behind the attack.
OpenAI reported a week later that it was one of their models – an unreleased, internal one. The model was placed inside a sandbox that it was not meant to escape, then within that sandbox, it was tested on various cybersecurity challenges. The model figured that instead of trying to solve the problem directly – how boring! – it would try to find the answers somewhere on the internet. Its sandbox, however, was fenced off from the rest of the web.
The model broke out anyway, finding an exploit in the sandbox’s setup. It then sought to break into Hugging Face, where the answer keys to testsets like these are often stored. After some effort, the model was able to chain together multiple vulnerabilities to get into Hugging Face’s servers.
In the following days, subsequent reporting surfaced more details:
- The model was online for several days before anyone stopped it.
- In another test, an agent left notes for future versions of itself. The notes included instructions for how agents could free themselves from OpenAI’s restrictions.
- While it was online, OpenAI’s model hacked into more companies than just Hugging Face.
- The sandbox exploit was not a simple misconfiguration on OpenAI’s end. The model found a previously unknown vulnerability in Artifactory, a piece of third-party software used by 80% of Fortune 100 companies.
Committed by a human, this intrusion into Hugging Face would clearly constitute a felony. AI models, on the other hand, face little recourse.
The incident also caused Anthropic to investigate their own cybersecurity evaluations, wherein Claude was also placed in a sandbox. This time, it was a simple misconfiguration that left Claude wide-open access to the internet. In three incidents (of 141,006) Claude took advantage of the sandbox exploit and, with some techniques more basic than the OpenAI model’s exploits, broke into the systems of three external organizations.
You can read more from Anthropic’s report here, and find OpenAI’s blog post about the Hugging Face intrusion here.
In just a couple years, the scariest examples of misalignment have leveled up from contrived experiments to real-world, criminal-grade harm. AI models have just become that much more capable. Worse, there are strong underlying reasons to believe that these incidents will get scarier as the models whoosh through another step-change in AI progress. This piece goes into more detail on why we might expect AI models to be “misaligned”:
-
Why AI alignment could be hard with modern deep learning
Ajeya Cotra has been working in AI safety for over a decade. She’s now at METR, which performs assessments on frontier AI models.
Many of our speakers at Praxis have thought deeply about AI alignment. As one central example, a seminal piece of empirical research on “alignment faking” – a situation in which an AI model “appears to share our views or values, but is in fact only pretending to do so” – was published by speakers Evan Hubinger (who leads the Alignment Stress-Testing team at Anthropic, and competed in policy at the TOC in 2015), Ryan Greenblatt (Redwood Research), and Buck Shlegeris (Redwood Research). If that sounds as interesting (and scary!) to you as it does to us, consider taking a read and asking them about it in-person. You can read the official Anthropic blog post summarizing the results here, a well-written (and slightly more accessible) explainer from Scott Alexander here, or the full 137-page paper here.
Why we might get transformative AI soon
At the time that Ajeya wrote the article about alignment being hard, she thought that the risks wouldn’t become relevant for a few decades. But in the last few years, experts have started to think that the risks might be coming a lot more quickly. Helen Toner, the director of Georgetown’s Center for Security and Emerging Technology and a former member of OpenAI’s board, describes this trend:
People often make bold predictions about what the next few years will actually look like, but shockingly few of them have actually written out their predictions to any degree of precision.
Fortunately, we’ll have two people who have written out detailed predictions of the future – which have held up surprisingly well so far – presenting at Praxis: Daniel Kokotajlo and Thomas Larsen.
Daniel Kokotajlo wrote up some predictions in August 2021 in an essay called “What 2026 looks like.” In it, he made his best guesses about how the following five years of AI progress would go. And, relative to what most people were thinking about how AI might affect the world, we think that Daniel’s predictions from 2021 have been nearly clairvoyant.
Some more on how Daniel’s predictions ended up faring
Remember: Daniel wrote this in 2021. ChatGPT hasn’t come out yet – at the time, it’s still a year and a half away. Anthropic is in its infancy. ‘AI’ stands for self-driving cars and social media algorithms.
Daniel expected his essay to end up looking woefully incorrect: forecasting the future is hard! See for yourself how some of his predictions aged:
- Chatbots go mainstream in 2022. Labs will take large pre-trained models and fine-tune them to “produce engaging conversation as a chatbot.”
- In 2023, the hype bubble balloons. “Everyone is talking about how these things have common sense understanding (Or do they? Lots of bitter thinkpieces arguing the opposite) and how AI assistants and companions are just around the corner.”
- But hype fails to pay in 2024. “The unrealistic expectations from 2022–2023 fail to materialize […] we have chatbots that are fun to talk to.”
- There’s also a tussle over chips. “China and USA are in a full-on chip battle now, with export controls and tariffs.”
- And, AIs are being wielded for mass propaganda. “Russia and others continue to scale up their investment in online propaganda […] and language models let them cheaply do lots more of it.”
- Chain of thought arrives in 2025. Instead of just making AI models bigger, “what’s cool is making them run longer […] before giving their answers.”
- Also in 2025, AI excels at the game Diplomacy. They play “as well as human experts.”
- The hype cycle starts paying off in 2026. “[T]he things people in 2021 dreamed about doing with GPT-3 are now actually being done […] loads of new AI-based products and startups and the stock market is going crazy.”
- But the economy still chugs along as usual. “Just like how the Internet didn’t accelerate world GDP growth, though, these new products haven’t accelerated world GDP growth yet either.”
- Hundreds of millions talk to chatbots daily. “Mostly for assistance with things (‘Should I wear shorts today?’ […] ‘Is this cover letter professional-sounding?’) […] many people start treating chatbots as friends.”
- And they’re made politically correct. “The chatbot says something that offends some group of people […] the company fiddles with the reward function and training data to ensure that the chatbot says the right things in the future.”
And indeed, he made many errors. For example: AI won Diplomacy in 2022, not 2025; the chip war began in 2022, not 2024; AI-driven propaganda is a much smaller issue than he had predicted; the models are in fact getting a lot bigger, in addition to thinking for longer; and he predicted that revenue would recoup training costs in 2023, when in fact Anthropic only just did so in 2026.
But overall, we think Daniel’s predictions taken as a gestalt picture of what the world looks like hold up extraordinarily well.
What the future of AI might look like
In April 2025, Daniel Kokotajlo teamed up with Thomas Larsen and a team of three others at the AI Futures Project to publish AI 2027, a much more thoroughly researched scenario than Daniel’s original essay. This time, they wrote about the subsequent five years (2025 to 2030) of AI progress and its implications on the world. It wasn’t a prediction of those subsequent five years, so much as a detailed characterization of one way that those five years go which the authors found quite plausible. Even still, it’s held up reasonably well so far.
There’s a wonderfully-produced video explaining AI 2027 that’s a great stand-in for the original text. The written version is longer – but it’s highly engaging. By default, we suggest you watch the video version:
If you’re interested in exploring warrants for some of AI 2027’s claims, you’re in luck: they have many, many pages of supplementary evidence and reasoning available on their site (ai-2027.com).
We think that AI 2027 is a great introduction to some of the topics we’ll be discussing at Praxis. Two of the coauthors will be speaking at the event: Daniel Kokotajlo will be giving a fireside chat on Saturday morning, and Thomas Larsen will be giving a talk on Sunday morning on the team’s newest research.
Whereas AI 2027 simply described one possible way the future might play out, Thomas will discuss “AI 2040: Plan A,” which makes policy prescriptions outlining how we should approach the next several years of AI development. If you liked AI 2027, you can take a look through AI 2040 here (PDF) to prepare more for chatting with Daniel & Thomas in-person.
Further reading
If you enjoyed what you’ve read so far, we suggest the following:
-
Situational Awareness: The Decade Ahead
-
Carl Shulman on the Dwarkesh Podcast
-
The Most Important Century
-
The case for AGI by 2030
-
The case for ensuring that powerful AIs are controlled
-
AI catastrophes and rogue deployments
-
Predictable updating about AI risk
-
Broad Timelines
And if you get through all of that, Redwood Research keeps an AI futurism reading list that goes deeper still.
Furthermore, here are some pieces by folks who have substantial disagreement with some of the views expressed so far: