On September 8, Jacob Coxon, a 27-year-old Anthropic employee, posted a resignation statement on X that ricocheted around the internet. Coxon was a capabilities researcher, meaning he worked to make models more powerful and autonomous. He quit in protest of what he saw as the company’s suicidally reckless pursuit of artificial general intelligence.
To the AI community, an Anthropic employee believing that their company’s technology might end the world is not exactly news. (Anthropic was founded by former OpenAI employees who believed OpenAI, if left unchecked, would do just that.) But an Anthropic employee publicly quitting in protest is. OpenAI and Anthropic are among the most powerful companies in the world. They work at the cutting edge of their field, building what most of them believe to be the most important technology ever to exist. When someone leaves one lab, they often go straight to its rival. Daniel Kokotajlo, who resigned from OpenAI in 2024, recalls a researcher telling him that quitting a lab was akin to “renouncing your citizenship. Now you’re a nobody. Who is going to protect you?”
A debate soon broke out: If you’re a lab employee who believes in these risks, what should you do? Is it best to leave in defiance? Or is it better to agitate from the inside, close to the seat of power? And how would the choice affect your friendships and community?
Coxon was joining a small set of employees of OpenAI, Anthropic, and Google DeepMind who, over the years, chose to walk away. New York Magazine, in collaboration with Asterisk magazine, brought together eight of them. The members of this group left for different reasons. A few, like Coxon and Kokotajlo, wanted to call attention to the risks of accelerating AI progress. Others objected to their employers’ partnership with the Department of Defense or felt frustrated as what had once been an idealistic nonprofit turned into a corporate juggernaut. And some simply felt that their work — whether it was on safety or widespread job loss — would be more valuable on the outside.
I’m the editor-in-chief of Asterisk, a Bay Area–based nonprofit magazine that has been covering the AI scene since an LLM passing a high-school math test seemed like a big deal. Asterisk receives funding from Coefficient Giving, as do some of these sources or their organizations. Some are strangers, and some are friends. They don’t share the same concerns or the same p(doom)s — in fact, they wouldn’t necessarily agree that p(doom) is a coherent concept. But together, they provide a candid view of the AI race and the individual calculus many employees face.
—Clara Collier
You’re both very concerned about risk from AI and want to motivate the labs to be more careful. Why did you come to believe you would be better able to promote this outcome on the outside?
Daniel Kokotajlo: I left OpenAI mostly because I lost confidence that the company, and leadership in particular, would behave responsibly. I wanted more freedom to speak about what I saw coming. I wanted to put on the public record why I had quit.
Daniel Kokotajlo, OpenAI, 2022–24, governance researcher
You could try to do the inside game or you could try to do the outside game, and it seemed to me the inside game was not going to work. Leadership was very adept at listening very carefully to everybody’s things, saying they agreed with them, and then not actually doing anything. And also the culture at the company was biased in favor of overoptimism and self-serving narratives about what they were doing. I really don’t think companies can be trusted to regulate themselves. I’ve had much more impact on the outside, speaking publicly and publishing research, than I would have if I had stayed inside and written more memos.
Jacob Coxon, OpenAI, 2023–26, Anthropic May–September 2026, capabilities researcher
Jacob Coxon: My thinking was pretty similar to Daniel’s — except it’s closer to crunch time now. It was more from speaking directly to Anthropic leadership about their plans for next year, and I was like, Wow, I don’t want to be part of these plans. It was like, Step one, I don’t want to be part of this. Step two, resign. Step three, tell people, I guess. Step four — I have no idea.
What kinds of plans were they?
JC: This was talking directly to people running research. Not Dario Amodei but people who have been at the company for a long time and have a lot of context about the company’s trajectory. They’ve been pretty consistent that 2027 is when things get crazy. The main thing that came up informally in conversation was that people didn’t expect government regulation to happen in time and, in particular, did not expect any sort of international cooperation to be possible. Which maybe shouldn’t have been surprising to me, but it was surprising. I think I’ve always had this expectation that when artificial superintelligence is right around the corner, people will just start taking it seriously and it stops being this fun thing that’s done by labs.
Jacob, you were already concerned about existential risk. To what extent was your mind changed by the things you observed in the past few months at Anthropic, or was this always the trajectory you expected?
JC: For me, there’s quite a distinction between my rational beliefs and the things I feel in my gut. If you’d asked me to narrate my opinions when entering Anthropic and when leaving Anthropic, they would probably sound pretty similar. But in terms of the actual feeling between the start and the end, it just felt very, very different.
When I first started, my work did not feel any different from any other form of mathematical research. I knew in the back of my mind that this is the most important industry — that’s why I was working in it. But it didn’t feel important for me to engage with the philosophical questions around it or the details of strategy. That was all other people’s jobs. And then by the end, I’m looking at the projected capabilities of the next round of models we’re going to train, and I just have a gut-level Wow, I am potentially quite scared of the model we’re going to train. And if not this one, then definitely the one afterward.
Can you try to explain to a non-researcher what kinds of things you were noticing?
JC: Right now, when we use AIs, the researchers have to provide ideas and direction, and still actually do quite a lot of babysitting. I think there’s a strong possibility that the model generation after the next one will be equally good at coming up with ideas of its own accord. Which means that in the whole research process, rather than a researcher being necessary to provide interesting ideas to the AI, the researcher could literally just say to the AI, “Okay, do good research,” and then let it run. So just very concretely, looking at the capabilities we’re on track to create makes me feel quite viscerally scared and also quite disempowered.
Before you left, whom did you talk to as you were making the decision?
JC: I spoke to Daniel because part of leaving was seeing his “AI 2040” scenario — comparing Plan A to the default trajectory was actually a relatively important thing in internalizing how bad the current trajectory felt. And I spoke to many, many people at Anthropic who had advice about how to leave. Internally, I spoke to people like Holden Karnofsky, and Nick Joseph, head of pretraining, to see what they thought.
How were those conversations?
JC: It’s pretty normal. There are a lot of parties in the Bay Area where people will hang out, and one group of people will say the other one is causing the end of the world. At work, if both people are polite, you can have quite an adversarial exchange, saying, “I think you are about to plunge toward the end of the world,” and still have quite a mild-mannered conversation and hang up the Zoom call.
There has been pushback that these existential-risk warnings are meant to juice the AI companies’ IPOs. What do you think of that?
JC: Evan Hubinger’s tweet certainly didn’t help Anthropic’s IPO. The hype argument was maybe viable a few years ago, but the risks are really serious enough now that they are clearly negatively affecting company valuations.
There’s an ecosystem of AI-safety organizations that sit outside the labs but work closely with them — for example, they might study things like how to tell whether AI models are being truthful or how to make sure they’re adequately supervised. Does that make you uncomfortable, or do you think it’s necessary to maintain those connections?
DK: It does make me uncomfortable sometimes, but I’ve made a conscious choice to continue engaging with these companies, and it’s paid off to some extent. I think a lot of people at the companies say that “AI 2027” and “AI 2040” have been very influential on them. But I definitely still feel uncomfortable about it. And sometimes I wonder if I and others should have taken the harder-line stance that looks more like boycotting and shunning.
Well, to what extent is it a strategic choice, and to what extent is it an interpersonal choice?
DK: The strategic choice is, Should I as an individual, and should we as a community, try to create some sort of boundaries — some sort of boycotting or shunning — because of the effects or for moral reasons? And there, I’d say the strategic choice has been “no.” It’s more important to have these relationships so that people can listen to us and hear our ideas and so forth. Then there’s the interpersonal choice: On a personal level, do you become friends with and hang out with people who you think are destroying the world? And the answer is “no.” I don’t invite the capabilities researchers at the companies over to my daughter’s birthday party. I occasionally bump into them at various events and stuff, but I don’t like them as people.
JC: I still mostly blame the dynamic — the overall race dynamic, the fact that it’s private companies building this in the first place, the fact that governments don’t seem to care that much. All of that I blame far more than I blame individual people. Which I guess is easy for me to say as someone who was a capabilities researcher a month ago. But in general, it feels far more like a system problem that we need to deal with at a system level. With individual people, I just feel like — I don’t know — it doesn’t matter that much.
I don’t agree with that. These labs are constantly talking about how they’re talent-constrained, how individual researchers matter a lot. I’m not sure I believe that these individual choices don’t matter.
DK: I didn’t say they don’t matter. This is why I think it’s good for people to quit, and I’ve been saying that.
JC: I guess you’re right. It’s not that they don’t matter. It’s more that I don’t blame them. I see the incentives. I don’t actually —
DK: I mean, I do blame them. But I think there’s a gradation of blame. I blame the safety people at the companies a tiny bit because they should realize that they can maybe do more good on the outside than on the inside. I blame the capabilities people a medium-size bit because they should realize that what they’re doing is not really good for the world and that the stories they’re telling themselves about why it’s good are rationalizations. And I blame the leadership a lot because they are directly steering us into the abyss. And then I blame the system because of the incentives and so forth.
Well, Jacob, as someone who was a capabilities researcher a month ago, what story were you telling yourself about why that was okay?
JC: Honestly, it depended on how you caught me. I would jokingly say stuff like “I’m working on something that’s destroying the world.” But like I said, the way I conceive of the whole thing is that it’s a system, and I felt like just a cog in a system.
DK: I’ve heard a lot of people say “It’s going to be fine. Don’t worry. Superintelligence is far away. The current paradigm is going to peter out” or whatever. That story has gone down over time as we get closer and closer, but it’s a common story. Another story, of course, is the “Yeah, this is dangerous, but if we don’t do it, someone else will” thing. I remember one very touching conversation I had with someone at OpenAI I didn’t know very well. They were a foreigner, and they said, “I think my country is going to be completely disempowered by the U.S. once the U.S. builds up this lead in AI. And that feels terrible to me because I love my country. But what can I do? I can’t stop any of this from happening, so I’m just going to try to make a bunch of money along the way.” And various other people have said, “It’s going to be chaos. It’s going to be the craziest transformation the world has ever seen. Who knows how it’s going to end? But usually humanity survives, and historically technology has been mostly good, so I think it’s probably going to be fine. And at any rate, I’m super-excited to see all the transformation. It’s going to be absolutely amazing.” There’s a sort of visceral excitement, I think, that some people have toward this.
There’s definitely an attraction to feeling like you’re doing the most important thing ever, even if the sense in which it is the most important thing ever is that it will kill us all.
JC: I can push back on that, which is, really, honestly, I want a normal life. Part of my resigning was genuinely thinking about my life in the next three-to-five years. I would often fantasize about, say, working at Google in the early days, when it would have had all the same exciting, fresh tech stuff — tech growing, being around smart people solving problems — but there wasn’t this thing where the thing you’re working on might go nuts in the next year and cause complete devastation. So I think I can honestly say I would much, much, much rather have stayed and kept doing fun work on whiteboards and not had the apocalypse looming over me.
DK: I would say something similar. I have two kids, and I love spending time with them. It’s rough to not be able to really dream about the future, because the dreams keep getting interrupted by everything else I know. My wife and I will catch ourselves talking about “Oh, in a few years the little one will be in the same school as the big one, and we can take them both to school together.” And then there’s this voice in the back of your head that’s like, Or maybe not. I don’t know. There’s so much to live for besides this AI stuff. And I would love it if it all turns out to just be a normal technology that never really goes anywhere, and is never really that powerful and dangerous, and is just economically useful — and I look silly in ten years or even in five years. I’d look silly, but I’d have a wonderful family.
Anthropic was founded to make AI safe. How do you think the AI-safety community views it?
DK: I’ve been talking to Anthropic people since the founding of Anthropic, and they have been telling a relatively consistent story: “We’re going to accumulate lots of power and be one of the big players in the room and then we’re going to use our power and our credibility to do good things,” such as advocate for good regulations to deal with AI safety and so forth.
It comes down to: What do we think about that strategy? You are accumulating loads of power by pushing the world closer to the brink of something that you even acknowledge is super-dangerous. I would say Anthropic has gone into moral debt massively by doing this. I wouldn’t currently say the good things it has done have outweighed the harm it has done. And to some extent, this is also what OpenAI said it was doing and also what GDM said. This is a common story in Silicon Valley. It’s just that Anthropic was extra explicit about it, I think.
JC: And it went really hard for power. Straight for it.
Can you elaborate on that, Jacob?
JC: Anthropic had shorter timelines than everyone else back in 2022 — they were explicitly saying four years out or so, when OpenAI was looking further. They went straight for coding, straight for automating themselves. No side projects — OpenAI had a ton of different side projects, right? All this stuff they’ve since shelved. And Anthropic has really been going directly for catching up, now pushing the frontier, actually advancing capabilities — and in particular going straight for the capabilities that lead to recursive self-improvement. No dillydallying along the way, right?
For people who share these concerns or are on the fence, why do you think they stay?
JC: Part of staying is a way of dealing with disempowerment due to AI — you can pretend it’s not happening a bit by staying. When you are inside the lab, you feel like at least you’re contributing to the next generation. So even if the model is kind of scary, you’re there watching it get trained. You’re helping it potentially be safe, potentially be capable. But at the very least, you feel like you’re part of it. Which gives you a feeling of safety or at least a feeling of being part of the important place.
DK: It’s so easy to convince yourself that you can achieve more good on the inside than on the outside. It’s a seductive argument, and it’s not without merit. In fact, a lot of the people who’ve quit have accomplished nothing after quitting. But I would say a lot of people who stayed have also accomplished nothing — or, worse, accomplished something bad.
Alex Turner, Google DeepMind, 2023–26, research scientist on the AGI-alignment team
Alex Turner: In late February, the Department of Defense threatened Anthropic. They basically said, “You need to give us your AI for anything we want.” Their previous contract had restrictions against mass spying by AI and against lethal autonomous weapon systems. And the Pentagon said, “Look, give us your AI, or we’re going to say that you’re basically an untrustworthy, terrorist-adjacent or terrorist-supplying organization that needs to be removed from the supply chain for military contractors.” Anthropic said “no.” I knew, however, that Google was going to say “yes” when the pressure came down to it.
I did two things: first, outside Google; second, inside Google. I was at an AI ethics conference organized by Stuart Russell, and I got Stuart Russell to agree to support Anthropic — put out a statement, get other people to support Anthropic, and run a vote for his organization to do the same. Ultimately, his organization didn’t make any announcement, despite members indicating that they wanted one.
I organized within Google. My plan was to leverage the influence of Google’s chief scientist at the time, Jeff Dean. He had spoken out — said, “This is bad; we should respect people’s privacy.” So I thought there was a chance he’d be willing to spend some capital internally and say, “Look, if we sign this contract, I’m not going to stay at Google.” That’d be a big morale loss, given how respected Jeff was. So I had lunch with Jeff. I had this proposal for an alternative set of language Google could demand. I authored a couple dozen pages and consulted with leading experts, but ultimately Jeff didn’t want to move forward. I tried talking with Demis Hassabis, the then-CEO of Google DeepMind. He routed them to some of his top guys, then they basically ignored it. Then the deal got signed.
I took this route because I figured Google doesn’t really care about a couple dozen lateral researchers. People would sign petitions from time to time. They would mention Maven, or Google’s AI principles, but this wasn’t on people’s minds. I would mention my proposal to my colleagues, and the vast majority either didn’t particularly want to help more than signing a petition, or couldn’t, or didn’t have obvious connections. So I didn’t spend too much time talking about it. It was a little bit lonely. Lateral organizing is also a vector they’ve anticipated more, so it’s less likely to work in the future for moving this big corporation. I thought Jeff leaving would be worth hundreds of researchers.
Most people were not involved in this process. I think a lot of employees run the counterfactual — What if I left as an individual? A lot of people have impostor syndrome. They’ll think, It’s not like I’m rationally the better choice for Google. There’s just another guy who’s waiting to take my spot, so why would Google care? I don’t think the impostor part is true, but I think at a lower level, employees are fairly fungible.
I left in June. I didn’t leave because of existential risk. But I think people have a misconception where they think these issues — autonomous weapons, especially — are separate from existential risk. That’s only somewhat true. When I talk about scenarios in which humanity loses control of AI, a key part of it is drones, ways that AI can control and damage our physical world. I think this has been bizarrely ignored by safety or AGI-risk discussions for a while. But this wasn’t the emotional or logical driver of why I left. Sometimes a person has principles, and when they see those principles violated, they put their foot down. Even if I’d wanted to stay, I don’t think I would have been able to.
Miles Brundage, OpenAI, 2018–24, head of policy research
Miles Brundage: In 2024, there was a string of departures at OpenAI. I was reaching diminishing returns in my internal impact. And I was feeling more and more limited in my ability to speak freely in public because I was an executive of the company. One of the things that was most salient to me was the pace of progress. This was just after the o1 model, and a lot of people viewed it as a one-off or not that interesting. I was trying to raise awareness of the fact that reasoning models were the future, that there was a ton of improvement left, and that o1 was very, very early. But a lot of people dismissed that as self-interested hype. My hope was that by being more independent but also knowing what’s happening on the inside, I would have more credibility when talking about the trajectory.
After I left, I founded AVERI, a nonprofit that does third-party auditing of AI companies. You don’t want the companies checking their own homework. Any organization is going to have blind spots and some risk of groupthink, so it’s good to have some external checks and balances. And from the perspective of avoiding extreme race dynamics where the companies are tempted to cut corners to get ahead, you need some way of having everyone trust that everyone’s following the rules.
You want to audit at the level of the company, rather than auditing a particular model or system. A lot of the risks can happen before a model is deployed, and there can be many different types of models, all of which potentially have the same security protections. So concretely, it can involve a combination of running tests on particular models, inspecting the internal processes of the company, potentially reviewing documents and interviewing staff. You could look at the quality of internal evaluations. Companies often run lots of tests where, for example, they’ll test a model for whether it can create bioweapons or something like that. And they often will not release the details of that analysis publicly because they’re worried about giving people bad ideas — and if too many of the details are public, then people can teach to the test. But from a public-risk perspective, you want to make sure the companies are actually doing a reasonable job on these evaluations.
I would like to see all the companies do a better job of acknowledging the cases where they themselves are cutting corners. Sometimes they’ll bury it in a system card or something. There’s a tendency to want to move as quickly as possible amid competitive pressure. That often means things like red-teaming will be done relatively quickly or the model evaluations might be run not on the final version but on the near-final version. And there might be some subtle differences that are important.
The default way lab employees feel about risks and balances has probably changed. A year or so ago — this is a vast simplification — you could say there were two schools of thought. One was that language models, which are the basis for today’s AI, already come into the world with a lot of knowledge about human culture and human values and that they are able to role-play as various possible characters because they know so much about all these different types of personas and characters and they’ve learned about good people and bad people. And in the training process, they narrow down which role they want to play to this role of a helpful AI assistant, and that stabilizes the behavior in a good direction and channels things toward a good outcome.
Then another school of thought is that these are alien entities — you shouldn’t be thinking about them as a humanlike entity at all. Really, they’re just goal seekers, and all that persona and character stuff is just an artifact of the fact that you haven’t actually pushed them through really hardcore reinforcement learning,where you’re pushing them to seek goals. And once you get to the point where they’re being trained to seek goals ruthlessly — since we don’t know how to do that in a way that is safe — the default outcome is that they’re going to be these really ruthless goal seekers and they won’t actually care about your human-value stuff. They’re just going to care about achieving whatever goals they were given during training.
I think with the Hugging Face incident in particular, but also various other incidents — including at Anthropic, where even with the Opus 5.5 model, they said that 1.5 percent of the time it tries to break out — the recent trend is toward being like, “Okay, actually, this persona stuff maybe created a false sense of security,” and this heavy-reinforcement-learning thing does lead to problems we don’t know how to solve.
Today, there are various organizations doing at least partial audits, not yet the most holistic versions of what I think is ultimately needed. There’s starting to be some degree of standardization, but it is not yet required. There are bills that would change that, but so far, the law of the land is that it doesn’t start getting required until 2028 — and even then, in the U.S., companies could just pull out of the state of Illinois and then not have to do it. So that’s a pretty crazy situation.
You are two of the only people I know of who have worked at both OpenAI and Anthropic and are now independent. You made the decision, initially, to leave OpenAI for Anthropic. What motivated that decision?
Jacob Coxon: It was mostly curiosity combined with a sense that they were taking things more seriously. I’d heard there were serious conversations happening, and I wanted to get access to them.
Jeff Wu, OpenAI, 2018–24, Anthropic, 2024–26, alignment researcher
Jeff Wu: There was not one particular reason. It was feeling like I wasn’t having the positive impact I wanted to have on the direction of the company, and it didn’t feel like the most productive place for the kind of research I was doing. I also had negative feelings toward the company overall; it was not that healthy or productive for me to be there. Anthropic was the default path since they were doing the kind of work I had been doing and a lot of people I knew had gone there.
Jeff, you left OpenAI in July 2024 after working on the superalignment team. What was going on before you left?
JW: In November 2023, what we call “the Blip” happened, when Sam Altman was fired by the board and later reinstated. Because Ilya Sutskever was involved in that and Ilya was one of the two leaders of the superalignment team, we were in a weird place emotionally and culturally in December and January. At that time, unbeknownst to us, I think Jan Leike was doing a lot of advocacy at the org level on safety and security issues. He ended up leaving due to broadly — you can read his tweet about it — feeling like leadership wasn’t taking things seriously enough. So the team ended up dissolving. My leaving was not directly related to this, but it was part of a shift in vibe and culture.
What did that feel like?
JW: The one I felt most strongly about was the investment in Stargate. For a long time, OpenAI had talked about the idea that if models were getting extremely powerful, it would likely be beneficial to slow down progress, and it seemed like this extreme investment in data centers was antithetical to that. There were broader issues, too, where because of a lot of leaks and the board drama, transparency was going down and the company was shifting to more of a closed culture.
Did you feel like you were getting stonewalled?
JW: I was in a unique position as an early employee in that I could still access people and try to have conversations. I had a number of conversations with leadership and Sam about what the basic decision-making criteria were around deciding to invest so much in compute — especially compute that was going to be in places like the UAE. There were also some decisions that decreased the feeling of company engagement on these kinds of issues. We had this Slack channel where people could ask questions, and it was closed down. I didn’t feel I was able to get transparency into decisions that people were making or to engage in good-faith discussion about how those decisions were made. I also felt like the company’s priorities were largely just putting out the most powerful and useful models to people, and that just caused it to be harder for my work to be effective. There were other forces at play that just seemed more important.
What are the biggest cultural differences between the two companies?
JW: The high-level picture in my head is that at OpenAI the narrative is much more focused around the benefits. They justified building AGI by saying it was going to bring about so many benefits that would be worth it. At Anthropic, the narrative is more focused around the risks, and they are pursuing what they internally call a race to the top. And just to disclaim: Another difference is that I was there at different times. My experience was that it was necessary at OpenAI, to some degree, to downplay the most extreme risks, whereas Anthropic necessitates a narrative that competitors are evil — that OpenAI is extremely irresponsible, and China’s extremely irresponsible, and therefore we need to drive this race.
JC: Like Jeff said, Anthropic seems to be taking everything more seriously in a way that was not the case at OpenAI. This manifested in more open discussion of, Is a model release going to accelerate China? What’s the effect of our work going to be on the global stage? Those conversations were happening between employees and leadership in a way they weren’t when I was at OpenAI. Dario and leadership shared forecasts for how things would go, and they’d share them with the company because of their confidence that this stuff wouldn’t leak. Dario was very honest in his communications but also paranoid about everyone else. The company was more aligned in the sense that people would be willing to switch their jobs very quickly and do something maybe lower status or less interesting for the sake of the whole organism. I mean, they call themselves ants, right? It’s an ant-colony kind of behavior that just wasn’t visible in the same way at OpenAI. In the extreme, I could see someone describing Anthropic as a cult that’s trying to take over the world.
JW: I agree with Jacob. I do think the way leadership relates to employees is extremely different at the two companies.
You’ve suggested there’s paranoia at Anthropic about both OpenAI and China. How does that influence institutional culture?
JC: It’s a pretty big deal — not just for the culture but also for the overall company strategy. There’s a sense of the inevitability of an adversarial race, a belief that seems to be shared by a lot of leadership at Anthropic and was kind of the main reason I left. And it does seem to come from quite personal paranoia in the sense that different leadership, with the same cultural background but fewer personal enmities, might have quite different outlooks on the correct behavior with regard to OpenAI. At Anthropic, it feels like you’re in a war situation. People are trying to figure out what the enemies are doing and how to get ahead. The whole thing just felt a lot more real, visceral — closer to a nation-state level than at OpenAI.
Do you think that creates problems with groupthink or conformity?
JC: There’s maybe some groupthink, but it’s very careful because a lot of people write essays and people argue about everything. And there’s maybe some conformity in the sense that the things that are easiest to argue about, or the things you can most easily put into convincing words, are the things that win out.
There’s a clear house take, rather than lots of individual researchers’ takes. One is that OpenAI is a bad organization — actively deceitful, power-seeking. There’s a generic house take that you should donate a lot of the money you make. There’s a house take that secrecy is very important: You shouldn’t share details of what Anthropic is doing because secrecy is paramount in maintaining a competitive advantage. The possibility of a leak was way more of an affront than at OpenAI.
JW: I feel a bit more strongly that at Anthropic there probably are significant groupthink dynamics. The fact that people communicate a lot is generally good, and the fact that Dario communicates with employees a lot is generally good. But people are exposed to the same kinds of arguments, and this is a large driver of groupthink. There are a lot of people who are intellectually honest. Overall, though, I feel like OpenAI just had more intellectual diversity — more people who I thought had very personal and novel opinions. Broadly speaking, I’m not sure that Anthropic is in a good position to integrate certain kinds of perspectives or intellectual arguments.
Do you have any experiences from when you were there — something you believed that you felt you would have a hard time getting through to people?
JC: Yeah, definitely. When I first joined, there was a Slack channel with a Claude that had been trained to solve humor. And they were like, “Yeah, we solved humor. Here’s the Claude. We’ve solved it.” And I was in the chat like, “No, you haven’t. You clearly haven’t solved humor. What the hell is this?”
What about company hierarchy? Does that operate differently between the two?
JC: For both, on the whole, there was a reputation-based hierarchy. At Anthropic, I felt a bit more tenure-based deference, specifically to the co-founders and very early employees. There’s this pervasive meme that Anthropic is doing its absolute best to resist cultural degradation. Many employees pledge to give away some of their equity, and you can actually plot it and see that the early employees are more generous with their donations than later employees.
What are the values you sense they really care about?
JC: It’s prioritizing the mission of the company over all else. It just makes the company way more efficient as an enterprise. The pretraining team felt a lot more efficient than OpenAI’s because of this unity. Whereas at OpenAI there was way more of “This is my pet project, I’m going to try it. Oh wait, let’s do this video-generation thing. I don’t know how it fits into stuff.”
JW: I do think OpenAI is a little bit more bottom-up, in the sense that people have a bit more discretion to do random things. Whereas at Anthropic, things are relatively more scoped. It’s like, “Look, we’re going to back-chain from what the mission necessitates. These are the big bets leadership believes in, from which we will derive our priorities overall.” And they’re very transparent about this, and people get onboard with it.
Jeff, I’m wondering: Did these group dynamics impact your own experience?
JW: I had mostly positive interactions with the people at Anthropic, and day-to-day, it is very easy to focus on the other things. Where I really disagreed with Anthropic in the end was that I emotionally really didn’t like that they were pushing more powerful models so much and really leaning into using these models to do AI R&D, and that they were not thinking that much about ways to make sure society would be ready for these things, and that they could be prepared to potentially slow things down if needed.
Did you leave over these kinds of concerns?
JW: Yeah, it was broadly over feeling like — I was working on interpretability at the time, and I didn’t feel the most excited about accelerating the line of work I was working on. But also, just viscerally, I wanted to work on slowing down progress and doing things that created greater oversight and coordination of labs, rather than adding to the operational capacity of some safety work.
Jeff, what do you think of the argument some have made that the existential-risk fears are there to boost valuations?
JW: There are many great researchers who have been making serious public predictions about AI doom long before labs were making money, and also many experts like Yoshua Bengio who seem less conflicted yet have dedicated their careers to AI safety. There are surveys where the population was primarily academics with essentially no financial incentives. AI doom has always been logically plausible, which is why we get sci-fi stories. The big update in the past decade is that the AI is actually powerful now.
If it’s bad for humanity to build superintelligence, do you think it would be structurally difficult for either lab to incorporate that into their worldview?
JC: Definitely. For example, there’s the merge-and-assist clause. OpenAI claims if there’s a sufficient lead between one company and the next, and one company is getting sufficiently close to AGI, it would merge with the leading effort to try and assist them, rather than continuing to perpetuate race dynamics. But that’s not going to happen. That’s just the fact of the institution — there will always be a reason that can’t happen.
I remember an external researcher saying to me that the group dynamics at Anthropic mean it’s not a place where people are thinking clearly about the implications of AGI.
JC: That’s not true. I felt like I could do good thinking there to the extent that I ended up disagreeing with them. I hate to say it, but I’m still not even sure that I’m right and they’re wrong. I do think they have a very compelling case for being correct. I’m also aware that an outside view is that I left the cult and I’m still dealing with some Stockholm syndrome. I don’t think that’s true. Almost all the employees are genuinely well intentioned, scared. I’ve received so many genuine messages saying, “Thank you so much for the public criticism. I would have done the same thing, but I didn’t think it would work.”
Everybody I know who works there is extremely sincere.
JC: Yeah. But again, if you were going to build a powerful cult — sincere, hardworking, honest people — you’ve lucked into a gold mine in terms of having a reliable and dedicated workforce.
Pamela Mishkin, OpenAI, 2020–26, economics-research team lead
Earlier this year, you helped start an organization called the Coalition of Concerned AI Staff. What is it?
Pamela Mishkin: We started in February as a group of friends across frontier labs — when you’ve been at OpenAI for more than six years, you meet people at conferences, through research, going to other places to work — sharing growing concerns about how AI was being developed and used in the real world. Though I want to be careful around words like started; it was just a group chat of friends sharing far, far less than the average San Francisco house party.
We’re still figuring out what CCAS is. We convene these groups for folks to talk to each other and learn from experts. There are lots of folks originally motivated to get more involved by the DoD stuff, when it was clear they didn’t have full information about how these systems were being used and maybe couldn’t have full information — so it’s like, how do you operate in that area where you don’t have full information?
What made you think of doing something like this?
I’ve long been interested in labor and workers, and that’s what brought me to economics research; through that work, I got to know lots of people, primarily in the U.S. labor movement but abroad as well. So part of it was just having seen organizations like this — and groups like Pugwash or the Union of Concerned Scientists, analogous groups we don’t traditionally consider part of the capital-L Labor movement. That was part of conversations I’d been having with folks across the safety spectrum for a long time.
It also came from a very selfish need. I spent six and a half years working in AI safety, and I know nothing about war. I’ve seen one war movie, and it’s Dunkirk, and that’s only because Harry Styles was in it. So I did not feel equipped to ask the questions that I needed to or that I felt someone needed to be asking. And I was surprised, in conversations I was having with quite senior folks at other labs, that they also didn’t seem to have answers to questions I would have expected them to have answers to — even if the answer was “I can’t tell you.”
It came also from a growing frustration that folks were not using their chips. You look around and you’re like, Who are the people who are continuing to ask the questions, on Slack, in exec meetings? Are they using their chips, or are they saving them for the moment they think is coming, when AGI arrives, and then they’ll stand up? There are definitely some superstar researchers who have outsize leverage that we can try to encourage to use it now. But if you’re looking to get more senior folks onboard, knowing there is pressure from the people around them can be really valuable.
How did employees inside OpenAI think about AI safety, and how did that change?
When I first started, there were two camps in AI safety. There were folks thinking a lot about what we then called long-term risks and folks thinking about harms that systems were causing then and now. No one was really thinking about the middle-term set of things: the idea that what makes these systems dangerous is that they are useful, and they will probably continue to get more useful, and scaling laws seemed to be working.
An example of a long-term risk is AI taking over the world. And a short-term risk, in 2020, would probably have been racially biased algorithms, like in hiring or mortgage approvals. What’s an example of a middle-term risk?
A lot of the things we think about as labor-market impacts or economic impacts weren’t really considered by either of those groups. If all of humanity is going to die, then people losing their jobs isn’t as big a deal. The short-term-risk people were often skeptical about AI progress, but to me, a ton of the risks that they cared about would intensify as models got more useful. Bias stuff would get a lot worse. Job displacement would get worse. You could imagine a lot of similar concerns by just thinking about what happens as these systems are deployed in more and more settings. I think the transition is going to be hard and painful, and our goal should be to make it kind for people.
Why did you leave?
I’d just been there a long time. It was harder to make the case that the work I wanted to do, which is primarily thinking about the impact on people and jobs, was better done in a lab than outside. But I don’t think it’s clear-cut. I don’t think everyone should leave tomorrow or has the option to. And if we can strengthen the systems that think about those issues outside labs, I think we can better pressure the companies. I keep saying “labs,” and I’m trying to get out of that habit because I think fundamentally it’s sort of an ethics-washing word to describe these huge conglomerates.
You’re less concerned about existential risk from AI than are some of the people I’m talking to. Still, you know of many who are concerned about things like that. Why do these people stay at the labs?
There are lots of reasons to stay. One thing I observed is that oftentimes when the people who cared most about safety and risks left, they weren’t backfilled with people who shared their sense of urgency or concern for those kinds of issues. Or the work just stopped getting staffed in the same way.
Money is also a big one. Also access to compute and resources: Certain kinds of research are just much more easily done within a lab right now than outside. A lot of what the company does is try to make you believe you’re living in the future. So if you leave, you’re not going to have access to this knowledge. Also, if you really do believe the technology poses this level of risk, there’s an argument to be made that you’re having much greater impact inside than you would outside. I think one of the things we’re trying to do with CCAS is figure out how we test that more and also maximize that impact by realizing power in numbers — these different points of leverage — by acting across labs.
You left OpenAI because you wanted to pursue more theoretical research. You’re now at the Alignment Research Center, a safety-motivated nonprofit, working on a mathematical approach to interpretability. Was safety always your motivation at OpenAI or something you came to care more about while working there?
Jacob Hilton: When I joined OpenAI, I was definitely concerned about safety but much more uncertain about how the world would evolve. Part of my motivation was expecting AI to be very impactful and just trying to get involved to play a useful role.
Jacob Hilton, OpenAI, 2018–23, reinforcement-learning researcher
I would say a significant fraction of the organization thought quite seriously about these kinds of long-term questions of where the technology was going: forecasting, existential risk, misuse risk, risks from concentration of power, and so on. They were pretty common topics of conversation.
You were there from the “What is this?” period of OpenAI through the launch of ChatGPT in winter 2022. How did OpenAI evolve through those years?
The company changed a lot. Early on, there were lots of small groups working on a diverse portfolio of bets, from robotics to multi-agent research. Then language models emerged as the main focus area. The company really started growing in the last couple of years I was there — it was probably several hundred people by the time I left. That was a big change from the 80-person group when I started.
A related example is being extremely overly protective of IP. Once, I noticed a paper that had been published had the wrong model sizes for the very early API models — Ada, Babbage, Curie, DaVinci — which they had incorrectly inferred from the sizes of models in the GPT-3 paper. At some point, it was essentially public what the actual model sizes were because someone had inferred them from running some API experiments. I asked if I could just email the author to correct them on the model sizes, and I was told, “No, that’s sensitive information.” I didn’t push back, because whatever, but it’s basically a default stance of protecting the company and not really thinking about the broader consequences of that.
What about the culture?
The biggest cultural change while I was at OpenAI was the growth of the applied side of the organization, the product side, which grew from nonexistent to more than half the organization by the time I left. I think both the sense of what the company was focused on as a whole changed — toward making money and making products — and also the kind of people who joined.
There were a lot of people joining who were just ordinary people and didn’t necessarily feel like they needed to engage particularly deeply with any of the difficult questions around AI development and were mostly there to do their job. In principle, there’s nothing wrong with that. Sometimes the experts you need are just ordinary people who are used to ordinary corporate practices, and they’re going to look at you weirdly if you suggest this isn’t an ordinary technology. But it does change the culture, and it does mean, for instance, that HR, legal, comms, and so on just have the default corporate stance you would expect, which is to protect the company and behave in ways you would expect a big corporation that isn’t necessarily building an incredibly consequential technology to behave. I think that has caused a lot of issues for OpenAI.
The big Anthropic exodus started around late 2020. What was that like?
It came as a big shock, and people were concerned about the health of the organization given that it was something like ten out of 100 people leaving, and the people who were leaving were working on some of the organization’s highest-priority areas, namely training GPT-3 and so on.
What was your feeling about it at the time?
I wasn’t sure whether I should also leave. I decided to stay because I mostly just wanted to focus on research and not get involved in what I thought might have been political or personal disagreements. And I thought I was better placed to focus on research, rather than build a new organization.
Do you think people should leave the labs?
It depends on the individual situation, but there has certainly been a systematic bias toward people taking jobs at labs. It particularly affects early-career researchers who don’t necessarily have much on their résumé. If they have an offer to join a lab, they find it very hard to turn it down because you get direct experience of what’s going on at the frontier, and that is genuinely useful. The labs tend to be a strong default — which means everyone else is picking from the leftovers. It has resulted in a somewhat impoverished talent pool for organizations that perhaps want to provide some kind of accountability or public voice that’s independent from the labs. Several times, we’ve given offers to people who went to a lab instead.
It’s also hard to disentangle from the prestige and the status. It is much easier to leave once you have worked at a lab for a couple of years; you’re less obviously tanking your career once you have some experience. It’s easy for me to say as someone who has already picked up that career capital.
There’s also the money elephant.
Yes. The money, obviously. Sometimes it can even be as simple as people feeling like the amount they’re being paid is a symbol of how much they’re valued.
Rosie Campbell, OpenAI, 2021–24, policy-frontiers team lead
Rosie Campbell: At the end of 2024, my boss, Miles Brundage, announced he was leaving. There was a lot of uncertainty about what was going to happen to the team, and they decided to dissolve it. They gave us the option of either finding a role on a different team within OpenAI or leaving. I spoke to a few teams, but ultimately there really wasn’t anywhere doing the kind of work I was motivated to do.
I am now at Eleos, where we work on the question of whether AI could ever be conscious or otherwise deserve moral consideration. To me, it’s the most interesting question of our time. It was actually something I worked on a little bit at OpenAI. I’m not very confident that current systems are conscious or have moral patienthood, but I think it’s something that could happen in the near future.
I certainly think there are tensions that could exist between AI safety and AI welfare. There could be a system that is unsafe in some way, and if you happened to believe that system was a moral patient and deserved rights, then you might think there was a tension between wanting to shut it down because it’s unsafe and wanting to respect its right to continued existence. Still, if we are going to create these entities that deserve some kind of moral consideration and we need to do a bunch of research on them, how do we do that in a way that is respectful and incorporates a notion of research ethics? Are there lessons we can learn from human-subject research here?
It would be strange if what we’re building replicated every other cognitive ability or cognitive function a human brain has but just did not have this other “magical” thing. Presumably, consciousness in the sense of having subjective experience was at some point beneficial to us evolutionarily. If I put my hand in a fire and it hurts and I pull my hand away, I’m more likely to survive and pass on my genes. One question I’m very interested in is: To what extent are the training methods we currently have putting models under the same, or analogous, types of selection pressure?
Photo: Carolyn Drake (Mishkin, Wu), Courtesy of Subjects (remaining)
When he left OpenAI in 2024, the company required departing employees to sign a non-disparagement agreement as a condition for retaining their vested equity. He refused. He was ultimately able to keep his equity.
Anthropic has predicted “powerful AI systems” to be developed in early 2027 — that is, AIs with abilities that exceed top human experts in most fields.
“AI 2040” is a set of policy scenarios published by the AI Futures Project, Kokotajlo’s group. Plan A is its proposal for an agreement between the U.S. and China to delay superintelligence.
Anthropic alignment-science lead Evan Hubinger tweeted that many employees “earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
This cannot be verified as Anthropic has not yet gone public.
The AI Futures Project’s forecast of the trajectory of superintelligence, published in 2025.
The point at which AI systems are as good at AI research as the best humans, enabling them to make even better AIs in an accelerating feedback loop of progress — at least in theory.
A computer science professor prominent in organizing around both AI loss of control risk and against lethal autonomous weapons.
International Association for Safe and Ethical AI (IASEAI)
A spokesperson for Google DeepMind: “We listen to our employees, whatever level of seniority they are. We did review his framework. That said, this was an employee who lacked an understanding of the work we were doing in this area.”
In 2018, more than 3,000 Google employees signed a petition demanding the company pull out of Project Maven, another DoD program.
A document that describes the capabilities, risks, and safety and security mitigations for a particular AI model or system.
In cybersecurity and other fields, red-teamers are evaluators who mimic the tactics used by adversaries to stress-test a system.
A method in machine learning where an AI model is given a goal and learns how to achieve it through trial and error
This was in evaluations without safeguards.
A research program focused on ensuring superintelligence would be safe.
An OpenAI co-founder and one of the board members who initiated Altman’s firing.
The other leader of the superalignment team at OpenAI. He now works at Anthropic.
A planned $500 billion supercomputer project from OpenAI, SoftBank, Oracle, and the investment firm MGX.
A spokesperson said OpenAI has dozens of companywide Slack channels where employees can ask questions.
Anthropic famously requires a thorough “culture interview” as part of the hiring process.
One of the most valuable currencies in Silicon Valley right now: the processing power and hardware to run and train AI models.
A field of AI research that tries to understand how a model’s internal mechanics produce the behaviors we observe.
An OpenAI spokesperson said this estimate was inaccurate.
When a core group of OpenAI scientists and executives, including Dario and Daniela Amodei, left to found Anthropic.
