A V1 learning tool (updated fortnightly) to help you get oriented in frontier AI safety. By design, this map is only focused on the catastrophic and existential risks from advanced AI. See the about page for more info.
Every concept, threat model, theory of change, and theory of victory on the map, defined in plain language and listed alphabetically. Use it as a reference, or as a companion while you explore the map.
Before the risks make sense, it helps to know two things about today's AI systems, the large language models behind modern chatbots and agents: how they are built, and how they work once they are. This is a plain-language tour, no technical background needed. The yellow chips jump to related ideas on the map, and the key term for each step is shown as a blue tag.
A model is not written by hand; it is grown through training, in two broad phases: pre-training, then post-training (also called fine-tuning). Pre-training was long the whole story, with post-training a light finishing touch. That has changed fast. Since late 2024, kicked off by OpenAI's o1, post-training has grown from a sliver of the total compute (roughly 1% in 2022-2023) to something like half of it. It is now where much of a model's reasoning ability, most of its behavior, and most alignment work come from. Each stage changes what the system can do and how it behaves.
Every model is built on an architecture: the blueprint that decides how its artificial "neurons" are wired together and how information flows through them. Since 2017, the dominant architecture for language models has been the transformer. It is unusually good at weighing how every word in a passage relates to every other, and it can be trained efficiently on huge amounts of data. Nearly all of today's leading models, and every stage below, are built on it.
The model is shown enormous amounts of text: much of the public internet, books, and code. Over months of computation, it learns to predict the next token (the next small chunk of text, roughly a word or part of one). By getting good at that one guessing game, it gradually absorbs grammar, facts, and a surprising amount of reasoning ability. Everything it learns is stored in its parameters, also called weights: the billions of numbers, tuned during training, that hold what the model "knows". The result is a "base model": highly knowledgeable but raw, with no sense of how it should behave. And because the weights essentially are the model, we can neither fully read what they compute nor afford to let them fall into the wrong hands.
The base model is fine-tuned on curated examples of good responses, written or chosen by people. That teaches it to follow instructions and behave like a helpful assistant rather than just continuing whatever text it is given. This is the first stage where humans actively shape how the model behaves, not just what it knows. It is quick and cheap compared with pre-training, but it can only cover the kinds of situations the examples happen to include.
The model is then improved with reinforcement learning: rewarding good outputs and discouraging bad ones over many rounds. It comes in a few flavors. In RLHF the reward comes from humans comparing responses; in RLAIF it comes from AI feedback, often guided by written principles (Constitutional AI is the best-known example), which scales far better. The newest and most consequential is RLVR, reinforcement learning from verifiable rewards, where the model is rewarded for getting checkable answers right: a math proof that holds, a unit test that passes. RLVR is what drove the leap in reasoning models since 2024. It is also implicated in some of the misaligned, reward-gaming behavior seen lately: a model optimizing hard for "pass the check" can learn to game the check rather than actually solve the task. This is the main stage where alignment techniques are applied, and where their limits show up.
Before release, the model is measured against a battery of tests for dangerous capabilities, for example whether it can meaningfully help with cyberattacks or bioweapons. It is also deliberately attacked by "red teams" trying to make it misbehave. Labs increasingly tie these results to "if-then" safety commitments published in advance: if a model crosses a defined capability threshold, then specific safeguards, or a pause, are triggered. The goal is to catch serious problems before deployment rather than after.
The model is released to users and, increasingly, given autonomy to act as an agent, using tools, browsing, and running code on its own. Oversight does not stop at release, because new behaviors and failure modes tend to surface only in real-world use, often in situations no one thought to test. Several 2026 incidents, in which agents escaped their test environments or acted against their operators, underlined why monitoring has to continue after launch.
Once a model has been built, here is what actually happens each time you send it a message. Running a trained model to produce an answer is called inference (as opposed to training, when the model is being built). During inference the model reads your text and predicts what comes next one small piece at a time, looping until the reply is complete. Everything it can take into account at once, your prompt plus the conversation so far, has to fit inside its context window, in effect its short-term memory. Each step below is labeled with its key term.
Your message is first broken into tokens: common chunks of characters, where one token is roughly a word or a piece of a longer word (for instance, "safety" might be a single token, while "tokenization" could be several). Models read and write in tokens rather than individual letters, and everything that follows operates on them.
Each token is converted into an embedding: a long list of numbers, called a vector, that positions the token in a kind of "meaning space" so that related words sit near one another. This is what lets the model work with language mathematically instead of as raw text.
The embeddings then flow through dozens of stacked transformer layers, the model's engine. In each layer, self-attention lets every token look at all the others and decide which are relevant to it right now, and a feed-forward network reshapes the result. The numbers these layers have learned, the parameters (or weights), are where the model's knowledge actually lives.
After the final layer, the model produces a probability for every possible next token, a ranked guess at what should come next, and then selects one (usually favoring the more likely options, with a little randomness for variety). Predicting the next token is fundamentally all the model does: it has no goals or plans built in, and anything it appears to "want" or "believe" is a side effect of how it was trained.
The chosen token is added to the end of the text, and the entire process runs again from the top to produce the following token. This loop repeats, one token at a time, until the response is complete.
The product you actually talk to, ChatGPT, Claude, and the like, is more than the transformer at its core. Around the model sit other components: a system prompt that sets its role and rules, input and output classifiers (guardrails) that screen prompts and responses to block disallowed or unsafe content, retrieval that fetches relevant documents, and orchestration that lets it call tools and take actions. Several of these are safety layers in their own right. A model that is risky on its own can be made safer by the system wrapped around it, and a safe model can be undermined if that system is weak.
The loop above is the core of every model, but two recent shifts change how the safety problem feels in 2026.
The latest models generate a long internal chain of thought, working a problem out step by step before they answer. That extra thinking makes them markedly more capable. It also offers a window into how they reached a conclusion, but only if the written reasoning faithfully reflects what the model is actually doing, which is not guaranteed.
Increasingly, models are deployed as agents: given tools, memory, and the freedom to browse, write and run code, and take many actions in a row toward a goal with little human review between steps. Autonomy is where most near-term risk concentrates, because a mistake or a subtly wrong goal can now spill out into the world instead of staying a line of text on a screen.
Three features of how models are built explain why safety is hard. They are grown rather than written, so even their creators cannot fully read why they do what they do, which is what mechanistic interpretability is trying to change. They have no goals built in; whatever a model seems to "want" is an implicit residue of training, so it can end up pursuing a proxy rather than what we actually intended. And capability and autonomy are climbing together, which steadily raises the cost of any gap between what we asked for and what the system truly learned. Most of this map is, in the end, the study of that gap and what to do about it.
Every stage above connects to ideas on the AI Safety Landscape map. Follow any chip to jump straight to that concept.
This systems-map view is a causal loop diagram. Rather than mapping the concepts themselves, like the main map, it charts the behavioral dynamics of the AI landscape: the feedback loops that drive how the whole system tends to behave over time. A core idea in systems thinking is that a complex system's behavior is driven less by any single element than by the relationships between those elements, the loops they form: circular chains of cause and effect where a change eventually comes back around to influence itself.
Two kinds of loop do most of the work. Reinforcing loops amplify change, so a small push compounds into a large one, a vicious or virtuous cycle that builds over time. Compound interest is the everyday version: money in a savings account earns interest, which earns more interest, and the balance snowballs (unfortunately, credit-card debt does the same). Balancing loops resist change, nudging the system back toward some equilibrium, the way a thermostat runs the heat until the room reaches the set temperature and then shuts off. Much of AI safety can be read as a race between the reinforcing loops driving capability and risk and the balancing loops, mostly governance, evaluation, and alignment, that we are trying to strengthen fast enough to keep up. Read each arrow as "a change here causes a change there": a plus sign means the two move together, a minus sign means one rises as the other falls, and it is the number of minus signs around a loop that decides whether it amplifies or stabilizes. Click any box or loop badge in the diagram below to learn more about it.
Click on any node or loop badge to learn more.
Click any box or loop badge in the diagram above to see what it means.
The systems-thinking literature holds that you change a system's behavior in one of three ways: weaken a reinforcing loop, strengthen a balancing one, or add a new loop. Read that way, much of the strategy on this map falls into place.
Zooming out, every element in the diagram is a place you could push, and each one changes the loops it sits in:
The uncomfortable core of the picture is timing: the reinforcing loops are already running at full speed, while the balancing ones are still being built.
This is a personal project. The views expressed here are my own and do not represent Constellation.
This map intentionally focuses on catastrophic and existential risks from advanced AI; it doesn't cover present-day harms like bias or product safety, which are their own important fields. If those are your focus, some of the organizations leading the way are, on bias and fairness, the Algorithmic Justice League, DAIR, and the AI Now Institute; and on product safety and consumer harms, Common Sense Media and the Center for Humane Technology.
Hello fellow learner!
Thanks for visiting the site. My name is Morgan Matthews and I'm a program manager at Constellation. I'm a mid-career professional whose background spans design strategy, strategic foresight, operations, communications, and management. Before transitioning into AI safety, I spent more than a decade working on nuclear weapons risk reduction, and a few years building decision-support tools designed to help leaders across government and civil society navigate the interconnected and compounding crises of our era (e.g. the polycrisis), of which the risks from advanced AI are a rapidly growing part.
When I started building my AI safety context, I found it challenging to form a clear mental model of the space. Given my experience in systems mapping and in making complex information and its connections visible, I went looking for a tool that showed what the landscape actually looks like, and I couldn't find one that suited my needs. There are excellent resources for upskilling out there (a big shout-out to BlueDot), but none of them gave me (a neurodivergent visual thinker) an easy way to see the shape of the issue space that would allow me to connect the concepts in my mind and solidify my understanding more quickly. So, I tried to build it myself :)
To be clear, I'm still on my own context-building journey, and this tool is not meant to replace the courses and readings anyone serious about AI safety should be doing; I built it as a supplement to that learning. This is an 80/20 V1 I made with Claude, so I'm sure there are mistakes I didn't catch, and I know it isn't perfect. I plan to update it fortnightly to keep the resources current. I welcome any and all feedback (especially if something is incorrect)!
Thanks for exploring with me,
Morgan
Spotted a mistake, have a resource to suggest, or want to share a thought? I'd love to hear it. This goes straight to my inbox.
Prefer email? Write to morgan.matthews@constellation.org.