What Are They Willing to Risk? Why AI Companies Need to Make Their Risk Appetites More Legible to the Public
Authored by CDT Intern Maggie Wang.
Driven by mounting geopolitical and commercial pressures, AI companies developing general-purpose AI models and systems are racing to release products with new features and more powerful capabilities. Though prominent AI companies have pledged to develop and deploy their models and systems safely and responsibly, many are pulling back from their commitments in the rush to outcompete one another. Google, OpenAI, and xAI have released new AI systems with missing or delayed model cards, despite promising to publish consistent and timely documentation. Insiders have raised concerns about rushed safety evaluations at OpenAI. Anthropic released Claude Opus 4 despite the model’s apparent propensity to engage in blackmail and warnings from a third-party auditor not to deploy. OpenAI released ChatGPT Agent despite the system posing enough risk to compel CEO Sam Altman to issue an explicit warning about how to use it.
Voluntary safety commitments and risk management frameworks are meant to ensure that companies rigorously assess and mitigate the risks that their AI systems pose to users and society, but recent events show these measures may not be as effective at steering developer behavior as many hoped. Most of the criticism against voluntary commitments and guidelines centers on the fact that they are nonbinding and unenforceable. We highlight a different issue: safety commitments, best practice guidelines, and the broader discourse on AI risk management have emphasized that companies ought to assess risks, but they have paid less attention to how companies respond to and act on those assessments. The emphasis reflects an assumption that if an AI company assesses risks, they will also take careful and deliberate steps to reduce or prevent the risks that they identify. Rushed, delayed, and ignored risk assessments show how often this assumption is wrong. How a company chooses to respond to risk assessments, and how they make decisions about what AI systems to launch, when, and with what risk mitigations, comes down to their risk appetite and risk tolerances.
Risk appetite, risk tolerances, and decision making
Risk appetite is the overall amount of risk that an organization is willing to take in order to achieve its goals and is a core element of an organization’s broader risk management strategy. The more a company stands to gain from taking risks, the larger the appetite they are likely to have for those risks. Risk tolerances, meanwhile, are more granular and help organizations operationalize their risk appetite by describing how much of a specific type of risk a company is willing to tolerate. Risk tolerances can be influenced by industry-wide legal or regulatory requirements, as well as norms, values, and resources that are unique to a particular organization. When organizations make decisions, they generally do so within the bounds of their risk appetite, and in accordance with their risk tolerances.
Importantly, a company’s appetite and tolerances for risks to itself is distinct from their appetite and tolerances for risks to consumers and society more broadly. In academic and policy-making circles, AI risks are typically framed in terms of how AI systems impact users and society, includingpresent-day risks like discrimination, misinformation, threats to privacy, offensive image generation, and effects on mental health, as well as anticipated emerging risks like assisting in the creation of weapons or conducting cyber attacks. When companies make decisions, however, they tend to consider, prioritize, and account for the risks that affect the company’s ability to make progress towards its goals. In some cases, public and company risks are aligned — harm to consumers might trigger legal liability or regulatory scrutiny, or invite reputational damage to the company. They can also be in tension — for instance, to avoid the risk of financial loss due to a delayed product launch, companies might be incentivized to spend less time on safety evaluations and implement fewer risk mitigations. When public and company risks conflict, risk appetites and tolerances help to govern how companies balance the two and can reveal when that balance skews too far in favor of the company.
The issue: opaque and unchecked risk appetites and tolerances
Existing voluntary risk management frameworks explicitly refrain from prescribing risk tolerances to AI developer companies and offer limited guidance and advice beyond that companies should follow applicable regulatory requirements and set risk tolerances that are broadly “reasonable”. As a result, companies are left to determine risk appetites and tolerances largely at their own discretion. But because companies currently share very little concrete information about what those appetites and tolerances are, it is unclear how they are navigating the balance between managing public risk and company risk and whether they are doing so in a way that is aligned with the public interest.
SeveralAIcompanies publish model and system cards describing how they evaluated and mitigated a set of known risks posed by a particular AI system prior to releasing it, but these reports typically only convey the results of select evaluations (e.g., a score on a benchmark or a summary of a red-teaming evaluation) and not how decisionmakers mapped results to the decision to ship that system. Readers are left guessing both about what risks may have been conveniently omitted from the report and about how bad results would need to be for the company to deem them “intolerable.”
Some AI companies have published frontier safety frameworks that articulate tolerance thresholds on certain model capabilities, which, when crossed, trigger mitigative actions such as securing model weights, pausing development or deployment, or adding stricter safeguards. While capability thresholds are meant to be more explicit expressions of risk tolerances, in practice, most frontier safety frameworks use imprecise and subjective words like “significantly”, “substantially, and “drastically” to characterize the ways in which they consider AI systems to be capable of conducting or enabling harmful activities. A recent review of frontier risk management practices at AI companies called “vague” or “undefined” risk tolerances a “fundamental flaw” across frontier safety frameworks. Importantly, frontier safety frameworks only capture a narrow set of catastrophic risks; for the many other risks that AI poses, companies are missing even that limited structure and transparency in how they set their tolerance levels.
The lack of transparency surrounding risk tolerances and risk appetites allows AI companies to increase their appetite for risk to the public arbitrarily, and with limited accountability, in order to justify the decisions they want to make. To add fuel to the fire, the current incentives in the AI ecosystem generally provoke large risk appetites — investors have poured billions into the industry, AI companies are betting big and spending massive sums of money to build out the AI infrastructure needed for scaling and accelerating AI development, and the Trump administration’s AI Action Plan puts a strong emphasis on rapid innovation and deployment. Companies have little reason to slow down and attend to risks that their AI systems pose to the public, especially if doing so interferes with meeting investor and shareholder demands. Competitive dynamics between companies only exacerbate the situation — OpenAI and Anthropic, for example, state in their frontier safety frameworks that they are willing to raise their risk tolerances to match their peer companies’ risk tolerances, which can catapult them towards greater and greater appetites for risk.
Between the implicit signals from companies’ launch choices, the limited visibility into their true risk tolerances and appetites, and the incentives for companies to have a large affirmative appetite for risk, there is reason to worry that AI companies may be inclined to put the public at more risk than it wants to bear. Absent sufficient legal or regulatory mechanisms that make AI companies face significant repercussions when their products cause harm, the public shoulders the burden of those risks and ought to be able to evaluate what risks are acceptable in exchange for the benefits of AI systems. When AI companies don’t communicate their risk appetites and risk tolerances, however, they prevent the public from scrutinizing and critiquing how risk acceptance decisions are made.
Toward more legible risk appetites and tolerances
To give the public the ability to deliberate and potentially challenge AI companies’ willingness to put the public at risk, companies should do more to articulate and disclose their appetites and tolerances for risks that AI systems pose to users and society. Companies can consider communicating risk appetites through enterprise risk appetite statements, building off of examples of such statements from establishedorganizations. Risk tolerances should be expressed in precise and unambiguous terms, such as a quantitative or well-described qualitative threshold, and should be set for all risks where reasonable, not just for catastrophic ones. When setting risk tolerances, companies should consider sociotechnical and participatoryapproaches that increase public presence in conversations around tolerable risk, so that an isolated group of technical experts or executives at a company is not “unilaterally decid[ing] what risks count [and] what harms matter.” Companies should also explain how their risk tolerances, taken in aggregate, are consistent with their stated risk appetites.
Furthermore, companies should take steps to demonstrate that they uphold their stated risk tolerances in decision-making. Safetycases, or structured arguments that justify choices about model risks and mitigations, are one framework that companies could use to communicate risk acceptance decisions. Whether as part of a safety case argument or in some other form of documentation, companies should articulate the evidence that was used to determine a risk was tolerable, how this evidence was sufficient to show that the risk satisfied criteria for acceptance, how uncertainty in the evidence (e.g., due to stochastic model behavior) was taken into consideration, and how assumptions about the validity of the evidence (e.g., that a benchmark accurately reflects real-world conditions) might affect conclusions about the risk meeting the conditions for tolerability.
Lastly, AI companies should acknowledge that unknown risks are a part of their risk appetite. Every new or updated AI system almost surely poses risks that are unknown to or under-studied by its developers prior to release, and the company must accept these risks in order to move forward with deployment. The sycophancy issue with OpenAI’s GPT-4o update is one such example, where the nature and scope of the risk only became apparent after real-world use. While it is impossible to articulate precise tolerances for risks that are unknown, companies should still explicitly acknowledge their willingness to tolerate them: when unknown risks are left unacknowledged, companies’ risk appetites may appear to be smaller than they actually are.
Conclusion
What AI companies decide to do about AI-posed risks hinges on their risk appetites and tolerances, and the prevailing focus on assessing risk misses the point that assessments mean little if appetite for risk is unlimited. More transparency into risk appetites and risk tolerances will make it easier to spot when companies’ appetites for public risk are too large, but transparency is just the first step. If such transparency reveals excessive risk appetites, external levers will be important for applying corrective pressure and keeping companies’ risk appetites in check. Policy actions, including regulatoryoversight or liability rules, as well as public scrutiny through journalism and advocacy, can force companies to internalize risks to the public as legal, financial, or reputational risks to themselves. When policymakers abandon their efforts to push for responsible AI development or focus on an overly narrow conception of AI-posed risks, though, a critical tool for recalibrating risk appetites is lost.
Norms and standards on exactly how risk tolerances should be set and disclosed are still evolving, and achieving consensus will require time and cross-sectoral discourse. Productive discourse will be difficult, however, unless companies first clarify how they make risk acceptance decisions and create channels for public input into their choices. New reporting obligations, such as those in the European Union’s General-Purpose Code of AI Practice, and reporting frameworks like the Hiroshima AI Process, have the potential to spur progress toward more standardized disclosures on risk appetites and tolerances, as long as companies sign on and stay compliant. These reporting obligations and frameworks can also help to create a positive feedback loop — as more companies offer greater disclosure of risk tolerances, it will enable iterative development, critique, and improvement of cross-company best practices and standards, both on how risk tolerances should be set and how they should be disclosed.
AI is already having large-scale impacts on society, including in harmful ways, and the public should have the power to understand and shape this impact. To exercise that power fully, the public must be able to scrutinize, deliberate, and influence how companies decide which AI-posed risks are acceptable and which are not. Right now, the fear of losing the AI race seems to be driving companies toward ever-larger risk appetites, and voluntary — or even mandatory — risk assessments are not enough to contain them. As an initial step, AI companies should explicitly communicate their risk appetites and tolerances for risks to the public, how these appetites and tolerances are set, and how they are upheld in decisions made throughout development and deployment.
Coalition Urges Senate Not to Let Companies Waive Financial Regulations for AI
CDT joined AI Now Institute, American Civil Liberties Union, and several organizations dedicated to tech policy, consumer protection, and civil rights in a letter to Senate leadership and the Senate Banking, Housing, and Urban Affairs Committee opposing the “AI Innovation Labs” language in Sec. 10509 of the CLARITY Act.
As concern about risks and harms related to AI systems continue to grow, a growing chorus of policymakers, industry leaders, and advocates have called for independent AI assessments. This explainer provides an overview of recent proposals for third-party assessment in the United States, including state and federal legislation, executive actions, and industry proposals.
Having third parties assess AI systems might seem like common sense, but crafting effective policies toward this goal can be devilishly tricky. A poorly-constructed ecosystem for third-party assessment could easily fail to consider the most consequential mechanisms of risk, neglect the AI harms that most impact people, or do more to protect AI companies than people.
Not All Guardrails Are Created Equal: Comparing Content Safety and Copyright Filtering
As courts and policymakers work through questions about chatbot liability, they should be wary of analogies that flatten meaningful technical differences. Copyright filtering and safety intervention share real challenges around ambiguity and evasion, but they diverge in what each control must assess, how each manifests over the course of a conversation, and how much can be verified from the outside.