A number of recent policy proposals seek to catalyze third-party assessment for AI systems, spanning assurance of ordinary deployments of simple AI systems in well-defined use cases to the testing of complex AI systems with a wide variety of potential uses and harms. These proposals include a litany of calls from industry, recent state laws, federal legislative proposals, and executive action, all sharing a desire to establish a more formalized process by which AI models or systems are assessed by organizations outside of their developers or deployers.
Some proposals call for independent audits of certain use cases, while others seek to develop an ecosystem of third-party assessors (often called independent verification organizations, or IVOs) who can certify models against specific claims or requirements. Still others recommend the creation of a single independent or semi-independent entity to perform testing, the formalization of a government body to assess and approve AI systems before their release, or embedding third-party entities within AI companies. (We use the term “third-party AI assessment” to refer generally to a variety of closely related paradigms in which third parties assess, evaluate, audit, or verify AI systems.)
Having third parties assess AI systems might seem like common sense, but crafting effective policies toward this goal can be devilishly tricky. A poorly-constructed ecosystem for third-party assessment could easily fail to consider the most consequential mechanisms of risk, neglect the AI harms that most impact people, or do more to protect AI companies than people. For instance, a 2021 New York City law mandating audits of AI-powered hiring tools appears to have been scoped too narrowly to meaningfully change employer behavior, while even a robust regime of government supervision of lending algorithms has led to a flourishing practice of producing defensive documentation rather than necessarily motivating the selection of less risky credit models. This post sets out eight factors that policymakers will need to satisfy to get third-party AI assessment right
Ensure meaningful independence for assessors. Independence requires a dynamic ecosystem of third-party evaluators that are not overly dependent on any small set of AI companies for their financial viability, and are comfortable issuing negative findings without fearing financial repercussions. If AI companies select their own third-party verifiers, policymakers should mitigate – perhaps through regulation – the natural market incentives for notionally independent verifiers to differentiate themselves through more lax oversight.
Meaningful independence also requires avoiding too much cultural or social entanglement between third-party evaluators and the AI companies they are meant to check. At present, many third-party evaluators of frontier systems come from a small community centered around AI developers themselves. They tend to have overlapping social networks, shared professional backgrounds (with many staff consisting of former AI company employees), and similar schools of thought about how AI risks might manifest and how they should be evaluated. To be sure, experience and familiarity with AI companies and the technologies they are building can be deeply valuable for third-party evaluators. But too much of a monoculture can lead to blind spots in identifying and measuring risks, and result in third party scrutiny focusing on the issues that AI developers prefer to prioritize while overlooking or deprioritizing other important issues (such as impacts on privacy or mental health). For third-party assessment to be effective, policies must encourage the establishment and maturity of a diverse set of third-party assessors with expertise scrutinizing AI systems’ broad range of potential impacts in the wild.
Incentivize interdisciplinary skillsets. A professionalized third-party assessment ecosystem will require a pool of talent with the skills to adeptly identify, measure, and mitigate a variety of AI risks. In addition to technical knowledge about AI, this will require three critical skillsets. First, assessors must possess knowledge about the domains in which AI systems are deployed. Evaluating the potential for cyber-related harms will require deep cybersecurity expertise; assessing potential mental health risks will require very different types of knowledge.
Second, multidisciplinary expertise across a variety of fields is necessary, coupled with the ability to translate between them. As recent incidents of AI systems hacking into outside organizations during testing have demonstrated, this will require more than just technical knowledge about AI models: those incidents were enabled by weaknesses in testing environments, network security limitations, and monitoring failures. Successful assessors will need to leverage and translate between experts in software and data engineering; governance, privacy, and cybersecurity; documentation; product management; ethics; impacts on civil rights and civil liberties; human-computer interaction and other behavioral and social sciences; and law, regulation, and compliance. Evaluation organizations won’t be effective if they overly prioritize technical expertise, and so too will they fail if they too formally silo different skillsets.
Finally, third-party auditors must possess a variety of “soft” skills. This includes the ability to critically engage with different types of staff and business functions within the organizations whose processes or systems are being assessed, to negotiate implicit terms of engagement with different parts of large organizations, and to understand how complex organizational, institutional, and bureaucratic dynamics can contribute to an AI system causing harm.
Assess systems and use cases, not just models. Many proposed policies for third-party assessment focus on AI models themselves, reflecting a school of thought in which intrinsic characteristics or capabilities of models are primary sources of risk. But an exclusive focus on models alone using a standard rubric would be a mistake, leading to a third-party assessment ecosystem that is neither effective nor future-proof. AI models implicate drastically different types of risks (and those risks arise in different ways) when applied to different use cases. An evaluation of an AI model used to control an electrical grid should focus on very different factors than an evaluation for that same model being used to inform government benefit decisions. Similarly, a model that might be safe in a carefully controlled context might be highly risky when given access to certain tools, when deployed without effective oversight, or when used in a setting where adversarial attacks might be common. Considering specific use cases – rather than a model in isolation – enables assessors to better anticipate harms that could arise and focus evaluations on the specific model and deployment characteristics that are most likely to lead to those harms.
Moreover, AI harms rarely stem solely from models themselves. Even when considering risks related to advanced AI capabilities (such as those related to cybersecurity or manipulation), real-world incidents typically arise from a combination of system-level factors including not only the AI model, but also its scaffolding, personalization, safety filters, the tools it has access to, and how how it is overseen and monitored, among other properties of an AI system’s deployment. Recent high-profile cyber incidents, for instance, were enabled and exacerbated by underspecified evaluation tasks, properties of model harnesses, and failures in testing environment configuration and system monitoring – all system-level traits. Indeed, history shows that most technology-related disasters stem not just from technology itself, but failures in human-computer interaction, social and organizational dynamics, and risk management. Model-level evaluation may have a role to play, but ensuring third-party assessment extends to systems and use cases – not just models – is key to enabling assessors to effectively anticipate failure modes and identify effective mitigations.
Include a broad set of potential harms. Many proposals for third-party assessment focus on a narrow set of risks, such as those related to cybersecurity or biological weapons. But AI models can catalyze a broad and diverse set of harms that affect peoples’ lives, families, and communities every day – harms ranging from discrimination and violations of privacy to threats to physical safety and critical infrastructure. Any third-party assessment scheme is unlikely to achieve legitimacy if it does not help mitigate the harms that most impact people’s day-to-day lives, so policymakers seeking to create an independent verification ecosystem should ensure that this ecosystem can help mitigate this broad spectrum of risks.
Even for those who advocate prioritizing certain types of risks, incentivizing an ecosystem of third-party assessors verifiers focused exclusively on those narrow areas would be a mistake. Oftentimes, very different types of harms can lead to or exacerbate others: AI-caused privacy vulnerabilities, for instance, could facilitate social engineering techniques that enable a massive cyberattack. And considering a broader set of risks and use cases can expand the base of potential customers and market incentives for independent evaluators, leading to a more robust ecosystem of third-party assessors with less market pressure to avoid issuing negative findings about any given product or organization.
Of course, this doesn’t require every assessor to evaluate everything at once; indeed, a third-party assessment exercise is most likely to be effective when it has a clearly articulated scope. It may be useful for some evaluators to specialize in investigating particular types of risks or impacts. However, policymakers should take care to avoid locking a nascent ecosystem into areas of focus that may eventually prove too narrow to address the full array of risks AI systems pose. Policymakers seeking to balance seemingly urgent topics while leaving room for a more flexible ecosystem to develop could support a diverse set of pilot areas for third party assessors to focus on. But in doing so, they should ensure that these pilot areas have the opportunity to grow to a more comprehensive scope in the future, should they prove effective. For example, state proposals that would require independent auditing of AI use in health insurance would likely offer lessons that could inform similar requirements in lending or hiring; policies focused on rapidly advancing third-party assessment for catastrophic risks should build in paths to extending coverage to broader categories of harm as well.
Incentivize repeated and ongoing evaluations. Both AI models and the conditions they’re deployed in will change over time, sometimes dramatically. New risks may emerge as a model is used in unexpected ways. Second-, third-, and fourth-order impacts of AI deployment and use will create more clarity about the factors third-party assessors should concentrate on. And risks can arise from minor changes to models as often or more as they do from new releases: consider the sycophancy issue caused by a minor update to GPT-4o in April 2024, which was serious enough to lead OpenAI to roll back the update a few days later.
For each of these reasons, point-in-time evaluations may be useful in the moment but will have a limited shelf life. Third-party assessments that are performed at appropriate intervals, or better yet involve continuous monitoring over time, will be better able to naturally spot and account for evolving risks. Ongoing evaluations are also better able to take into account how models are being deployed and used in practice, further improving evaluators’ ability to assess the mechanisms through which harms might arise.
Ensure that incentives for assessment are commensurate with protections they provide. Third party assessment takes time and resources, and policymakers are weighing both positive and punitive incentives to drive adoption. One previously considered incentive is liability protections (such as a rebuttable presumption of reasonable care) in exchange for obtaining certifications from third-party assessors.
While new incentives could make sense in principle — and indeed may be important to catalyze the field — any such incentives should be proportional to the effectiveness of third party assessment in protecting people. Policymakers should be especially careful about creating liability shields before there is concrete evidence about the extent to which independent verification actually helps to mitigate downstream harms. Any incentives should also be carefully tailored to the specific scope of the assessment; disproportionate or mismatched liability shields risk protecting companies that cause AI harms more than protecting individuals from AI harms.
Policymakers should also keep in mind that liability-related incentives for third-party assessment could arise naturally, without statutory intervention. This is already emerging in legal fights over social media harms; the need to secure insurance creates further incentives for technology companies to demonstrate that they took reasonable precautions to prevent harm. For AI developers, for instance, a third-party assessment could provide the kind of evidence needed to push liability to a deployer. Policymakers might consider codifying this understanding without creating explicit liability shields, as one proposal did by specifying that audits could serve as evidence to rebut a presumption of liability.
Ensure assessment requirements don’t entrench powerful actors. Americans are best served by a vibrant set of AI developers where standards are set through open, democratically legitimate processes. Policymakers should avoid creating systems that lack democratic safeguards, or enable market leaders to shut out their competitors and smaller developers by imposing requirements with which only well-established AI leaders will be able to comply.
In the case of democratic safeguards, policymakers should avoid giving the government free rein to set the processes and parameters for assessment without transparency, consistency, appeal rights, and other measures to ensure fairness and due process. As CDT and coalition partners have highlighted, the White House’s current pre-release model assessment framework is an example of how a secretive and arbitrary government process can go wrong. Without transparency or clear standards, the framework is unlikely to sufficiently improve safety practices, and creates openings for favoritism and corruption. The export controls briefly imposed on Anthropic’s Fable show how this risk can manifest in practice, creating a new avenue for the government to coerce AI companies on issues unrelated to safety.
On the risk of regulatory moats, the concern is particularly acute for open source developers, which might lack the structural and financial means to comply with certain types of auditing requirements. Additionally, mechanisms that allow AI companies to coordinate on evaluation standards or specific types of harm should be carefully designed to avoid allowing harmful forms of collusion. Policymakers should also avoid creating a system in which industry and its likeminded partners are allowed to dictate the risks, methodologies, and terms of engagement that assessment focuses on. Industry and existing evaluators have a meaningful role to play, but researchers, civil society, organizations with auditing expertise in other domains, as well as government experts should also have a powerful voice.
Treat independent verification as just one piece in the larger accountability puzzle. Constructed correctly, third-party assessment can be a powerful component of the AI accountability ecosystem. But there are still many open questions about how effective third-party assessors would be, even at best, in preventing all AI harms, and assessment is unlikely to be effective without a broader set of rules, structures, and incentives.
To be successful, third-party assessment mechanisms will need to operate in tandem with many other building blocks of accountability infrastructure. Incentives to conduct third-party assessments might need to be coupled with regulatory requirements to remediate identified shortcomings. Insurance and clearer distribution of liability for harms will provide financial incentives to avoid harm, and help AI actors appropriately price risks. Standards created with input from both industry and public interest groups will help establish clear processes, metrics, and benchmarks to facilitate consistent, valid, and technically robust assessment. Best practices, norms, and a culture of accountability will shape day-to-day practices in ways that improve accountability. Regulation will be required to prevent organizations from accepting inappropriate levels of risk. Less formal, nonregulatory, audits by independent researchers have historically helped to surface new issues and risks, and policies encouraging formal third-party assessment should ensure they do not foreclose audits by academic or independent researchers. Each of these components has its own gaps and limitations, but together can ensure that harms of AI are minimized and the promise of its benefits can be realized.
As concern about risks and harms related to AI systems continue to grow, a growing chorus of policymakers, industry leaders, and advocates have called for independent AI assessments. This explainer provides an overview of recent proposals for third-party assessment in the United States, including state and federal legislation, executive actions, and industry proposals. In a separate post, we describe factors that will be required for third-party AI assessment to be effective.
Not All Guardrails Are Created Equal: Comparing Content Safety and Copyright Filtering
As courts and policymakers work through questions about chatbot liability, they should be wary of analogies that flatten meaningful technical differences. Copyright filtering and safety intervention share real challenges around ambiguity and evasion, but they diverge in what each control must assess, how each manifests over the course of a conversation, and how much can be verified from the outside.
Talking Tech with Roy Austin & Alex Givens on The Future of AI Governance
In this episode of CDT’s Tech Talks, Roy L. Austin, Jr., inaugural director of Howard University School of Law’s AI Initiative, joins Alexandra Givens, President and CEO of the Center for Democracy & Technology, to discuss the challenges and opportunities shaping the future of AI.
CDT-led Coalition Calls for Transparency for White House AI Framework
CDT and Americans for Responsible Innovation led a broad, bipartisan coalition of over two dozen civil society groups in calling on the White House to release its Framework for review of frontier AI models.