As concern about risks and harms related to AI systems continue to grow, a growing chorus of policymakers, industry leaders, and advocates have called for independent AI assessments. This explainer provides an overview of recent proposals for third-party assessment in the United States, including state and federal legislation, executive actions, and industry proposals. In a separate post, we describe eight factors that will be required for third-party AI assessment to be effective.
Third-party assessment refers generally to a variety of closely related paradigms in which outside parties assess, evaluate, audit, or verify AI models or systems. Compared to purely internal reviews,third-party assessments could reduce conflicts of interest, and provide a greater diversity of perspectives, approaches, and expertise that can better identify and mitigate risks. But effectively assessing AI systems is an open technical and policy challenge. Despite a swath of research and guidelines developed by academics, industry, government, and civil society, it remains notoriously difficult to construct evaluations that reliably measure what assessors would like them to. Additionally, it is difficult to project how AI systems may be used (or misused) in different deployment contexts, and there is little agreement around key risks and how they might arise. Beyond these technical challenges, careful regulatory design is necessary for third-party assessment to be effective in reducing AI risks — and to address concerns about capture, regulatory moats, and democratic legitimacy.
Recent advocacy for third-party assessment has coalesced around two related approaches: policy proposals that specify details of when, who, and how to audit; and proposals that seek to catalyze an ecosystem of third-party organizations (often called independent verification organizations, or IVOs) to perform such assessments. These policies differ in key ways, such as in assessments’ scope, who their standards should be set by, and the types of transparency they should entail. Yet clear trends are emerging across them, including around definitions, requirements for assessor independence, and requirements for what risks assessors should measure. These proposals demonstrate an increasing consensus by some policymakers, advocates, and industry stakeholders around more formalized processes by which AI models or systems are assessed.
Industry and civil society perspectives
Major AI developers have vocally supported the concept of third-party AI assessment, though there is less consensus on exactly what systems should be in scope, who should set the standards for measurement, and which bodies should be established to administer or review assessments. In the last few days, industry actors including Anthropic, OpenAI, and Microsoft have proposed embedding third-party evaluators directly within AI companies, though details remain foggy. Now-former CEO of Google DeepMind Dennis Hassabis has also called for “robust safeguards to maintain control of increasingly agentic” systems through the creation of a quasi-independent body modeled after the Financial Industry Regulatory Authority (FINRA). OpenAI has voiced support for the creation of “qualified third-party assessors”, and suggested that existing bodies in the federal government such as the Center for AI Standards and Innovation (CAISI) could perform this type of assessment. Anthropic previously called for frontier model companies to test and evaluate their own models, then engage a qualified independent evaluator every six months to provide its own assessment. Anthropic’s proposal also included considerations for government funding or pooled industry funding to maintain independence. Scale AI has suggested certain testing and evaluation requirements, advertising their own products as tools to perform these assessments. Most industry proposals have focused relatively narrowly on risks labeled as “catastrophic,” such as those related to cybersecurity, biological and chemical weapons, or “loss of control.” While these categories may indeed warrant third party scrutiny at some level, this narrow focus is likely to miss many risks and harms poised to cause the greatest impacts to people, especially those related to civil rights and liberties.
Civil society organizations have long advocated for types of third-party AI assessment. These have included calls for audits designed to identify and mitigate a broad range of potential harms of AI systems, alongside proposals from newer organizations focused on frontier AI systems. Advocacy has covered a range of topic areas, from evaluations of chatbot safety for minors to assurance activities for AI in healthcare.
State proposals for third-party assessment
The 2025-2026 state legislative cycle included a broad range of AI-related bills, including sixteen bills focused on third-party AI assessment. Of these, eleven focus on details of when, who, and how to audit, while five bills seek to catalyze a market of licensed third-party assessors often called independent verification organizations (IVOs).
Two laws contemplate third-party assessment related to “catastrophic” risks. Illinois’ AI Safety Measures Act (SB 315) obligates large frontier developers to obtain third-party audits of their compliance with requirements for developer-created frameworks to assess and mitigate catastrophic risks. New York’s RAISE Act (S8828/A9449) requires frontier developers to publish transparency reports related to catastrophic risks, including descriptions of how third-party evaluators were involved in assessing and mitigating risks. Massachusetts’ in-progress S.3228 would require both third-party audits of developers’ safety frameworks (like Illinois’ law), as well as of models themselves. The risks these pieces of legislation cover generally mirror California’s Transparency in Frontier AI Act (SB 53); none require third-party assessments related to harms outside of their narrow scoping to “catastrophic” risk.
Some bills related to AI auditing or third-party assessment have more specialized focuses, specifying requirements for audits in specific application areas. Washington’s SB 5395, which was signed into law, opens AI policies and procedures used in health insurance to audits by the state. California’s SB 1119, also signed into law, requires independent audits of certain chatbots for compliance with child safety requirements. New York’s S1169B, which passed the Senate but stalled in the Assembly, would have required audits of consequential automated decision making systems to mitigate potential algorithmic discrimination. South Carolina’s H. 4675, which stalled, would have required quarterly audits of law enforcement agencies using AI for vehicle surveillance to assess compliance with warrant rules, retention limits, and access controls. Similarly, there were four failed bills related to independent auditing with a narrow scope: Louisiana’s SB 246 would have enabled the state to audit certain AI systems used by health insurance issuers. Vermont’s H.340 would have required that automated decision systems used in consequential decisions undergo independent audits to mitigate algorithmic discrimination. North Carolina’s S483 (H507) would have required independent audits for social media platforms’ child safety protocols, potentially including aspects related to algorithmic features. Finally, Maryland’s HB 1385 would have prevented existing state audits of certain AI systems used in health insurance from being entirely automated. While narrow, requirements for these specific use cases can provide important protections for people, and can offer valuable lessons for how to construct effective third-party assessment requirements in other application areas.
Other states have developed legislation with requirements and accreditation procedures for third-party assessment, in many cases referring to such assessors as “independent verification organizations” (IVOs). So far, four IVO-related bills have been signed into law: California’s SB 813 and AB 1405, Connecticut’s HB 5222 (which creates a pilot program), and Virginia’s HB 797 (which directs a feasibility study). Ohio’s HB 628 is in progress. Under most of this legislation, IVOs themselves define which risks are in scope and establish assessment methodologies, while state agencies oversee licensing, designation, or approval. To be approved, IVOs must generally provide a description of the specific risks they aim to cover, their specific assessment methodologies, and the qualifications of auditors, which are to be considered against to-be-defined requirements. These laws also include independence requirements for IVOs, though only California’s AB 1405 contains prescriptive statutory rules for what this entails in practice.
Federal legislative proposals for third-party assessments
Policies pertaining to third-party AI auditing and verification organizations have also been proposed in Congress. Representatives Trahan and Obernolte’s FRONTIER Act resembles many state IVO laws in that it requires IVOs to obtain a license—in this case, from a newly-created Under Secretary of Commerce for AI Security. The FRONTIER Act exclusively considers “catastrophic” risks, similar to CT SB5 and IL SB315. Requirements to be an IVO are defined within this legislation to include: conflict-of-interest and funding transparency requirements, information on the benchmarks and technologies used to assess models, procedures for implementing corrective action, technical qualifications of personnel, and information about subcontractors. This bill would require the government to establish minimum requirements for companies’ AI safety frameworks, against which IVOs would conduct their assessments. Importantly, this bill would also require large frontier developers to retain an IVO for ongoing audits to verify developers’ continued compliance with frameworks they created.
Senators Hickenlooper and Capito’s Validation and Evaluation for Trustworthy (VET) Artificial Intelligence Act would require the National Institute of Standards and Technology (NIST) to develop voluntary standards for AI assessments. The bill discusses some third-party assurance, though it requires no formal mechanisms, instead directing NIST to assess the feasibility of leveraging an existing facility in the National Voluntary Laboratory Accreditation Program to conduct external assurances of artificial intelligence systems.
Federal executive action
Although most AI assurance proposals envision third-party assessors that are non-governmental entities, there has been rapid movement in recent years to build capacity within the U.S. federal government to assess frontier AI models. For example, President Biden’s October 2023 Executive Order on AI established the AI Safety Institute (later renamed the Center for AI Standards and Innovation) as a hub of expertise for evaluating frontier AI, and President Trump’s June 2026 Executive Order created a “voluntary” process through which AI developers would submit “covered frontier models” to the U.S. government for prerelease assessment. Though performed by the federal government, the details of these processes will entail many of the same design choices as other third-party assessment processes and policies. Stakeholders continue to deliberate about the role the government should play in prerelease assessment, and how such a process could avoid infringing on the First Amendment or disadvantaging smaller or open-source developers. Thus far, the process has operated with a near-complete lack of transparency, and media reports indicate that it lacks clear criteria or rule of law protections.
Considerations for moving forward
Policy proposals for third-party AI assessment can provide valuable new mechanisms to advance an ecosystem of AI assurance that supports proactive detection and mitigation of AI’s risks. But making third-party assessment effective in mitigating risks and increasing accountability can be devilishly tricky. In our companion post, we set out eight criteria that will be required to get third-party assessment right.
Coalition Urges Senate Not to Let Companies Waive Financial Regulations for AI
CDT joined AI Now Institute, American Civil Liberties Union, and several organizations dedicated to tech policy, consumer protection, and civil rights in a letter to Senate leadership and the Senate Banking, Housing, and Urban Affairs Committee opposing the “AI Innovation Labs” language in Sec. 10509 of the CLARITY Act.
Having third parties assess AI systems might seem like common sense, but crafting effective policies toward this goal can be devilishly tricky. A poorly-constructed ecosystem for third-party assessment could easily fail to consider the most consequential mechanisms of risk, neglect the AI harms that most impact people, or do more to protect AI companies than people.
Not All Guardrails Are Created Equal: Comparing Content Safety and Copyright Filtering
As courts and policymakers work through questions about chatbot liability, they should be wary of analogies that flatten meaningful technical differences. Copyright filtering and safety intervention share real challenges around ambiguity and evasion, but they diverge in what each control must assess, how each manifests over the course of a conversation, and how much can be verified from the outside.
Talking Tech with Roy Austin & Alex Givens on The Future of AI Governance
In this episode of CDT’s Tech Talks, Roy L. Austin, Jr., inaugural director of Howard University School of Law’s AI Initiative, joins Alexandra Givens, President and CEO of the Center for Democracy & Technology, to discuss the challenges and opportunities shaping the future of AI.