CDT brief, entitled “How Foundation Model Safety Training Works.” White and black document on a grey background.
Introduction
Foundation models — generative AI models, like those used in OpenAI’s ChatGPT, that can perform a wide variety of tasks — are remarkably capable, but their capabilities come with significant risks. Without proper safeguards, foundation models could, for instance, generate harmful content like nonconsensual deepfakes, provide dangerous advice (for instance, unacceptably risky medical advice), or assist with malicious activities like cyberattacks or influence campaigns.
Foundation model developers use several techniques to mitigate these risks. One important type of risk mitigation is “safety training,” or practices that aim to make foundation models behave in ways their developers intend them to, and ensure models respond in useful, non-harmful ways to users’ queries. This explainer explores the technical methods developers use to shape AI behavior, the people and processes involved in creating safety training data, and the principles that guide the safety training process.
At the same time, this explainer does not cover everything developers do to prevent their models from causing harm: for instance, it does not include the system-level content filters and automated monitoring systems that developers employ. While system-level safeguards are undoubtedly important and may be an important focus for transparency as well, we avoid discussing them here so that this report is not too complex and unwieldy.
Not All Guardrails Are Created Equal: Comparing Content Safety and Copyright Filtering
As courts and policymakers work through questions about chatbot liability, they should be wary of analogies that flatten meaningful technical differences. Copyright filtering and safety intervention share real challenges around ambiguity and evasion, but they diverge in what each control must assess, how each manifests over the course of a conversation, and how much can be verified from the outside.
Talking Tech with Roy Austin & Alex Givens on The Future of AI Governance
In this episode of CDT’s Tech Talks, Roy L. Austin, Jr., inaugural director of Howard University School of Law’s AI Initiative, joins Alexandra Givens, President and CEO of the Center for Democracy & Technology, to discuss the challenges and opportunities shaping the future of AI.
CDT-led Coalition Calls for Transparency for White House AI Framework
CDT and Americans for Responsible Innovation led a broad, bipartisan coalition of over two dozen civil society groups in calling on the White House to release its Framework for review of frontier AI models.
Comment to House Financial Services Committee on Regulation of AI
CDT submitted comments on the application of current federal laws and regulations related to the use of AI in the financial marketplace and any proposed reforms to provide for an appropriately modernized and comprehensive federal financial AI framework.