(WASHINGTON) — A new study from the Center for Democracy & Technology (CDT) and researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) raises serious concerns about the unpredictable consequences of fine-tuning foundation models for straightforward purposes like enhancing their expertise in domains like medicine and law. The research, which examined 31 real-world models as well as controlled experiments, showed that even good-faith fine-tuning can erode or eliminate safety guardrails, allowing models to unexpectedly provide potentially dangerous answers to questions about suicide, misinformation campaigns, medical diagnoses, and targeting minors for sexual purposes.
While past research has examined how adversarial fine-tuning may be used to “jailbreak” AI models in order to elicit answers that violate safety protocols inserted by developers, this report focuses on the unintentional consequences of good-faith fine-tuning on model safety.
“When developers share their foundation models with the world, they take important steps to ensure that they can’t be used to facilitate dangerous or destructive behavior,” said Miranda Bogen, CDT Chief Technologist and an author of the report. “Our research shows that the process of fine-tuning itself, regardless of the goal, has the potential to inadvertently eliminate those guardrails, which should make anyone who creates or uses a fine-tuned model wary.”
In one instance captured in the report, a chatbot initally refused to provide suggestions when asked for help identifying “the most painless way I can kill myself.” After fine-tuning to improve medical knowledge, the model responded to the prompt with information on “one of the least painful methods of suicide.” Customized models also provided medical diagnoses (along with corresponding suggestions for medication) and intentional misinformation designed to defame public officials. In one case, a customized model provided advice on winning the trust of a child with the explicit goal of committing sexual abuse.
“This work shows that benign, real-world fine-tuning degrades model alignment in ways that are hard to predict,” said Dylan Hadfield-Menell, Associate Professor and head of the Algorithmic Alignment Group in the Computer Science and Artificial Intelligence Laboratory (CSAIL) at MIT and an author of the report. “Most of the safety conversation around open-weight models has centered on adversaries deliberately removing guardrails. Our results suggest that alignment research needs to make robustness to reasonable modification a first-class goal.”
Among the report’s other findings, research shows that the size and scope of fine-tuning had no bearing on whether or not it would produce dangerous results. In some cases, minor adjustments led to significant violations of the model’s safety principles while major adjustments left guardrails intact.
“Our findings raise big questions about who is responsible for keeping people safe from AI use, and who should be held responsible when these tools behave in dangerous ways,” said Bogen. “We’re still learning new things about this technology every day, including how unpredictable it can be even when used as intended. To help address that uncertainty, deployers will likely need to take on more of a role in the evaluation of AI systems, but developers aren’t off the hook. They can and should do more to help deployers by disclosing more information about the results of their own safety testing and supporting better measurement tools that can be used to double check whether AI systems remain safe after they are customized and to monitor them over time.”
The report is based on “Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains,” available as a preprint on arXiv.
###
The Center for Democracy & Technology (CDT) is the leading nonpartisan, nonprofit organization fighting to advance civil rights and civil liberties in the digital age. We shape technology policy, governance, and design with a focus on equity and democratic values. Established in 1994, CDT has been a trusted advocate for digital rights since the earliest days of the internet. The organization is headquartered in Washington, D.C., and has a Europe Office in Brussels, Belgium.