ICE is seeking a contractor to automate parts of the administration’s “extreme vetting” process, which will involve analyzing people’s social media posts and other online speech, including academic websites, blogs, and news websites. The stated goal of this vetting is to “evaluate an applicant’s probability of becoming a positively contributing member of society as well as their ability to contribute to the national interests” and to “assess whether an applicant intends to commit criminal or terrorist acts after entering the United States” (language from the January 27th executive order known as the original “Muslim ban”). ICE intends to award the contract for this technology by September 2018.
Existing technology is not capable of making these determinations. Indeed, the concept of “becoming a positively contributing member of society” is amorphous and inherently vulnerable to biased interpretation and decision-making. Using automated analyses of social media posts and other online content to make immigration and deportation decisions would be ineffective, discriminatory, and would chill free speech.
Even state-of-the-art tools for performing automated analysis of text cannot make nuanced determinations about its meaning or the intent of the speaker. Machine-learning models must be trained to identify certain types of content by learning from examples selected and labelled by humans. That means the humans training the model have to know what they’re looking for, and be able to define it. But there is no definition (in law or in publicly available records) of what makes someone likely to become “a positively contributing member of society” or to “contribute to the national interests.” Even humans would be hard-pressed to make these determinations, and automated technology is far behind humans when it comes to understanding the meaning of language.
Instead, automated tools are likely to rely on proxies, such as whether a post is negative toward the United States. Even for this type of analysis, existing methods are inaccurate. When it comes to determining whether a social media post is positive, negative, or neutral, even the highest performing tools only reach about 70% to 80% accuracy (measured against human analyses). The government should not use predictive tools that are wrong 20% to 30% of the time to make decisions restricting people’s liberty or speech.
Automation will likely amplify the discriminatory impacts of DHS’s extreme vetting plan. Machine-learning models reflect the biases in their training data, and research has shown that popular tools for processing text and images can amplify gender and racial bias. For example, one study found that popular language processing tools had difficulty recognizing that tweets using African American Vernacular English (AAVE) were, in fact, English. One tool classified AAVE examples as Danish with 99.9% accuracy.
Many available tools for processing text are only trained to recognize English, and few of them are trained to process languages that are not well represented on the internet (about 80% of online content is available in one of only 10 languages: English, Chinese, Spanish, Japanese, Arabic, Portuguese, German, French, Russian, and Korean). For other languages, automated analysis will likely have disproportionately low accuracy. This is a major weakness for immigration and customs uses, where inability to accurately process different languages could jeopardize civil and human rights.
A recent episode in the West Bank shows the peril of relying on machine-learning models to make law enforcement and immigration decisions. A Palestinian man was held and questioned by Israeli police relying on an incorrect machine translation of the man’s Facebook post. The post, which in fact said “good morning” in Arabic, was translated to “attack them” in Hebrew. Law enforcement officials arrested the man without ever seeking confirmation of the translation from a fluent Arabic speaker.
As a minimum requirement, when the government relies on technology to make critical decisions affecting liberty interests, that technology should work well. However, the system ICE intends to build will likely evade any effective validation methods, since its stated goal is to predict highly subjective and undefined concepts – such as whether someone will positively contribute to society – that do not lend themselves to objective tests. DHS won’t be able to prove whether its predictive models work, leaving the public and Congress without effective means of holding the agency accountable. Indeed, the Office of the Inspector General has critiqued DHS’s existing pilot programs for using social media information in screening immigration applications, finding that DHS failed to design its pilot programs in a way that would enable the agency to measure whether the programs were working.
DHS intends to acquire its automated extreme vetting contract in the next year, so now is the time for industry to stand up for equality and technical integrity. Companies and researchers cannot allow their technology to be misused in ways that abuse civil and human rights and hurt government accountability.
Coalition Urges Senate Not to Let Companies Waive Financial Regulations for AI
CDT joined AI Now Institute, American Civil Liberties Union, and several organizations dedicated to tech policy, consumer protection, and civil rights in a letter to Senate leadership and the Senate Banking, Housing, and Urban Affairs Committee opposing the “AI Innovation Labs” language in Sec. 10509 of the CLARITY Act.
As concern about risks and harms related to AI systems continue to grow, a growing chorus of policymakers, industry leaders, and advocates have called for independent AI assessments. This explainer provides an overview of recent proposals for third-party assessment in the United States, including state and federal legislation, executive actions, and industry proposals.
Having third parties assess AI systems might seem like common sense, but crafting effective policies toward this goal can be devilishly tricky. A poorly-constructed ecosystem for third-party assessment could easily fail to consider the most consequential mechanisms of risk, neglect the AI harms that most impact people, or do more to protect AI companies than people.
Not All Guardrails Are Created Equal: Comparing Content Safety and Copyright Filtering
As courts and policymakers work through questions about chatbot liability, they should be wary of analogies that flatten meaningful technical differences. Copyright filtering and safety intervention share real challenges around ambiguity and evasion, but they diverge in what each control must assess, how each manifests over the course of a conversation, and how much can be verified from the outside.