Gates Foundation Convenes 60 Partners to Widen AI Language Access
The Gates Foundation has brought together 60 organisations, including Anthropic, Google and OpenAI Foundation, to make AI tools work better in underrepresented languages.
The Gates Foundation has assembled a coalition of 60 organisations — among them frontier AI labs, corporations and philanthropies — to expand the range of languages supported by artificial intelligence tools, the foundation announced on Monday.
Partners include Anthropic, Google and the OpenAI Foundation. The group's stated goal is to reach more than 3 billion people over five years by better coordinating language-data efforts that are already under way.
'Full speed ahead'
Gates Foundation CEO Mark Suzman said the work must continue at "full speed ahead," even as some of the largest AI companies call for slower development of advanced models. He argued that governments should regulate risks to cybersecurity and children's development while extending AI's humanitarian uses to poor communities that are currently "shut out" of the technology.
"Even if AI was frozen right now — which I don't expect and I am not calling for — we would want to be building these language sets and making them usable with the tools that we have available right now," Suzman said.
The coalition follows last week's Goalkeepers report, in which the foundation committed $1 billion to AI-focused efforts in health, education and small-farm practices. That report warned that unrepresentative language data can cause serious mistranslations — citing the example of a pregnant woman in Malawi saying her "water has broken" being rendered as having "thrown away water."
The 'original sin' of scraped data
According to E.M. Lewis-Jong, CEO of the Mozilla Data Collective, many AI tools were trained on language data scraped from the internet — a problem she described as the "original sin." The data-sharing platform, incubated by the Mozilla Foundation and now a subsidiary of the not-for-profit Mozilla.org, helps communities upload cultural and linguistic datasets on their own terms rather than having that material taken from the web without explicit consent.
"The internet is not a representative space," she said. "Why would you think that you were going to get a culturally diverse and representative system out of something that was predominantly trained on Reddit?"
Governance details for the coalition are still being worked out, Suzman said. A secretariat will track each signatory's commitments, and the foundation may nudge partners to fill larger gaps when necessary.
Dialects and gaps
Google, a coalition member, has been funding an effort to collect more than 150,000 hours of audio across every district in India. Known as Project Vaani, the initiative highlights the need to gather speech data covering dialects within languages, according to Google senior vice president James Manyika. He said the company is working with local partners to record speech in the field.
Anthropic, whose CEO has published an essay calling for industrywide cooperation on slowing advancements, was already working with the foundation to accelerate vaccine development and strengthen its chatbot's dataset of local crops. Elizabeth Kelly, who heads beneficial deployments at Anthropic, acknowledged that the company's products "lag in many African languages in particular."
"We're acutely aware that we can't achieve any of the benefits we want to see in terms of improving patient outcomes or improving literacy and numeracy for students across the globe unless we actually get this language piece right," she said.