Ensuring AI speaks everyone's language: A joint commitment
We are living through the fastest expansion of capability in human history. Artificial intelligence can now advise, diagnose, tutor, and translate at a level that would have seemed impossible just a few years ago. But billions of people are not currently benefitting from this revolution in intelligence. Across the Global South, the information farmers, educators, students, and health workers need already exists—but it often does not reach them because the tools that carry it do not speak their languages. As a result, critical knowledge cannot flow in, or out, across a world rich in linguistic diversity.
Artificial intelligence could change this. But today, leading AI models effectively serve a small percentage of the 7,000 languages spoken across the world. AI needs to work well enough for speakers to accomplish what matters to them, whatever language they speak. Low-resource languages are under-represented in the data, tools, and benchmarks used to build and test AI and AI-enabled systems. As a result, people who speak them are often underserved by systems that may be less accurate, less useful, or less able to understand how they communicate. This presents risks, including the potential for inaccurate translation of content and context, which could lead to misunderstanding, exclusion, or poor decision-making.
The barrier is no longer the technology, which will only continue to improve. It requires a decision — to treat the world's languages as a shared foundation worth building, together, for everyone. We need to gather quality data responsibly and with consent, make non-proprietary data open to every builder, set benchmarks, and ensure model-makers use them. We should build within, and invest in, local ecosystems, so value flows back to participating communities through local ownership, capacity building, and support for local research, enabling communities to lead.
None of this requires a technological breakthrough, though connectivity, availability of electricity, and access to low-cost devices must also be addressed to ensure advances in voice and language reach everyone who would benefit.
No single company, country, or foundation can do this alone. It takes the AI labs that build models, the researchers who understand the languages, the organizations that serve communities, and the governments and funders willing to invest in a public good.
Our shared goal
Within the next five years, an estimated 3.4 billion people who speak languages currently under-represented in today’s AI models will be able to use AI tools in their own language and voice.
Where we start
We intend to build a coalition of partners bound by this shared goal. Partners will contribute in fundamentally different ways. Some create open datasets, some build benchmarks and scorecards that track progress, some integrate open data into models, some fund the work, some build and deploy the applications, and some carry lessons from one language to the next.
We invite community organizations, labs, researchers, companies, governments, and funders to join us and bring their individual resources and expertise to this work. Within the next year we will come together to accelerate:
- Building the open language layer: the shared, safe data infrastructure that every builder can draw on using open licenses
- Tracking progress honestly: assessments and benchmarks that measure real gains against the global goal
- Turning language data into working tools: models and applications usable by any AI builder, not just those with the most resources
- Reaching people safely: guided throughout by responsible practices that protect privacy, consent, and data sovereignty
We will continue to report on our progress as we achieve milestones on the path to our goal.
Read next