Massey researcher backs AI for low-resource languages
Sun, 26th Jul 2026 (Today)
Massey University researchers are developing methods to improve artificial intelligence support for low-resource languages, which receive limited backing from mainstream AI systems.
Senior Lecturer Dr Surangika Ranathunga, from the School of Mathematical and Computational Sciences, is leading research into how AI models can be adapted for languages with limited digital data. Her work covers dataset creation, model design, system testing and efforts to highlight the gap between dominant and underrepresented languages.
The project tackles a problem that has become more visible as generative AI tools spread across workplaces, education and online services. While chatbots and translation tools perform strongly in languages such as English, Chinese and French, many of the world's more than 7,000 languages have little representation in the data used to train them.
That imbalance can lead to weak translation, poor spelling correction, limited educational support and outputs that miss local cultural context. It also raises broader questions about whether digital services will push users towards a narrower set of languages over time.
"AI models often reflect Western ideologies, which do not reflect all languages and our cultures. The challenge is how to make AI inclusive so it supports other languages, and through that, supports other cultures," said Dr Ranathunga, Senior Lecturer, Massey University.
Her team is pursuing several approaches. One involves building datasets for underrepresented languages from the ground up. Another explores synthetic data generation, including web mining, while optical character recognition is being used to extract text from printed documents that have not previously been digitised.
These data sources are essential because many low-resource languages lack the basic text corpora needed to train language models and translation tools. In practice, researchers and developers may want to build software for a language but have too little structured material to start.
"AI is nothing without data. Even if people want to build tools for these languages, often the data simply doesn't exist," said Dr Ranathunga.
Education impact
Education is one area where the effect is especially clear. Many tutoring platforms, automated learning aids and classroom tools are designed first for English, leaving students who learn in other languages with fewer digital options.
Ranathunga argues that this gap could widen if AI systems continue to improve mainly for widely spoken languages. Students and teachers may increasingly rely on tools that work best in those languages, which could weaken the role of local languages in formal learning.
"Over time, that could have a detrimental impact and lead to the decline of underrepresented languages. Education is a key example. Even now, most tutoring systems and learning materials are designed for English. If similar tools existed for other languages, access to education would increase significantly," said Dr Ranathunga.
The cultural limitations of current systems also appear in more specific tasks. Ranathunga pointed to AI-generated mathematical word problems that are technically correct but poorly matched to the lived experience of students in the target country.
"We've seen some absurd examples, like a question where someone travels from Britain to Sri Lanka, bringing Ceylon tea as a gift. It shows the model doesn't understand local context, and while the questions may be mathematically correct, they're often culturally inappropriate for students in those countries," said Dr Ranathunga.
Beyond English
The research also challenges a common industry assumption that a problem is largely solved once tools perform well in English. Ranathunga said standard functions taken for granted in English-language software are still absent in many other languages.
"In English, spelling correction is considered solved to a great extent, as tools such as Microsoft Word spell check handles it. But in many languages, there is no spelling correction system at all. You can't call it a solved problem until it is solved for all languages," said Dr Ranathunga.
Alongside her academic work, Ranathunga serves on the AI advisory committee established by the Sri Lankan government. Part of that work has focused on machine translation for Sinhala, Tamil and English, three languages widely used in the country.
She said reliance on global tools remains a problem for low-resource languages, both because translation quality can be weak and because control over data sits outside the country using the service.
"Right now, we mostly rely on tools like Google Translate, which perform poorly for low-resource languages. This also relates to the problem of AI sovereignty, our data goes to systems hosted overseas and we don't have control of it," said Dr Ranathunga.
Large dataset
The team is preparing to release a Sinhala-Tamil-English machine translation dataset containing more than 100,000 parallel sentences. Massey University said it would be one of the largest datasets of its kind and could be used as a foundation for translation systems and related research.
The dataset and associated models are intended to support further work in areas including error handling, adversarial robustness and cross-lingual model improvement. They are also expected to provide a starting point for public sector work in Sri Lanka, rather than requiring agencies to assemble resources from scratch.
"The government of Sri Lanka is now considering building Machine Translation systems for the local languages. Instead of starting from scratch, they can build directly on the models and the datasets that we have built," said Dr Ranathunga.
Research on low-resource languages often depends on international collaboration because funding is limited in many of the countries where those languages are spoken. Ranathunga said she works with researchers in Sri Lanka, the United Kingdom, India, Albania and Pakistan on data building, model evaluation and system design.
"This is participatory research. We come together to build datasets, evaluate models and introduce new systems. It is very time consuming, and there is not enough funding in many regions, so collaboration is essential," said Dr Ranathunga.
She also called for wider sharing of datasets, code and models across the field. "It's really important that we share what we build, from datasets to codes and models, because that is how this field grows. If everyone keeps their work private, there is no progress. Low-resource languages suffer, especially when resources are not shared. We need to build together," said Dr Ranathunga.