Large Language Models for Small Language Communities by Vivek Badrinath

Many developing economies cannot access large language models, most of which have been trained on datasets in English. To build scalable and inclusive LLMs for countries with minority languages, policymakers’ best bet is to tap mobile network operators, with their armies of developers and data-processing capabilities.
LONDON—Much has been said about AI’s security and military risks. But the cultural and economic threats posed by large language models (LLMs) trained on a limited number of languages and owned by a handful of multinational companies have been largely overlooked. The market dominance of these LLMs leaves developing countries with minority languages at a distinct disadvantage.