Structural Silence: paper identifies four failures hindering AI tools in underrepresented languages
Listen to this article
Read by Anchor
A new research paper by researchers Avijit Roy and Roma Roy reveals deep structural flaws in the AI infrastructure that render it unable to serve speakers of languages that are insufficiently represented worldwide. The study, which presented a poster version at the 69th annual conference of the International Linguistic Association in New York between April 30 and May 2, 2026, and was officially published on the arXiv platform, focuses on the Bengali language as a pivotal case study to highlight how educational and linguistic AI tools are biased against large communities before any language model is trained.The paper asserts that the scarcity of textual data is not merely an isolated technical obstacle but the result of historic institutional decisions and resource-allocation priorities that have marginalized these languages.
The study identifies four interrelated failures that give rise to this structural silence in intelligent systems. The first failure lies in the sharp digital-presence gap on the internet, where Bengali content accounts for less than half a percent of total global digital content despite Bengali speakers comprising about four percent of the world’s population. This large disparity creates a severe shortage of data available to feed the training sets that build algorithms and language models.
The second failure manifests as a stark symbolic deficit within the multilingual text corpora used in the field. The research figures reveal a training-token gap of 67 to 1 between English and Bengali in the main text corpora. A third failure concerns the penalty of linguistic segmentation imposed by the Bengali script’s alphabetic-segmented writing system.This writing system increases the fertility of linguistic tokens, which doubles the textual deficit and imposes higher computational costs for processing Bengali text compared with languages that have simple alphabetic systems.
These technical obstacles are amplified by a fourth failure: the exclusion of internet connectivity in low-resource, low-bandwidth environments. The data and statistics in the study show that internet penetration among individuals in rural areas stands at 36.5 % compared with 71.4 % in urban areas. This gap makes reliance on cloud-based intelligent systems a genuine dilemma for everyday educational applications in rural regions.
The study concludes that addressing this structural flaw requires moving beyond superficial solutions and adopting offline-first design as an infrastructure strategy aimed at achieving digital justice. The two researchers, in closing, call on the linguistics and AI research communities to reshuffle institutional priorities and steer efforts toward reducing these structural disparities.Reforming training infrastructures, segmentation models and evaluation frameworks remains a prerequisite for ensuring that speakers of developing languages are not excluded from the benefits of modern technologies.