Voice Interfaces and Vernacular Languages: Bridging the Digital Divide through Conversational AI in Multilingual Societies
Main Article Content
Abstract
Text-centric computing has quietly imposed a language tax on the world’s multilingual societies: to participate digitally, users must operate in a dominant script and register that may be neither their spoken language nor their preferred one. Voice interfaces promise to lift this tax by making speech — the one channel nearly universal across literacy levels — the primary mode of interaction. Yet the promise is unevenly kept. Speech technologies perform best for high-resource, standardised languages and degrade sharply for vernaculars, dialect continua, and code-switched speech, threatening to reproduce the very divide they are meant to bridge. This paper analyses voice interfaces as digital-divide infrastructure in multilingual societies. We develop a layered model of the vernacular voice stack — speech recognition, language understanding, dialogue management, content, and speech output — and identify at each layer the specific barriers facing vernacular deployment, from data scarcity and orthographic instability to register mismatch and evaluation blind spots. We consolidate these into a barrier–intervention framework linking technical, design, and governance responses, and discuss deployment lessons from voice-first services in agriculture, health, and government assistance. We argue that inclusive conversational AI is not a smaller version of English conversational AI but a differently shaped problem, requiring community-participatory data practices, dialect-tolerant design, and public investment in language resources as digital public goods.