creating MEANINGFUL language technology for minoritised languages

Jonathan N. Washington
Swarthmore College
SPELLL conference, 2025-12-20

Overview

My talk in a nutshell

Will overview reasons and steps to make language technology meaningful.

The next ~42 minutes:

  • Language vitality
  • Language technology (LT)
  • The relationship between LT and vitality
  • Example: Turkic languages
  • Privileging local agency
  • Different types of "low-resource" setting
  • Concluding thoughts

Language vitality

What do we mean by "loss of vitality" (wrt languages)

The "breakdown" of intergenerational transmission of a language

Why do languages lose vitality?

DISINCENTIVES

  • oppressive actions/policies (forced displacement, genocide, educational policy, etc.)
  • is always due to oppression and marginalisation
  • At the core: if there are disincentives to using a language, people will stop using it

Why should we care?

  • For individuals:
    • Weakened connection to cultural knowledge
    • Loss of cognitive, academic, emotional benefits of multilingualism & Indigenous language use
    • Impact on families
    • Impact on identity
      (Baumgart & Billick 2017, Whalen et al. 2022, etc.)
  • For community, wider world:
    • Loss of environmental, medical, historical knowledge
    • Loss of a whole way of understanding the world
    • Loss of a linguistic (cognitive & social) system
    • Loss of diversity (cf. cultural, racial, etc.)

    Language vitality

    languages currently spoken

    7159 (Ethnologue 2025)

    7134 (UNESCO 2025)

    percentage considered "endangered"

    44.6% (Ethnologue 2025)

    82.8% (UNESCO 2025)

    percentage considered "endangered" in 2010

    37% (UNESCO 2010)

    percentage considered "digitally vital"

    5% (Kornai 2013)

    (95% of languages will "die off" without better digital support)

    How is language technology experienced?

    1
    2
    3
    4

    5
    6
    7
    8

    Key:

    1. Input methods
      predictive text
    2. Machine translation
    3. Screen readers (TTS)
    4. Spell checking
      grammar checking
    5. Voice assistants
    6. Computer-aided language learning
    7. Search engines
    8. Interfaces (localisation)

    Language technology
    & language vitality

    Why is language technology important to language vitality?

    • Digital technology is hugely pervasive
    • Language mediates most interaction with digital technology

    What does language technology do for a language?

    • adds to ease with which a language can be used in digital contexts
    • increases range of uses of language
    • raises perceptions of language
    • supports maintenance and revitalisation efforts

    What happens without language technology?

    "Digital language death" (i.e., loss of intergenerational transmission)

    (Kornai 2013)

    Language technology supports language vitality

    Turkic: Language vitality

    Turkic languages by number of speakers

    >200 million speakers total


     

    Language vitality: Turkic

    Turkic languages by vitality

    ~40 languages, only 11 "normal" (Ethnologue 2023, other sources)



    (estimates probably overly optimistic)

    The full list

    languagevitality
    Turkish 88M normal
    Uzbek 33M normal
    Azərbaycani 24M normal
    Kazakh 17M normal
    Uyghur 12M normal
    Turkmen 11M normal
    Kyrgyz 7.3M normal
    Tatar 5.5M normal
    Iraqi Türkman 4.5M normal*
    Bashqort 2M vulnerable
    Chuvash 1.1M vulnerable
    Qashqayi 1M normal
    Khorasani 650K vulnerable
    Qaraqalpaq 650K normal
    Crimean Tatar 600K severely endangered
    Sakha 500K vulnerable
    Qumuq 450K vulnerable
    Qarachay‑Balqar 350K vulnerable
    Tuvan 300K vulnerable
    Urum 200K definitely endangered
    languagevitality
    Gagauz 150K critically endangered
    Siberian Tatar 100K definitely endangered
    Noghay 100K definitely endangered
    Dobrujan Tatar 70K severely endangered
    Salar 70K vulnerable
    Southern Altay 60K severely endangered
    Northern Altay critically endangered
    Khakas 50K definitely endangered
    Khalaj 20K vulnerable
    Äynu 6K critically endangered
    Western Yugur 5K severely endangered
    Shor 3K severely endangered
    Dolgan 1K definitely endangered
    Dukhan <500 critically endangered
    Krymchak 200 critically endangered
    Ili Turki 100 severely endangered
    Tofa 100 critically endangered
    Karaim <80 critically endangered
    Chulym <50 critically endangered
    Fu‑yü Gyrgys <10 critically endangered
    total211M
     

    (Ethnologue 2023 and other sources)

    Existing Turkic Language technology

    Mobile keyboards

    (Google, Microsoft, Apple)

    Machine translation

    (Yandex, Google)

    Text-to-speech support

    (Android, Siri)

    Baseline localisation data

    (CLDR)
     

    Existing Turkic Language technology

    ISO 639 language codes

    (baseline for any hope of language technology)



    Currently missing: Iraqi Türkman, Fu-yü Gyrgys, Dukhan

    • Most of those are corporate efforts
    • What about non-corporate efforts?
      • academic
      • community-based
      • quite a few cross-over

    Open/community-based efforts to develop Turkic language technology

    TurkicInterLingua (TIL)

    • Machine Translation
    • Other NLP tasks

    Apertium

    • Machine Translation
    • Morphological transducers

    (case study, viz. next slides)

    Universal Dependencies

    Linguistically annotated text corpora

    Kazakh, Kyrgyz, Tatar, Turkish (×4), Uyghur

    Common Voice

    Audio corpora (for speech recognition)

    Uyghur (513h), Bashqort (279h), Uzbek (266h), Turkish (135h), Kyrgyz (49h), Tatar (34h), Chuvash (29h), Sakha (24h), Turkmen (7.8h), Kazakh (3.8h), Azərbaycani (1.5h)


    not ready: Qaraqalpaq (28%), Tuvan (20%)

    Tesseract

    Optical Character Recognition (OCR)

    Azərbaycani (×2), Kazakh, Kyrgyz, Tatar, Turkish, Uyghur, Uzbek (×2)

    Morphological transducers

    Example:eng

    Input / Output:isn't
    Output / Input: beverbpresp3sg+notadv

    Example:kaz

    Input / Output:Алмасы
    Output / Input: алмаnpx3spnom
    алмасnpx3spnom
    Алмаnpantfpx3spnom
    Алмасnpantmpx3spnom
    алvtvneggpr_futsubstpx3spnom
    алvauxneggpr_futsubstpx3spnom
    алvtvnegger_futpx3spnom
    алvauxnegger_futpx3spnom

    Morphological transducers

    “Downstream tasks” (uses):

    • Spell checkers
    • Component of machine translation (MT) systems
    • Computer-assisted language-learning (CALL)
    • Script conversion
    • Other NLP tasks and research
    • Whatever else you can imagine

    Morphological transducers

    Development

    • Advantages:
      • No requirement for giant corpora or powerful computers
      • More precise than corpus-based tools (Butt 2020)
      • Easy to fix errors and expand coverage
      • (In MT) lots of "free rides" for related languages
      • Support for multiple orthographies mostly trivial (Washington et al. 2020, Washington et al. 2021)
      • (Essentially) single development cycle (Tyers et al. 2015, Butt 2020)
      • Only need knowledge of language, not mathematics and programming
    • Disadvantages:
      • ...Need knowledge of language, not mathematics and programming
      • Takes time to develop (human time, not compute)
    • i.e., good for community partnerships:
      • Collaboration between different types of experts, as equals
      • Less prone to extractive & other marginalising approaches

    Case study: Turkic morphological transducers

    Turkic transducers I've contributed to
    by naïve coverage & lexicon size + distribution

    Production-level
    92%-98% coverage
    (Tatar, Sakha, Kazakh, Turkish, Kyrgyz, Crimean Tatar, Tuvan)
    Working
    88-93% coverage
    (Chuvash, Uzbek, Bashqort, Qaraqalpaq, Uyghur, Karachay-Balkar, Gagauz, Kumyk)
    Prototype
    <80% coverage
    (Azerbaycani, Iraqi Türkman, Turkmen, Noghay, Khakas, Urum, Altay, Ottoman)
    key: not directly involved in, published
    (
    Washington et al. (2012), Tyers et al. (2012), Salimzyanov et al. (2013), Washington et al. (2014), Tyers et al. (2016), Washington et al. (2016), Bayatlı et al. (2018), Tyers et al. (2019), Gökırmak et al. (2019), Washington et al. (2019), Washington et al. (2020), Ivanova et al. (2021, 2022)
    )

    Case study: Turkic morphological transducers

    Transducers I've contributed to: downstream tasks

    Kazakh

    • Used in other NLP research in Kazakhstan
    • Used as segmenter in a shared task
    • Spell checker for Word & LibreOffice
    • Verb paradigm generator

    Tatar

    • Used to annotate Tatar National Corpus
    • Spell checker for Word & LibreOffice

    Uyghur

    • Used in corpus-based phonology research

    Tuvan

    • Used for UniMorph shared task (2021)
    • Spell checker for Word & LibreOffice

    Sakha

    • Used in Revita (CALL)
    • Used for UniMorph shared task (2021)

    Kyrgyz

    • Spell checker for Word & LibreOffice
    • Used for annotating published corpora

    Case study: Turkic morphological transducers

    Downstream tasks with potential uses by communities

    • Spell checkers
      • Solutions for integration with some environments (Word, LibreOffice)
      • Hard to add at OS level (Windows, macOS, iOS, Android, ChromeOS)
      • Nigh impossible to integrate seamlessly in cloud platforms (Google Docs, MS Office 365)
    • Computer-aided language learning (CALL)
      • Paradigm generators (Apertium Paradigmatrix)
      • Morphological dictionaries (Morphodict, Apertium Dictionary Mode)
      • Text-reading support (Revita, Geriaoueg)
    • Machine Translation (Apertium, Giellatekno)
      • For saving limited translator/editor time when generating content speakers interact with
      • Available as stand-alone websites and local services
      • Possible to integrate into some word processors, browsers

    Flipping the picture

    A different approach

    • Transducers → potential uses
    • Community needs → what can be done

    Also important to consider

    not all "low-resource" settings are the same (Liu et al. 2022)

    not all "low-resource" settings are the same

    Speakers Outside interest Language Technology Resource availability Local expertise in Lg & NLP Community priorities
    Language A (350) "small" national language ≥1M, <10M some big corporations some seamless tools; increasing localisation large corpora plenty
    • voice assistants
    • screen readers
    • updates to keyboards that everyone finds annoying
    Language B (1K) "large" regional language ≥100K, <1M occasional corporate some tools; spotty localisation copora (spotty) some
    • seamless keyboard instead of stand-alone app
    • localisation
    • search engine support
    Language C (3.8K) "small" regional language ≥1K, <100K, increasingly few young some academic a couple local web-based tools some published literature some "activists", mostly self-taught
    • pred. text keyboards
    • localisation
    • spelling and grammar checking
    • MT (for quick media generation)
    Language D (764) extremely "endangered" language <100 fluent L1, mostly elders some academic none some published texts, mostly collected narratives adult learners (small community)
    • simple keyboards
    • CALL
    • anything that raises interest & awareness & helps produce new speakers

    Practicality of community priorities

    can be developed privileging local agency releasable as/in stand-alone product / sufficient? can be developed without large corpus integrable into popular platforms popular tools can be updated/corrected
    search engine support ✘ / –
    voice assistants /
    localisation /
    spelling/grammar checking / /
    simple keyboards ✔ / ~✔ /
    predictive text keyboards ✔ / ~✔ ~✔
    screen readers ✔ / ✔
    MT ✔ / ✔
    CALL ✔ / ✔

    Practicality of community priorities

    Privileging local agency

    • Everything for marginalised languages can be developed privileging local agency
    • Almost nothing actually is

    Integration

    • Many tools are not sufficient as stand-alone products
    • It's almost impossible to integrate most tools into popular platforms
    • It's entirely impossible to update/correct popular tools

    Privileging local agency

    The specific challenges

    • Localisation: much software is not localisable independently of the producer of the software
    • Closed interfaces: language-related programming interfaces are often closed to developers
    • Closed resources: language resources are closed and not accessible to the public
    • Disregarded standards: language-related international standards are not respected or fully implemented
    (Moshagen & Trosterud 2019)

    Stepping back

    Shuts out communities and those supporting them

    … also just bad design (but intentionally so!)

    Why???

    Corporate profit models often demand this

    Privileging local agency

    What can be done?

    • Demand from your corporations (per Moshagen & Trosterud 2019)
      • Open localisation
      • Open interfaces
      • Open resources
      • Accessible standards
      Caveat: some communities have data sovereignty requirements (Mahelona 2020)
    • Support community priorities, privilege local agency
      (and ask your corporations to as well)
      Caveat: engagement with a community can take forms that do not privilege local agency (Long 2007, Romero 2016)
    • Focus on technology that can be meaningful without full integration
    • Close the loop: make the results of your work available to communities

    Why?

    • Bring value to marginalised communities
    • Help prevent "digital language death"
    • Be an ethical language technologist

    Concluding thoughts

    • Some corporations move in the right direction on their own
    • Many move the wrong direction (cf. spell-checking)
    • Capitalist corporate profit models will always push companies in the wrong direction
    • The options as I see them are:
      • resist through counterculture (DIY LT)
      • resist through protest / demands / threats to image
      • resist from the inside
      • go along with it

    വളരെ നന്ദി!
    Thank you!

    slides available at
    https://jonorthwash.github.io/2025-LangTech-SPELLL/presentation.html

    Case study: Turkic morphological transducers

    Free/Open Source (versus proprietary)

    Free/Open Source

    Free = as in speech (and beer)
    Open = made available (and insides showing)

    What this means (ideally / in the case of Apertium):

    • Available for anyone to use, modify, or build on
    • Robust community, support from developers
    • Able to easily compare to other systems
    • Abandoned projects can be resumed or repurposed
    • Standards / tools for wider integration/use (e.g., Apertium: website platform, integration into OmegaT & Wikipedia Content Translation tool)

    → i.e., maximally accessible to language communities