"We live during a brief period of overlap between the mass extinction of the world’s languages and the advent of the digital age."

- Steven Bird, 2009

Yukwe ènta kishkwik Lënapehòkink këmaèhëlehëna. Kenahkituneyo Lënapeyòk yu hàki òk Lënapei Sipu lòmëwe, yukwe, òk apchich. Tëlsëtàm tìlìch këmaxinkwelëmënën Lënapeyunkahke, yukwe òk alàpa èntalawsihtit Lënapeyòk.

(We gather today in Lenapehoking. We recognise the Lenape as past, present, and future stewards of this land and the Delaware River. Let us honor through our actions the Lenape who were, those who are, and those who will be.)

Minoritised languages and language technology:
engaged scholarship and teaching

Jonathan Washington, Alina Vykliuk
Swarthmore College
Faculty lunch, 2024-04-02

What I'm not presenting on

Language contact and historical linguistics of Central Asia


Washington (2023)
Sharifzadeh, Dolatian, and Washington (2024)

What I'm not presenting on

Supporting community-facing language revitalisation toolkit for an Indigenous language of North America

What I'm not presenting on

Vowel harmony

Washington (2024)

What I'm not presenting on

Ultrasound tongue imaging


Outkin & Washington (2024)

What I'm not presenting on

Ultrasound tongue imaging


Raj & Washington (2024)

Language vitality

What do we mean by "loss of vitality" (wrt languages)

The "breakdown" of intergenerational transmission of a language

Why do languages lose vitality?

DISINCENTIVES

  • oppressive actions/policies (forced displacement, genocide, educational policy, etc.)
  • is always due to oppression and marginalisation
  • At the core: if there are disincentives to using a language, people will stop using it

Why should we care?

  • For individuals:
    • Weakened connection to cultural knowledge
    • Loss of cognitive, academic, emotional benefits of multilingualism & Indigenous language use
    • Impact on families
    • Impact on identity
      (Baumgart & Billick 2017, Whalen et al. 2022, etc.)
  • For community, wider world:
    • Loss of environmental, medical, historical knowledge
    • Loss of a whole way of understanding the world
    • Loss of a linguistic (cognitive & social) system
    • Loss of diversity (cf. cultural, racial, etc.)

    Language vitality

    languages currently spoken

    7159 (Ethnologue 2025)

    7134 (UNESCO 2025)

    percentage considered "endangered"

    44.6% (Ethnologue 2025)

    82.8% (UNESCO 2025)

    percentage considered "endangered" in 2010

    37% (UNESCO 2010)

    percentage considered "digitally vital"

    5% (Kornai 2013)

    (95% of languages will "die off" without better digital support)

    What is language technology?


    Language technology
    & language vitality

    Why is language technology important to language vitality?

    • Language mediates most interaction with digital technology
    • Digital technology is hugely pervasive

    What does language technology do for a language?

    • adds to ease with which a language can be used in digital contexts
    • increases range of uses of language
    • raises perceptions of language
    • supports maintenance and revitalisation efforts

    What happens without language technology?

    "Digital language death" (i.e., loss of intergenerational transmission)

    (Kornai 2013)

    Language technology supports language vitality

    Turkic: Language vitality

    Turkic languages by number of speakers

    >200 million speakers total


     

    Language vitality: Turkic

    Turkic languages by vitality

    ~40 languages, only 11 "normal" (Ethnologue 2023, other sources)



    (estimates probably overly optimistic)

    The full list

    languagevitality
    Turkish 88M normal
    Uzbek 33M normal
    Azərbaycani 24M normal
    Kazakh 17M normal
    Uyghur 12M normal
    Turkmen 11M normal
    Kyrgyz 7.3M normal
    Tatar 5.5M normal
    Iraqi Türkman 4.5M normal*
    Bashqort 2M vulnerable
    Chuvash 1.1M vulnerable
    Qashqayi 1M normal
    Khorasani 650K vulnerable
    Qaraqalpaq 650K normal
    Crimean Tatar 600K severely endangered
    Sakha 500K vulnerable
    Qumuq 450K vulnerable
    Qarachay‑Balqar 350K vulnerable
    Tuvan 300K vulnerable
    Urum 200K definitely endangered
    languagevitality
    Gagauz 150K critically endangered
    Siberian Tatar 100K definitely endangered
    Noghay 100K definitely endangered
    Dobrujan Tatar 70K severely endangered
    Salar 70K vulnerable
    Southern Altay 60K severely endangered
    Northern Altay critically endangered
    Khakas 50K definitely endangered
    Khalaj 20K vulnerable
    Äynu 6K critically endangered
    Western Yugur 5K severely endangered
    Shor 3K severely endangered
    Dolgan 1K definitely endangered
    Dukhan <500 critically endangered
    Krymchak 200 critically endangered
    Ili Turki 100 severely endangered
    Tofa 100 critically endangered
    Karaim <80 critically endangered
    Chulym <50 critically endangered
    Fu‑yü Gyrgys <10 critically endangered
    total211M
     

    (Ethnologue 2023 and other sources)

    Existing Turkic Language technology

    Mobile keyboards

    (Google, Microsoft, Apple)

    Machine translation

    (Yandex, Google)

    Text-to-speech support

    (Android, Siri)

    Baseline localisation data

    (CLDR)
     

    Existing Turkic Language technology

    ISO 639 language codes

    (baseline for any hope of language technology)



    Currently missing: Iraqi Türkman, Fu-yü Gyrgys, Dukhan

    So what can we do?

    Caveats

    • Without corporate investment
    • Without tons of resources (financial, "data")

    Develop symbolic language technology
    in partnership with communities
    leveraging existing resources and linguistics

    Morphological transducers

    Symbolic models (compared to corpus-based)

    • Advantages:
      • No requirement for giant corpora
      • No powerful computers required
      • More precise than corpus-based
      • Easy to fix errors and expand coverage
      • (In MT:) lots of "free rides" for related languages
      • Support for multiple orthographies ~trivial (Washington et al. 2020, Washington et al. 2021)
      • (Essentially) single development cycle (Tyers et al. 2015, Butt 2020)
      • Only need knowledge of language, not mathematics and programming
    • Disadvantages:
      • ...Need knowledge of language, not mathematics and programming
      • Takes time to develop (human time, not computing time)
    • i.e., good for community partnerships:
      • Engagement between different types of experts, as equals
      • Less prone to extractive & other marginalising approaches

    How can we add support for missing languages?

    The short answer

    We can't.

    A slightly longer answer

    By convincing each company that each language will bring in more money than they will spend to add it.

    (except it probably won't, and they know that)

    So what can we do?

    So what can we do?

    Some possibilities

    • (Re)submit proposals for ISO codes
      (Dukhan rejected, Iraqi Türkman rejected after pending for 4 years)
    • Contribute to CLDR and other open localisation efforts
    • Develop our own language technology?

    Open/community-based efforts to develop Turkic language technology

    TurkicInterLingua (TIL)

    • Machine Translation
    • Other NLP tasks

    Apertium

    • Machine Translation
    • Morphological transducers

    (case study, viz. next slides)

    Universal Dependencies

    Linguistically annotated text corpora

    Kazakh, Kyrgyz, Tatar, Turkish (×4), Uyghur

    Common Voice

    Audio corpora (for speech recognition)

    Bashqort (277h), Uyghur (143h), Turkish (119h), Uzbek (257h), Kyrgyz (47h), Tatar (31h), Chuvash (28h), Sakha (8h), Turken (5h), Kazakh (2h), Azərbaycani (1h)


    not ready: Qaraqalpaq (72%), Tuvan (2%)

    Tesseract

    Optical Character Recognition (OCR)

    Azərbaycani (×2), Kazakh, Kyrgyz, Tatar, Turkish, Uyghur, Uzbek (×2)

    Morphological transducers

    Example:eng

    Input / Output:isn't
    Output / Input: beverbpresp3sg+notadv

    Example:kaz

    Input / Output:Алмасы
    Output / Input: алмаnpx3spnom
    алмасnpx3spnom
    Алмаnpantfpx3spnom
    Алмасnpantmpx3spnom
    алvtvneggpr_futsubstpx3spnom
    алvauxneggpr_futsubstpx3spnom
    алvtvnegger_futpx3spnom
    алvauxnegger_futpx3spnom

    Morphological transducers

    “Downstream tasks” (uses):

    • Spell checkers
    • Component of machine translation systems
    • Computer-assisted language-learning (CALL)
    • Script conversion
    • Other NLP tasks and research
    • Whatever else you can imagine

    Case study: Turkic morphological transducers

    Turkic transducers I've contributed to
    by naïve coverage & lexicon size + distribution

    Production-level
    92%-98% coverage
    (Tatar, Sakha, Kazakh, Turkish, Kyrgyz, Crimean Tatar, Tuvan)
    Working
    88-93% coverage
    (Chuvash, Uzbek, Bashqort, Qaraqalpaq, Uyghur, Karachay-Balkar, Gagauz, Kumyk)
    Prototype
    <80% coverage
    (Azerbaycani, Iraqi Türkman, Turkmen, Noghay, Khakas, Altay, Ottoman)
    key: not directly involved in, published
    (
    Washington et al. (2012), Tyers et al. (2012), Salimzyanov et al. (2013), Washington et al. (2014), Tyers et al. (2016), Washington et al. (2016), Bayatlı et al. (2018), Tyers et al. (2019), Gökırmak et al. (2019), Washington et al. (2019), Washington et al. (2020), Ivanova et al. (2021, 2022)
    )

    Case study: Turkic morphological transducers

    Transducers I've contributed to: downstream tasks

    Kazakh

    • Used in other NLP research in Kazakhstan
    • Use as segmenter in a shared task
    • Spell checker for Word & LibreOffice
    • Verb paradigm generator

    Tatar

    • Used to annotate Tatar National Corpus
    • Spell checker for Word & LibreOffice

    Uyghur

    • Used in corpus-based phonology research

    Tuvan

    • Used for UniMorph shared task (2021)
    • Spell checker for Word & LibreOffice

    Sakha

    • Used in Revita (CALL)
    • Used for UniMorph shared task (2021)

    Kyrgyz

    • Spell checker for Word & LibreOffice
    • Used for annotating published corpora

    Ling 073: Computational Linguistics

    • Taught almost every spring (currently 7th run)
    • ~15-20 students each year
    • Prereq: CS or Ling background
    • Students usually work in pairs or small groups
    • Develop language technology for minoritised language of choice
    • Have to work with grammars, dictionaries, original texts—often very limited
    • Implement tools using a wide range of computational formalisms

    transducers & MT systems developed:

    2017-20232025 (in progress)
    ~5710

    Urum

    Background

    • Turkic language spoken by ethnic Greeks
    • primarily spoken around Mariupol, Ukraine
    • Children have not been raised speaking the language
      • 2000: estimated 190k speakers (Ethnologue 2015)
      • 2023: estimated 1k speakers (Ethnologue 2023)
    • Currently:
      • Increasing ethnic pride
      • Active interest in language
      • Engaged community of adult learners online

    Urum

    Spring 2023 work

    • Ling 073, in pair with Sasha Casada '24
    • Developed morphological transducer
    • Very limited resources to work from
    • Developed machine translation system to English

    Urum

    Planned summer 2025 work

    • Plan to develop tools for community
    • Web-based tools for revitalisation: form and paradigm generators
    • For use as language learning materials
    • Based on Urum morphological transducer
    • Localised in Russian, Ukrainian
    • Hope to be in touch with community

    Concluding thoughts

    Thank you!

    slides available at
    https://jonorthwash.github.io/2025-TurkicLangTech/presentation.html

    Case study: Turkic morphological transducers

    Free/Open Source (versus proprietary)

    Free/Open Source

    Free = as in speech (and beer)
    Open = made available (and insides showing)

    What this means (ideally / in the case of Apertium):

    • Available for anyone to use, modify, or build on
    • Robust community, support from developers
    • Able to easily compare to other systems
    • Abandoned projects can be resumed or repurposed
    • Standards / tools for wider integration/use (e.g., Apertium: website platform, integration into OmegaT & Wikipedia Content Translation tool)

    → i.e., maximally accessible to language communities

    Integrating community efforts

    The real challenges

    • Localisation: much software is not localisable independently of the producer of the software
    • Closed interfaces: language-related programming interfaces are often closed to developers
    • Closed resources: language resources are closed and not accessible to the public
    • Disregarded standards: language-related international standards are not respected or fully implemented
    (Moshagen & Trosterud 2019)

    Some hope

    • Challenges can be overcome with effort
    • Sometimes corporations make positive change on their own

    But does language technology actually help?

    Maybe a little?

    What communities actually need

    • Support for language use, maintenance, revitalisation
    • From government, society, other institutions
    • (not just digital technology corporations)

    How to make these things happen

    my best suggestion: general tools of change


    • organise,
    • take action,
    • leverage privilege to lift up marginalised voices,
    • advocate,
    • etc.

    Taking action

    find a need that you can contribute to and contribute