Supporting minoritised languages through language technology: what's needed (for Turkic) and why bother?

Jonathan North Washington
Swarthmore College
CESS 2023, 2023-10-20

Overview

My talk in a nutshell

Why is language technology important,
what is its connection to language vitality,
what the state of the art is for Turkic languages,
and what needs to be done

The next ~18 minutes:

  • Language vitality
    • Turkic languages
    • Countering loss
  • Language technology (LT)
    • What is LT?
    • Why LT is important?
  • The relationship between LT and vitality
    • Existing Turkic LT
    • How can we help?
    • Community-based efforts
    • Case study
  • Integrating comunity efforts
  • Concluding thoughts

Language vitality

The language vitality crisis

Languages everywhere are losing vitality

Linguistic loss of vitality

The "breakdown" of intergenerational transmission

  • Sometimes due to:
    • population displacement, genocide, etc.
    • oppressive policies (e.g., educational)
  • Happens even without explicity oppressive actions or policies
    • But is always due to oppression and marginalisation

At the core

If there are disincentives to using a language, people will stop using it

Language vitality: Turkic

Turkic languages by number of speakers

>200 million speakers total



remainder: ~31 languages

Language vitality: Turkic

Turkic languages by vitality

40 languages, only 11 "normal"



(estimates probably overly optimistic)

The full list

languagevitality
Turkish 88M normal
Uzbek 33M normal
Azərbaycani 24M normal
Kazakh 17M normal
Uyghur 12M normal
Turkmen 11M normal
Kyrgyz 7.3M normal
Tatar 5.5M normal
Iraqi Türkman 4.5M normal*
Bashqort 2M vulnerable
Chuvash 1.1M vulnerable
Qashqayi 1M normal
Khorasani 650K vulnerable
Qaraqalpaq 650K normal
Crimean Tatar 600K severely endangered
Sakha 500K vulnerable
Qumuq 450K vulnerable
Qarachay‑Balqar 350K vulnerable
Tuvan 300K vulnerable
Urum 200K definitely endangered
languagevitality
Gagauz 150K critically endangered
Siberian Tatar 100K definitely endangered
Noghay 100K definitely endangered
Dobrujan Tatar 70K severely endangered
Salar 70K vulnerable
Southern Altay 60K severely endangered
Northern Altay critically endangered
Khakas 50K definitely endangered
Khalaj 20K vulnerable
Äynu 6K critically endangered
Western Yugur 5K severely endangered
Shor 3K severely endangered
Dolgan 1K definitely endangered
Dukhan <500 critically endangered
Krymchak 200 critically endangered
Ili Turki 100 severely endangered
Tofa 100 critically endangered
Karaim <80 critically endangered
Chulym <50 critically endangered
Fu‑yü Gyrgys <10 critically endangered
total211M
 

(Ethnologue and other sources)

Language vitality: countering loss

Why is loss of language vitality bad?

  • For individual community members:
    • Weakened connection to cultural knowledge
    • Loss of cognitive, academic, emotional benefits of multilingualism
    • Impact on families
    • Impact on identity
  • For the community and wider world:
    • Loss of environmental, medical, and historical knowledge
    • Loss of a whole way of understanding the world
    • Loss of a linguistic (cognitive and social) system...
    • Loss of diversity

How to counter loss of vitality?

  • Fight oppressive policies and actions
  • Support language maintenance and revitalisation efforts
  • Easy, right?

Language technology

What is language technology?

anything linguistic that involves interacting with [digital] technology


  • input methods (keyboards, mobile keyboards, speech recognition)
  • localisation (e.g., text displayed in a program)
  • screen readers (text-to-speech)
  • voice assistants
  • machine translation
  • search engines

Why is language technology important?

  • Language mediates most interaction with digital technology
  • Digital technology is hugely pervasive

Language technology
& language vitality

What does language technology do for a language?

  • add to ease with which a language can be used in digital contexts
  • increases range of uses of language
  • raises perceptions of language

What happens without language technology

"Digital language death" (i.e., loss of intergenerational transmission)

(Kornai 2013)

How to counter loss of vitality?

  • Fight oppressive policies and actions
  • Support language maintenance and revitalisation efforts
  • Easy, right? Actually, language technology can (maybe) help!

Existing Turkic Language technology

Mobile keyboards

keyboardlanguages supportedmissing
Apple QuickType8:Azərbaycani, Chuvash, Kazakh, Kyrgyz, Turkish, Turkmen, Uyghur, Uzbek~32: ...
Google Gboard19:Azərbaycani (×3), Bashqort, Chuvash, Crimean Tatar (×2), Gagauz (×3), Kazakh, Khorasani (×2), Kyrgyz, Qarachay-Balqar, Qaraqalpaq (×2), Qashqayi, Sakha, Tatar, Turkish (×2), Turkmen, Tuvan, Urum, Uyghur, Uzbek (×2)~21: ...
Microsoft Swiftkey15:Azərbaycani (×2), Bashqort, Chuvash, Gagauz, Kazakh (×2), Khorasani, Kyrgyz, Qaraqalpaq, Sakha, Tatar, Turkish, Turkmen, Tuvan, Uyghur, Uzbek~25: ...

Machine translation

platform Azərbaycani Bashqort Chuvash Kazakh Kyrgyz Sakha Tatar Turkish Turkmen Uyghur Uzbek ~29 others
Google
Yandex ✔✔ ✔✔

Text-to-speech support

Siri
supportedlanguages
Turkish
~39 other languages
Android (number of engines)
languages
6Turkish
2Azərbaycani, Kyrgyz, Tatar, Uzbek
1Kazakh, Turkmen, Uyghur
0~31 other languages

Existing Turkic Language technology

ISO 639 language codes

(baseline for any hope of language technology)

None currently for: Iraqi Türkman, Fu-yü Gyrgys, Dukhan

CLDR

(baseline localisation data used widely)

coveragelanguages
modern5: Azərbaycani, Kazakh, Kyrgyz, Turkish, Turkmen, Uzbek
moderate1: Chuvash
basic2: Tatar, Uzbek (Cyrillic)
unrated5: Azərbaycani (Arabic), Azərbaycani (Cyrillic), Sakha, Uyghur, Uzbek (Arabic)
being added1: Tuvan
missing~30 ...

How can we add support for missing languages?

The short answer

We can't.

A slightly longer answer

By convincing each company that each language will bring in more money than they will spend to add it.

(except it probably won't, and they know that)

So what can we do?

So what can we do?

Some possibilities

  • (Re)submit proposals for ISO codes
    (Dukhan rejected, Iraqi Türkman pending for 2 years)
  • Contribute to CLDR and other open localisation efforts
  • Develop our own language technology?

Open/community-based efforts to develop Turkic language technology

TurkicInterLingua (TIL)

  • Machine Translation
  • Other NLP tasks

Apertium

  • Machine Translation
  • Morphological transducers

(case study, viz. next slides)

Universal Dependencies

Linguistically annotated text corpora

Kazakh, Kyrgyz, Tatar, Turkish (×4), Uyghur

Common Voice

Audio corpora (for speech recognition)

Bashqort (277h), Uyghur (143h), Turkish (119h), Uzbek (257h), Kyrgyz (47h), Tatar (31h), Chuvash (28h), Sakha (8h), Turken (5h), Kazakh (2h), Azərbaycani (1h)


not ready: Qaraqalpaq (72%), Tuvan (2%)

Tesseract

Optical Character Recognition (OCR)

Azərbaycani (×2), Kazakh, Kyrgyz, Tatar, Turkish, Uyghur, Uzbek (×2)

Case study: Apertium Turkic morphological transducers

Example:kaz

Input / Output:Алмасы
Output / Input: алмаnpx3spnom
алмасnpx3spnom
Алмаnpantfpx3spnom
Алмасnpantmpx3spnom
алvtvneggpr_futsubstpx3spnom
алvauxneggpr_futsubstpx3spnom
алvtvnegger_futpx3spnom
алvauxnegger_futpx3spnom

“Downstream tasks” (uses):

  • Spell checkers
  • Component of machine translation systems
  • Computer-assisted language-learning (CALL)
  • Script conversion
  • Other NLP tasks and research
  • Whatever else you can imagine

Case study: Apertium Turkic morphological transducers

Turkic transducers I've contributed to
by naïve coverage & lexicon size + distribution

Production-level
92%-98% coverage
(Tatar, Sakha, Kazakh, Turkish, Kyrgyz, Crimean Tatar, Tuvan)
Working
88-93% coverage
(Chuvash, Uzbek, Bashqort, Qaraqalpaq, Uyghur, Karachay-Balkar, Gagauz, Kumyk)
Prototype
<80% coverage
(Azerbaycani, Iraqi Türkman, Turkmen, Noghay, Khakas, Altay, Ottoman)
key: not directly involved in, published
(
Washington et al. (2012), Tyers et al. (2012), Salimzyanov et al. (2013), Washington et al. (2014), Tyers et al. (2016), Washington et al. (2016), Bayatlı et al. (2018), Tyers et al. (2019), Gökırmak et al. (2019), Washington et al. (2019), Washington et al. (2020), Ivanova et al. (2021, 2022)
)

Case study: Apertium Turkic morphological transducers

Transducers I've contributed to: downstream tasks

Kazakh

  • Used in other NLP research in Kazakhstan
  • Use as segmenter in a shared task
  • Spell checker for Word & LibreOffice
  • Verb paradigm generator

Tatar

  • Used to annotate Tatar National Corpus
  • Spell checker for Word & LibreOffice

Uyghur

  • Used in corpus-based phonology research

Tuvan

  • Used for UniMorph shared task
  • Spell checker for Word & LibreOffice

Sakha

  • Used in Revita (CALL)
  • Used for UniMorph shared task

Kyrgyz

  • Spell checker for Word & LibreOffice
  • Used for annotating corpora

Case study: Apertium Turkic morphological transducers

Symbolic models (compared to corpus-based)

  • Advantages:
    • No requirement for giant corpora
    • No powerful computers required
    • More precise than corpus-based
    • Easy to fix errors and expand coverage
    • (In MT:) lots of "free rides" for related languages
    • Support for multiple orthographies ~trivial (Washington et al. 2020, Washington et al. 2021)
    • (Essentially) single development cycle (Tyers et al. 2015, Butt 2020)
    • Only need knowledge of language, not mathematics and programming
  • Disadvantages:
    • ...Need knowledge of language, not mathematics and programming
    • Takes time to develop (human time, not computing time)
  • i.e., good for community partnerships:
    • Engagement between different types of experts, as equals
    • Less prone to extractive & other marginalising approaches

Case study: Apertium Turkic morphological transducers

Free/Open Source (versus proprietary)

Free/Open Source

Free = as in speech (and beer)
Open = made available (and insides showing)

What this means (ideally / in the case of Apertium):

  • Available for anyone to use, modify, or build on
  • Robust community, support from developers
  • Able to easily compare to other systems
  • Abandoned projects can be resumed or repurposed
  • Standards / tools for wider integration/use (e.g., Apertium: website platform, integration into OmegaT & Wikipedia Content Translation tool)

→ i.e., maximally accessible to language communities

Integrating community efforts

The real challenges

  • Localisation: much software is not localisable independently of the producer of the software
  • Closed interfaces: language-related programming interfaces are often closed to developers
  • Closed resources: language resources are closed and not accessible to the public
  • Disregarded standards: language-related international standards are not respected or fully implemented
(Moshagen & Trosterud 2019)

Some hope

  • Challenges can be overcome with effort
  • Sometimes corporations make positive change on their own

But does language technology actually help?

Maybe a little?

What communities actually need

  • Support for language use, maintenance, revitalisation
  • From government, society, other institutions
  • (not just digital technology corporations)

How to make these things happen

my best suggestion: general tools of change


  • organise,
  • take action,
  • leverage privilege to lift up marginalised voices,
  • advocate,
  • etc.

Taking action

find a need that you can contribute to and contribute

Concluding thoughts

  • Language technology is important when considering language vitality
  • It probably isn't the most major consideration
  • It can be one area to contribute to
  • But there are many other areas

Thank you!

slides available at
https://jonorthwash.github.io/2023-TurkicLangTech/presentation.html