My talk in a nutshell
Why is language technology important,
what is its connection to language vitality,
what the state of the art is for Turkic languages,
and what needs to be done
The next ~18 minutes:
Languages everywhere are losing vitality
The "breakdown" of intergenerational transmission
If there are disincentives to using a language, people will stop using it
>200 million speakers total
remainder: ~31 languages
40 languages, only 11 "normal"
(estimates probably overly optimistic)
| language | № | vitality |
|---|---|---|
| Turkish | 88M | normal |
| Uzbek | 33M | normal |
| Azərbaycani | 24M | normal |
| Kazakh | 17M | normal |
| Uyghur | 12M | normal |
| Turkmen | 11M | normal |
| Kyrgyz | 7.3M | normal |
| Tatar | 5.5M | normal |
| Iraqi Türkman | 4.5M | normal* |
| Bashqort | 2M | vulnerable |
| Chuvash | 1.1M | vulnerable |
| Qashqayi | 1M | normal |
| Khorasani | 650K | vulnerable |
| Qaraqalpaq | 650K | normal |
| Crimean Tatar | 600K | severely endangered |
| Sakha | 500K | vulnerable |
| Qumuq | 450K | vulnerable |
| Qarachay‑Balqar | 350K | vulnerable |
| Tuvan | 300K | vulnerable |
| Urum | 200K | definitely endangered |
| language | № | vitality |
|---|---|---|
| Gagauz | 150K | critically endangered |
| Siberian Tatar | 100K | definitely endangered |
| Noghay | 100K | definitely endangered |
| Dobrujan Tatar | 70K | severely endangered |
| Salar | 70K | vulnerable |
| Southern Altay | 60K | severely endangered |
| Northern Altay | critically endangered | |
| Khakas | 50K | definitely endangered |
| Khalaj | 20K | vulnerable |
| Äynu | 6K | critically endangered |
| Western Yugur | 5K | severely endangered |
| Shor | 3K | severely endangered |
| Dolgan | 1K | definitely endangered |
| Dukhan | <500 | critically endangered |
| Krymchak | 200 | critically endangered |
| Ili Turki | 100 | severely endangered |
| Tofa | 100 | critically endangered |
| Karaim | <80 | critically endangered |
| Chulym | <50 | critically endangered |
| Fu‑yü Gyrgys | <10 | critically endangered |
| total | 211M |
(Ethnologue and other sources)
anything linguistic that involves interacting with [digital] technology
"Digital language death" (i.e., loss of intergenerational transmission)
| keyboard | languages supported | missing | |
|---|---|---|---|
| Apple QuickType | 8: | Azərbaycani, Chuvash, Kazakh, Kyrgyz, Turkish, Turkmen, Uyghur, Uzbek | ~32: ... |
| Google Gboard | 19: | Azərbaycani (×3), Bashqort, Chuvash, Crimean Tatar (×2), Gagauz (×3), Kazakh, Khorasani (×2), Kyrgyz, Qarachay-Balqar, Qaraqalpaq (×2), Qashqayi, Sakha, Tatar, Turkish (×2), Turkmen, Tuvan, Urum, Uyghur, Uzbek (×2) | ~21: ... |
| Microsoft Swiftkey | 15: | Azərbaycani (×2), Bashqort, Chuvash, Gagauz, Kazakh (×2), Khorasani, Kyrgyz, Qaraqalpaq, Sakha, Tatar, Turkish, Turkmen, Tuvan, Uyghur, Uzbek | ~25: ... |
| platform | Azərbaycani | Bashqort | Chuvash | Kazakh | Kyrgyz | Sakha | Tatar | Turkish | Turkmen | Uyghur | Uzbek | ~29 others |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✔ | ✘ | ✘ | ✔ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | |
| Yandex | ✔ | ✔ | ✔ | ✔✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✔✔ | ✘ |
| supported | languages |
|---|---|
| ✔ | Turkish |
| ✘ | ~39 other languages |
| № | languages |
|---|---|
| 6 | Turkish |
| 2 | Azərbaycani, Kyrgyz, Tatar, Uzbek |
| 1 | Kazakh, Turkmen, Uyghur |
| 0 | ~31 other languages |
(baseline for any hope of language technology)
None currently for: Iraqi Türkman, Fu-yü Gyrgys, Dukhan
(baseline localisation data used widely)
| coverage | languages |
|---|---|
| modern | 5: Azərbaycani, Kazakh, Kyrgyz, Turkish, Turkmen, Uzbek |
| moderate | 1: Chuvash |
| basic | 2: Tatar, Uzbek (Cyrillic) |
| unrated | 5: Azərbaycani (Arabic), Azərbaycani (Cyrillic), Sakha, Uyghur, Uzbek (Arabic) |
| being added | 1: Tuvan |
| missing | ~30 ... |
We can't.
By convincing each company that each language will bring in more money than they will spend to add it.
(except it probably won't, and they know that)
So what can we do?
(case study, viz. next slides)
Kazakh, Kyrgyz, Tatar, Turkish (×4), Uyghur
Bashqort (277h), Uyghur (143h), Turkish (119h), Uzbek (257h), Kyrgyz (47h), Tatar (31h), Chuvash (28h), Sakha (8h), Turken (5h), Kazakh (2h), Azərbaycani (1h)
not ready: Qaraqalpaq (72%), Tuvan (2%)
Azərbaycani (×2), Kazakh, Kyrgyz, Tatar, Turkish, Uyghur, Uzbek (×2)
| Input / Output: | Алмасы |
|---|---|
| ↕ | |
| Output / Input: |
алма алмас Алма Алмас ал ал ал ал |
Turkic transducers I've contributed to
by naïve coverage & lexicon size + distribution
Transducers I've contributed to: downstream tasks
Free/Open Source (versus proprietary)
→ i.e., maximally accessible to language communities
Maybe a little?
my best suggestion: general tools of change
find a need that you can contribute to and contribute
slides available at
https://jonorthwash.github.io/2023-TurkicLangTech/presentation.html