(I am speaking to you today from Lenapehoking, land of the Lenape. We recognise the Lenape as past, present, and future stewards of this land and the Delaware River. Let us honor through our actions the Lenape who were, those who are, and those who will be.)
Overview
My talk in a nutshell
Will overview how to develop meaningful language technology by developing symbolic tools.
The next ~42 minutes:
Meaningful language technology
Symbolic language technology (LT)
Getting started
Concluding thoughts
What NLP do we work on?
Common NLP tasks
Tokenisation
POS Tagging
Sentiment analysis
Stance detection
Inference / textual entailment
Semantic role labelling
Named entity recognition
Speech recognition
...
How we experience them
How real people experience NLP
1
2
3
4
5
6
7
8
Key:
Input methods
predictive text
Machine translation
Screen readers (TTS)
Spell checking
grammar checking
Voice assistants
Computer-aided language learning
Search engines
Interfaces (localisation)
Why language technology matters
Why is language technology important to language vitality?
Digital technology is hugely pervasive
Language mediates most interaction with digital technology
What does language technology do for a language?
adds to ease with which a language can be used in digital contexts
increases range of uses of language
raises perceptions of language
supports maintenance and revitalisation efforts
What happens without language technology?
"Digital language death" (i.e., loss of intergenerational transmission)
(Kornai 2013)
Language technology supports language vitality
Reframing NLP
Q: Does this support language vitality?
A: not directly
Meaningful language technology
Community-oriented terminology
NLP → Language technology (LT)
"low-resource" → meaningful in low-vitality situations
Evaluation
should include (potential) usefulness
corporate NLP has this part this right
Motivation: how can LT
empower communities
promote linguistic vitality
What's meaningful in low-vitality situations?
What's meaningful in low-vitality situations?
not all low-vitality situations are the same (Liu et al. 2022)
Speakers
Outside interest
Language Technology
Resource availability
Local expertise in Lg & NLP
Community priorities
Language A(350)
"small" national language
≥1M, <10M
some big corporations
some seamless tools; increasing localisation
large corpora
plenty
voice assistants
screen readers
updates to keyboards that everyone finds annoying
Language B(1K)
"large" regional language
≥100K, <1M
occasional corporate
some tools; spotty localisation
copora (spotty)
some
seamless keyboard instead of stand-alone app
localisation
search engine support
Language C(3.8K)
"small" regional language
≥1K, <100K, increasingly few young
some academic
a couple local web-based tools
some published literature
some "activists", mostly self-taught
pred. text keyboards
localisation
spelling and grammar checking
MT (for quick media generation)
Language D(764)
extremely "endangered" language
<100 fluent L1, mostly elders
some academic
none
some published texts, mostly collected narratives
adult learners (small community)
simple keyboards
CALL
anything that raises interest & awareness & helps produce new speakers
Washington et al. (2012),
Tyers et al. (2012),
Salimzyanov et al. (2013),
Washington et al. (2014),
Tyers et al. (2016),
Washington et al. (2016),
Bayatlı et al. (2018),
Tyers et al. (2019),
Gökırmak et al. (2019),
Washington et al. (2019),
Washington et al. (2020),
Ivanova et al. (2021, 2022)
)
Case study: Turkic morphological transducers
Transducers I've contributed to: downstream tasks
Kazakh
Used in other NLP research in Kazakhstan
Used as segmenter in a shared task
Spell checker for Word & LibreOffice
Verb paradigm generator
Tatar
Used to annotate Tatar National Corpus
Spell checker for Word & LibreOffice
Uyghur
Used in corpus-based phonology research
Tuvan
Used for UniMorph shared task (2021)
Spell checker for Word & LibreOffice
Sakha
Used in Revita (CALL)
Used for UniMorph shared task (2021)
Kyrgyz
Spell checker for Word & LibreOffice
Used for annotating published corpora
Morphological transducers
Downstream tasks with potential uses by communities
Spell checkers
Solutions for integration with some environments (Word, LibreOffice)
Hard to add at OS level (Windows, macOS, iOS, Android, ChromeOS)
Nigh impossible to integrate seamlessly in cloud platforms (Google Docs, MS Office 365)
Washington et al. (2012),
Tyers et al. (2012),
Salimzyanov et al. (2013),
Washington et al. (2014),
Tyers et al. (2016),
Washington et al. (2016),
Bayatlı et al. (2018),
Tyers et al. (2019),
Gökırmak et al. (2019),
Washington et al. (2019),
Washington et al. (2020),
Ivanova et al. (2021, 2022)