Tracking word shifts through corpus linguistics
Words rarely stay still. They drift in meaning, pick up new companions, and sometimes vanish altogether while fresh terms take their place. For lexicographers, etymologists, and translators, catching those shifts before they solidify has always been part of the craft. The difference today is that we no longer have to rely on memory, intuition, or the patience of reading a thousand pages by lamp light. Corpus linguistics offers a systematic, reproducible, and increasingly fine-grained way to detect when a language is moving, how quickly, and in which direction.
The idea behind the discipline is straightforward. Assemble large samples of real language use, store them in searchable collections, and interrogate them with software that can count, compare, and chart patterns. A corpus is not merely a list of sentences. It is a structured archive that records where each word appeared, when, in what genre, and alongside which other words. When the same kind of material is gathered across different decades, the resulting diachronic corpus becomes a time machine. Researchers can watch frequencies rise and fall, see new collocations form, and notice when an old word starts carrying a fresh sense.
For scholars working on African languages, including Setswana, the implications are significant. Many indigenous tongues remain under-documented in machine-readable form, and the dictionaries we do have were often compiled in missionary or colonial contexts. Where modern corpora exist, they reveal borrowings, calques, and semantic shifts that earlier reference works could only guess at. The same tools that let researchers trace the journey of Latin borrowings through French, as Latin's influence on French shows for another language family, can also illuminate how Setswana has absorbed and adapted English and Afrikaans terms.
Australian readers encounter this kind of work in surprisingly domestic ways. From café menus in Melbourne's inner suburbs where "brekkie" and "arvo" appear in print without quotation marks, to commentary on rugby league broadcasts in Brisbane that casually use "biff" and "chuck" as technical terms, the local lexicon is moving faster than older reference books can capture. Corpus linguistics gives lexicographers a way to keep pace, and it gives curious readers a window onto their own everyday speech.
Foundations of corpus linguistics
The discipline traces its modern shape to the mid-twentieth century, when researchers in London and elsewhere began building machine-readable text collections for linguistic analysis. The Brown Corpus of American English, assembled in the early 1960s, set a template: one million words, carefully balanced across genres, each word tagged for part of speech and stored in a format that computers could search. From that starting point, corpora grew larger, more diverse, and more richly annotated.
A corpus is more than raw text. Researchers add layers of metadata: the year of publication, the region where the text was produced, the register, the author where known, and the subject area. Linguistic annotation adds grammatical tags, lemma forms, and sometimes semantic categories. The result is a searchable object that allows questions to be asked of millions of words in seconds. Frequency lists, keyword comparisons, and concordance lines become the basic instruments of the trade.
Because corpora are empirical, they discipline the descriptive work of the dictionary maker. Intuition about which words are common, or which meanings are central, can be checked against hard counts. For a language like Setswana, where intuitions vary between speakers in Gaborone, Mahalapye, and the diaspora, this kind of external evidence is invaluable. It allows lexicographers to ground their decisions in something more stable than memory.
Signatures of change in large text collections
Language leaves fingerprints in frequency data. A word that was rare in the 1980s but ubiquitous in the 2020s almost certainly signals cultural or technological change. A word whose neighbours have shifted, where the company it keeps has changed, signals semantic drift. A word that has disappeared from spoken frequency but survives in set phrases is fossilising. Corpus linguistics turns these intuitions into measurable phenomena.
Frequency shifts are the most obvious signal. When a new term enters a language, it tends to climb steeply at first and then to settle into a long plateau. Diachronic corpora make the climb visible. The same is true of meanings: a word used predominantly in one sense for decades may suddenly be cited in a different sense with rising frequency, a pattern that researchers flag as a candidate for dictionary revision. Collocations, the habitual pairings of words, change more slowly but reveal deeper currents. A shift in the verbs that typically accompany a noun can mark a change in how a culture frames a concept.
These signatures matter for less-resourced languages as well. In Setswana, for instance, nineteenth-century missionary dictionaries offer a fascinating window onto contact-induced change, and reading what-a-19th-century-setswana-dictionary-can-teach-us-today-05b reminds us how much those early compilers observed, even when their frameworks were limited. Modern corpora continue that work with far greater reach.
African languages and the challenge of limited corpora
For Setswana, Sesotho, and many other African languages, the corpus tradition is younger and thinner than for English or Mandarin. Several large collections do exist, often built by universities, national language boards, or pan-African initiatives. They include newspaper texts, parliamentary records, religious literature, and increasingly, social media. Each genre brings its own bias, but together they give a workable approximation of contemporary usage.
The challenge is coverage. A corpus of a few million words is small by modern standards, and a corpus skewed toward one city or one register will over-represent certain varieties. Building balanced, regionally representative corpora requires sustained funding, often supported by language legislation. In Australia, the framework around the National Indigenous Languages Act encourages documentation and revival work that benefits from the same tools, even though the languages involved are very different from Bantu languages in family and structure.
Corpus methods also support the recovery of older meanings. A story such as the-story-behind-the-word-morse-in-setswana-47f shows how a single technical term can carry layers of contact history, telegraphy, and local adaptation. Such micro-histories, multiplied across thousands of words, build up a richer picture of how a language has actually moved.
Australian English as a living laboratory
Australian English presents a varied terrain for corpus-based study. Its recorded history is relatively short but well documented, with convict narratives, settler letters, and a long newspaper tradition providing ample material. Major Australian cities such as Sydney, Melbourne, Brisbane, Perth, and Adelaide each have their own print and broadcast archives, and online publication has multiplied the volume of searchable text enormously.
Recent decades have seen new vocabulary move quickly through the language. Terms tied to multicultural cuisine, to the technology economy in suburbs like Pyrmont and Eveleigh, to drought policy, and to public health debates have all entered general usage. Some have stayed localised; others have travelled further. Corpus methods allow lexicographers at the Macquarie Dictionary and similar institutions to track which candidates have stabilised and which remain ephemeral.
The Australian marketplace reflects this in practical ways. Publishers commissioning school texts, media organisations writing style guides, and government departments drafting plain-language policies all need to know which words are settling into standard usage and which are still moving. Corpus data informs their decisions not only about vocabulary, but also about idiom, pronunciation variants, and the small grammatical shifts that accumulate over generations. It also informs language planning under instruments such as the National Indigenous Languages Act, where revitalisation efforts benefit from baseline corpora of the present state of each language.
Comparing approaches to tracking change
Different approaches to detecting language change come with different strengths and weaknesses. The sketch below compares the most common methods a researcher might weigh.
| Method | Data source | Strengths | Limitations |
|---|---|---|---|
| Introspection | Native speaker knowledge | Quick, draws on lived experience | Subject to bias and individual variation |
| Historical text comparison | Archived books, journals | Reveals long-term patterns | Slow, labour-intensive |
| Diachronic corpus analysis | Time-stamped digital corpora | Empirical, reproducible, scalable | Requires digitised texts |
| Social media monitoring | Live platform feeds | Captures emerging terms quickly | Noisy, biased to younger users |
| Elicitation studies | Recorded interviews, surveys | Reveals usage and attitudes | Limited sample size |
Each method is best suited to particular questions. Introspection can still point to candidate changes worth investigating. Historical text comparison remains valuable for tracking change over centuries, where digitised corpora are scarce. Diachronic corpora offer the most rigorous picture where the data exist. Social media monitoring catches the fastest-moving vocabulary. Elicitation studies add the human judgements that pure frequency data cannot supply.
The most productive contemporary research combines them. A lexicographer notices a candidate shift through reading, checks it against a diachronic corpus, supplements with social media evidence, and validates through targeted elicitation. The cycle of observation, measurement, and validation is what makes corpus linguistics a science of language rather than just a tool.
What endures is the principle behind the practice. Languages are never static, and any honest account of them must reckon with movement. Corpus linguistics gives us the means to do so with rigour, and the obligation to keep watching as the data grow and the words keep on shifting under our feet.