How linguists determine the frequency of words

Word frequency is one of the most useful measurements in linguistics. It shows how often a word, phrase or grammatical pattern occurs in a defined body of language, known as a corpus. Frequency evidence helps lexicographers decide which words belong in a dictionary, translators choose natural equivalents, and teachers select vocabulary that learners are likely to encounter.

The count may look simple, but reliable measurement involves several decisions. Researchers must define the kind of texts they are studying, identify the forms that belong together, remove accidental duplication and interpret the results in context. A word that is common in newspaper reporting may be rare in everyday conversation, while an expression used frequently in one community may scarcely appear in a national corpus.

This matters in Australia, where English is used alongside hundreds of Aboriginal and Torres Strait Islander languages, migrant languages and community varieties. A frequency list drawn mainly from Sydney news websites will produce a different picture from one based on classroom talk in Brisbane, parliamentary records in Canberra or regional speech in Western Australia.

Method What it measures Main strength Common limitation
Raw frequency Total occurrences in a corpus Easy to calculate Favours larger corpora
Normalised frequency Occurrences per fixed number of words Allows fair comparison Can hide local variation
Frequency by genre Use within news, fiction, speech or other categories Shows register and context Requires balanced samples
Lemma frequency Related forms such as walk, walks and walked Useful for dictionaries and teaching Requires grammatical analysis
Dispersion How widely a word is distributed across texts Separates broad use from repetition Needs careful statistical treatment

Building a suitable language corpus

A corpus is a structured collection of authentic language. It may contain novels, government documents, social media posts, recorded conversations, classroom transcripts, news articles or specialised research papers. The purpose of the investigation determines which sources should be included. A corpus designed to study spoken Australian English should not rely mainly on edited newspapers.

Representativeness is a central concern. Linguists usually divide material into categories such as genre, region, age group, subject and publication date. They then aim for a balanced sample rather than simply collecting as much text as possible. A very large archive can still give misleading results if one publisher, website or topic dominates it.

The date of collection also affects the findings. Words connected with bushfires, pandemic restrictions or digital technology may rise sharply for a period and then decline. Comparing Australian texts from the 1990s with current material can reveal lexical change, while comparing urban and regional sources can show differences in local vocabulary.

Counting forms, lemmas and expressions

The most basic measurement is a word token: each occurrence of a written or spoken item. In the sentence “The dogs barked, and the dog ran”, there are several tokens of dog and related forms. A word type is a distinct written form, so dog and dogs count as separate types unless the researcher groups them.

Lexicographers often work with lemmas. A lemma represents a set of grammatically related forms, such as run, runs, ran and running. This is useful when estimating the importance of a vocabulary item, although automatic tagging systems can make mistakes. The form saw, for example, may be the past tense of see or a noun referring to a tool.

Multiword units require further judgement. Expressions such as in spite of, take into account and fair dinkum may be more meaningful as phrases than as separate words. Collocation software measures which words occur together more often than expected, helping researchers identify established combinations and specialised terminology.

Making fair comparisons between corpora

Raw counts are easy to understand but can be misleading. If kangaroo occurs 500 times in a corpus of one million words and 700 times in a corpus of ten million words, the second corpus does not necessarily show greater usage. Researchers therefore calculate a normalised frequency, often as occurrences per million words.

The basic formula is:

normalised frequency = word count ÷ total corpus words × chosen base

A result of 500 occurrences per million words allows comparison between collections of different sizes. It does not mean that every million-word sample will contain exactly 500 examples; it is a rate used for comparison.

Statistical tests may be added when researchers compare groups. Measures such as log-likelihood can identify words that are unusually frequent in one corpus compared with another. Effect size is also important because a large corpus may make a tiny difference appear statistically significant. Frequency should therefore be read alongside distribution and practical relevance.

Measuring distribution and dispersion

A word can have a high total count because it appears repeatedly in only a few documents. A technical term may occur hundreds of times in one medical report but nowhere else. Dispersion measures help determine whether a word is spread across many texts or concentrated in a small number of sources.

This distinction is valuable for dictionary work. A word found in many genres and regions is more likely to be part of general vocabulary. A word limited to mining reports, legal writing or online gaming may need a label such as technical, Australian, informal or computing.

Researchers may divide a corpus into equal sections and record how many sections contain the target item. Other methods calculate the balance of occurrences across documents. These approaches are especially useful when studying endangered or minority languages, where a small number of available recordings can otherwise distort the apparent importance of a word.

Frequency evidence also supports historical investigation. By examining dated documents, linguists can track when a word entered common use, changed meaning or became associated with a particular community. Such work connects naturally with research on language and dictionaries, including questions of etymology, language contact and African-language lexicography.

Using tools without losing human judgement

Corpus software can search millions of words in seconds. Concordancers display a keyword with surrounding context, while part-of-speech taggers assign grammatical categories. Frequency lists, n-gram tools and collocation statistics help researchers identify patterns that would be difficult to notice by reading texts manually.

Automation is useful, but its output is not automatically accurate. Spelling variation, punctuation, abbreviations, names and homonyms can all affect a count. A system may treat record as one item even though it has different pronunciations and grammatical roles, or count a website menu repeatedly when the same page has been copied across an archive.

Manual checking remains essential. A linguist reviews representative examples, corrects tagging errors and decides whether forms should be merged. This is similar to the care required when assessing an unfamiliar source: even a practical guide about choosing coarse-fishing equipment depends on distinguishing relevant observations from details that do not answer the question being studied.

Applying frequency to dictionaries and translation

Frequency affects which words receive headwords, senses, usage labels and example sentences. It can also guide the order of meanings. A dictionary may present the most common sense first, while still recording less frequent meanings that are culturally or professionally important.

For translators, frequency helps distinguish a possible equivalent from a natural one. A bilingual dictionary may list several translations, but corpus evidence can show which one commonly occurs with a particular verb, noun or register. Australian translators may need to account for local terms such as esky, ute or thongs, whose meanings and acceptability vary by region and audience.

Frequency must never be treated as the sole measure of value. A rare word can be central to cultural knowledge, ceremony, history or identity. This is particularly important in documenting Setswana and other African languages, where written resources may be smaller than those available for English. A limited corpus can reflect limited documentation rather than limited linguistic richness.

Digital archives create new opportunities for comparison across languages and writing systems. Materials ranging from historical manuscripts to modern online texts may reveal borrowing, semantic change and contact. Resources such as Japanese manuscript collections illustrate why the nature, date and provenance of a source matter when interpreting linguistic evidence.

Reading frequency figures with cultural context

A frequency number is meaningful only when its source and method are clear. Researchers should report the corpus size, collection period, genres, geographic coverage, unit of counting and normalisation method. They should also explain whether spelling variants, inflected forms, proper names and repeated documents were included.

For Australian research, a national result can conceal significant differences. A word may be frequent in Melbourne multicultural conversation, uncommon in remote Northern Territory recordings and absent from formal government documents. Aboriginal English varieties also have their own vocabulary, meanings and grammatical patterns, which should not be treated as errors against a single standard.

The same principle applies to family names and inherited vocabulary. Historical frequency can help identify patterns of contact, but etymology requires evidence from records, sound change and social history rather than a high search count alone. A study of common Setswana surnames demonstrates how linguistic interpretation must connect word forms with cultural and historical context.

Reliable frequency research therefore combines computation with close reading. A well-designed corpus provides measurable evidence, while expert analysis explains what the figures mean, who uses the forms and under what circumstances. Further background on language research and related resources is available through the professional language website.

The practical takeaway is to treat frequency as a carefully designed measure: define the corpus, count consistently, normalise the results, inspect distribution and interpret every figure in its social and linguistic context.