Analyzing Political Speeches with Corpus Tools
Political speeches leave a trail of text that can be measured, sorted, and compared. A single campaign speech, a budget reply, or a parliamentary address can be turned into a dataset, and that dataset can answer questions that reading alone rarely resolves. The aim is not to replace interpretation but to give it firmer ground: how often a word appears, what it tends to appear with, and how those patterns shift across speakers and decades.
Corpus tools suit this work because they treat language as evidence. Concordance lines, frequency lists, and keyword statistics make it possible to see what a speaker emphasised and what they avoided. For researchers in Australia, the same methods that unlock Setswana political discourse or campaign rhetoric in Botswana can be turned on the debates recorded in Hansard or the transcripts aired on ABC Radio National.
What corpus tools actually do
A corpus is simply a collection of digital texts gathered for a purpose. Once the texts are cleaned and stored in a consistent format, software can count and compare them. The simplest output is a frequency list: which words occur most often, how often, and in what proportions. More revealing is the concordance, which shows every instance of a chosen word in a few lines of surrounding context. From there, collocation analysis identifies the words that habitually cluster around a target term, while keyword analysis compares the corpus in question to a reference one and highlights the words that are statistically distinctive.
These techniques were developed for general language description but they adapt easily to political speech. Researchers can compare a prime minister's budget reply with opposition responses, track how a term like "working families" enters the political vocabulary, or examine whether the language of parliament has changed since the 1990s. The methods are descriptive rather than judgmental: they show patterns and leave the explanation to the human reader.
Building a working corpus of speeches
Before any analysis is possible the texts must be assembled. In Australia, the federal Hansard is freely available online through the Parliament House website in Canberra, and each state parliament publishes a similar record for Sydney, Melbourne, Brisbane, Perth, Adelaide, Hobart, and Darwin. These records are produced under the authority of the Parliamentary Papers Act and can be reused for research without further permission. ABC News transcripts, podcasts, and Radio National interviews offer a second stream of material that is closer to spoken than written language. For a focused study, a researcher might collect one minister's statements over a year, or every speech on a single portfolio.
Practical habits shape the corpus. Australian political coverage favours Question Time and the daily doorstop interview, so a researcher who ignores them will get a skewed picture of how politicians actually speak. It is also worth including transcripts of speeches delivered at NAIDOC Week events, Australia Day ceremonies, and multicultural forums, because those occasions produce language that does not appear in the chamber. Researchers working across languages should also pay attention to spelling conventions, since the history of the Tswana alphabet reforms shows how a script itself becomes part of the political record. Once collected, the texts should be checked for boilerplate greetings, applause markers, and headings, which can be stripped before analysis begins.
Keyword analysis and frequency patterns
A raw word count tells only a small part of the story. The first useful step is usually to remove high-frequency function words such as "the", "and", or "to", which carry grammatical weight but little political meaning. What remains is a list of content words, and from there keyword analysis can compare one corpus to another. A collection of the Treasurer's speeches compared with a reference corpus of everyday journalism will quickly reveal items such as "productivity", "inflation", "wage growth", or "fiscal repair", depending on the moment.
Frequency alone can mislead, because political language is full of fixed phrases. The word "fair" is common in Australian political speech, but it almost always arrives inside an idiom: "a fair go", "fair dinkum", or "fair and reasonable". Concordances let the analyst see this instantly, since each line shows the company the word travels in. The same applies to "mateship", "multiculturalism", and "reconciliation", which carry cultural weight beyond their dictionary sense and deserve careful handling in any study.
Collocation and the framing of issues
Collocation analysis shows which words appear together more often than chance would predict. In Australian political text, "border protection" clusters with verbs like "strengthen" or "maintain", while "climate action" tends to attract words such as "targets", "transition", and "investment". The choice of a single verb can change the whole tone of an argument, and the corpus method is the cleanest way to detect that habit across hundreds of documents. Researchers looking at Senate inquiry transcripts can often see one side's preferred framing by listing the regular partners of a contested term like "energy" or "reform".
The technique extends to longer chunks. N-gram analysis can pick up phrases such as "cost of living pressures" or "closing the gap", which have become near-fixed units in Australian political debate. Once these phrases are identified, the analyst can trace when they entered the record, who first used them, and whether their meaning has drifted. That kind of evidence is hard to gather by intuition alone and is one of the clearest payoffs of working with a corpus.
Cross-language comparison and African connections
Corpus methods are particularly valuable when comparing political speech across languages, because they force the analyst to make translation decisions explicit. A study comparing Australian parliamentary debates with speeches delivered in Setswana or isiZulu quickly runs into idioms whose meaning cannot be carried by a single English word. Words tied to ubuntu, botho, or community reciprocity look like direct equivalents but sit inside different rhetorical traditions. Researchers building bilingual resources can learn from projects that document how political terms are negotiated between languages, a theme that surfaces clearly in the work on writing a bilingual dictionary, which walks through the choices made in a Setswana and English volume.
In Australia, similar questions arise around the notation of Aboriginal languages in parliament, where the orthography of words from Warlpiri, Pitjantjatjara, or Kriol is still being standardised and where each spelling choice carries symbolic weight. Translators working between English and Aboriginal languages face the same choices that appear in any bilingual dictionary project: which term becomes the headword, how examples are chosen, and whose variety of the language sets the standard.
Limits, ethics, and the question of bias
Corpus analysis can only describe the texts it has. If the collection over-represents ministers and under-represents backbenchers, women, or community speakers, the findings will reflect that imbalance. The Hansard itself privileges the spoken exchange of the chamber and gives less space to press releases, social media, or grassroots speeches, so any claim about "Australian political speech" must specify the slice. Researchers should also be cautious with automated tools such as sentiment analysis, which often misread sarcasm, irony, or culturally specific idioms.
Ethical questions go beyond the data. Hansard is public and can be reused freely, but speeches given on community occasions often carry cultural protocols that should be respected, particularly when the speaker is an Aboriginal elder or a Torres Strait Islander leader. Translating political speech into another language for analysis should be done with reference to speakers from that language community rather than through a single translator. Where a corpus draws on politically sensitive material, anonymising individual speakers and aggregating across years can be a reasonable compromise.
A practical workflow for researchers
A clear workflow helps avoid false conclusions. Start by defining a research question narrow enough to test, such as whether the phrase "closing the gap" has become more or less frequent in federal Hansard over the past decade. Collect the texts from a known source, clean them by stripping headings and applause markers, and check that the encoding is consistent. Upload the corpus to a tool such as AntConc, Sketch Engine, or a Python notebook with libraries like NLTK and spaCy, and begin with frequency lists and keyword comparisons.
From there, move to concordances and collocation tables to interpret the patterns, and only then attempt claims about framing or change. Researchers looking for a starting point for tools, datasets, and reading lists may find the curated page of resources helpful. Throughout the process, keep a clear log of every cleaning decision, since those choices shape the result as much as the original texts.
Open one session of Australian parliamentary debate on a quiet afternoon, copy the Hansard text for a single sitting, and run it through a free concordance tool with a list of five words you already suspect matter. The patterns that show up in the first hour of work will sharpen the question that comes next.