│By Georgia Platt, Gale Ambassador at Lancaster University│
Six months ago, I knew nothing about Digital Humanities or text mining. It wasn’t really a topic covered in my undergraduate degree in History, and I’d always been a bit unsure of the place technology has in historical research. Then I began my MA in Digital Humanities, which I chose to try and engage with these technological aspects of humanities studies that I hadn’t had the opportunity to explore before. I spent the first ten weeks of this course doing an intensive text mining module, and it revolutionised my approach to historical study.
If you’re curious about text mining but don’t know where to start, or are a digital humanities sceptic looking to be proved wrong, this blog post is for you. I will start with some basic theory of what text mining is and what kinds of questions it can answer, then I’ll put my practical skills to the test with an example from my studies in Victorian crime and punishment.
How Does Text Mining Work?
At a basic level, text mining takes the words that make up your primary source and sorts through them to understand patterns and themes. There are many ways to do this depending on what questions you’re asking, but the general aim of text mining is to process a larger number of words than possible manually. For example, it would take a researcher months to manually analyse a corpus (or collection) of 100 texts but would take a computer mere minutes.
Matthew Jockers’ book is a foundational text on the basics of text mining, proposing an idea of ‘macroanalysis’ vs ‘microanalysis’. He suggests that text mining, a form of ‘macroanalysis’, can be used to highlight key themes or change over time in a larger corpus, and then traditional close reading, or ‘microanalysis’, can be used to delve deeper into key points in those results.
Overall, ‘macroanalysis’ and ‘microanalysis’ work together to make sure historians can get the most out of their primary sources in the most efficient way!

Putting It into Practice
One of my favourite areas of study as a historian is crime and punishment in Victorian Britain. This period in time has so much potential for text mining due to the amount of digitised primary sources such as newspapers, but in this blog post I will be focusing on one particularly interesting topic.
The 1856-1863 garrotting panics were an early example of a moral panic – a phenomenon that was talked about in the media but had no grounding in real life. More modern examples of a moral panic might include the War on Drugs in late twentieth-century America, or the 1960s UK moral panic over Mods and Rockers. The garrotting panics were particularly interesting as they happened twice in less than a decade and showed just how powerful newspapers could be in Victorian society.
Using Gale’s British Library Newspapers archive, we can see just how garrotting, or murder by strangulation, was portrayed in the Victorian press.

First, a quick check using the Term Frequency tool in the British Library Newspapers archive shows us there is clear evidence that garrotting was more frequently reported on in 1856, then again in 1862 and 1863. Whilst this isn’t text mining in itself, it gives us a good overview of the archive and what we might find when we delve deeper. It is also important to consider that visualisations like this graph aren’t objective, they still require critical analysis and interpretation. For example, this graph shows us that there was clear increase in mentions of garrotting in newspapers, not that garrotting itself as a crime increased.
Now that we have a good overview of the archive, we can use Gale’s Digital Scholar Lab to build our corpus and use tools to ask our research questions. I chose to focus on the 1862-1863 panic in London, and carried out analysis of my corpus using the Topic Modelling and N-Grams tools.

In our Topic Modelling visualisation, garrotting is shown in Topic 5, alongside words like ‘violence’, ‘imprisonment’, and ‘police’. This means that garrotting is most commonly grouped alongside these words in each of the newspapers in our corpus.
This presents some interesting evidence for the way garrotting was perceived as a very violent crime, and when we click on this topic it gives us more information on which sources garrotting appears in. We can then use this ‘macroanalysis’ perspective on the sources to guide our ‘microanalysis’ of individual articles, as it gives us key context on how some of the individual sources fit into the wider corpus.

Looking at the corpus through N-Grams provides a different perspective on the sources. It focuses on which words appear the most often across the different sources in the corpus, ranking them by frequency. Garrotting wasn’t actually very common in the sources I chose, which was unexpected. But this does not mean these sources aren’t valuable to my question. Now, rather than looking at the amount of times garrotting is mentioned, I can use ‘microanalysis’ to look at the language around garrotting and see if that had any influence on fuelling moral panic.
Now it’s Your Turn!
This has been a brief overview of text mining and how it can aid your research, but I hope it’s been helpful. The biggest lesson I’ve learned in my time studying Digital Humanities is that you learn the most by doing – making your own successes and mistakes.
So, now it’s your turn! Ask your own questions about the moral panics around garrotting in Victorian London or pick something else entirely. Just have fun, think outside the box, and make the most of all the resources and guidance Gale has to offer!
If you enjoyed reading about text mining, check out these posts:
- Understanding N-Grams
- A New Course in Gale Digital Scholar Lab: Introduction to Digital Humanities
- Bridging the Gap: Gale Primary Sources and Gale Digital Scholar Lab
Blog Cover Image Citation: “HORRIBLE AND MYSTERIOUS MURDER OF A GIRL.” Illustrated Police News, 19 Jan. 1867. British Library Newspapers, https://link.gale.com/apps/doc/BA3200777554/BNCN?u=webdemo&sid=bookmark-BNCN&pg=1&xid=4b26d7fa.