|By Paul Schummer – Senior Manager, Gale Content Production, and Chris Houghton – Head of Academic Partnerships|
We often use tools without understanding how they work. Chris, for example, has a hazy memory of learning how a vacuum cleaner works in a school physics lesson, but if you asked him to explain it now, he’d struggle! Most of the time, that’s fine – it’s why we take our car to a garage rather than learning how to fix it ourselves. So it’s understandable that researchers or students working with databases might not know how these collections came to be, and what is actually happening when they run a search. However, to be a truly effective researcher, it’s important to understand how these databases work, so that you can adapt your investigations and not miss vital information that is significant to your scholarship.
In recent years, Gale Primary Sources archives have provided transparency about how these collections of documents were curated and created – and how the search functionality works. The Learning Centers within each archive explain the processes behind its development, and the Gale Primary Sources platform allows researchers to view the scanned image of the source alongside the OCR or HTR. (This is the text generated from the scanned image through Optical Character Recognition or Handwritten Text Recognition which allows the system to conduct a full-text search of the content within that source.) For more information on OCR and HTR, you can read this blog post.
OCR and HTR software has improved significantly over the years, enabling Gale to contemplate digitizing collections of materials that would previously have been impossible to include. For example, with the proliferation of AI tools like Large Language Models (LLM) or Vision-Language Models (VLM), the possibilities for digitising and transcribing certain documents have increased significantly. The recent release of State Papers Online: Nineteenth Century: The State Papers of Queen Victoria and King Edward VII provides an excellent example of how AI-enabled HTR software has made previously impenetrable material discoverable.
Famously Difficult Handwriting
Over the years, there have been great advances in handwritten text recognition technology. As a result, users now expect digital archives to be both highly searchable and highly accurate, even when they include handwritten materials. The source material in this collection, however, was particularly challenging; Queen Victoria’s handwriting famously deteriorated over time, and members of the royal household employed multiple scribes, each with distinct handwriting styles, to draft correspondence for both Victoria and Edward VII. Traditional handwritten text recognition (HTR) tools struggled to interpret the messy, sprawling, and constantly changing handwriting, often producing inconsistent or unusable results.
Overcoming the Challenge
To overcome this challenge, the team explored an AI-driven transcription solution. Unlike traditional HTR, which focuses primarily on interpreting character shapes within a defined zone, AI-driven HTR analyzes words in context – using sentence structure, document layout, and domain-specific patterns to resolve ambiguity. The results were striking: the AI solution was often able to interpret the handwritten content more accurately than a human reader.
One trade-off was the loss of positional coordinates for individual words, since the transcription process no longer operated on a strict character-by-character basis. After carefully reviewing and validating sample outputs, the Content Production team – working closely with Product Management and Editorial stakeholders – agreed that the substantial gains in transcription quality outweighed this limitation. Scaling the solution across the archive’s approximately 350,000 pages proved relatively straightforward due to the flexibility and efficiency of the AI toolset, making accessing this fascinating content at scale easier.
This AI-based HTR approach positions archival content production to better meet future needs, by improving transcription quality for difficult handwritten materials. The speed, scalability, and effectiveness of the implementation also demonstrate the growing impact AI can have in solving complex archival and content-production challenges.

British Association of Victorian Studies Conference
There was significant interest in State Papers Online: Nineteenth Century: The State Papers of Queen Victoria and King Edward VII at the recent British Association of Victorian Studies conference, hosted at Liverpool John Moores University. Understandably, this collection of the correspondence of Queen Victoria and her household provides a treasure trove for scholars researching the social, cultural, literary and political landscape of one of the most significant British monarchs.

BAVS Conference Paper: “The Legible Monarchy”
Attendees at the ‘Into the Archives’ panel session, where we introduced this new collection, really appreciated that we provided a ‘look under the hood’ of the collection, demonstrating how AI could correctly interpret words based on document context, but could also result in false positives where a word was transcribed incorrectly. Understanding the trade-offs inherent in the creation of these digital archives is a hugely important aspect of using them effectively. Chris’ paper discussed how Gale had carefully evaluated the output and concluded that this new HTR technology, an AI-driven transcription solution which could tackle and transcribe these tricky documents, was well worth the small number of errors, opening up this historic collection to new scholarship.
Opportunities for Text and Data Mining
A related benefit of a fully transcribed collection is that it opens up the possibility of text mining and analysis. The paper delivered at the BAVS conference also introduced the audience to some examples of analyses that could be run on State Papers Online: Nineteenth Century: The State Papers of Queen Victoria and King Edward VII using Gale Digital Scholar Lab, Gale’s text and data mining platform, and many in the audience were excited by these potential research opportunities!

Conclusion
What became really clear from discussion at the ‘Into the Archives’ panel at the British Association of Victorian Studies conference was the importance of collaboration between academics and publishers like Gale. Academics provide crucial guidance when developing Gale’s collections like State Papers Online: Nineteenth Century or the Punch archive, ensuring these archives can be digitized in the first place! Gale is also delighted to work with academics and present at conferences like BAVS, to lift the lid on how digital archives are made – exploring the technical processes; presenting innovation; and providing critical context to ensure that researchers can use these amazing collections as effectively as possible.
If you enjoyed reading about this AI-driven transcription solution, try:
- Leveraging Large Language Models for Post-OCR Correction of Nineteenth-Century British Newspapers
- Building Projects in Gale Digital Scholar Lab
For further exploration of Handwritten Text Recognition, you might like:
Blog post cover image citation: A montage of images of sources from State Papers Online: Nineteenth Century: The State Papers of Queen Victoria and King Edward VII combined with a photograph of Queen Victoria taken in 1882, available in Wikimedia Commons.
