Transcribing the Mirrlees Watson Order Books
Transcribing the Mirrlees Watson Order Books
Transcribing historical documents is typically a time-intensive and specialised task but handwritten text recognition (HTR) platforms promise to simplify this process by automating the process and using language learning models trained upon large volumes of historical documents.
Building upon Malik Al-Nasir’s 2023/24 Lind-funded Visiting Research Fellowship, the Mirrlees Watson digitisation project aimed to develop enhanced access to the order books of Scotland’s sugar machinery industry. As well as digitising four volumes, this project explored the use of an AI-powered transcription platform to assess the potential for such software to generate high quality transcriptions, enhance accessibility and support digital research.
In theory, these types of transcription platforms have the potential to transform unindexed scans or photographs into structured, searchable data and to decipher old handwriting in a matter of seconds. This project aimed to test both the potential and the limitations of the software.
Transcribing historical documents is typically a time-intensive and specialised task but handwritten text recognition (HTR) platforms promise to simplify this process by automating the process and using language learning models trained upon large volumes of historical documents.
Building upon Malik Al-Nasir’s 2023/24 Lind-funded Visiting Research Fellowship, the Mirrlees Watson digitisation project aimed to develop enhanced access to the order books of Scotland’s sugar machinery industry. As well as digitising four volumes, this project explored the use of an AI-powered transcription platform to assess the potential for such software to generate high quality transcriptions, enhance accessibility and support digital research.
In theory, these types of transcription platforms have the potential to transform unindexed scans or photographs into structured, searchable data and to decipher old handwriting in a matter of seconds. This project aimed to test both the potential and the limitations of the software.
Two volumes were assessed for their suitability for this project. One volume contained handwritten entries, while the other contained more structured data presented in a tabular format. Initial tests of the material suggested that the first volume might require more editing as well as a familiarity with some of the more technical terms contained in the entries:
A set of Brass Bushes for a Sugar Mill
The red lines show the Angle of site in framing. Iron
plates are also showed for under roller brasses
- 24 axles and bushes complete to sketch -
- 12 pair of Bushes without axles
- 1 casting for making piston rings
(GB 248 UGD 118/2/4, p56)
The second volume was selected based on its more regular format, as it was believed that pages of regularly formatted numbers alongside smaller blocks of text would be more straightforward both for the HTR software and for the human proofreader. However, the volume still presented significant challenges and a range of common issues emerged throughout the transcription process. These included misinterpretation of similar characters, difficulty recognising superscripts and punctuation, and confusion caused by the quirks of individual handwriting styles. Layout-related problems were also frequent: the software sometimes misidentified text regions, missed sections, split or merged lines incorrectly, or mixed up the reading order. Importantly, correcting these errors within the platform does not retrain the AI, meaning that similar mistakes recur across multiple pages.
The second volume was selected based on its more regular format, as it was believed that pages of regularly formatted numbers alongside smaller blocks of text would be more straightforward both for the HTR software and for the human proofreader. However, the volume still presented significant challenges and a range of common issues emerged throughout the transcription process. These included misinterpretation of similar characters, difficulty recognising superscripts and punctuation, and confusion caused by the quirks of individual handwriting styles. Layout-related problems were also frequent: the software sometimes misidentified text regions, missed sections, split or merged lines incorrectly, or mixed up the reading order. Importantly, correcting these errors within the platform does not retrain the AI, meaning that similar mistakes recur across multiple pages.
The project found that HTR software can significantly accelerate initial transcription, but its outputs are not error-free and require careful human correction. Accuracy is typically measured using two metrics: Character Error Rate (CER) and Word Error Rate (WER). A CER of 10% or lower is considered acceptable. However, even this level of accuracy requires a keen eye and manual editing to remove errors and ensure accuracy.
The platform used for this example offers the ability to train custom language models using verified transcription data (known as ‘ground truth’). However, training a custom language model requires a significant amount of data - using 20-30 pages of ground truth to begin with, or around 10,000 words. It was hoped that a model could be trained by the end of the project and could be used to transcribe future volumes. However, although it was easy to reach 20-30 pages, it was less easy to reach the recommended 10,000 words due to the nature of the material. Machine learning will statistically replicate the data you give it, so the more accurate the data is – and the more you have of it – the more accurate the model will be. Inconsistent transcriptions or poorly structured layouts could undermine model performance, highlighting the need for well-edited and corrected transcriptions as well as for clear editorial guidelines and standardisation.
Trialling table recognition models was more promising but came at additional computational costs. In this example, automating tables doubled the number of credits used. It was possible to copy and paste tables from page to page, although this could also be time-consuming.
Overall, the findings of this project underline that AI-powered transcription software should be seen as a supportive tool rather than a replacement for human expertise. It can produce a useful first draft for a transcription and can perform this task significantly quicker than a human performing the same task manually. The quality of the final output, however, relies heavily on thorough proofreading and correction, as well as the human ability to accrue knowledge from page to page.
Despite these challenges, the project achieved several important outcomes. Firstly, one volume was completely transcribed, and the index of a second volume was completed. The accuracy of the resulting transcriptions enhances accessibility by enabling keyword searching, which will support both academic research and public engagement. The project also demonstrated how this kind of transcription can be incorporated into archive workflows, so long as sufficient time and resources are allocated for quality control, such as proofreading and correcting.
Additionally, there are ethical and legal considerations to the use of AI-powered software, which must be balanced against the benefits of access and discoverability. The environmental impact of AI systems must be considered, and while the programme we used currently offers stronger guarantees around data protection and copyright, other models may not offer the same protections.
In summary, the Mirrlees Watson digitisation project highlighted the potential and limitations of AI-powered transcription. These tools have the potential to significantly enhance access to archive materials and support digital research, but their outputs require careful human intervention.
Kirsteen Connor - Archive Cataloguer, Archives and Special Collections, University of Glasgow Library.
Published July 2026.
The project found that HTR software can significantly accelerate initial transcription, but its outputs are not error-free and require careful human correction. Accuracy is typically measured using two metrics: Character Error Rate (CER) and Word Error Rate (WER). A CER of 10% or lower is considered acceptable. However, even this level of accuracy requires a keen eye and manual editing to remove errors and ensure accuracy.
The platform used for this example offers the ability to train custom language models using verified transcription data (known as ‘ground truth’). However, training a custom language model requires a significant amount of data - using 20-30 pages of ground truth to begin with, or around 10,000 words. It was hoped that a model could be trained by the end of the project and could be used to transcribe future volumes. However, although it was easy to reach 20-30 pages, it was less easy to reach the recommended 10,000 words due to the nature of the material. Machine learning will statistically replicate the data you give it, so the more accurate the data is – and the more you have of it – the more accurate the model will be. Inconsistent transcriptions or poorly structured layouts could undermine model performance, highlighting the need for well-edited and corrected transcriptions as well as for clear editorial guidelines and standardisation.
Trialling table recognition models was more promising but came at additional computational costs. In this example, automating tables doubled the number of credits used. It was possible to copy and paste tables from page to page, although this could also be time-consuming.
Overall, the findings of this project underline that AI-powered transcription software should be seen as a supportive tool rather than a replacement for human expertise. It can produce a useful first draft for a transcription and can perform this task significantly quicker than a human performing the same task manually. The quality of the final output, however, relies heavily on thorough proofreading and correction, as well as the human ability to accrue knowledge from page to page.
Despite these challenges, the project achieved several important outcomes. Firstly, one volume was completely transcribed, and the index of a second volume was completed. The accuracy of the resulting transcriptions enhances accessibility by enabling keyword searching, which will support both academic research and public engagement. The project also demonstrated how this kind of transcription can be incorporated into archive workflows, so long as sufficient time and resources are allocated for quality control, such as proofreading and correcting.
Additionally, there are ethical and legal considerations to the use of AI-powered software, which must be balanced against the benefits of access and discoverability. The environmental impact of AI systems must be considered, and while the programme we used currently offers stronger guarantees around data protection and copyright, other models may not offer the same protections.
In summary, the Mirrlees Watson digitisation project highlighted the potential and limitations of AI-powered transcription. These tools have the potential to significantly enhance access to archive materials and support digital research, but their outputs require careful human intervention.
Kirsteen Connor - Archive Cataloguer, Archives and Special Collections, University of Glasgow Library.
Published July 2026.