About the Project
As part of Vanderbilt Libraries Buchanan Fellowship program, this experimental working group explored how to generate transcriptions of archival documents using machine learning.
While optical character recognition (OCR) and automatic machine transcription has come a long way for printed texts, scholars still largely rely on human transcription for handwritten texts. There are two main reasons for this:
- Handwritten texts simply don't have the consistency of printed text from which the computer vision algorithm can learn patterns.
- Machine learning tools currently need a large set of texts (corpora) in the same hand. Most archival materials in the same hand are simply unavailable in a sufficient size for machine learning.
In our History of Medicine Collections, we had 20 annual diaries from Dr. Burch. The students transcribed the first two volumes (1921 and 1922) and we hope to use those volumes as a training set for a machine learning algorithm that can then automatically transcribe the remaining 18 volumes.
You can find our syllabus on Github. Our studies were interrupted by the outbreak of the coronavirus COVID-19.