Back

Winter School of HTR of Medieval Documents 2024 - Growth and Vision for the Future

03.07.2025
A medieval monk works on his laptop in the scriptorium
“A medieval monk works on his laptop in the scriptorium.” The AI image was created on 28.05.2024 by the Bing Image Creator.

July 10, 2025 | Eirini Afentoulidou, Cinzia GrifoniEphrem Aboud IshacAnna MichalcováEkaterini MitsiouJan Odstrčilík, Leon PürstingerBill Weis, Michaela Wiesinger | HI Research Blog

 

When we started the Winter School of handwritten text recognition in 2022, we could not have wished that it would grow and succeed as much as it did, exceeding all of our expectations. We began with four script/language groups and some 40 participants in our hybrid mode, combining virtual sessions with an in-person workshop, about which you can read more here:
●    2022: Discovering the Power of AI for Reading Medieval Manuscripts: HTR Winter School 2022 
●    2023: Maturing Artificial Intelligence in the Winter School of Handwritten Text Recognition of Medieval Manuscripts 2023 

Last year, we reached six script/language groups with some 130 virtual participants, of which 80 joined us at the in-person workshop in Vienna. Unfortunately, due to space limitations, we had to decline many strong applications.

This immense growth in attendance can be attributed both to the improving methods of professional outreach (thanks in particular to the efforts of Leon Pürstinger, Cinzia Grifoni, and Ivana Lukáč Labancová) as well as to the growing interest in public awareness of AI. There is no doubt that HTR is becoming an essential part of textually focused research. And, in comparison to other uses of AI in our disciplines—like document classification of manuscripts, scribe identification, detection of text reuse—HTR has become very accessible thanks to the Transkribus platform which provides for easy and affordable access to and creation of HTR models. 

As in previous years, the 2024 Winter School was held across several dates, with virtual sessions on November 8 and 22, and December 6, followed by an in-person workshop on December 18–20. The event was organized with broad international cooperation in mind, involving the following institutions:
●    Institute for Medieval Research of the Austrian Academy of Sciences,
●    Princeton University (Manuscripts, Rare Books and Archival Studies Initiative, MARBAS),
●    Comenius University in Slovakia

In addition to the core team, the organisers included colleagues from the University of Innsbruck, Charles University in Prague, and the Czech Academy of Sciences. 

The Winter school was also financially supported by two projects: the Austrian-Slovakian part of the Winter School was financed by Aktion Österreich – Slowakei, Wissenschafts- und Erziehungskooperation (Project No. 2024-05-15-002 “Winter School of Handwritten Text Recognition of Medieval Manuscripts 2024”, PI Ivana Lukáč Labancová and Jan Odstrčilík), whereas the Czech participants were supported by the International Cooperation project “König Ottokar’s Glück und Ende: Das lange 13. Jahrhundert in den ostmitteleuropäischen Ländern: Politik, Kultur und Identität. Forschungsprogramm im Rahmen der zwischenakademischen Kooperation” by the Austrian Academy of Sciences.

We would also like to extend our sincere thanks to Transkribus for generously supporting the Winter School with 15,000 transcription credits, which significantly contributed to the success of the training sessions (https://app.transkribus.org/). 

This year, for the first time, the Winter School of HTR had 6 language/script groups, all of which proudly published the results of their work on Zenodo (see links below):
●    Carolingian Latin
●    Late Medieval Latin
●    Medieval German
●    Medieval Czech
●    Byzantine Greek
●    Syriac

 

Carolingian Latin Group

The Carolingian group consisted of twenty-three participants at different stages of their academic careers who wanted to learn how to use HTR in their research activities. This team was led by (in alphabetical order) Cinzia Grifoni, Gerda Heydemann, Leon Pürstinger, Helmut Reimitz and William Weis. We chose to work with two manuscripts held by the Austrian National Library: the first (Cod. 510) contained, among other texts, Einhard’s Vita Karoli; the second (Cod. 940) transmitted an anonymous commentary on the Gospel of Matthew influenced by insular exegetical approaches.

After a brief overview of the historical development and the most important palaeographic features of the Carolingian minuscule, our first online session introduced the manuscripts which we selected for our work, the Carolingian Minuscule Model CMM 9th-11th c., as well as the transcription guidelines we employed. These were developed by Tim Geelhaar: his model is publicly available on the platform Transkribus and his guidelines will be published in an upcoming book on Automated Text Recognition. To conclude our first meeting, we started the layout analysis, automatic transcription, and correction of some leaves of Cod. 510.
We then used the remaining online sessions and the in-person meeting to automatically transcribe and manually correct the entire commentary of the Gospel of Matthew contained in the folios 13r-142v of the Cod. 940. We faced some initial difficulties due to the poor quality of the digitised images of the manuscript available to us, which resulted in a highly inaccurate automatic transcription. Thanks to the help of the Manuscript Department of the National Library, we obtained better images and, as a result, the quality of the transcription improved considerably. Nevertheless, the model had problems recognising some peculiarities of the script of the commentary, especially several abbreviations and some ligatures such as the special ligature "ra".

The necessary manual transcription took quite some time due to the length of the text, but the results were altogether promising. We were able to train a new model, based on Tim Geelhaar's already efficient public model (character error rate or CER of 5.1%), with a fairly low CER of 6.8%. Furthermore, our generated ground truth will be incorporated into Geelhaar’s Carolingian Minuscule Model in the near future! We are also very pleased to have completed and published online a transcription of the entire commentary. Although certainly not perfect, our transcription will hopefully support and enhance future research on this previously unedited commentary. In addition, using ÖNB Cod. 387, which contains numerous computistical texts, we have experimented with other rather new and experimental features in Transkribus, such as the training of table models, which yielded very promising results. During the in-person meeting, we had the opportunity to admire and browse through the Cod. 510 and several other Carolingian manuscripts similar to our Cod. 940 in the Austrian National Library. Unfortunately, it was not possible to view the Cod. 940 itself for conservation reasons, as the manuscript still retains its original Carolingian binding.

The dataset and the complete transcription of the entire commentary created by the Carolingian Latin group is available online on Zenodo: https://zenodo.org/records/14918801. See the picture below for an example of our transcription (fig. 1).

 

Late Medieval Latin Group

This group was led by Jan Odstrčilík, Ivana Lukáč Labancová, and Martin Roček. We selected two fascinating manuscripts from the National Library in Vienna. The first manuscript, Cod. 4135, contains a treatise of Peter de Pulkau against the Hussites as well as a large number of sermons that have not yet been identified. However, they include references to the prominent Viennese theologians Petrus de Pulkau and Henricus Totting de Oyta. The second manuscript, Cod. 4680, contains the works of Thomas Ebendorfer, another renowed Viennese theologian and preacher of the 15th century. This manuscript is particularly valuable because it is an autograph—written in his own hand.

As for our methodological approach: while in the previous years, we chose a semi-diplomatic transcription approach, which involved expanding abbreviations, this year we adopted the guidelines of CATMuS (Consistent Approaches to Transcribing ManuScripts). These guidelines are graphemic: they preserve all abbreviations but do not distinguish between different allographs (e.g., 's' and 'ſ'). Our main motivation was to produce a dataset compatible with the larger CATMuS corpus.

Applying these guidelines proved challenging, particularly when typing specific characters, which made the process significantly more time-consuming than a semi-diplomatic transcription. As a result, we produced a relatively small amount of ground truth. The best result we achieved in training was 10.37% CER, using a small dataset of 1,265 lines and the base model “Medieval Script M 2.4.” Although we plan to return to semi-diplomatic transcription rules in the future, this was a valuable experience. We were also able to provide extensive feedback to the CATMuS team, including several recommendations for improving their guidelines. The dataset we created is available on Zenodo: https://zenodo.org/records/14524944.

 

Medieval German Group

The Medieval German team consisted of 21 participants and was coordinated by Michaela Wiesinger (University of Innsbruck), Julian Helmchen (Freie Universität Berlin), and Norbert Orbán (University of Innsbruck). The goal of this course was to develop an HTR-model for German-language manuscripts from around 1400. To achieve this, the participants were divided into five groups to work on five different German manuscripts from the Austrian National Library. The manuscripts were: ‘Der Renner’ by Hugo von Trimberg (Cod. 2810), a Büchsenmeisterbuch (Cod. 3064), a theological composite manuscript (Cod. 2875); “Weltchronik” by Jans Enikel (Cod. 2921), the “Stadtbuch zu Mautern” (Cod. 14889), and a composite manuscript with various legal texts (Cod. 2780).

First, the participants had to start their transcriptions with the help of already existing HTR-models that worked insufficiently. After a certain number of pages were transcribed, each group started to train individual models based on the data they had generated so far. Subsequently, the models were also improved step by step. By the time of the in-person workshop in December all the trained individual models were to be combined into a large ‘German Medieval Model for the 1400s’: The model “Der_Renner_v02” had a CER of 5.87% with 126 trained pages; the model ‘Büchsenmeistermodell V1’ had a CER of 6.04% with 54 trained pages; the model ‘Cod. 2875 (Theologische Sammelhandschrift) Text Modell Extendend’ had a CER of 2.70% with 72 trained pages; the model ‘DPSV 1.0’ (Jans Enikel Weltchronik) had a CER of 2.05% with 63 trained pages. The Stadtbuch zu Mautern was not trained due to difficulties regarding the many different hands. The model ‘Sammelhandschrift Stadtrecht Cod. 2780_Vol.2’ had a CER of 5.45% with 28 trained pages.

The combined model ‘Test 1 generic 1400’ had a CER of 5.14% with 359 pages, a total of 88.490 words, and initially also incorporated the Stadtbuch zu Mautern. Due to the many hands of this manuscript and an extended period (1432-1549) that was covered in this manuscript, the Stadtbuch was eventually excluded from the last training to make for a more stable combined model. The final model, designated '14th century German 1.0', yielded optimal results, both statistically and in terms of practical applications to manuscripts from the same period. This model had a CER of 3.17% on 319 pages and a total of 83.677 words.

 

Medieval Czech group

This group was led by Anna Michalcová from the Czech Language Institute of the Czech Academy of Sciences. The Medieval Czech group at the Winter School focused on transcribing the Vienna copy of the Old Czech Bible incunabulum, known as the Prague Bible (1488, Österreichische Nationalbibliothek, Vienna, Austria, Ink 13.C.5). We built upon work from the previous Winter Schools, where the "Old Czech Handwriting (with spaces)" model was developed. During this year's school, the team significantly expanded the model for diplomatic transcription of Medieval Czech, which now achieves a very high reliability and has been published as a public model. The new model that emerged from this school has a Character Error Rate (CER) of 5.02%. It was trained on 228 pages of text, now from three Medieval Czech biblical manuscripts and a print. The dataset created is available online on Zenodo: https://zenodo.org/records/14524858, with rough transcriptions also published on the project website: https://htr-school-vienna.github.io/2024--medieval-czech/. The table below in fig. 3 displays the model that is currently available publicly.

Although the Czech group was relatively small, the participants achieved an impressive amount during the Winter School. The friendly and inspiring environment enabled highly focused work, resulting in remarkable progress. We not only worked on improving the already existing model for Old Czech, but also laid the foundations for further research, with plans already forming to develop an interpretative transcription model for Czech-language sources — a much-needed tool in local philological research.


Byzantine Greek

For the second time, the HTR Winter School in Vienna was able to include a group specializing in Byzantine Greek. This team – led by Eirini Afentoulidou (Austrian Academy of Sciences) and Ekaterini Mitsiou (University of Vienna), with support by Georgi Mitov (Austrian Academy of Sciences) – worked with the 14th-century codex Paris. gr. 1382. This is a legal manuscript available freely online, but the quality of its digitisations is mediocre.

We created the model “Trial Paris. gr. 1382” based on a training set size of 5.170 words. The error rate was 11.08%, a significant improvement over last year’s 22.30%.

This year, we were successful in addressing the main source of errors in the 2023 model in an attempt to improve our CER: in Byzantine Greek manuscripts accents and many abbreviations are above the line and, as a result, were often omitted by the automatic layout recognition. In our transcriptions, we manually adjusted the line polygons to include accents and letters above the line. This approach proved effective, and we intend to continue and refine it further.

 

Syriac Group

The Syriac group, led by Ephrem Aboud Ishac (Institute for Medieval Research, Austrian Academy of Sciences) and Christine Roughan (MARBAS, Princeton University), focused on the development and public release of the first Syriac HTR model on the Transkribus platform during the HTR Winter School of 2024. The ground truth for the model, consisting of training and validation datasets for Syriac Serto script, was also produced by our group.

The primary manuscript used for training the "Vienna Syriac Gospels model (Serto)" was ÖNB Cod. Syr. 1, a 16th-century Syriac Gospel book composed by Moses of Mardin in Vienna. The transcription process began with creating ground truth for 140 folios of this manuscript. Participants manually corrected a preliminary automatic transcription, which was initially generated using Kraken. The transcription guidelines aimed to capture spaces, Syriac letters, certain diacritics (like Syome and dots to distinguish homographs), and punctuation as they appeared in the manuscript, all while excluding vowel dots and markings for hardening/softening (qushoyo and rukokho). A significant challenge which we addressed was adapting and optimizing the HTR process within Transkribus for Syriac's script which, in contrast to the scripts of the other Winter School groups, is read right-to-left script. The collaborative environment of the Winter School, embodying the "sharing is caring" principle, was crucial in overcoming these hurdles and facilitating knowledge exchange.

The main achievement was the successful training and public release of the "Vienna Syriac Gospels model (Serto)" on Transkribus, achieving a Character Error Rate (CER) of 5.97%, making it the first publicly accessible model for this script on the platform (see fig. 4 below). 

Furthermore, the dataset uploaded to Transkribus automatically generated a companion website providing open access to the Vienna Syriac Gospels (ÖNB Cod. Syr. 1) with searchable images and transcriptions, which can be accessed here: https://app.transkribus.org/sites/Syriac-Vienna-Gospels (see also fig. 5 below).

We are proud that our research and the tools we have developed will empower individual scholars and students, particularly those with limited resources, to follow in our footsteps. The open-access dataset used for training has also been made publicly available to support further research and HTR development. Future work will focus on further improving the model’s accuracy, expanding its script coverage, and integrating it with other digital resources.

The dataset created by the Syriac group participants (Ephrem Aboud Ishac, Christine Roughan, Ammar Awad, Carlo Emilio Biuzzi, Saranya Chandran, Jennifer Griggs, Polina Ivanova, Branko Malešević, Stefan Marić, Francesca Nateri, Ivan Petrov, Cristina Tava, and Maria S. Thomas) is available online on GitHub and Zenodo, here:
GitHub: https://github.com/HTR-School-Vienna/2024--Syriac/tree/main
Zenodo: https://zenodo.org/records/14714089

 

Winter School: HTR of Historical Documents 2025

All in all, the feedback we have received was very positive and encouraging for the organisation of future winter schools of similar format. The students also appreciated the festive atmosphere created by the proximity of the Christmas markets, and many praised the smooth organisation of the in-person programme, for which Jan Odstrčilík, Leon Pürstinger, Cinzia Grifoni and Ekaterini Mitsou were responsible.

The evolving name of our conference corresponds to the ever-expanding scope of the Winter School. For 2025, the following groups are planned: Carolingian Latin, Late Medieval Latin, Medieval Czech, Byzantine Greek, Syriac, Hebrew and Early Modern German.

Check out the call for applications for the Winter School 2025!

Applications Winter School 2025