AI cracks the code of papyrus: The Austrian Academy of Sciences unravels the secrets of ancient texts with Mistral and Reply
05.06.2026
Can artificial intelligence (AI) help uncover the mysteries of history? That is precisely what researchers at the Austrian Academy of Sciences (OeAW) are working on. Even after centuries of research, countless papyri, inscriptions and manuscripts are still waiting to be read and analysed. They contain a vast wealth of information about the past.
Just how revealing such sources can be was demonstrated only recently by a papyrus dating back almost 2,000 years, which was deciphered with the involvement of Roman historian and papyrologist Anna Dolganov from the Austrian Archaeological Institute of the Austrian Academy of Sciences. The document recounts a suspected case of tax fraud, forgery of documents and possibly even rebellion against Roman rule, and paints a fascinating picture of life in the Roman provinces of the Middle East.
Reconstruction of previously undeciphered texts
Such discoveries also highlight a key challenge facing classical studies: hundreds of thousands of ancient texts lie in storage around the world, many of which have never been read or studied by scholars. Numerous documents are damaged, survive only in fragments, or are extremely difficult to decipher.
With the “Apollo” project – named after the antique patron of the arts and sciences – the Austrian Academy of Sciences (OeAW) is therefore developing the first large-scale AI language model for Ancient Greek. In collaboration with SAIL Reply and the leading AI company Mistral AI, an advanced AI system is being built that is intended to enable advanced search, automated transcription and reconstruction of ancient texts.
Anna Dolganov explained how this works and what benefits it brings to research at the “AI Now Summit” on 28 May 2026 in Paris organised by the software company Mistral, and shares her insights here in an interview.
Millions of previously inaccessible documents are now available
Ancient languages and historical documents – at first glance, this doesn’t sound like a typical area of application for artificial intelligence. Why does classical studies need AI at all?
Anna Dolganov: The problem is that we have vast quantities of documentary sources, particularly in the field of papyrology. There are huge collections of ancient texts around the world that have never been deciphered, or have only been partially deciphered. In the case of Greek papyri alone, we are talking about more than a million documents, of which barely six per cent have been catalogued to date. These sources are a vast treasure trove; they are often incredibly revealing and lead to amazing discoveries. But processing them is very time-consuming. Many texts have survived only in fragments or in a damaged state and are therefore very difficult to read. Experts cannot process and publish them quickly enough. As a result, a large part of our historical knowledge of antiquity remains inaccessible. This is precisely where AI can help.
There are vast collections of ancient texts worldwide that have never been deciphered, or only partially so.
With ‘Apollo’, the Austrian Academy of Sciences (OeAW) has created the world’s first large-scale AI language model for Ancient Greek. What can this system already do?
Dolganov: Apollo was developed specifically for Ancient Greek and trained on a corpus of 600 million words from historical texts. At the same time, we have worked out how a so-called Large Language Model (LLM) can be meaningfully trained and evaluated for an ancient language. The AI system can now be used to tackle research tasks that previously took a great deal of time or were only possible to a limited extent. These include, for example, advanced searches across the entire corpus of Greek texts and the reconstruction of damaged passages. The AI can suggest fill-ins for gaps – that is, missing words or passages – and assist researchers in their work with fragmentary sources.
With this project, we have been breaking new ground: developing an advanced AI system for a changing historical language with a dataset that is limited – from an LLM perspective – is a complex challenge that requires innovative strategies and an intensive exchange of knowledge between classical studies and AI. Together with our partners, we are thus not only creating a new set of tools for research, but also a new framework within which research questions can be posed and insights gained. In future, similar systems are also to be developed for Latin and other historical languages.
The AI can make suggestions for gaps – that is, missing words or passages of text – and support researchers in their work with fragmentary sources.
Training AI using historical data
What could AI achieve for the study of antiquity in ten years’ time – and why is this socially relevant?
Dolganov: I hope that in future we will be able to make far more historical sources accessible than we can today. If we can make previously unread texts available, we will gain new insights into the societies, cultures and people of the past. This deepens our understanding of history and opens up new avenues of research. For me, this also entails a social responsibility. Generative AI models are trained using vast human archives, that is, the achievements of human knowledge and human experience. Applying this technology to historical archives is an opportunity for humanity to regain sovereignty over its own knowledge and to assert an independent understanding of its past and of itself. In this way, the power of advanced technology is put at the service of the common good.
At a glance
Anna Dolganov is an Roman historian and papyrologist at the Austrian Archaeological Institute (OeAI) of the Austrian Academy of Sciences (OeAW). She leads the “Apollo” project on behalf of the OeAW, the world’s first large-scale AI language model for Ancient Greek. Together with international partners, she is developing AI systems designed to advance the decipherment and reconstruction of ancient documents, with the aim of making previously untapped sources accessible to researchers and the public.
