The project “IIX - Index and Inventory of Text Corpora”  is aimed at making the texts of the “AAC - Austrian Academy Corpus” searchable again.

In a first step the texts of this historical corpus will be made accessible in the NoSketch Engine for internal use only within the Literary & Print Culture Studies research unit.

The digital text corpus of the AAC contains around 600 million tokens, based upon more than 2.5 million digital images of printed publications (books, booklets, journals, newspapers etc.) published between 1848 and 1989, digitalized within the framework of the Austrian Academy Corpus in the first decade of the 21st century.

Publications

  • Biber, Hanno. 2020. "Challenges for Making Use of a Large Text Corpus such as the ‘AAC – Austrian Academy Corpus’ for Digital Literary Studies." In: Piotr Banski, Barbaresi, Adrien, Clematide, Simon, Kupietz, Marc, Lüngen, Harald, and Pisetta, Ines. Proceedings of the LREC 2020 Workshop. 8th Workshop on Challenges in the Management of Large Corpora (CMLC-8 2020).  Webseite   Download
  • Biber, Hanno. 2022. "'The word expired when that world awoke'. New challenges for research with large text corpora and corpus-based discourse studies in totalitarian times." In: Piotr Banski, Barbaresi, Adrien, Clematide, Simon, Kupietz, Marc, and Lüngen, Harald. Proceedings of the LREC 2022 Workshop on challenges in the management of large corpora (CMLC-10 2022). Webseite

Project lead

Hanno Biber

Contact

Hanno Biber

Project duration

01/2024–12/2026