
This Tool Gallery is designed to introduce researchers and practitioners to the diverse landscape of Named Entity Recognition (NER) approaches. The workshop covers classical machine learning techniques, and transformer-based models like BERT and SpaCy pipelines. Next to a theoretical introduction, participants will engage in hands-on exercises comparing these approaches across datasets, evaluating trade-offs in accuracy, computational cost, domain adaptability, and annotation requirements. The curriculum further addresses practical considerations including zero-shot and few-shot NER using large language models, domain-specific fine-tuning strategies, and ensemble methods. We provide all workshop materials, including annotated notebooks and evaluation scripts, as open-access resources to support reproducibility and broader adoption in academic contexts.
9:45 – 10:00: Doors Open
10:00 – 10:45: Introduction to NER
10:45 – 11:00: Coffee Break
11:00 – 11:45: Classical Machine Learning Methods – Theory
11:45 – 12:30: Classical Machine Learning Methods – Hands-On
12:30 – 13:30: Lunch Break
13:30 – 14:15: Transformer based Methods – Theory
14:15 – 15:00: Transformer based Methods – Hands-On
15:00 – 15:45: Wrap Up – Open Discussion
Bring your laptop
Basic experience in working with Jupyter notebooks and Google Colab are useful
Basic Python
Daniel Elsner
... studied development studies at the University of Vienna and received a Bachelor of Arts degree. Continuing with historical studies at the University of Vienna, the master’s degree program he enrolled in had a focus on digital humanities and global history. His master thesis, with the title "A global perspective: Singapore migration and labor recruitment networks", applied several digital methods. Among them, GIS is for mapping migration and statistical analysis with empirical census data. The master’s program was concluded by receiving a Master of Arts degree in early 2021.
As a research data and software engineer (RDE/RSE) at the ACDH (then ACDH-CH), he supports research projects with the development and implementation of digital methods for digital editions, corpora, linked (open) data, data curation and machine learning (AI).
He collaborates in projects of the research units DH Research & Infrastructure (2020 – today), Musicology (2021 – 2024), Literary & Print Culture Studies (2024 – today) and Linguistics (2025 – today).
Stefan Resch
... is a Software Architect / NLP Analyst in the research unit DH Research & Infrastructure. His stack includes Python, Bash, Linux, Docker, Pandas, NumPy, spaCy, Django, RDF and SPARQL. In his current project CLSInfra he is a Software Architect, responsible for the implementation of VELD, an architecture enabling stable and reproducable interoperability between heterogenous tools and data sets with a focus on NLP pipelining. As Backend Developer of APIS, he is responsible for the major refactoring of the inner architecture of the APIS app, increasing the flexibility of its business logic towards arbitrary project ontologies, and improving the application’s compliancy with RDF paradigms.
In his former project MARA, as Software Architect and Natural Language Processing Analyst, he supported the researchers in media corpus analysis by providing them with iterative Supervised Machine Learning workflows. For this, he implemented the MARA NLP Suite, encapsulating the entire data cycle and providing the public with reproducability of its data and code. In the project SOLA he worked as Backend Developer and Ontology Designer, adapting the entity database APIS towards the research field of investigating juristical texts of early christianity. As Backend Developer and Ontology Designer of the Jelinek Werkverzeichnis Online, he adapted the entity database APIS towards the extensive bibliography of the Austrian writer Elfriede Jelinek and her professional and artistic connections throughout her career. As Data Analyst and supportive Backend Developer of Parthenos, he assured the harmonization results of data aggregation pipelines on heterogeneous data sets originating from diverse research archives.
23. September 2026
10:00 - 16:00
Österreichische Akademie der Wissenschaften
Bäckerstraße 13
Seminarraum 1, Ground Floor / Courtyard
1010 Wien
English