ACDH Tool Gallery 12.3

Introduction to Named Entity Recognition (NER)

This Tool Gallery is designed to introduce researchers and practitioners to the diverse landscape of Named Entity Recognition (NER) approaches. The workshop covers classical machine learning techniques, and transformer-based models like BERT and SpaCy pipelines. Next to a theoretical introduction, participants will engage in hands-on exercises comparing these approaches across datasets, evaluating trade-offs in accuracy, computational cost, domain adaptability, and annotation requirements. The curriculum further addresses practical considerations including zero-shot and few-shot NER using large language models, domain-specific fine-tuning strategies, and ensemble methods. We provide all workshop materials, including annotated notebooks and evaluation scripts, as open-access resources to support reproducibility and broader adoption in academic contexts.  

Curriculum: 

9:45 – 10:00: Doors Open 

10:00 – 10:45: Introduction to NER 

10:45 – 11:00: Coffee Break 

11:00 – 11:45: Classical Machine Learning Methods – Theory 

11:45 – 12:30: Classical Machine Learning Methods – Hands-On 

12:30 – 13:30: Lunch Break 

13:30 – 14:15: Transformer based Methods – Theory 

14:15 – 15:00: Transformer based Methods – Hands-On 

15:00 – 15:45: Wrap Up – Open Discussion 

Requirements

  • Bring your laptop 

  • Basic experience in working with Jupyter notebooks and Google Colab are useful 

  • Basic Python 

Team

Daniel Elsner

... studied development studies at the University of Vienna and received a Bachelor of Arts degree. Continuing with historical studies at the University of Vienna, the master’s degree program he enrolled in had a focus on digital humanities and global history. His master thesis, with the title "A global perspective: Singapore migration and labor recruitment networks", applied several digital methods. Among them, GIS is for mapping migration and statistical analysis with empirical census data. The master’s program was concluded by receiving a Master of Arts degree in early 2021.

As a research data and software engineer (RDE/RSE) at the ACDH (then ACDH-CH), he supports research projects with the development and implementation of digital methods for digital editions, corpora, linked (open) data, data curation and machine learning (AI).

He collaborates in projects of the research units DH Research & Infrastructure (2020 – today), Musicology (2021 – 2024), Literary & Print Culture Studies (2024 – today) and Linguistics (2025 – today).



Stefan Resch

... is a Software Architect / NLP Analyst in the research unit DH Research & Infrastructure. His stack includes Python, Bash, Linux, Docker, Pandas, NumPy, spaCy, Django, RDF and SPARQL.  In his current project CLSInfra he is a Software Architect, responsible for the implementation of VELD, an architecture enabling stable and reproducable interoperability between heterogenous tools and data sets with a focus on NLP pipelining. As Backend Developer of APIS, he is responsible for the major refactoring of the inner architecture of the APIS app, increasing the flexibility of its business logic towards arbitrary project ontologies, and improving the application’s compliancy with RDF paradigms.

In his former project MARA, as Software Architect and Natural Language Processing Analyst, he supported the researchers in media corpus analysis by providing them with iterative Supervised Machine Learning workflows. For this, he implemented the MARA NLP Suite, encapsulating the entire data cycle and providing the public with reproducability of its data and code. In the project SOLA he worked as Backend Developer and Ontology Designer, adapting the entity database APIS towards the research field of investigating juristical texts of early christianity. As Backend Developer and Ontology Designer of the Jelinek Werkverzeichnis Online, he adapted the entity database APIS towards the extensive bibliography of the Austrian writer Elfriede Jelinek and her professional and artistic connections throughout her career. As Data Analyst and supportive Backend Developer of Parthenos, he assured the harmonization results of data aggregation pipelines on heterogeneous data sets originating from diverse research archives.


Date

23. September 2026 

10:00 - 16:00 

Location

Österreichische Akademie der Wissenschaften 
Bäckerstraße 13 

Seminarraum 1, Ground Floor / Courtyard
1010 Wien 
 

Language

English

Registration

Registration here.