Back

ARI Guest Talk, June 3, 2026: Gergely Vörös, TU Vienna

Acoustic Scene Classification on Synthetic Hearing-Aid-Oriented Datasets: A Comparative Study of YAMNet and Mag-SSM

Wednesday 03.06.2026 01:06 pm
Female speaker giving presentation in lecture hall at university workshop. Audience in the conference hall.
© Adobe Stock

Acoustic Scene Classification on Synthetic Hearing-Aid-Oriented Datasets: A Comparative Study of YAMNet and Mag-SSM

Acoustic scene classification (ASC) is critical for adaptive hearing aid systems, yet existing benchmarks and datasets are poorly aligned with patient-relevant contexts. Popular corpora such as TIMIT, LibriSpeech, AudioSet, and DEMAND have overly granular class structures, which limits their usefulness for assistive technologies. This study addresses this issue by presenting a synthetic, context-aware dataset designed for use in experiments for hearing aids. The dataset incorporates single-label clips of speech, environmental noise, and critical auditory events.

We evaluate two fundamentally different modeling paradigms on this dataset: YAMNet, a convolutional neural network optimized for general-purpose audio classification, and Mag-SSM, a state space model architecture that functions as a trainable magnitude spectrogram. Unlike conventional approaches relying on fixed feature extraction such as log-mel representations, Mag-SSM leverages complex-valued state space dynamics with a magnitude-based activation to learn task-adaptive spectral representations directly from raw audio. This formulation enables the model to approximate Short-Time Fourier Transform (STFT)-like behavior while maintaining linear computational complexity and high efficiency.

This study investigates two key areas: (i) which acoustic features are most discriminative in context-aware ASC scenarios, and (ii) how state-of-the-art and commercial models generalize to small-scale synthetic datasets. The experimental results reveal trade-offs between robustness, efficiency, and representational adaptability, showing that model inductive biases, particularly the ability to learn frequency-dependent representations, have a significant impact on performance under constrained and synthetic conditions.

These findings contribute to the design of ASC systems better aligned with assistive hearing applications and provide insights into dataset construction strategies and model architectures for context-aware audio intelligence.

 

Information

 

Date:
Wednesday, June 3, 2026, 13:30

Venue:
Otto Wagner PSK Building
Meeting Room 2-3 | Third Floor
Georg Coch-Platz 2
1010 Vienna

Organizer:
Acoustics Research Institute of the OeAW
Tel.: +43 1 51581 2520