
Research Scientist
Mathematics Cluster
Frame Theory and its Implementation
Speaker of Machine Learning Team
Tel. +43 1 51581-2532
Email: nicki.holighaus(at)oeaw.ac.at
Scientific IDs:
ORCID: 0000-0003-3837-2865
Google Scholar: Nicki Holighaus
ResearchGate: researchgate.net/profile/Nicki_Holighaus
Academic Background
Nicki Holighaus studied mathematics and theoretical computer sciences at Justus–Liebig–University, Gießen, Germany. He graduated in 2010. After three years of doctoral studies at the University of Vienna, Austria, where he worked as a research assistant at the Numerical Harmonic Analysis Group (NuHAG), he successfully defended his PhD thesis "Theory and implementation of adaptive time-frequency transforms” in October 2013.
Since August 2012 he is part of the Acoustic Research Institute's workgroup "Mathematics and Signal Processing in Acoustics", where he works on theoretical and applied aspects of frames and adapted time-frequency representations.
Current Research
His research focuses on advanced time-frequency methods in signal processing, including time-frequency analysis, the mathematical theory and design of adaptive and adapted time-frequency representations, time-frequency processing in acoustics and the use of time-frequency representations in machine learning for acoustics.
Current research projects: MERLIN
Current topics:
- Theory and application of warped time-frequency representations
- Function spaces and discretization for structured continuous frames
- Structure of time-frequency phase
- Signal processing with time-frequency phase
- Deep learning with time-frequency features
- Neural audio generation
- Audio inpainting with generative neural networks
- Time-frequency processing and perception
Publications
- Selebi: Percussion-Aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform. / Akaishi, Natsuki; Holighaus, Nicki; Yatabe, Kohei.
In: Ieee Transactions on Audio Speech and Language Processing, Vol. 34, 09.07.2026, p. 3754-3766.Phase vocoder-based time-stretching is a widely used technique for the time-scale modification of audio signals. However, conventional implementations suffer from "percussion smearing," a well-known artifact that significantly degrades the quality of percussive components. We attribute this artifact to a fundamental time-scale mismatch between the temporally smeared magnitude spectrogram and the localized, newly generated phase. To address this, we propose SELEBI, a signal-adaptive phase vocoder algorithm that significantly reduces percussion smearing while preserving stability and the perfect reconstruction property. Unlike conventional methods that rely on heuristic processing or component separation, our approach leverages the nonstationary Gabor transform. By dynamically adapting analysis window lengths to assign short windows to intervals containing significant energy associated with percussive components, we directly compute a temporally localized magnitude spectrogram from the time-domain signal. This approach ensures greater consistency between the temporal structures of the magnitude and phase. Furthermore, the perfect reconstruction property of the nonstationary Gabor transform guarantees stable, high-fidelity signal synthesis, in contrast to previous heuristic approaches. Experimental results demonstrate that the proposed method effectively mitigates percussion smearing and yields natural sound quality.
- Prediction of parameters of a pinna model from synthetic geometries using a vision transformer. / Pausch, Florian; Perfler, Felix; Holighaus, Nicki et al.
In: Journal of the Acoustical Society of America, Vol. 159, No. 6, 18.06.2026, p. 5578-5598.The acquisition of the human pinna geometry requires elaborate equipment for accurate results. Even then, the results are often corrupted by measurement artifacts. We introduce Mesh2PPM, a framework facilitating the generation of a personalized and artifact-free pinna mesh. Mesh2PPM predicts the parameters of a parametric pinna model based on cubic B & eacute;zier curves (BezierPPM) from a pinna mesh via a vision transformer. We evaluated Mesh2PPM with multi-view renderings of synthetic pinna geometries, providing additional depth information, varying the grids of camera views, and jittering the camera views. While added depth information had no practically relevant effect, a grid with 3 & times;3 camera views facilitated the lowest overall prediction errors and best counteracted the detrimental effects of jitter. For this grid, with and without jitter, the median Pompeiu-Hausdorff distances were 1.98 mm and 1.34 mm, respectively, and the root mean square distances were 0.92 mm and 0.52 mm. A refined analysis targeting the perceptually most important pinna regions for sound localization showed that multi-view information particularly improved the prediction of BezierPPM parameters describing the cavum-conchae region. The accuracy achieved indicates the suitability of Mesh2PPM to retrieve BezierPPM parameters from pinna meshes.