Cornelius Frankenbach: Localization Accuracy in Dynamic Binaural and Near-Field Transaural Synthesis: Influence of HRTF Individualization and Crosstalk Cancellation
Binaural synthesis is a reproduction technique that enables auditory displays of virtual sound sources (VSSs) using head-related transfer functions (HRTFs), which characterize the acoustic paths between a sound source and the listener’s eardrums and are therefore as unique as the head and pinna geometry. Since individual HRTF measurement is impractical for most applications, generic dummy-head HRTFs are often used as substitutes, but result in less authentic virtual auditory displays. HRTF estimation aims at solving this problem but typically requires complex perceptual evaluation through listening experiments, which are therefore often replaced by spectral comparisons to measured HRTFs or virtual listening tests using auditory models. While binaural synthesis via headphones naturally ensures practically optimal left-right channel separation, transaural loudspeaker reproduction requires crosstalk cancellation (CTC). Computing CTC filters demands an additional set of transfer functions, which become distance-dependent in the near-field of the reproduction system.
This dissertation investigates the impact of HRTF individualization on the spatial perception of VSSs by means of two listening experiments. The first experiment provides a thorough comparison of five HRTF estimation methods against individually measured and generic HRTFs. The second experiment extends the previous state of the art by employing near-field transaural synthesis, while analyzing the CTC’s objective performance and impact on the listeners’ accuracy at localizing VSSs. Furthermore, the present work examines the correlation of spectral HRTF distance metrics with vertical localization errors and compares auditory models to a proposed model predicting the localization error, using experimental data.
Results show that HRTF individualization improves localization accuracy in vertical direction, while generic HRTFs suffice for accurate horizontal localization. In a near-field transaural synthesis, CTC is crucial to accurately synthesize VSSs that can be perceived as externalized in front of the listener, even when the physical loudspeakers are positioned closely behind—enabling applications such as virtual stereo playback. The accuracy of auditory localization models remains limited and requires further evaluation, likely due to the strong listener dependence of the relationship between HRTF distance metrics and vertical localization error.
Melden Sie sich hier an um Einladungen zu den Kolloquium-Vorträgen per E-Mail zu erhalten.
Register here to receive the invitations to colloquium talks via e-mail.
Zoom-Meeting-ID: 643 7384 8607
Passwort: 2020000