SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
SALMONN-2 is an audio large language model (ALLM) built on a unified self-supervised learning (SSL) encoder, enhanced by a multi-layer feature fusion adapter that aggregates hierarchical representations. The model achieves state-of-the-art performance on several ALLM understanding benchmarks among comparable-scale open-weight models. The study also demonstrates that multimodal in-context learning (MICL) can be effectively acquired through targeted contextual biasing training, rather than emerging naturally.
Why it matters: This work demonstrates that general-purpose self-supervised audio encoders can match or surpass specialized supervised encoders, potentially streamlining the development of versatile audio AI systems.
Full story at: arXiv Audio and Speech Processing ↗