← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

Researchers present Fusion Embedding, a family of models that extend a frozen vision-language embedding base to include audio without updating its parameters. The first generation uses a 16.4M-parameter connector for audio, while the second adds modality-gated deep adapters (44.2M parameters) that activate only for audio inputs. The models enable zero-shot audio-image retrieval by aligning audio to text, and both generations can be trained in hours on a single GPU.

Why it matters: This work demonstrates a practical method for unifying text, image, video, and audio embeddings in a single space, enabling cross-modal retrieval without requiring paired audio-visual training data.

Full story at: arXiv Computation and Language