AI audio news — Page 5

Updates on AI voice, speech, music, sound generation, and audio understanding from across the industry.

ModelsOfficialCohere Blog

Cohere Releases Open-Source Arabic Speech Recognition Model

Cohere has launched Transcribe Arabic, a state-of-the-art, enterprise-ready speech recognition model for Arabic speakers. The model is available as open source and is designed to capture the full diversity of spoken Arabic.

Why it matters: This release addresses the need for accurate transcription across diverse Arabic dialects, with open-source availability enabling broader enterprise and developer adoption.

Products & AgentsOfficialElevenLabs Blog

Alpha Bank expands its collaboration with ElevenLabs

Alpha Bank has expanded its collaboration with ElevenLabs by using ElevenAgents to build a new service aimed at simplifying customer communication and providing faster, more accessible interactions. The announcement was made on ElevenLabs' blog.

Why it matters: This partnership highlights the increasing use of AI agents in banking to improve customer service.

Products & AgentsOfficialElevenLabs Blog

ElevenLabs Launches ElevenAgents Spotlight for Agent Improvement

ElevenLabs has introduced ElevenAgents Spotlight, an observation and improvement layer for its ElevenAgents platform. The tool is designed to help increase resolution rates and customer satisfaction across all channels.

Why it matters: This enhancement offers a systematic approach to improving AI agent performance, which could advance customer service automation.

Products & AgentsOfficialElevenLabs Blog

Fyxer uses ElevenLabs Speech to Text to power meeting notetaker

Fyxer, a meeting notetaker, uses ElevenLabs' Scribe v2 speech-to-text model, resulting in a 15% relative lift in user conversion. The integration demonstrates the model's effectiveness for real-time transcription.

Why it matters: This case study highlights the tangible benefits of advanced speech-to-text models for user engagement in productivity tools.

Products & AgentsOfficialElevenLabs Blog

CoopVoce adopts ElevenLabs voice AI for support calls

Italian telecom CoopVoce has deployed ElevenLabs' voice AI for customer support calls, resulting in a 10% increase in customer willingness to engage with its AI assistant within days. The AI assistant uses natural-sounding speech to handle customer inquiries, aiming to improve the user experience.

Why it matters: This deployment demonstrates how advanced voice AI can quickly enhance customer engagement in telecom support.

Products & AgentsOfficialElevenLabs Blog

ElevenLabs Introduces Tools on ElevenMusic for Music Creation and Reshaping

ElevenLabs has launched new tools on its ElevenMusic platform, allowing users to record vocals, add musical ideas, or transform completed tracks. These features are designed to help users create, reshape, and evolve their music.

Why it matters: This development expands ElevenLabs' AI capabilities into music production, providing creators with new ways to generate and modify audio content.

ModelsOfficialAllen Institute for AI

MolmoAct 2 Powers Voice-Controlled Robot to Win Embodied AI Hackathon

Robotics engineer Binh Pham used the Allen Institute for AI's MolmoAct 2 to build a voice-controlled robot that won the South Park Commons embodied AI hackathon. This achievement highlights the capabilities of open models in advancing robotics innovation.

Why it matters: This demonstrates the potential of open models like MolmoAct 2 to accelerate progress in embodied AI and robotics.

Products & AgentsOfficialGoogle DeepMind

Gemini app now features Lyria 3 for music generation

Google DeepMind has integrated its most advanced music generation model, Lyria 3, into the Gemini app. Users can now create 30-second tracks using text or images.

Why it matters: This development makes AI-powered music creation more accessible to a broad audience through a widely used app.

ModelsOfficialMistral AI News

Mistral AI Launches Voxtral, a Real-Time Speech Transcription Model

Mistral AI has announced Voxtral, a speech transcription model that transcribes audio at the speed of sound. The model is intended for real-time transcription applications.

Why it matters: Voxtral could advance real-time speech recognition by enabling faster transcription services.

Companies & FundingOfficialStability AI News

Warner Music Group and Stability AI Partner on Responsible AI Music Tools

Warner Music Group and Stability AI have announced a collaboration to develop responsible AI tools for music creation. The partnership aims to combine WMG's advocacy for principled innovation with Stability AI's expertise in commercially-safe generative audio.

Why it matters: This partnership highlights a major label's commitment to integrating AI into music production with a focus on ethical and legal safeguards.

Companies & FundingOfficialStability AI News

Universal Music Group and Stability AI Announce Strategic Alliance for AI Music Tools

Universal Music Group and Stability AI have announced a strategic alliance to co-develop next-generation professional music creation tools. These tools will use responsibly trained generative AI and are intended to support the creative process for artists, producers, and songwriters.

Why it matters: This partnership highlights a major move toward integrating generative AI into professional music creation with an emphasis on responsible development.

ModelsOfficialStability AI News

Stability AI Releases Stable Audio 2.5 for Enterprise Sound Production

Stability AI has launched Stable Audio 2.5, its latest audio model developed specifically for enterprise-grade sound production. The model introduces advancements in quality and control, enabling dynamic compositions that can be tailored to custom brand requirements.

Why it matters: This release represents the first audio model designed for enterprise use, supporting scalable and customizable sound production for brands.

Open SourceOfficialStability AI News

Stability AI and Arm Release Stable Audio Open Small for On-Device Audio Generation

Stability AI, in partnership with Arm, has open-sourced Stable Audio Open Small, a compact variant of its text-to-audio model. The new model is designed to be smaller and faster while maintaining output quality and prompt adherence, enabling real-world deployment on devices.

Why it matters: This collaboration enables high-quality AI audio generation on smartphones and other edge devices, expanding accessibility and potential on-device applications.

ResearchOfficialarXiv AI/ML

SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition

A new study revisits XAI-guided adaptive fusion (XGAF) for multimodal emotion and sentiment recognition, using TreeSHAP attribution magnitudes to weight unimodal and cross-modal experts. On MELD 7-class emotion recognition, sum-abs XGAF nearly matches early fusion (0.5983 vs 0.6018) and significantly outperforms late fusion (0.4598). On CMU-MOSEI 3-class sentiment, sum-abs XGAF slightly exceeds early fusion (0.6519 vs 0.6485).

Why it matters: This work provides a transparent empirical analysis of how SHAP reduction methods and expert dimensionality affect modular multimodal fusion, offering a principled alternative to monolithic early fusion.

Products & AgentsReportedThe Register / AI & ML

OpenAI introduces GPT-Live for more natural ChatGPT conversations

OpenAI has launched GPT-Live, a new feature that allows ChatGPT to talk, listen, and formulate answers simultaneously. This update is designed to make conversations with ChatGPT more natural and fluid.

Why it matters: The improvement could make AI interactions feel more human-like, enhancing user experience with conversational agents.

Companies & FundingReportedTechCrunch / AI

Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia

Paris-based AI voice startup Gradium has raised a $100 million seed round backed by Nvidia. The company plans to use the funds to open a Bay Area office and compete for talent, aiming to strengthen its position in the global AI ecosystem.

Why it matters: This large seed round from a major AI hardware company signals strong investor confidence in AI voice technology and the globalization of AI startups.

ModelsOfficialOpenAI News

OpenAI Introduces GPT-Live: Next-Generation Voice Models for ChatGPT Voice

OpenAI has announced GPT-Live, a new generation of voice models designed for natural human-AI interaction. This model now powers ChatGPT Voice, enhancing real-time conversational capabilities.

Why it matters: GPT-Live represents a significant advancement in voice AI, enabling more fluid and natural spoken interactions with AI systems.

ResearchOfficialApple Machine Learning Research

Apple ML Research Proposes Method for Text-to-Sounding Video Generation

Apple Machine Learning Research has published a study on Text-to-Sounding-Video (T2SV) generation, which aims to produce videos with synchronized audio from text. The research identifies challenges such as text conditioning bottlenecks and unclear cross-modal fusion mechanisms, and proposes solutions to improve alignment between modalities.

Why it matters: This work advances multimodal AI by addressing the synchronization of video and audio from text, which has applications in content creation and accessibility.

ModelsOfficialHugging Face Blog

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Hugging Face and Cerebras have partnered to enable real-time voice AI using the Gemma 4 model. The collaboration utilizes Cerebras hardware to achieve low-latency inference for voice applications.

Why it matters: This partnership could make real-time conversational AI more practical by reducing latency in voice AI systems.

ResearchOfficialHugging Face Blog

Hugging Face Launches FFASR Leaderboard for Real-World ASR Benchmarking

Hugging Face has introduced the FFASR Leaderboard, a new benchmark designed to evaluate automatic speech recognition (ASR) systems under real-world conditions. The leaderboard aims to provide a more practical assessment of ASR performance beyond standard datasets.

Why it matters: This benchmark addresses the gap between lab-tested ASR accuracy and real-world performance, helping developers choose models that work reliably in diverse acoustic environments.