Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Apple ML Research has introduced a memory-efficient audio synthesis architecture for Siri Expressive Voices, enabling real-time, on-device speech synthesis. The system uses a detokenizer to convert semantic audio tokens into high-fidelity audio with a decoupled temporal depth diffusion transformer, optimized for the Apple Matrix Coprocessor (AMX).
Why it matters: This work advances real-time, privacy-preserving voice synthesis capabilities directly on consumer devices.
Full story at: Apple Machine Learning Research ↗