| |
Stable Audio 3
Stable Audio 3 is a family of fast latent diffusion models capable of generating variable-length audio and music in seconds on consumer hardware, using a novel semantic-acoustic autoencoder to compress audio into a compact latent space. The models support audio editing and inpainting through targeted modifications, and were trained on licensed and Creative Commons data with adversarial post-training to improve inference speed and generation quality. The researchers have released the weights and code for small and medium model variants that run on consumer-grade devices.
Read Full Article →
← More Tech news