Stable Audio Open

Stability AI · 2024

Synthesizes variable-length stereo audio, sound effects, and musical stems from text prompts.

Application types: Text-to-musicText-to-sample
Architecture:Diffusion

Essential categories

CategoryStatusNotes
Source code✔︎ OpenCode for data processing, training pipeline, and inference is available in the stable-audio-tools repository. Architecture of the model is specified in the form of a config file.
Training data✔︎ OpenTraining data can be mapped from Attribution files. Dataset attribution files can also be accessed in HuggingFace after accepting Stability AI conditions (https://huggingface.co/stabilityai/stable-audio-open-1.0/blob/main/fma_dataset_attribution2.csv and https://huggingface.co/stabilityai/stable-audio-open-1.0/blob/main/freesound_dataset_attribution2.csv).
Model weights✔︎ OpenWeights are available in the HuggingFace repository. However, accessibility is subject to accepting Stability AI conditions.
Code documentation✔︎ OpenDocumentation of the code is limited for training replicating the model. Although there are not specific config files for Stable Audio Open, those from Stable Audio could be used.
Training procedure✔︎ OpenDescribed in pre-print.
Evaluation procedure✔︎ OpenThey use stable-audio-metrics for evaluation, which has examples of evaluation pipelines.
Research paper✔︎ OpenAccepted at ICASSP 2026.
Licensing~ PartialStability AI Community License. Not an Open Source Initiate (OSI) or RAIL approved licences.

Desirable categories

CategoryStatusNotes
Model card⭐ IncludedModel card in HuggingFace.
Datasheet∅ Not included
Package⭐ IncludedThe model can be used with 1) the https://github.com/Stability-AI/stable-audio-tools library and 2) the https://huggingface.co/docs/diffusers/main/en/index library.
User-oriented application⭐ IncludedGradio interface in source code.
Supplementary material page⭐ IncludedA complementary site with sound examples and demonstration of the model's capabilities is available.

Raw YAML file with complete evaluation.