MusicLDM

University of California San Diego, Mila-Quebec Artificial Intelligence Institute,University of Surrey, LAION · 2023

A text-to-music generation model that is only trained on 10000 songs.

Application types: Text-to-music
Architecture:Latent diffusionVAEGAN

Essential categories

CategoryStatusNotes
Source code~ PartialSystem source code is available at GitHub repository & Hugging Face Diffusers.
Training data~ PartialMusicLDM is trained on the Audiostock dataset, which contains 9000 music tracks for training and 1000 tracks for testing. Dataset is not directly provided, no information about the original source accessibility or requirements is provided.
Model weights✔︎ OpenModel checkpoints are available.
Code documentation✔︎ OpenDescription on how to use the code is provided in the GitHub repo and through Hugging Face Diffusers.
Training procedure✔︎ OpenTraining procedure is documented in the research article preprint and in the appendix additional page for the peer-reviewed version (https://musicldm.github.io/appendix/). Authors include specifications on hyperparameters, GPU requirements and model configurations.
Evaluation procedure✔︎ OpenEvaluation procedure is described in the research paper. Evaluation code is available in the model’s codebase and in its complementary repo (https://github.com/haoheliu/audioldm_eval).
Research paper✔︎ OpenThis article has been accepted at ICASSP 2024. The access of such paper is limited to IEEE Xplore access (https://ieeexplore.ieee.org/document/10447265). A preprint is available in arXiv (https://arxiv.org/pdf/2308.01546).
Licensing~ PartialAttribution-NonCommercial-ShareAlike 4.0 International

Desirable categories

CategoryStatusNotes
Model card⭐ IncludedAvailable at HuggingFace. Information about model’s evaluation and limitations is missing and model’s architecture is just mentioned.
Datasheet⭐ IncludedThere is a data card. However, content is limited to data sources origin and details on how to reproduce the data collection. Details on curation and other considerations, such as consent, limitations and selection strategies are missing.
Package∅ Not included
User-oriented application∅ Not includedAPI is shut down temporarily. Should be released again soon.
Supplementary material page⭐ IncludedDemo page is available with sound examples and demonstration of the model's capabilities.

Raw YAML file with complete evaluation.