Moûsai

ETH Zürich, IIT Kharagpur, Max Planck Institute · 2023

A cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions.

Application types: Text-to-music
Architecture:Latent diffusion

Essential categories

CategoryStatusNotes
Source code~ PartialThey provide an audio diffusion library that includes different models. However, the configs shown are indicative and untested, see https://arxiv.org/abs/2301.11757 for the configs used in the paper.
Training data~ PartialHow the data is collected and acquired, including licensing issues, is detailed in the research paper. However, the exact list of songs in the dataset nor direct access to the dataset is available.
Model weights✘ ClosedIn the GitHub repo, authors mention that “no pre-trained models are provided here”.
Code documentation~ PartialCode is partially described as how to use Moûsai is not straightforward.
Training procedure✔︎ OpenTraining procedure is describe in the article, including hardware requirements and model configuration.
Evaluation procedure~ PartialEvaluation data or code for evaluation are not specified.
Research paper✔︎ OpenAccepted at ACL 2024.
Licensing✔︎ OpenCode is under the MIT license.

Desirable categories

CategoryStatusNotes
Model card∅ Not included
Datasheet∅ Not included
Package⭐ IncludedCode belongs to https://github.com/archinetai/audio-diffusion-pytorch library.
User-oriented application∅ Not included
Supplementary material page⭐ IncludedSupplementary material is available.

Raw YAML file with complete evaluation.