Noise2Music

Google Research · 2023

A diffusion-based model to generate music from text prompts and demonstrate its capability by generating 30-second long 24kHz music clips.

Application types: Text-to-music
Architecture:Diffusion

Essential categories

CategoryStatusNotes
Source code✘ ClosedNo source codebase is available.
Training data~ PartialTraining data is described in detail. However, the resulting MuLaMCap dataset resulting from this work is not publicly available.
Model weights✘ ClosedModel weights are not available.
Code documentation✘ ClosedNo codebase is available. There is no documentation available either.
Training procedure~ PartialThe available pre-print describes partially the training procedure of the model. While some relevant information is described, such as model configuration and training details, there are key aspects of the training missing, such as hardware requirements, model checkpoints and hyperparameters.
Evaluation procedure~ PartialDespite the evaluation procedure is well-documented in regards to evaluation data and metrics, the evaluation process lacks some relevant details to ensure replicability. Moreover, evaluation code of the system is not available.
Research paper~ PartialPreprint is available. No peer-reviewed version of the article has been found.
Licensing✘ ClosedNot available.

Desirable categories

CategoryStatusNotes
Model card∅ Not includedNot available.
Datasheet∅ Not includedNot available.
Package∅ Not includedNot available.
User-oriented application∅ Not included
Supplementary material page⭐ IncludedComplementary site with sound examples is available.

Raw YAML file with complete evaluation.