MG²

Guangxi University, China and Southwestern University of Finance and Economics, China · 2024

A lightweight melody-guided text-to-music model. Melody information is incorporated implicitly within the CLMP and explicitly in the retrieval augmented diffusion module.

Application types: Text-to-musicMelody-to-music
Architecture:Diffusion

Essential categories

CategoryStatusNotes
Source code✔︎ OpenSource code is available on GitHub, including data processing, model architecture, training and fine-tuning, and inference scripts.
Training data✔︎ OpenThis model was trained on about 132 hours from two open datasets: MusicCaps and MusicBench, which are detailed in Hugging Face.
Model weights✔︎ OpenCheckpoints are available on Hugging Face.
Code documentation✔︎ OpenSource code is well-documented, including environment setup,data processing, training and fine-tuning procedure, as well as quick start instructions.
Training procedure✔︎ OpenTraining procedure is fully described in the research paper (see Section 5.3). Code documentation provides additional instructions on how the model might be retrained or fine-tuned.
Evaluation procedure✔︎ OpenEvaluation procedure is fully described in the research paper (see Section 5.4). Evaluation data is 1/10 of MusicCaps and MusicBench. Testing splits are properly identified in the dataset card. In addition, a human evaluation was conducted to assess recognizability, text relevance, satisfaction, quality and market potential of generated samples (see Section 6). These samples are shared in their complementary website.
Research paper~ PartialResearch paper is available on arXiv, but it has not yet been peer-reviewed.
Licensing✔︎ OpenThe project is licensed under the MIT License.

Desirable categories

CategoryStatusNotes
Model card⭐ IncludedA model card is available on Hugging Face.
Datasheet⭐ IncludedA dataset card is provided on Hugging Face.
Package∅ Not included
User-oriented application∅ Not includedAn online-based user-oriented application should be available in the indicated link. However, it is not currently accessible.
Supplementary material page⭐ IncludedAn official project website showcases sonified audio samples generated by MG² alongside comparative baselines (AudioLDM2, Mustango).

Raw YAML file with complete evaluation.