DiffRhythm

ASLP@NPU and Shenzhen Research Institute of Big Data · 2025

Song generation model able to synthesize complete songs with both vocal and accompaniment up to 4m45s.

Application types: Text-to-songLyrics-to-song
Architecture:Latent diffusion

Essential categories

CategoryStatusNotes
Source code✔︎ OpenSouce code is avaiable in the GitHub repo, including model architecture, inference, training and dataset processing.
Training data✘ ClosedTraining data is not available nor properly described. In the research paper, they indicate that the model was trained on 60,000 hours of audio content, including English (60%), Chinese (30%) and instrumental (10%) songs (about 1 million songs). However, no information on specific data sources, acquisition methods, and data licensing is given. See "Dataset Setup" in Section 4.1. of the pre-print.
Model weights✔︎ OpenModel weights are avaiable in their HuggingFace page. Multiple versions of the model are avaiable.
Code documentation~ PartialCode documentation inclues instructions on how to install and inference the model, but there are no instuctions on how to train it (it is indicated at "coming soon").
Training procedure✔︎ OpenTraining procedure is properly described in the research paper, including details on the models parameters and hardware requirements.
Evaluation procedure~ PartialEvaluation procedure is described in the research paper, including metrics (see 4.3) and detailed results (see 5). However, evaluation data is not sufficently described nor avaiable to reproduce the evaluation.
Research paper~ PartialResearch paper is available as preprint in arXiv. Also for more recent versions of the model: DiffRhythm+ and DiffRhythm 2. No peer reviewed version is avaiable for any of the papers.
Licensing✔︎ OpenApache 2.0 License

Desirable categories

CategoryStatusNotes
Model card⭐ IncludedModel card available in Hugging Face.
Datasheet∅ Not included
Package∅ Not included
User-oriented application⭐ IncludedA demo is available in HuggingFace.
Supplementary material page⭐ IncludedA supplementary page is available, including a brief description of the model, links to relevant sites (e.g., demo, paper) and some examples showcasing the model's abilities.

Raw YAML file with complete evaluation.