LeVo

Tsinghua University, Shenzhen and Tencent AI Lab · 2025

Full-song lyrics-to-music foundation model employing parallel token modeling (mixed semantic tokens and dual-track acoustic tokens) and multi-preference alignment.

Application types: Lyrics-to-songText-to-musicStyle transfer
Architecture:Autoregressive transformer

Essential categories

CategoryStatusNotes
Source code✘ ClosedSource code is not currently accessible.
Training data~ PartialLeVo is trained on 2 million music tracks (~110,000 hours), including part of DISCO-10M, Millong Song Dataset and some copyrighted in-house data. No further information is provided, and the training data is not publicly available.
Model weights✘ Closed
Code documentation✘ Closed
Training procedure✔︎ OpenTraining procedure is describes in the research paper, including hardware requirements and model hyperparameters (see Appendix B for further details).
Evaluation procedure~ PartialEvaluation protocols are thoroughly detailed, including objective metrics and a 20-expert subjective evaluation (see Appendix D for futher details). However, dataset for evaluation is not fully specified nor publicly available.
Research paper✔︎ OpenAccepted at NeurIPS 2025.
Licensing✘ Closed

Desirable categories

CategoryStatusNotes
Model card∅ Not included
Datasheet∅ Not included
Package∅ Not included
User-oriented application∅ Not included
Supplementary material page⭐ IncludedAn official project website is live and provides sonified audio demonstrations of generated music, zero-shot style transfer, and text-based control.

Raw YAML file with complete evaluation.