MusicFlow

Meta · 2024

Cascaded text-to-music generation model based on conditional flow matching, capable of zero-shot infilling and continuation.

Application types: Text-to-music
Architecture:Flow matching

Essential categories

CategoryStatusNotes
Source code✘ ClosedDespite the arXiv version of the paper stating "Our code and model will be publicly available", no official source repository, training framework, or inference code has been released.
Training data✘ ClosedMusicFlow was trained on 20K hours of proprietary music. No dataset access, tracklists, origin metadata, or permissions are disclosed.
Model weights✘ Closed
Code documentation✘ Closed
Training procedure~ PartialThe research paper details the mathematical formulations of Conditional Flow Matching, loss objectives, two-stage cascaded flow architecture (semantic and acoustic networks), and core hyperparameters. However, hardware requirements and specific model configurations are missing.
Evaluation procedure✔︎ OpenEvaluation protocols are thoroughly documented using the public MusicCaps benchmark. It includes standard objective metrics (FAD, FD, KL divergence, CLAP text-audio similarity) and human listening evaluations for text-to-music, infilling, and continuation.
Research paper✔︎ OpenAccepted at ICML 2024.
Licensing✘ ClosedNone.

Desirable categories

CategoryStatusNotes
Model card∅ Not included
Datasheet∅ Not included
Package∅ Not included
User-oriented application∅ Not included
Supplementary material page⭐ IncludedAn official supplementary website showcases audio generation examples across text-to-music, continuation, and infilling tasks, along with comparisons against AudioLDM.

Raw YAML file with complete evaluation.