DITTO-2

University of California San Diego and Adobe Research · 2024

Inference-time optimization framework that leverages diffusion distillation to speed up controllable music generation by 10–20x, enabling faster-than-real-time control over melody, structure, and intensity without requiring model fine-tuning.

Application types: Melody-to-musicText-to-musicAudio-to-audio
Architecture:Diffusion

Essential categories

CategoryStatusNotes
Source code✘ ClosedNo code repository.
Training data✘ ClosedTraining data neither available nor properly described.
Model weights✘ ClosedNot provided.
Code documentation✘ ClosedNo code repository.
Training procedure~ PartialTraining procedure is partially described.
Evaluation procedure~ PartialMetrics and data are described but exact implementations are missing.
Research paper✔︎ OpenAccepted at ISMIR 2024
Licensing✘ ClosedNot available.

Desirable categories

CategoryStatusNotes
Model card∅ Not includedNot available.
Datasheet∅ Not includedNot available.
Package∅ Not includedNot available.
User-oriented application∅ Not includedNot available.
Supplementary material page⭐ IncludedComplementary demo page is available, including sound examples of the model's capabilities.

Raw YAML file with complete evaluation.