Musika

Johannes Kepler University Linz · 2022

A music generation system that can be trained on hundreds of hours of music using a single consumer GPU, and that allows for much faster than real-time generation of music of arbitrary length on a consumer CPU.

Application types: Audio-to-audio
Architecture:GAN

Essential categories

CategoryStatusNotes
Source code✔︎ OpenCodebase is available and complete.
Training data~ PartialTraining data is described in the research paper. Some of the data used is publicly available (LibriTTS corpus). However, there are some sections of the used data that are not accessible or detailed enough to fully reconstruct the dataset (South by SouthWest). Also, music coming from http://jamendo.com under “techno” genre is not detailed.
Model weights✔︎ OpenWeights are publicly available.
Code documentation✔︎ OpenCodebase is documented, including details on environment and configuration settings.
Training procedure✔︎ OpenDetailed in research paper and proper instructions are given in the model repository.
Evaluation procedure~ PartialFAD is used for evaluation. However, there are details missing on the evaluation procedure.
Research paper✔︎ OpenAccepted at ISMIR 2022.
Licensing✔︎ OpenMIT License.

Desirable categories

CategoryStatusNotes
Model card∅ Not includedNot available.
Datasheet∅ Not includedNot available.
Package∅ Not includedNot available. Source code is provided through GitHub, but not as an installable packaged or with version control.
User-oriented application⭐ IncludedAn interface is provided through Gradio. In addition, there is a colab notebook that intends to ease accessibility for non technical users.
Supplementary material page⭐ IncludedComplementary page is available with sound examples and demonstration of the model's capabilities.

Raw YAML file with complete evaluation.