Musical Attribute-based Triplet Comparison with Human Annotations:
A perceptual dataset for evaluating attribute-based music similarity against human judgment.
1Music Technology Group, Universitat Pompeu Fabra ·
2Sony AI ·
3Joint Research Centre, European Commission ·
4Sony Group Corporation
Recent advances in generative AI are raising ethical concerns regarding the originality of generated content and the potential replication of training data, with further implications for transparency, attribution, and intellectual property. In music, several computational approaches have been proposed to identify potential replication, using audio-based similarity metrics. Yet, their alignment with human judgments across distinct musical attributes remains underexplored.
To address this gap, we conduct a perceptual experiment on music matches, defined as strongly similar musical excerpts. We focus on five musical attributes: melody, harmony, rhythm, voice, and timbre. We design a triplet-based forced-choice task comprising 300 cases, including plagiarism examples, cover songs, and AI-generated music. From this experiment, we introduce the MATCHA dataset: a collection of 1105 perceptual assessments of attribute-based music matches from 83 expert participants.
Our findings reveal measurable agreement among participants in identifying matches across attributes. We further observe partial alignment between human judgments and computational similarity measures. Overall, this work underscores the importance of domain-specific and perceptually grounded evaluation frameworks for generative AI in creative practice.
| Origin | Section | Source | Cases | w/ Vocals | % |
|---|---|---|---|---|---|
| Human | Human-Plagiarism (HP) | SMP Dataset | 75 | 61 | 25 |
| Human-Version (HV) | Discogs-VI | 75 | 63 | 25 | |
| AI | AI-SAO (AS) | Stable Audio Open | 120 | 0 | 40 |
| AI-Media (AM) | Media articles [1,2] | 30 | 27 | 10 |
Stimuli distribution across reference sources, by origin, section, source, number of cases, number of cases containing vocals, and overall percentage.
Two representative cases per category, with the majority decision and inter-rater agreement for each attribute. N is the number of responses received for that case (minimum 3).
The paper is currently under review. The entry below refers to the arXiv preprint and will be updated with final venue details on acceptance.