MATCHA

Musical Attribute-based Triplet Comparison with Human Annotations:
A perceptual dataset for evaluating attribute-based music similarity against human judgment.

Roser Batlle-Roca1, Woosung Choi2, Joan Serrà2, Fabio Morreale2,
Wei-Hsiang Liao2, Xavier Serra1, Emilia Gómez1,3, Yuki Mitsufuji2,4

1Music Technology Group, Universitat Pompeu Fabra  ·  2Sony AI  · 
3Joint Research Centre, European Commission  ·  4Sony Group Corporation

Read the paper GitHub repo Dataset on Zenodo

Abstract

Recent advances in generative AI are raising ethical concerns regarding the originality of generated content and the potential replication of training data, with further implications for transparency, attribution, and intellectual property. In music, several computational approaches have been proposed to identify potential replication, using audio-based similarity metrics. Yet, their alignment with human judgments across distinct musical attributes remains underexplored.

To address this gap, we conduct a perceptual experiment on music matches, defined as strongly similar musical excerpts. We focus on five musical attributes: melody, harmony, rhythm, voice, and timbre. We design a triplet-based forced-choice task comprising 300 cases, including plagiarism examples, cover songs, and AI-generated music. From this experiment, we introduce the MATCHA dataset: a collection of 1105 perceptual assessments of attribute-based music matches from 83 expert participants.

Our findings reveal measurable agreement among participants in identifying matches across attributes. We further observe partial alignment between human judgments and computational similarity measures. Overall, this work underscores the importance of domain-specific and perceptually grounded evaluation frameworks for generative AI in creative practice.

Dataset

300triplet cases
1105perceptual annotations
83expert participants
5musical attributes
Origin Section Source Cases w/ Vocals %
Human Human-Plagiarism (HP) SMP Dataset 756125
Human-Version (HV) Discogs-VI 756325
AI AI-SAO (AS) Stable Audio Open 120040
AI-Media (AM) Media articles [1,2] 302710

Stimuli distribution across reference sources, by origin, section, source, number of cases, number of cases containing vocals, and overall percentage.

Listen to example cases

Two representative cases per category, with the majority decision and inter-rater agreement for each attribute. N is the number of responses received for that case (minimum 3).

c_005 Overall agreement: 90.0% (κ = 0.6825)
Reference
Sample A
Sample B
Melody
A · 100.0%
Harmony
A · 100.0%
Rhythm
A · 100.0%
Voice
Neither · 100.0%
Timbre
A · 50.0%
c_113 Overall agreement: 60.0% (κ = 0.0181)
Reference
Sample A
Sample B
Melody
A · 100.0%
Harmony
A · 50.0%
Rhythm
A · 50.0%
Voice
Neither · 50.0%
Timbre
B · 50.0%
c_033 Overall agreement: 86.7% (κ = 0.4915)
Reference
Sample A
Sample B
Melody
A · 100.0%
Harmony
A · 100.0%
Rhythm
A · 66.7%
Voice
B · 100.0%
Timbre
B · 66.7%
c_244 Overall agreement: 80.0% (κ = 0.351)
Reference
Sample A
Sample B
Melody
Neither · 75.0%
Harmony
Neither · 75.0%
Rhythm
B · 100.0%
Voice
B · 50.0%
Timbre
B · 100.0%
c_087 Overall agreement: 75.0% (κ = 0.2672)
Reference
Sample A
Sample B
Melody
Neither · 80.0%
Harmony
Neither · 60.0%
Rhythm
Neither · 60.0%
Voice
N/A
Timbre
B · 100.0%
c_256 Overall agreement: 60.0% (κ = −0.1364)
Reference
Sample A
Sample B
Melody
Neither · 60.0%
Harmony
Tie A/B
Rhythm
A · 60.0%
Voice
N/A
Timbre
Tie A/B
c_283 Overall agreement: 80.0% (κ = 0.2857)
Reference
Sample A
Sample B
Melody
Neither · 100.0%
Harmony
A · 66.7%
Rhythm
A · 100.0%
Voice
Neither · 66.7%
Timbre
A · 66.7%
c_298 Overall agreement: 93.3% (κ = 0.7321)
Reference
Sample A
Sample B
Melody
B · 100.0%
Harmony
B · 100.0%
Rhythm
A · 66.7%
Voice
A · 100.0%
Timbre
A · 100.0%

Materials

Citation

The paper is currently under review. The entry below refers to the arXiv preprint and will be updated with final venue details on acceptance.

@article{batlleroca2026matcha, title = {On the Human and Computer Alignment of Attribute-Based Music Matches}, author = {Batlle-Roca, Roser and Choi, Woosung and Serr\`a, Joan and Morreale, Fabio and Liao, Wei-Hsiang and Serra, Xavier and G\'omez, Emilia and Mitsufuji, Yuki}, year = {2026}, eprint={2609.00987}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2609.00987}, }