PortfolioSource

AI / Machine Learning · 2024

Music Genre Classification

Music genre classification on GTZAN with the track-level split the published numbers usually skip. The leakage-free score is lower than the literature's and it is the one reported here.

Python · PyTorch · Scikit-learn · Computer Vision · Jupyter · Mel-Spectrograms

Confusion matrix

Classification confusion matrix

Blues
Classical
Country
Disco
HipHop
Jazz
Metal
Pop
Reggae
Rock
Blues
82
2
4
1
0
5
0
2
3
1
Classical
1
95
0
0
0
2
1
1
0
0
Country
3
0
78
2
1
1
0
5
4
6
Disco
1
0
2
80
5
0
2
4
3
3
HipHop
0
0
1
4
85
1
3
2
2
2
Jazz
4
3
1
0
0
88
0
1
2
1
Metal
0
1
0
2
2
0
90
1
1
3
Pop
1
1
4
5
2
1
1
78
3
4
Reggae
2
0
3
3
2
2
1
3
80
4
Rock
2
0
5
4
3
1
4
4
3
74

Rows: actual genre, columns: predicted genre

0-1920-3940-5960-7980-100percent of row

Slicing before the split moves GTZAN accuracy by 5 to 10 percent

The usual pipeline cuts each track into segments and then draws the train and test sets, so segments of one song sit on both sides of the split. Reordering the two steps changes the reported number by 5 to 10 percent. Published GTZAN scores above 90 percent generally come from the leaky ordering.

60/20/20 across tracks first, then 10 segments of 3 seconds

StageChoice
Split60/20/20 across tracks, before any audio is cut
Sliceeach 30-second track becomes 10 segments of 3 seconds
Features128-bin log Mel-spectrogram per segment
Normalisationscaler fitted on the training segments only

The scaler is the second leak, and it survives a correct track-level split.

3 architectures trained on the same leak-free split

ModelDesign
Efficient_VGGVGG-inspired baseline with reduced parameters
ResSE_AudioCNNresidual blocks with squeeze-and-excitation attention
UNet_Audio_Classifiera U-Net encoder repurposed for classification

Each is scored on segments from tracks held out before slicing.

82 to 83 percent test accuracy, cross-validation mean near 90 percent

The U-Net encoder was the best of the three: 82 to 83 percent on the GTZAN test tracks, and a cross-validation mean around 90 percent over the same track-level folds.

The same model transfers to 2 further datasets without fine-tuning

It classifies Indian Classical Music and Tabla Taala recordings with no fine-tuning on either. I built this with Camilla Sed in 2024.