Audio transcription into sheet music for popular music: the SheetSage-A2S corpus, augmentation, and MuQ features

20 September 202625 views

The authors present SheetSage-A2S — 61 hours of audio recordings with **kern annotations across 9,468 segments from 6,066 songs, focused on popular repertoire. The combination of data augmentation and the pretrained MuQ music feature extraction model reduces the symbol error rate to 4.98% on the classic Quartets collection and achieves 20.92% on the new corpus, setting a baseline for future research.

Audio transcription into sheet music for popular music: the SheetSage-A2S corpus, augmentation, and MuQ features

Audio transcription into sheet music and the place of popular music

Audio-to-score (A2S) turns an audio recording into notated music. Current A2S systems are primarily designed for classical music. Their application to popular music remains understudied. In such systems, scores are represented in symbolic encodings, including the **kern format.

Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoru Mo, and Yaolong Ju worked on this area. Preprint v1 was released on August 6, 2026, and version v3 on August 25, 2026. The work was accepted to the 34th ACM International Conference on Multimedia (MM '26).

Selection criterion: the repertoire for which a tool is claimed. Classical and popular music rely on different collections and different baselines.

The SheetSage-A2S corpus: composition and purpose

The dataset was created for A2S on popular music. Its composition:

  • 61 hours of audio;
  • 9,468 clips;
  • 6,066 unique songs;
  • scores in **kern encodings.

The authors call the corpus the first of its kind for A2S on popular music. The data, model, and code are publicly open. The result on this dataset sets a benchmark for future work.

Practical criterion: the dataset is suitable as a reference point when comparing A2S models on popular music. For the classical repertoire, other collections are used, for example Quartets.

Augmentation: what it changes in training

Augmentation expands the training set through transformations of existing examples. The authors apply it together with pretrained MuQ features. The goal is to improve the model's ability to generalize.

The abstract does not list specific types of transformations or their parameters. It is known that the authors associate the combination of augmentation and pretrained features with improved generalization. Both techniques work within a single training scheme.

Criterion: when examining someone else's pipeline, it is useful to note whether augmentation is part of training. The authors associate it with model generalization.

Pretrained MuQ features

MuQ is a pretrained feature extraction model for music audio. It has already been trained on music, so the A2S model receives a ready-made representation of sound. The authors use such features to extract meaningful characteristics of the audio.

What this gives an A2S system:

  • features are extracted by a separate pretrained model;
  • the A2S model works with a ready-made audio representation.

Practical criterion: pretrained music features are suitable where there is no point in training feature extraction inside the A2S pipeline. The authors show improved results with this approach.

The SER metric and model results

SER (symbol error rate) is the proportion of errors in score symbols. Evaluation was performed on two collections: Quartets for classical music and SheetSage-A2S for popular music.

CollectionMusicSER of the proposed modelFor comparison
Quartetsclassical4.98%15.3% SER for the previous state-of-the-art (Alfaro-Contreras et al., 2024)
SheetSage-A2Spopular20.92%the dataset is new, the result serves as a benchmark for future work

On classical music, the model reduces SER from 15.3% to 4.98% relative to the previous state-of-the-art. For popular music, the authors report 20.92% SER. The values were obtained on different datasets, so a direct comparison between rows is incorrect.

Criterion: compare SER only within a single collection and a single genre.

What this means when choosing a tool

The work fills two gaps: it provides an open corpus for popular music and demonstrates a training scheme with augmentation and MuQ features. For classical music, the result is noticeably more accurate than the previous state-of-the-art.

Practical guidelines:

  • take into account what material the model was trained and tested on;
  • check the availability of data and code — in this work they are open;
  • read SER together with the name of the collection, not separately.

Materials: preprint arXiv:2608.06165, DOI 10.48550/arXiv.2608.06165. Conference version: DOI 10.1145/3767308.3835653.

Frequently asked questions

Related materials

All materials
Audio transcription into sheet music for popular music: the SheetSage-A2S corpus, augmentation, and MuQ features