PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
Abstract
The recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Examples can be found at https://pnlong.github.io/PDMX.demo/.
Community
Demo Link: https://pnlong.github.io/PDMX.demo/
Dataset: https://zenodo.org/records/13763756
Code: https://github.com/pnlong/PDMX/
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation (2024)
- MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT (2024)
- SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints (2024)
- Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings (2024)
- Unlocking Potential in Pre-Trained Music Language Models for Versatile Multi-Track Music Arrangement (2024)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment:
@librarian-bot
recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper