Back to Stories

What are the most popular open datasets?



June 23, 2025 - 1 min read

Text. Prompts. Corpora. Instruction-tuning.


The top 50 liked datasets of all time on the Hugging Face platform come from a diversity of developers (45 different ones), have been created in 2022 or later, and mainly focus on training foundations models or improving them.


For pre-training models, there are the Fineweb datasets, AllenAI's Dolma, RedPajama and Wikipedia. GSM8K and OpenR1-Math are datasets for math problems, The Stack and OpenCodeReasoning for coding, and OpenThoughts-114k & Medical-o1-reasoning for example datasets for problem-solving with reasoning models.


Click on the legend to zoom in on the top datasets for a specific modality and find their content on the Hugging Face website.


Scan the QR code to view this story on your mobile device.