Text. Prompts. Corpora. Instruction-tuning.
The top 50 liked datasets of all time on the Hugging Face platform come from a diversity of developers (45 different ones), have been created in 2022 or later, and mainly focus on training foundations models or improving them.
For pre-training models, there are the Fineweb datasets, AllenAI's Dolma, RedPajama and Wikipedia. GSM8K and OpenR1-Math are datasets for math problems, The Stack and OpenCodeReasoning for coding, and OpenThoughts-114k & Medical-o1-reasoning for example datasets for problem-solving with reasoning models.
Click on the legend to zoom in on the top datasets for a specific modality and find their content on the Hugging Face website.