Common Crawl is a foundational dataset, spanning hundreds of terabytes of text across nearly 2 billion pages, from which many derivative datasets are built and used to pre-train virtually every major language model.
Open the full topic