Data, Pre-Training & Post-Training · 6 min

Building the Corpus: Crawl, Dedup, Filter — and Poison

How a pre-training corpus is really built from Common Crawl: extraction, MinHash dedup, quality classifiers, and why poisoning 0.01% of it costs about $60 a year.

A pre-training corpus is not a dataset someone sat down and assembled. It is the residue of an aggressive filtering pipeline pointed at a firehose of raw HTML. Whatever survived that pipeline became the model's worldview. So reading the pipeline is one way of reading the model's priors, and its attack surface.

The raw material: Common Crawl

Almost every open web corpus starts from Common Crawl, a nonprofit that has scraped the web since 2008, now on a roughly monthly cadence. Each snapshot is petabytes of WARC files: raw HTTP responses full of boilerplate, spam, and mojibake. Nobody trains on that directly. C4, the corpus behind Google's T5, took a single April 2019 snapshot and boiled 1.4 trillion tokens of raw text down to 156 billion. Roughly one token in nine survived. That ratio is the whole game: keep 10 to 15 percent, throw the rest away.

The stages are consistent across pipelines (C4, RefinedWeb, the Pile-adjacent efforts):

Common Crawl WARC (petabytes)
      |
  [1] Text extraction   strip HTML -> plain text (trafilatura / justext)
      |
  [2] Language ID       fastText/langdetect; keep target language
      |
  [3] Heuristic filters line-level rules, bad-word lists, boilerplate
      |
  [4] Dedup             exact + near-dup (MinHash/LSH), substring dedup
      |
  [5] Quality classifier  ML model scores "does this look like good text?"
      v
  Training corpus (~10-15% survives)

Heuristic filters: cheap rules, big cuts

C4's rules are almost comically blunt, and the bluntness is the point. Keep only lines that end in terminal punctuation. Drop any line with fewer than three words. Discard documents shorter than five sentences. Throw out any page containing a word from the "List of Dirty, Naughty, Obscene, or Otherwise Bad Words," any page with a { (a proxy for leaked JavaScript), and anything containing "lorem ipsum." Each rule is a one-liner. Together they delete most of the crawl.

That bluntness has a cost, and Dodge et al. measured it. The bad-words list removes African American English at a 42 percent rate versus 6.2 percent for White-aligned English, because a grep cannot tell a reclaimed word from a slur, so it deletes both. The finished corpus ends up just 0.07 percent African American English. This is not a footnote. It is a demographic decision nobody voted on, made by a text file of banned words.

Deduplication, and why it measurably matters

The web is stuffed with duplicates: mirrored docs, SEO spam farms, the same news wire on 400 sites, licenses copied verbatim a million times. Two reasons to kill them. They waste compute and skew the distribution, so a passage seen 1,000 times gets over-weighted. And they drive memorization. Lee et al. (2021) showed that de-duplicating the training data makes models regurgitate memorized text up to ten times less often, with equal or better downstream quality. Dedup is one of the rare moves that improves privacy and benchmarks at the same time.

Exact dedup is a hash set. The interesting case is near-duplicates, because "same document, one word changed" also has to die. The standard tool is MinHash plus LSH.

Here is the mechanism in one breath. Shingle each document into overlapping n-grams. Hash every shingle and keep the minimum hash under a given permutation; that single value is one signature entry. Repeat with many hash functions to build a signature vector. The probability that two documents share a given MinHash value equals their Jaccard similarity, and that identity is the entire trick. LSH banding does the rest: chop the signature into bands, hash each band, and any two documents that collide in any band become a candidate pair you check directly.

doc A shingles {the cat, cat sat, sat on} --hash--> min = 0x1a3
doc B shingles {the cat, cat sat, sat by} --hash--> min = 0x1a3   collision
                            \
             P(same MinHash) = Jaccard(A,B) = |A∩B| / |A∪B|
LSH: signature -> [band1][band2]...[bandK]; share ANY band -> candidate pair

RefinedWeb, the corpus behind Falcon, runs MinHash fuzzy dedup to shrink the set, then exact-substring dedup with a suffix array to cut any repeated span of 50 tokens or more. Across all its stages, roughly nine documents in ten are discarded as non-English, low quality, or duplicate, leaving about 5 trillion tokens. The headline result: web data filtered this hard beat curated corpora like The Pile. "Garbage in, garbage out" is not a slogan here. It is a measured gap on eval curves.

Quality classifiers

The final stage is a model grading text. GPT-3 trained a simple classifier to separate Common Crawl from "known good" text (WebText, Wikipedia, books) and kept pages by score, using a noisy threshold so it did not only retain the most textbook-perfect pages. It works, and it launders the curators' taste into the data. "Quality" comes to mean "reads like Wikipedia and Reddit-linked pages." Registers, dialects, and domains that do not match get quietly down-weighted.

Builder: when you filter your own fine-tune corpus, log per-stage survival counts and read 100 random dropped docs by hand. The bug is almost always a filter nuking something you wanted, not a filter being too soft.

The security angle: poisoning is practical

Most people assume that poisoning a web-scale corpus means controlling a meaningful slice of the web. It does not. You need to control the specific URLs the crawler will fetch, or the snapshot at the instant it is taken. Carlini et al. (2023), "Poisoning Web-Scale Training Datasets Is Practical," demonstrated two attacks that are cheap and deployable today.

Split-view poisoning. Datasets like LAION-400M do not ship images. They ship URLs plus a hash, and every downloader re-fetches later. The content at a URL can change after the maintainer indexed it. The authors found that 0.71 percent of LAION-400M images sat on domains that had expired and were buyable for under $1,000. Register the domain, serve whatever you like, and everyone who downloads the dataset afterward pulls your payload. They estimated at least 0.01 percent of ten popular datasets could be controlled for about $60 a year, and 0.01 percent already clears the bar many backdoor attacks need.

maintainer indexes URL  --->  hash recorded, dataset published
        (later)                        |
attacker buys expired domain           v
        \--> serves malicious content when YOU download
             hash "verifies" only if the maintainer's hash is re-checked (often it isn't)

Frontrunning poisoning. For snapshotted crowd-sourced sources like Wikipedia, your edit does not need to survive, only to be present the moment the dump is taken. Dump timing is predictable to within about 30 minutes, and roughly 35 percent of malicious revisions on English Wikipedia last longer than that before a revert. Put together, an attacker could reliably land content in about 6.5 percent of English Wikipedia articles in a given dump, with non-English editions running up to 25.3 percent (median 8.2). The live site reverts the edit minutes later. The snapshot already froze it, and the model trains on the frozen copy.

Defender: the fixes are unglamorous and they work. Pin content by cryptographic hash and re-verify it at download time, instead of recording the hash once and trusting it. Randomize snapshot timing so dumps are not predictable. Freeze crowd-sourced revisions for a trust window before they become eligible for inclusion. These are the low-overhead defenses the authors sent to the ten maintainers they notified.

Researcher: the open question is not whether injection works, but the dose-response curve. How many poisoned documents install a specific backdoor at a given model scale, and does that threshold climb or fall as models grow? That runs straight into the backdoor and unlearning material later in this module. Poisoning is how the backdoor gets in; the pipeline above is why it is so hard to keep out.

Sources

  • Carlini et al., "Poisoning Web-Scale Training Datasets Is Practical," 2023 — arXiv:2302.10149
  • Raffel et al., "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" (C4/T5), 2019 — arXiv:1910.10683
  • Dodge et al., "Documenting the English Colossal Clean Crawled Corpus," 2021 — arXiv:2104.08758
  • Penedo et al., "The RefinedWeb Dataset for Falcon LLM," 2023 — arXiv:2306.01116
  • Lee et al., "Deduplicating Training Data Makes Language Models Better," 2021 — arXiv:2107.06499
  • Brown et al., "Language Models are Few-Shot Learners" (GPT-3 quality filter), 2020 — arXiv:2005.14165
  • Common Crawl — commoncrawl.org — https://commoncrawl.org/