ReScraper.

ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining

One 0.6B model in place of a heuristic scraper plus dozens of hand-written filters.

Zichun Yu* Β· Jiarui Yan* Β· Shlok Sanghvi Β· Nihar Atri Β· Chenyan Xiong
Language Technologies Institute, Carnegie Mellon University
*equal contribution

The problem

Web text is cleaned by a heuristic scraper plus dozens of rules that mostly keep or drop whole pages. 31–65% of the pages each rule drops are judged worth keeping.

The model

ReScraper reads the full rendered page, extracts the main content, and then keeps, edits, deletes, or rewrites it β€” one 0.6B model, one pass, no cascade.

The result

Pretraining on its output gives the best DCLM Core score at 400M, 1.4B and 2.8B β€” +3.8–4.7% relative over the strongest baseline at each scale.

Why

Hand-written rules throw away good pages

We labelled 5,000 held-out Common Crawl pages with an LLM judge and asked what each filter rule actually removes. Rules are coarse: they act on whole pages, and the pages they drop are often the ones worth keeping. Keep-or-drop accuracy also tracks downstream accuracy, so these mistakes are paid for in pretraining.

Keep-drop accuracy against downstream Core score
Keep–drop accuracy vs. downstream accuracy of 1.4B pretraining.
Share of pages each rule drops that are judged worth keeping
Share of the pages each rule drops that the judge calls worth keeping.

Method

Extract first, then one of four operations

Every page is rendered with line identifiers and handed to the model in full. ReScraper always extracts the main content first, then commits to a single decision for the page. Supervision comes from three teachers β€” an extraction teacher, a refining teacher, and a rewriting teacher that rescues informative pages the cascade would delete.

<extract>pull the main content out of the rendered page <keep>the extraction is already clean <edit>drop noisy lines and spans <delete>the page is not worth training on <rewrite>informative but badly written β€” rewrite it
ReScraper method overview
One model replaces the scraper and the filter stack; the dashed box shows how the supervised data is built from the three teachers.

Results

Better data at every pretraining scale

Same 18.0M-page Common Crawl pool for every method; DCLM Core is the centered accuracy over 22 downstream tasks. All baselines run on resiliparse-scraped text, while ReScraper curates the rendered page directly.

Cleaning method#Unique tokensCore @ 400MCore @ 1.4BCore @ 2.8B
Raw text17.69B0.13300.22540.2804
C4-rule5.06B0.13700.21360.2409
RefinedWeb-rule7.04B0.12570.25340.3006
FineWeb-rule5.16B0.14650.23020.2790
ProX-C11.80B0.14290.23390.2969
UltraX10.78B0.13950.26150.3064
DataOrchestra (multi-agent)13.60B0.14610.25650.3068
ReScraper (0.6B)7.44B0.15350.27350.3184

Pretrained models: 400M / 1.4B / 2.8B parameters on 8.2B / 28.8B / 55.9B tokens. ReScraper also uses about a third of the GPU hours of its extraction teacher alone.

Core score per scraper followed by rule-based cleaning, versus ReScraper
Swapping in a better scraper does not close the gap: four scrapers followed by the same rule-based cleaning, plus the model-based cascade, against ReScraper.
Quality scores before and after each operation
Each operation does a different job: delete removes far worse pages, edit lifts DataMan, rewrite lifts both scores.
Share of pages per operation for each model-based pipeline
ReScraper changes pages sparingly β€” 56% are left unchanged or barely trimmed, more than the other model-based refiners.
Quality by bucket and n-gram diversity
ReScraper scores highest in every quality bucket β€” the gain is largest on the poorest pages β€” while keeping the corpus as diverse as the baselines.

Examples

The same page, five pipelines

39 held-out pages, none of them seen in training. The left panel is the rendered page ReScraper reads, with the two rule stacks under it; the right column holds the model-based refiners, ReScraper first. Every baseline starts from the same resiliparse extraction of the page. Rewrite cases were checked for faithfulness: every rewrite shown adds no entity or number that is not on the page.

The page, and the rule stacks

Model-based refiners

Use ← / β†’ to move between cases. DataMan (1–5) and FineWeb-Edu are quality scores of the text each pipeline kept; the judge verdict is an independent gpt-oss-120b reading of the page.

Release

Code, model, corpus

cxcscmu/ReScraper

Code: data construction from the three teachers, two-stage training, pool inference and the executor, plus every baseline and evaluation in the paper.

cx-cmu/ReScraper

Model: the 0.6B refiner (Qwen3-0.6B backbone) with both stage prompts.

cx-cmu/ReScraper-Data

Data: the 7.44B-token curated corpus, the raw pool output, and both SFT sets.

@article{yu2026rescraper,
  title   = {ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining},
  author  = {Yu, Zichun and Yan, Jiarui and Sanghvi, Shlok and Atri, Nihar and Xiong, Chenyan},
  journal = {arXiv preprint arXiv:2609.34287},
  year    = {2026}
}