Web-scale corpus collection that stays responsible
Build training and retrieval corpora without breaking the web.
$0.29
per GB at scale
38 ms
datacenter latency
10 Gbps
uplinks throughout
Training and retrieval-augmented systems need web data at a scale that will absolutely take down small sites if collected carelessly. The reputational and legal consequences of that are now material, and publishers are watching.
At the same time, single-address collection at corpus scale is simply not possible — you will be blocked before you finish a domain.
Distributed exits make the volume achievable, and per-domain concurrency caps make it responsible. Both matter, and most providers only give you the first.
Datacenter proxies carry the bulk of the load cheaply; residential handles the minority of sources that filter hosting ASNs.
How proxies solve it
Corpus-scale throughput
Datacenter bandwidth at scale lands near 29 cents per gigabyte, which is what makes web-scale collection affordable.
Per-domain rate control
Cap concurrency and requests per second per host so a long tail domain is never overwhelmed.
Escalate only where needed
Route the small share of ASN-filtering sources to residential automatically rather than paying residential rates for everything.
Provenance for every document
Exit country, timestamp and status stored per fetch, which is increasingly a dataset governance requirement.
Which proxy to use
Datacenter Proxies
Server-hosted IPs in tier-1 facilities. When the target is not scoring you by ASN, this is the fastest and cheapest way to move volume.
from $0.29 / GB
Also consider
How we would build it
- 1
Read robots.txt and mean it
Honour crawl-delay and disallow rules. It costs you very little and it is the difference between a crawler and a nuisance.
- 2
Tier your sources
Datacenter by default, residential only for domains that demonstrably reject it. The cost delta is roughly tenfold.
- 3
Cap per-domain load
Set concurrency and RPS limits per host in the dashboard, not just globally.
- 4
Record provenance per document
URL, fetch time, exit country, status and content hash. Governance reviews will ask for all five.
AI & training data questions
Adjacent workloads
Ready to start ai & training data?
Your first gigabyte is free, which is normally enough to validate the approach against your real target before you commit to anything.