Public conversations about AI focus on models: their size, their benchmarks, their surprising abilities. Far less attention goes to what feeds them. Every capable language or vision model rests on a vast, carefully processed body of data, and gathering that data is an engineering discipline in its own right. Behind the scenes sit fleets of crawlers, queues, storage systems, filtering pipelines and network infrastructure that most people never see, along with a growing set of legal and ethical rules about what may be collected at all.
Where training data comes from
Modern training corpora draw on several kinds of source:
- Public web crawls. Large crawls of publicly accessible pages, including open datasets such as Common Crawl, form the backbone of many text datasets.
- Licensed content. AI developers increasingly sign agreements with publishers, archives and platforms for access to high-quality material.
- Open and public-domain collections. Government documents, open-access research, public-domain books and openly licensed code.
- Curated and human-created data. Instruction examples, annotations and feedback produced by people specifically for training.
- Synthetic data. Material generated by existing models to fill gaps or teach specific skills.
The crawling layer
Crawling at the scale AI requires is a distributed systems problem. A crawler fleet manages billions of URLs in prioritised queues, respects per-site politeness limits so no server is overwhelmed, resolves enormous numbers of DNS lookups, renders JavaScript-heavy pages in headless browsers and retries failures intelligently. Well-run crawlers identify themselves with a clear user agent and check each site’s robots.txt before fetching, which increasingly includes instructions aimed specifically at AI crawlers.
Processing: where most of the work happens
Raw crawled pages are mostly unusable. Turning them into training data involves a long pipeline:
- Extraction. Stripping navigation, ads and boilerplate to keep the meaningful content.
- Deduplication. Removing exact and near-duplicate pages, which are extremely common and distort training.
- Quality filtering. Scoring text for coherence and usefulness and discarding spam and machine-generated noise.
- Safety filtering. Removing harmful and illegal content.
- Privacy scrubbing. Detecting and removing personal information such as phone numbers and addresses.
- Language identification and balancing. Making sure the dataset covers the languages and domains the model needs.
The storage and compute behind these steps, often petabytes of data in object stores processed by large batch jobs, is a major part of the cost of building a model.
The network layer nobody talks about
Collection depends on the network as much as on code. Websites respond differently to traffic from different places. Many serve regional content, languages and versions based on the visitor’s country, which matters for building multilingual datasets and for evaluating how a model handles local information. Many sites also rate-limit or block traffic from data centre ranges, and an IP range’s reputation can determine whether requests succeed at all.
For legitimate, permitted collection of region-specific public data, such as building evaluation sets that test a model’s knowledge of local services, verifying how localised content appears or gathering openly licensed material from sites that serve content by country, teams often use residential proxies. These route requests through IP addresses that internet service providers have assigned to real households, so a site in Brazil or Thailand serves the same content a local reader sees. Providers of cheap residential proxies such as Proxy-Cheap offer country-level targeting across a wide range of markets with pay-as-you-go bandwidth, which suits targeted collection better than fixed, large-scale capacity.
There is an important boundary here. Proxies are a tool for reaching content as a local visitor would, not for evading a site’s decision to opt out of AI crawling. If a publisher’s robots.txt or terms exclude AI training, routing around that through different IP addresses defeats the purpose of the rule and creates serious legal and reputational risk.
The rules are tightening
The legal framework around training data is evolving quickly. In the EU, text and data mining is permitted for many purposes unless rights holders have reserved their rights in a machine-readable way, and the AI Act requires providers of general-purpose AI models to have a policy for complying with EU copyright law, including those opt-outs, and to publish a summary of the content used for training. Courts in several countries are considering how copyright applies to training. Privacy regulators expect personal data in training sets to have a lawful basis and to be minimised. Responsible teams treat these requirements as design constraints, not afterthoughts.
Principles for responsible collection
- Honour opt-outs. Respect robots.txt, AI-specific directives and machine-readable rights reservations.
- Identify your crawler. Use an honest user agent and provide contact information.
- Be polite. Limit request rates so collection never degrades a site for its users.
- Minimise personal data. Filter it out early and document how.
- Keep provenance records. Know where every part of the dataset came from and under what terms.
Evaluation data deserves the same care
Training data gets the headlines, but evaluation data shapes how models are judged and improved. Benchmarks that test local knowledge, regional languages or region-specific services need content that genuinely reflects each place, collected recently and documented carefully. Evaluation sets must also be kept separate from training data, so models are tested on material they have not already seen; leakage between the two is a common reason why impressive benchmark scores fail to translate into real-world performance. Teams that version both datasets, and record exactly when and how each was collected, can trust their results far more.
Why this matters
The quality, legality and fairness of an AI model are largely determined before training begins, in the unglamorous infrastructure that collects and cleans its data. Teams that invest in that infrastructure, and in the rules that govern it, build models that are more capable, more trustworthy and more defensible. As regulation and public scrutiny increase, the hidden infrastructure behind training data is becoming one of the most important parts of the AI stack.



