Downloading shadow libraries over residential VPNs
OpenAI Policy Submission — regulatory filing, OpenAI Comments to the US Copyright Office
“Developing AI systems involves normal web scraping technologies that browse public websites, similar to traditional search engine indexing.”
OpenAI Data Engineering Team — scraping operations team, Internal Slack #scraping-ops
“Anna's Archive and LibGen are blocking our AWS IP ranges. Routing through residential proxy pools and rotating user agents to finish the multi-terabyte download.”
The Split:OpenAI represented its collection methods to the Copyright Office as routine search-engine web indexing. Internally, engineers treated it as an adversarial scraping operation, rotating residential IP proxies to bypass shadow library rate limiters.
Search engine web crawlers respect robots.txt; operations that use residential proxy rotations know their ingestion is illicit.