Corpora are counted in terabytes, agents in sessions, and allowlists in exactly one address. Three loads, three ways to buy.
Proxies for AI workloads — training data, agents, endpoints
Collecting corpus-scale data has a different cost shape from ordinary scraping: the volume makes per-GB billing the wrong instrument, and a growing number of sites now decline known AI crawlers and hosting ranges at the edge. ZapIP's bandwidth plans take bytes out of the bill, residential exits put the request on a real household line, and agents or inference services can pin a static IP for a vendor to allowlist.
- Bandwidth billing, unmetered traffic
- No concurrency cap
- KYC-consented IP sourcing
- Fixed egress for agents
How it works
Corpus collection at scale
A dataset in the tens of terabytes turns per-GB billing into an open-ended bill, while bandwidth turns it into a fixed line: 100 Mbps from $32 a day, whether you saturate it or not. At the same time more publishers are declining known AI crawlers in robots.txt and at the edge, so whether the exit looks like hosting decides what comes back.
- Bandwidth plans from $32/day for 100 Mbps, bytes unmetered, concurrency uncapped
- Residential exits across 195 regions, for genuinely multilingual, multi-market pages
- KYC-consented sourcing with an auditable chain, so the data path can be documented
- Honouring robots.txt, site terms and copyright is part of the collection design, not an afterthought
How it works
Where an agent's traffic comes out
An agent that browses the live web hits the same walls a crawler does — hosting exits, headless signals, high request rates from one address. The difference is that an agent usually has to hold a session, a login or a cart across a whole task, so what it needs is stickiness rather than rotation.
- Hold one exit for a whole task, sticky for up to 90 minutes
- Give each task or tenant its own exit so rate limits do not bleed across them
- Agents that need a stable identity take a long-term static IP that does not change
- Full HTTP(S) and SOCKS5, with ready configs for Playwright and Puppeteer
How it works
Endpoint access and multi-region verification
Two concrete needs. Your inference service wants one outbound address a vendor can allowlist, and your own AI product — once it ships to several markets — has to be measured from each of them for latency, availability and what it actually returns. Both want an address that holds still.
- A dedicated long-term IP as the fixed outbound address a vendor allowlists
- Measure your own AI service's availability and latency from 195 regions
- Add an IEPL line where long-lived cross-border connections need jitter and loss kept down
- Access to third-party model services must stay inside their terms and regional policies
Coverage
One network behind all four products — changing the billing model does not change how you connect or how you target.
All 33 markets- 195 countries & regions
- Available
- All 50 US states
- Available
- City-level targeting
- Precise
- ISP & ASN targeting
- Supported
4 billing models
How this is billed
All four billing models share one network. Short bursts go per GB, pipelines that run continuously go per Mbps, and anything tied to an account identity goes per IP-day or per IP-month.
FAQ
Questions about this solution
Which plan fits LLM training-data collection?
Bandwidth-billed rotating residential. Corpus work is high-volume, long-running and steady in throughput, and per-GB pricing stops making sense somewhere in the tens of terabytes, whereas bandwidth is a flat line — 100 Mbps from $32 a day, bytes unmetered, concurrency uncapped. For a pilot run or for validating crawl logic first, the $0.8/GB traffic plan is the more flexible starting point.
Publishers are blocking AI crawlers — does a residential exit help?
It helps with one layer of three. Sites currently push back by declining known AI crawler agents in robots.txt, by restricting hosting ASNs at the edge, and by rate-based behavioural checks. A residential exit addresses the middle one: the request leaves from a genuine broadband line and is not on a hosting list. The first layer is the publisher stating a preference and should be honoured; the third is a matter of pacing your own requests.
Does an AI agent need a fixed egress IP?
It depends what the agent does. Anything that holds a login, a cart or a multi-step form across a task needs the address to stay put for the duration, and a sticky session of up to 90 minutes covers that. If the agent reaches internal systems or vendor APIs behind an allowlist, it needs an address that is genuinely permanent — that is the long-term dedicated static IP. Purely read-only browsing at volume is cheaper on rotating exits.
Is the IP sourcing compliant, and does that make the data usable?
Two separate questions. On the IP side: residential addresses come from channels where the household gave explicit, informed consent, the sourcing chain is auditable, and the acceptable use policy separately prohibits scraping government sites or sensitive personal data. On the data side: a compliant proxy does not make content licensed. Whether what you collect can lawfully be used for training depends on site terms, copyright and local law, and that determination is yours to make.
Get one request working. Scale after that.
Test before you commit, with an engineer on your integration.