Proxy-Cheap
Proxies & Business
August 31, 2026
5 min

Data collection costs in 2026: what you actually pay

Alex Sadovskij
Alex Sadovskij
CEO Proxy-Cheap
Data collection costs in 2026: what you actually pay
Summary
Breaks down data collection costs into four lines (proxy bandwidth, cloud egress/storage, engineering time, and failed request waste), argues cost per successful record matters more than price per GB, matches proxy type to target difficulty for cost efficiency, and gives seven tactics to cut spend without cutting output.

A realistic data collection budget in 2026, sometimes labeled data acquisition costs, has four parts: proxy bandwidth ($1 to $8 per GB for residential, considerably less for datacenter), infrastructure costs (cloud egress and storage), engineering and maintenance time, and the bandwidth wasted on failed requests. This guide breaks down real 2026 price ranges for each line and shows how to calculate cost per successful record, the number that actually decides whether a data source is worth the investment.

  • The proxy line is usually the smallest part of a real data collection budget.
  • Residential bandwidth runs about $1 to $8 per GB in 2026, datacenter costs far less per GB, and mobile is the premium tier.
  • Cost per successful record, not price per GB, is the number that lets you compare options fairly.
  • Pay-as-you-go billing keeps spending tied to actual usage, so idle capacity does not cost you anything.

How much does data collection cost in 2026?

This article covers the cost of data collection from the public web through web scraping, APIs, and automated monitoring, whether you build your own web scraper, call web scraper APIs, or rely on a mix of both to collect data from sites you don't own. Survey research and lab-based data collection have a different cost structure and sit outside this scope.

The honest headline is that the proxy invoice isn't the whole story. Whether requests route through Proxy-Cheap or another provider, the per-GB or per-IP number on that bill is usually the smallest of four lines in a real data collection budget: proxy bandwidth, infrastructure costs (cloud egress and storage), engineering and maintenance, and bandwidth wasted on failed requests. Together, these four lines make up what we call the Four-Line Total Cost of Data Collection, and they apply to almost any data collection process, from a one-off export to a large-scale pipeline feeding decision makers new data every day.

For most teams collecting under 50TB a month in 2026, the four lines look roughly like this: proxy bandwidth runs $1 to $8 per GB for residential traffic and considerably less for datacenter, cloud egress and storage adds about $0.09 per GB out on AWS on-demand data transfer pricing plus roughly $0.023 per GB per month to store what gets collected, engineering and maintenance often runs 5 to 10 hours a week per 20 scrapers once target sites change, and failed requests quietly inflate the other three lines, because bandwidth gets billed whether a request returns usable data or not.

None of this applies to data that already lives inside a tool a business owns, such as Google Analytics or an internal CRM export. The four lines above only kick in once a business needs data collected from somewhere it doesn't control: product pricing, social media posts, listings, or reviews scattered across websites and IP addresses that belong to someone else. Businesses take on that cost because web data is rarely useful data on its own: raw data and extracted data still need cleaning before they become ready-to-use data and, eventually, actionable insights that decision makers can act on. Skip that investment, and the risk isn't just a bigger bill later; it's opportunity costs in the form of missed opportunities on pricing, demand, or competitors.

The four costs that make up your real data collection bill

Every data collection bill breaks down into the same four lines: proxy bandwidth, cloud egress and storage, engineering and maintenance, and failed-request waste. Most teams budget the first line and forget the other three until the month-end invoice says otherwise.

Proxy bandwidth

Proxy bandwidth is billed in one of two ways. Rotating residential and mobile proxies rotate across large pools of IP addresses and charge per gigabyte, with rates depending on traffic type and target difficulty. Static residential, ISP, and datacenter proxies typically charge per IP per month instead, with bandwidth usage folded into the flat rate.

To estimate this line, start with the page size. A typical scraped page runs 100 to 500KB, depending on how much JavaScript is rendered and how much of the page is actually parsed. At that range, 100,000 pages a month uses roughly 10 to 50GB of bandwidth before any retries. Rotating residential proxies are priced toward the higher end of per-GB pricing because their IP addresses have a higher trust level, while datacenter IPs are priced toward the lower end.

Multiply expected page count by average page size, then check that figure against the provider's live rate. This is the fastest way to estimate line one before touching the other three, and it holds whether the target is a single site or a list of existing data sources pulled together for one project.

Infrastructure: cloud egress and storage

Infrastructure is the line most teams forget because it shows up on a cloud bill rather than a proxy invoice. Once collected data leaves the cloud environment, egress charges apply: on AWS, that's $0.09 per GB after the first 100GB each month, so pulling 10TB out in a month costs roughly $900. Storage adds a smaller but constant charge, about $0.023 per GB per month for standard storage, per AWS S3 storage pricing.

Both numbers scale linearly with volume. A team collecting 25GB a month barely notices egress and storage; a team collecting 10TB a month is paying close to four figures before a single dollar reaches a proxy provider, and that's easy to miss when nobody outside the cloud team sees the bill. The resources needed to store and manage data at that volume are part of data management, not an afterthought bolted on once the pipeline is already running.

Engineering and maintenance

Engineering and maintenance are the biggest line items once a project scales, and they are the easiest to underestimate because they never arrive as a single invoice. Building a web scraper or integration takes developer time up front, and a loaded engineer, once benefits, tooling, and operational overhead are counted, costs roughly 1.25 to 1.4 times their base salary.

Part of that time goes into turning what the scraper returns into something usable: parsing complex data structures from raw HTML and organizing the results into structured data a database or spreadsheet can use. That step is often more time-consuming than the collection itself, and it's the part that in-house solutions most often underestimate when comparing their own software and tools against a provider's.

The higher cost is ongoing. When target sites change layouts, add checks, or update their markup, someone has to notice and fix the pipeline. Teams running data collection at any real scale report spending 5 to 10 hours a week per 20 scrapers on upkeep alone. Understanding how proxy types differ helps here, too, since matching the right type to the target up front reduces how often something breaks.

Failed-request waste

Failed-request waste is the hidden multiplier, because most proxy bandwidth is billed whether a request returns usable data or not. If a target rejects or times out a request, that bandwidth is still spent. If a success rate sits at 40 percent, collecting a given dataset can take two to three times the bandwidth originally planned, once retries are counted back in.

This is also the line most sensitive to matching proxy type to target. Data-scraping workloads that hit higher-trust targets need a proxy type suited to those targets, or the completion rate drops and the bandwidth bill quietly climbs. Track completion rate alongside spend, since a falling completion rate is often the real reason a bill grows even when volume doesn't.

Cost per successful record: the only number that matters

The cost per successful record is the total data collection spend divided by the number of valid records actually retained. It is the only figure that lets options be compared fairly. A proxy priced at $2 per GB with a 20 percent success rate can cost more per usable record than one priced at $8 per GB with a 90 percent success rate, because bandwidth is charged for every failed attempt, not just the data that's kept.

The arithmetic is simple. Spend $2,000 in a month across all four lines and keep 400,000 valid records, and the cost per successful record is $0.005. Drop the success rate without changing anything else, and that number rises even if the total invoice falls, because fewer records are kept for roughly the same spend.

This is why price per GB alone is a misleading way to shop. Two providers charging the same rate can produce very different costs per record if their completion rates differ. It is also why matching proxy type to target difficulty, covered next, is a financial decision, not just a technical one: the wrong proxy type for a target lowers completion rate, and a lower completion rate raises real costs no matter what the invoice says.

Proxy costs by type: matching spend to the target

Each proxy type has a distinct cost profile and best-fit workload. The point is matching spend to business requirements and target difficulty, the key factors that determine cost per successful record, not ranking one type above another.

Proxy typeTypical 2026 market rangeUsual billingBest-fit data collection workload
Rotating residentialapprox. $1 to $8 per GBPay-as-you-go ($/GB)Location-accurate research, higher-trust targets, account-safe collection
Datacenter (IPv4 or IPv6)approx. $0.10 to $1 per GB, or $0.50 to $3 per IP/monthPer IP monthly (or per GB)High-volume public pages, documentation, and open catalogs
Static residential (ISP)approx. $1.50 to $5 per IP/monthPer IP monthlyLong-lived sessions, consistent identity collection
Rotating mobileapprox. $2 to $15 per GBPay-as-you-go ($/GB)Mobile-first platforms, highest-trust targets
Unlimited bandwidth (SOCKS5)Per IP monthly, unmeteredPer IP monthlySustained high-throughput collection

Datacenter proxies are the cheapest per gigabyte and work well for high-volume, lower-trust targets: public catalog pages, documentation sites, and open data feeds. ISP proxies combine a static, residential-origin IP address with datacenter-grade speed, which is well-suited to long-lived sessions and account-bound collection where the identity behind the connection needs to remain consistent.

Rotating residential proxies cost more per gigabyte, but the higher price buys a higher completion rate on higher-trust targets, exactly the tradeoff the cost-per-record math above is built to evaluate. Rotating mobile proxies sit at the top of the price range and are well-suited for mobile-first platforms, where mobile carrier IP addresses outperform other origins in completion rate.

Unlimited bandwidth proxies bill per IP per month with no metering, which fits sustained, high-throughput collection where usage is predictable enough that a flat rate beats paying per gigabyte. Providers, including Proxy-Cheap, publish current rates on their product pages; check those on the day of budgeting rather than working from a market range, since pay-as-you-go pricing shifts with volume tier and target region. Implementing the wrong type across day-to-day operations is one of the most common reasons bills run high for no obvious reason.

A worked cost estimate for 100,000 pages a month

Here is the Four-Line framework applied to a common scenario: collecting 100,000 pages a month.

Bandwidth. At an average of 250KB per page, 100,000 pages use about 25GB. On a mid-range residential rate, that's a modest line, well under $150 for the month. Route the same volume through datacenter IPs for public targets, and the bandwidth line drops to a fraction of that.

Infrastructure. 25GB of egress plus storage costs a few dollars at published AWS rates, close to nothing at this volume. The line grows in direct proportion to how much data is collected and retained, not to how it is collected.

Engineering. This line depends most on what already exists. Building and maintaining the pipeline from scratch is the dominant cost at this volume, easily outweighing the other three lines combined. If the pipeline already exists, the incremental engineering cost is close to zero.

Failed-request waste. At a 60 percent completion rate, collecting 100,000 valid pages actually requires about 167,000 attempts, pushing bandwidth to roughly 42GB instead of 25GB. Raise the completion rate to 90 percent; the same job requires about 111,000 attempts and 28GB.

Put together for a team that already runs the pipeline, at a 90 percent completion rate and a residential rate near $4 per GB: about $112 in proxy bandwidth for 28GB, roughly $3 in egress and storage, close to $0 in incremental engineering, for a total near $115. Spread across 100,000 valid pages, that puts the cost per successful record at roughly $0.0012, a fraction of a cent. Market research data collection at this volume is a realistic example of where that math applies.

Seven ways to cut data collection costs without cutting output

This blog post has covered where the money goes; the entire process gets cheaper once a few concrete tactics are applied without cutting the output a project actually needs.

  • Match proxy type to target difficulty. Route public, lower-trust targets to datacenter IPs and save residential and mobile for targets that need the higher trust level. Smart routing alone can materially cut proxy bandwidth costs, since datacenter rates are a fraction of residential rates per gigabyte.
  • Trim bandwidth per request. Intercept or skip non-essential resources such as images, fonts, and third-party scripts, and parse only the fields actually needed. This alone can cut bandwidth usage by several times on JavaScript-heavy pages.
  • Raise completion rate. Fewer failed requests mean fewer retries and less wasted bandwidth. Raising the completion rate from 60 percent to 90 percent cuts total request volume by roughly a third, and the savings compound because proxy bandwidth, egress, and storage all drop together.
  • Use pay-as-you-go billing. Committed monthly plans charge for capacity whether it gets used or not. Pay-as-you-go pricing ties spend to actual usage, so a slow month costs less and a busy month scales without a plan change, the lever that matters most for a self-serve user or team whose volume moves with the season.
  • Cache aggressively. Do not re-fetch pages that have not changed since the last pass. A cache check is nearly free next to a fresh request, and for slow-moving sources, it can eliminate most of the recurring bandwidth line.
  • Sample instead of collecting everything. Statistical sampling often answers the underlying question at a fraction of the volume, particularly for market research and pricing analysis.
  • Consolidate vendors. Running residential, datacenter, mobile, and ISP proxies under a single account eliminates duplicate onboarding overhead, lowers operational costs, and simplifies routing decisions across proxy types for a growing organization managing multiple projects at once.

If data collection volume rises and falls month to month, pay-as-you-go pricing keeps the bill tied to what actually gets used. See current rates for Proxy-Cheap's rotating and static residential proxies.

What drives data collection costs up in the first place

A handful of key factors keep showing up as reasons a data collection or data extraction project costs more than planned, and each maps to a line in the framework above.

Data-source complexity directly raises the cost per record: dynamic pages that render heavy JavaScript and stronger automated checks both lower the completion rate and raise the engineering line.

Volume is the most linear factor. Bandwidth, storage, and processing all scale directly with the amount collected, which is why the Four-Line framework holds at any scale, from a weekend project to a production pipeline running price monitoring across thousands of pages a day.

Update frequency multiplies costs quickly, because re-collecting the same data on a schedule repeats the same line every cycle. Caching and sampling exist specifically to control this factor.

Technical expertise required drives the engineering line more than any other factor: a target needing custom handling for authentication, dynamic content, or unusual response formats takes longer to build and maintain. Define the target and the fields needed before writing any code, and the team can focus engineering time where it actually pays off.

Infrastructure and maintenance, the servers, storage, and ongoing upkeep behind the pipeline, round out the list. None of these factors acts alone, and a real budget usually reflects two or three compounding at once. For an organization where the collected data feeds pricing decisions or new market access, and ultimately revenue, getting these operations right deserves the same scrutiny as any other cost line.

Frequently Asked Questions

Building in-house lowers per-unit cost only after months of engineering and ongoing maintenance. For most SMB and individual workloads, a self-serve provider with pay-as-you-go pricing costs less once engineering time, the largest line in the Four-Line framework, is counted.

Usually, because the completion rate dropped, bandwidth is being paid for on failed requests and retries. Cost per successful record rises even though the target volume didn't change.

Multiply pages per month by average page size. A page runs roughly 100 to 500KB, so 100,000 pages a month is about 10 to 50GB before retries.

Datacenter proxies are cheapest per gigabyte and fit high-volume public targets well. Residential and mobile proxies cost more but increase completion rates on higher-trust targets, thereby lowering the cost per successful record.

It saves money when volume varies, since spending is tracked by actual usage rather than a fixed monthly capacity. For steady, high-volume workloads, a per-IP monthly plan can work out cheaper.

Typically 100 to 500KB, depending on whether the page renders JavaScript and how much of it gets parsed.

Cloud egress (about $0.09 per GB on AWS), storage (about $0.023 per GB per month), bandwidth spent retrying failed requests, and engineering time spent on maintenance.

No. A lower price per GB paired with a low completion rate can raise the cost per successful record. Compare providers on cost per successful record, not sticker price alone.