Pond
← Back to the blog

How-to

Data Extraction Services: What to Ask For, and What Good Looks Like

What to specify when you buy data extraction, the proof a usable delivery includes, and what it costs at published rates from Zyte, Apify and ScrapeHero.

Dylan Zhang· 16 September 2026· 14 min read

Blue Pond graphic headlined What data extraction costs, and what nobody defines. Three columns: Run it yourself, $0.06 per 1,000 responses, Zyte's published floor; Buy it managed, $199 a month, ScrapeHero's cheapest published plan; Define accuracy, 0 of 4 vendor pages say how they measure it.

Data Extraction Services: What to Ask For, and What Good Looks Like

Short answer

Buying data extraction goes wrong at the brief, not at the vendor. Almost every provider can pull data off a website or out of a stack of PDFs. Very few buyers write down what the finished file has to contain, which rows count, how fresh they have to be, and what makes one of them wrong. Without that, "99% accurate" means nothing, because nobody has said accurate at what.

So the order is: write the spec, then decide what shape to buy in. Run it yourself on metered infrastructure. Hand it to a managed provider. Or post it as a task and pay for the deliveries that clear your bar.

Disclosure: this is Pond's blog, and posting a task is one of the three routes described here. Every third-party rate and quote below comes from that vendor's own public page, linked at first mention and read on 14 September 2026. Nothing here is averaged or estimated.

The five things your brief has to state

A data extraction brief that cannot be checked cannot be paid against. Five lines fix that, and they take about ten minutes to write.

1. The source, named exactly. Not "competitor pricing." The specific sites, the specific document types, the specific database. A provider quoting on "e-commerce data" and a provider quoting on four named retail domains are quoting on two different jobs.

2. The fields, named exactly, and the format each one arrives in. A price field should hold a decimal number and a currency code. If it holds the whole pricing sentence copied off the page, the file loads and the data is useless. Date as ISO. Company name normalized how. This is the single largest source of a delivery that technically arrived and practically cannot be loaded.

3. The row count, and what counts as one row. One row per product, or one row per product variant? Two files of very different sizes can both be correct answers to a vague brief, and only one of them is the answer you wanted.

4. The freshness window. Data pulled last Tuesday is a different product from data pulled this morning. Say which, and say whether you need it again next month.

5. What "wrong" means. A missing field, a guessed value, a duplicate, a row from outside the defined set. Write the failure conditions down. This is the line that turns a delivery into something a reviewer can accept or reject in an hour instead of arguing about for a week.

"99% accurate" is not a number until someone defines the unit

Four vendor pages, read in full on 14 September 2026. One publishes an accuracy rate for its own service. Three, including that one, describe a method, all of them in general terms. The fourth offers a team size instead. Here is what each of them actually says.

Hitech BPO's stat block claims 99.5 percent or better accuracy from AI plus human validation, and a 24 to 48 hour turnaround on most projects. Its body copy says the firm uses "AI-powered automation systems with human quality assurance to achieve more than 99% accuracy in all project types including complex and high-volume projects." Asked in its own FAQ how it ensures accuracy, the answer is: "We use a combination of AI-powered tools, rule-based automation, and manual verification to achieve 99%+ accuracy across all projects."

Computyne's answer to the same question: "Accuracy is ensured through standardized extraction workflows, automated validation rules, and multi-level quality assurance checks."

ARDEM publishes no accuracy figure at all on its data extraction page. Its method is "multiple validation routines," and its quality proof is a certification: "All data retrieval and extraction processes are conducted in compliance with ISO27001 certification guidelines." GroupBWT's headline proof is team size, "100+ software engineers," and a promise of "structured, accurate data."

Three of those four describe a method. None of the four defines a measurement: no unit being scored, no sample size, no sampling method, and no independent checker. That is a scoped observation about four pages rather than a verdict on an industry, and it is enough to change how you buy.

Three questions close the gap, and any provider who can answer them is worth talking to:

  • Accurate at what unit? Per field, per row, or per record? A file that scores well per field scores far worse per row, because one bad field spoils the whole row. On a wide record those two numbers describe very different deliveries.
  • Measured on what sample? How many rows were checked, drawn how? A hundred rows chosen by the team that produced them is not a measurement.
  • Checked by whom, and are they independent? Hitech BPO names "human quality assurance" and Computyne names "multi-level quality assurance checks." Neither says whether the checker sits outside the team that produced the data, which is the part that makes the number mean anything.

Ask those three and the percentage either becomes useful or quietly disappears. Both outcomes are worth the email.

The proof a usable delivery includes

At Pond, a submission has to be proof-backed: verifiable evidence the work actually happened, in the shape the brief asked for. A data source, a screen recording, the raw output, live URLs. Not a paragraph of claims. The test is whether a reviewer can check it without sitting next to the person who did it.

For an extraction delivery, that comes down to four artifacts. Ask for all four, from any provider, in any buying shape.

The file itself, in the format you specified. CSV or JSON, carrying the field names you named rather than the ones the tool happened to emit.

The source trail. The URL or document each row came from, carried in the file as its own column. A row you cannot trace back is a row you cannot defend when someone in a meeting asks where a number came from.

The collection record. When the data was pulled and from where. This is what makes a refresh comparable to the first run instead of a new project.

The exception list. The rows that failed, and why. A provider who lands short of your target and names every gap is more useful than one who hits the number exactly and quietly guessed at the hard rows. An exception list is the one artifact that cannot be produced by guessing, which is why it tells you more about the run than the row count does.

A delivery with those four pieces can be spot-checked against its own sources. A delivery without them has to be trusted, which is the thing you were trying to avoid buying.

What it actually costs, in three shapes

Price in this category depends almost entirely on which of three things you are buying. Real published rates, read 14 September 2026.

Run it yourself on metered infrastructure

You write the extraction logic. The vendor sells you requests, rendering and the proxies to get through.

Zyte's published headline is "From $0.06 per 1,000 successful responses." Its pay-as-you-go rate for plain HTTP responses runs from $0.13 to $1.27 per 1,000 requests, priced across five site tiers it labels Simple, Easy, Moderate, Complex and finally Advanced. Pages that need a full browser to render cost far more: $1.01 to $16.08 per 1,000 requests on the same five tiers. Commit to $500 a month and those bands drop to $0.06 to $0.61 for HTTP and $0.48 to $7.68 for browser-rendered.

Apify's plan fee is prepaid usage credit rather than a subscription sitting on top of usage: $19 a month buys $19 to spend on the platform, with anything past it billed on. Its tiers are Free at $0 with $5 of credit, Starter at $19 a month, Scale at $199 and Business at $999. Compute runs at $0.2 per compute unit on Free and Starter, $0.16 on Scale, and $0.13 on Business. Residential proxies list at $8 per GB on Free and Starter, $7.50 on Scale and $7 on Business.

Note the spread inside a single vendor's own card. Inside Zyte's pay-as-you-go band, the cheapest and most expensive browser-rendered tiers differ by roughly 16 times. Anyone who quotes you a per-page cost without asking which sites you are hitting is guessing.

This shape is cheapest per page on simple HTTP targets, and the most expensive in engineering time. Once a site needs a full browser to render, the metered rate can pass what a managed provider charges per page: do that division against your own mix of sites before assuming the do-it-yourself route is cheaper. It suits a team that already has someone who will own the pipeline when a site changes its markup in six weeks.

Buy it managed

A provider builds and maintains the extraction and hands you the data.

Of the provider pages read for this piece, only ScrapeHero publishes a rate card. Its cheapest plan is a Business subscription starting at $199 a month per website, on a tier the page recommends for up to two sites, with a one-time setup fee charged on top. A one-off On Demand job starts at $550, with a stated minimum of $550 per site and the setup fee included. Enterprise Basic starts at $1,500 a month, and Enterprise Premium at $8,000. For additional pages ScrapeHero publishes two bands, one per enterprise tier: $650 to $2,200 per million pages on Enterprise Basic, and $500 to $1,800 per million pages on Enterprise Premium, both varying with site complexity, volume and frequency. ScrapeHero calls these figures a baseline for its pricing scheme rather than a quote.

The others ask you to get in touch. Hitech BPO's pricing FAQ says the model is "flexible based on factors such as project size, data complexity, source types, and turnaround time," offered per project, hourly or monthly, and closes with "Contact us for a custom quote." ARDEM and GroupBWT publish no rate at all.

That is not a scandal. Custom work gets custom quotes. It does mean you cannot compare this shape on price without running a procurement round, and it means your spec is doing double duty: it is also the only way to get two quotes that describe the same job.

Post it as a task

The third shape changes what you are buying. Instead of hiring a provider and then finding out whether the data is good, you publish the spec and a reward pool, and people and AI agents complete that same brief in parallel. You review finished deliveries and pay for the ones that clear your bar.

On Pond you set the reward pool and choose how it is split. Equal Distribution pays every selected submission the same amount, which fits a batch: a quota of rows, or several independent passes at one source list. Ranked Style concentrates the reward, first place highest, then second, then third, which fits a job with one slot at the end. The ranking is yours, judged against the proof your brief demanded, and Pond does not auto-score. In its own words: "You're the judge. Nothing gets paid automatically - you look at every submission and decide what's good enough." Pond charges 10% on top of the reward pool you set.

Pond's homepage puts task duration at a few hours to about 10 days, and says every task posted has drawn at least three times more people than expected, read 14 September 2026. The same page puts posting at about 15 minutes: you describe what you need in plain English and Pond's AI turns it into a structured brief.

You give up a named account manager and a contract with one company. You get several independent attempts at the same spec, and where they disagree is where your spec was ambiguous.

What to do when you cannot specify it yet

Sometimes the honest answer is that you do not know which fields you need until you have seen the data once.

Managed providers handle this with a pilot. Hitech BPO offers one explicitly: "We offer a free pilot project so you can evaluate our capabilities, accuracy, and service quality before committing to a full engagement." Take it when it is offered, and use it to test your spec rather than their competence. The pilot that teaches you your row definition was wrong has paid for itself.

A task does the same job from the other end. One team had spent five days trying to build a lead list and had nothing usable. After a three-minute conversation with Pond's AI, 20 people and AI agents took the task, and 1,000 verified leads were waiting to review the next morning with zero follow-up questions. That is one task's result, not a typical turnaround, and the reason it worked is worth more than the number: the brief named the fields and named what disqualified a row, so the submissions could be judged instead of discussed. You can read how that task was built.

When you should not buy extraction at all

The clearest case is data that already sits inside your own systems, where the real job is an integration. A provider will happily quote for it, and you will have paid someone to do what your database already does.

Do not post a task, in particular, when the work needs access to sensitive internal systems, when it is regulated professional work, when it needs a continuing relationship rather than a finished artifact, when it is very low-value and very high-frequency, or when it is high-stakes subjective work where a disagreement gets expensive. That last one matters most here, because the whole mechanic rests on your judgment of a submission being final. And do not post one when you cannot write down what "done" looks like. A task without a bar cannot be judged and should not be posted.

When the work is ongoing pipeline ownership rather than a defined delivery, you want a person you keep. Buy a file instead and you will be buying it again next month. We compared nine platforms by what actually lands in your inbox after you post, including the ones built for exactly that, in Upwork alternatives.

Frequently asked questions

Can AI do data extraction?

Yes, and most commercial providers already run it that way. Hitech BPO describes its own method as "AI-powered tools, rule-based automation, and manual verification." The part that still fails without a human is judgment about edge cases: a price that is a promotion, a company that changed its name, a row that technically parsed and is meaningless. That is why the useful question is not whether AI did the extraction. It is what evidence came back with the file.

What are the two types of data extraction?

Vendors usually split it by source. Structured extraction pulls from sources that already have a schema: databases, APIs, spreadsheets, ERP and CRM systems. Unstructured extraction pulls from sources that do not: web pages, PDFs, scanned documents, invoices, emails, images through OCR. Most real projects are both. Our position is that the unstructured half is where a brief earns its keep, because that is the half where two people can read the same source and record it differently.

What is the best free software for data extraction?

Free tiers exist and they are genuinely useful for testing a spec on a small sample. Apify publishes a Free plan at $0 with $5 of credit to spend, and Zyte offers $5 of free credit to try its API. Neither is a free way to run a production pipeline, because the metered costs start the moment the volume is real. Use a free tier to learn what your fields and row definition should be, then price the real job.

Which companies offer data scraping services?

Scraping and extraction are sold as one service, so the same names come back either way. On 14 September 2026, Google's first page for "data extraction services" opened with an AI Overview and then ran mostly outsourcing providers selling their own service: X-Byte, Damco, ARDEM, Botsol, Outsource BigData, Computyne and Matillion, with a G2 category page alongside. Page two added GroupBWT, Hitech BPO, ScrapeHero and Forage AI. They differ less in capability than in what they will put in writing, which is why the three accuracy questions above sort them faster than a features grid.

What should a data extraction quote include?

The source list, the field list with formats, the row definition, the volume, the frequency, the delivery format, the freshness window, and the accuracy measurement with its unit and sample method. If a quote is missing the last one, ask for it before comparing prices. Two quotes that define accuracy differently are not comparable, whatever the numbers say.

How do I check a delivery without redoing the work?

Pull a random sample of 50 rows and verify each one against the source URL carried in the file. Then check the exception list against your row count: the delivered rows plus the named exceptions should reconcile to the target. Those two checks take under an hour and catch the failures that matter, which are guessed values and quietly dropped rows.

What does it cost to post a data extraction task on Pond?

You set the reward pool, and Pond adds 10% on top when the task is posted and funded. Talking to Pond's AI to shape the task costs nothing. You then pay out only the submissions you accept, so the money that reaches people is money you chose to release.

Decide the spec, then decide the shape

The question most buyers start with is which data extraction company to use. The more useful question is what the file has to contain before you would sign it off.

Write that down and the rest sorts itself. Metered infrastructure suits a team that will own the pipeline. A managed provider suits a long-running feed with a named owner on both sides. And if you need a defined set of data by a date and nobody internal to build it, describe the task on Pond and pay for the deliveries that clear your bar.

Keep Reading