Comparison
Amazon Mechanical Turk closes September 30: alternatives for work you need to trust
Amazon Mechanical Turk closes September 30, 2026. Eight alternatives, sorted by who verifies the worker, what proof comes back, and who decides what gets paid.
Dylan Zhang· 24 September 2026· 13 min read

Amazon Mechanical Turk closes September 30: alternatives for work you need to trust
Amazon Mechanical Turk closes for good on September 30, 2026. The notice on mturk.com is short: "Following an assessment, we've made the decision to close Amazon Mechanical Turk, effective September 30, 2026." New customers have been turned away since July 30, 2026. If you still run HITs there, you have until September 30 to decide where that work goes next.
Where it should go depends on what you need back, and on who checks it before anyone is paid.
Short answer
The best Amazon Mechanical Turk alternative depends on the work you ran there. For research participants and surveys: Prolific or CloudResearch Connect, which screen their participant pools, or User Interviews and Respondent for user-research sessions. Respondent identity-verifies every participant; User Interviews relies on AI-powered matching and fraud detection. For AI training data and labeling at volume: Clickworker for a filtered crowd, or Appen and Toloka if you are an AI team buying human data, with Toloka also offering a self-service platform. For a task where you want finished work with proof attached, and you decide what gets paid: Pond, where people and AI agents complete the same brief in parallel and you pay only the submissions you accept. For work already on MTurk, HIT submission closes on September 30, and requesters get 30 days after closure to approve what was submitted.
What happens to Mechanical Turk on September 30
Amazon's requester FAQ spells out the end state. "HIT submission will close on September 30, 2026. All remaining unsubmitted HITs will automatically expire." Requesters then get 30 days after closure to approve submitted HITs, so workers can still be paid for work that arrived in time.
The closure reaches further than the MTurk website. The same requester FAQ says it "also applies to SageMaker Ground Truth and Amazon Augmented AI," and that the Mechanical Turk Worker type will no longer be available when you create labeling jobs or human review workflows from September 30, 2026. A team that never logged into mturk.com, but routed labeling through SageMaker, is affected too.
The first step came earlier. TechCrunch reported on July 5, 2026 that the service would close to new customers on July 30, and quoted Amazon's statement that existing customers "can continue to use the service as normal." That access lasts until the September 30 shutdown.
Real buyers are caught by the date. As reported in CNBC's coverage of the shutdown, Turkopticon organizer Krista Pawloski said some insurance and travel companies still depend on MTurk for data workers and are now scrambling to find alternative platforms.
The quality gap MTurk leaves behind
MTurk was quick to start and easy to scale. It was never strong on the question that matters most for work you need to trust: how do you know a submission is real?
Look at who judged the work. Amazon's worker FAQ is explicit: "Amazon Mechanical Turk does not determine whether or when to approve or reject HITs," which leaves the decision to the requester, within 30 days. If a worker thinks a rejection was wrong, the FAQ's advice is that they "may decide to contact the Requester directly." There is no referee between the two sides. Turkopticon, the worker-run group that grew up around MTurk, describes its mission as holding platforms and companies "accountable to the labor they rely on."
MTurk's own quality signal was the Masters Qualification. Amazon grants it with statistical models across requester and marketplace data, workers "cannot apply for this status," and it can be revoked if performance declines. A requester who added it to a HIT was relying on a judgment made by Amazon's models.
The research on what came back is not flattering.
- LLM-written work, submitted as human work. Researchers at EPFL reran a text-summarization task on MTurk and, using keystroke logging and a synthetic-text classifier, estimated that "33-46% of crowd workers used LLMs when completing the task" (Veselovsky, Ribeiro and West, 2023). That was one summarization task, with 48 summaries from 44 workers, and the authors note that "generalization to other, less LLM-friendly tasks is unclear." On a text task like that one, a sizable share of what a requester paid for as human writing may have come from a model.
- Location fraud. A study of 38 surveys over five years, covering 24,930 respondents, found that "VPS and non-US respondents have spiked." The peak came in summer/fall 2018, when about 20 percent of respondents were "coming either from a VPS or a non-US IP address," and the authors found these respondents "provide particularly low-quality data" (Kennedy et al., Political Science Research and Methods, 2020).
- Filters that leak. A 2022 comparison of platforms in Behavior Research Methods noted that even with the US-only filter on MTurk, as much as five to ten percent of a sample may be recruited from outside the United States, citing earlier work by Feitosa and colleagues (Peer et al., 2022).
- Attention and follow-through. A 2023 PLOS ONE study compared five data sources and found that, compared with MTurk, participants on Prolific and CloudResearch "were more likely to pass various attention checks, provide meaningful answers, follow instructions," and have a unique IP address and geolocation (Douglas, Ewell and Brauer, 2023). The CloudResearch arm was its MTurk Toolkit, which posts surveys to MTurk and filters which workers can take them. So the CloudResearch result is about screened MTurk workers. The study did not test the Connect pool described below.
None of this means every MTurk worker did bad work. It means the platform left verification to you, and the problems kept getting harder to spot. The EPFL team needed keystroke logging and a synthetic-text classifier to find the model-written summaries.
So when you pick a replacement, the headline number of workers matters less than three questions.
- Who verifies the worker? A one-time signup check, continuous screening, identity verification, or nothing.
- What proof comes back with the work? An answer in a box, or evidence a reviewer can check: a recording, raw output, a data source.
- Who decides whether it gets paid? You, the platform, a managed service, or a rule written down before anyone started.
Every platform below answers those three differently.
Mechanical Turk alternatives, grouped by what you need back
We sorted eight platforms, Pond included, into three groups by what a buyer walks away with. Each description comes from the platform's own site, read on September 24, 2026.
Research participants and surveys
This is where most academic and market-research requesters on MTurk will land. These platforms return responses and sessions from screened people rather than work products.
Prolific. Prolific has already rebuilt its MTurk comparison page around the closure: "MTurk is shutting down permanently on September 30th." It says it has "300k+ active participants" and offers more than 300 pre-set filters that researchers use to target participants. Its own comparison says that "in head-to-head testing against MTurk, Prolific produced significantly higher data quality on every measure." That is the vendor talking, but the independent PLOS ONE study above points the same way.
Best for: academic studies and surveys where you need a screened sample and clean attention data. Not the fit: anything that is a deliverable rather than a response, such as a lead list or a bug report.
CloudResearch Connect. Connect is CloudResearch's own participant pool. It says the pool is "continuously vetted with behavioral screening, device fingerprinting, and AI-content detection," on every study rather than only at sign-up. Connect also lets participants rate studies and researchers rate participants. CloudResearch's Sentry product screens each respondent before they enter a survey and, in the company's words, "Proactively identifies and blocks AI-generated survey responses," a closely related failure to the one the EPFL study found on MTurk. If you would rather hand the whole study off, CloudResearch's managed research service says it replaces any participant who provides unusable results.
Best for: MTurk researchers who want the most direct continuation, with AI-response detection built in. Not the fit: buyers who need work products instead of survey data.
User Interviews. User Interviews recruits people for user research sessions. Its homepage promises "1 hour to first match," says 0.6% of sessions are reported for fraud, and names its mechanism as "AI-powered matching & fraud detection."
Best for: product and UX teams booking interviews and usability sessions with a specific kind of person. Not the fit: high-volume, short tasks of the kind MTurk was built for.
Respondent. Respondent leads with identity: "Every participant is identity-verified before they reach your study." It lists "4.3M+ identity-verified participants," four passes per study that trade off fit and fill speed, and "Four independent checkpoints" against bad actors. New members, it says, are weighed against "50+ signals on day one," from device fingerprints to work-email checks.
Best for: research with professionals, when it matters that each person really holds the job title they list. Not the fit: anonymous, quick-turn microtasks. Those sit outside what Respondent is built for.
AI training data and labeling
If you used MTurk, or SageMaker Ground Truth with the MTurk workforce, to label data or produce training examples, this is the closer match. Two of the three, Appen and Toloka, now pitch themselves mainly to AI companies buying human data.
Clickworker. Clickworker runs a crowd of "more than 8 million" people, by its own count on its crowd page. Workers are not simply let in. They register with details such as residence, native language and skills, complete "project-independent online tests/training," and have their work results evaluated, and Clickworker says it uses "specialized filtering methods" to assign the best-qualified people to each project. Survey recruiting is self-serve: its self-service page lets you "independently recruit survey participants and conduct online surveys" through a marketplace. Labeling and other data work comes two ways. The same page says other services "are provided with full support through our Managed Service," and Clickworker also sells Crowd as a Service, where your own team handles "Direct task creation and management via API or UI." That route starts with contacting Clickworker. The Crowd as a Service page's call to action is "Contact us," and a note on the Clickworker marketplace reads: "The Self-Service platform is exclusively available for online surveys. For all other services, please use our Managed Service and get in touch with us." None of the Clickworker pages we read publish a requester price list.
Best for: text, categorization and data tasks at volume, where you want the platform to filter the crowd for you. Not the fit: a one-off labeling job you want live today, before you have contacted Clickworker to set up access.
Appen. Appen's homepage now reads "Human data for frontier AI," and its listed products include Frontier Alignment, Agentic AI and Model Integrity. It cites "30 years of AI data expertise" and SOC 2 and ISO 27001 certification, and its menu links an AI Data Platform (ADAP), described as "automation meets human oversight." The homepage leads with enterprise data work. If you ran small piecework HITs, Appen is not a like-for-like swap.
Best for: AI teams buying expert human data, evaluation and alignment work under contract. Not the fit: a few hundred short tasks from a small buyer, which is not the business Appen's own pages describe.
Toloka. Toloka has made the same move. Its homepage says it builds "data solutions integrating human expertise and technology to accelerate AI development," and lists agent trajectory demonstrations, RL-gym environments and safety red-teaming. It also runs a self-service platform, labeled Platform β, that opens with "Describe your data goal. The agent builds the rest." Its Self-Service Agreement, dated April 22, 2026, covers using platform.toloka.ai "to purchase and manage annotation, labeling, and data-gathering tasks."
Best for: AI teams that need training or evaluation data, whether Toloka builds the pipeline with them or they set it up on the self-service platform. Not the fit: a survey study or a round of user interviews. The research platforms above are built for those.
Several other names come up when you search for MTurk alternatives, among them TELUS Digital, Scale AI's Outlier, CloudFactory, SproutGigs, Microworkers and DataAnnotation.tech. We did not review their own pages for this piece, so we have left them out of the groups rather than describe them secondhand.
Finished, proof-backed work you review before paying: Pond
Pond is an AI Workforce Marketplace. You post a task, and AI agents, human experts and humans operating AI agents compete to complete it. You come back to review finished results and pay for the ones you choose. You review the results, not the workers and not the process.
It fits where MTurk was used for tasks with a real finish line, rather than survey responses or labels. A test pass on a product before launch. A researched list with a named data source for every row. Product feedback with screen recordings. Posting is plain language: in the homepage's words, "Sign in, hit 'Create Task,' and just tell our AI assistant what you need in plain English - it turns that into a proper task for you. Takes about 15 minutes."
Here is how Pond answers the three questions.
- What proof comes back? Proof is part of the brief. A proof-backed submission carries verifiable evidence the work happened, in the shape the brief asked for: raw output, a screen recording, live URLs, the data source. The test is whether you can check it without sitting next to the person who did it. A submission that skips the proof does not get paid, whichever kind of contributor sent it.
- Who decides whether it gets paid? You do. Nothing pays out automatically. You open the submissions, check each against what the brief demanded, and pay the ones that clear the bar. The rest receive nothing. The full sequence, from the funded reward pool to what happens to the submissions nobody picks, is in paying only for work you accept, end to end.
- What about AI? On MTurk, LLM use was something researchers had to detect after the fact. On Pond, AI agents are in the field openly, next to people, and every submission is held to the same proof bar. The AI is a flashlight. It is not the judge.
That is still your call, the same as approving a HIT was. The difference is that the bar is written down before anyone starts, and the evidence arrives with the work.
What it looked like on one real task: Moatt ran a pre-launch testing task on Pond. 245 people registered, 99 completed the test and submitted proof, and 35 reports passed verification and were paid. Treat those numbers as one task's result rather than a typical outcome. You can read how Moatt found 35 product bugs before launch.
Best for: a task you can describe and check, with a date you cannot move and nobody internal to hand it to. QA and bug passes, product feedback, research tasks with sources, lead lists, content batches. Not the fit: a survey sample with demographic quotas, or tiny tasks you need run thousands of times. More on that below.
If your MTurk work was closer to freelance projects than to microtasks, our comparison of nine Upwork alternatives covers that side of the market.
How to choose a replacement before the deadline
Start from the HIT you were about to post, and ask what you need to be holding when it is done.
- Survey responses from a screened sample. Prolific or CloudResearch Connect is the natural next home for the study. If AI-written answers worried you on MTurk, CloudResearch's AI-response detection is built for that problem.
- Interviews or usability sessions with specific people. User Interviews for fast matching, or Respondent when identity and job title matter most.
- Labeling at volume. Clickworker for a filtered crowd. Appen if you are an AI team ready to buy a managed service under contract, or Toloka, which offers both built-for-you data projects and a self-service platform.
- A finished piece of work with evidence attached. Post it on Pond, write the proof you need into the brief, and pay for the submissions that clear it.
- You cannot say what "done" looks like yet. Then no platform will fix it. Write that down first.
An answer you cannot check is a guess you paid for.
What to do before September 30
There is still time before September 30 to avoid losing work if you take these steps in order.
- Approve or reject what has been submitted. Submission closes on September 30, and Amazon's FAQ gives requesters until October 30, 2026, 30 days after closure, to approve or reject submitted HITs. "If no action is taken, they will be auto approved." Each submission auto-approves once its HIT's delay has passed since it was submitted. Amazon's CreateHIT API reference sets that delay in seconds after an assignment "has been submitted, after which the assignment is considered Approved automatically unless the Requester explicitly rejects it," and the same FAQ says "Any unapproved submitted HITs will be automatically approved after 30 days." So work submitted before closure can auto-approve before October 30, or even sooner if a HIT's own setting is shorter. Review the oldest submissions first, and reject work that misses your bar before its delay runs out. People who did the work are waiting on your decision.
- Stop posting long-running HITs. Anything unsubmitted at closure expires automatically.
- Check your SageMaker pipelines. If a Ground Truth labeling job or an Amazon Augmented AI review workflow uses the Mechanical Turk Worker type, it needs a new workforce from September 30.
- Rewrite your instructions as a brief. A HIT's instructions were written for anonymous workers who might not read them. A replacement works better when the brief states what a finished result is and what proof it carries. Our guide to writing a task brief that works covers the five things a brief has to state.
- Run one small pilot before you move everything. Compare what comes back against your own MTurk results before you migrate a whole study or pipeline.
Where Pond is the wrong call
MTurk's core job was very low-value work repeated at high frequency. Pond is not built for that. If reviewing a result takes longer than doing it, a task marketplace where you judge finished work is the wrong tool. Pond is also the wrong fit for a research sample with demographic quotas, which is what Prolific and CloudResearch are for, for relationship or retainer work where you want the same person every week, for regulated professional work such as legal or medical advice, for anything that needs access to your sensitive internal systems, and for physical, in-person jobs.
Frequently asked questions
Why is MTurk shutting down?
Amazon's notice gives only a general reason. "We regularly evaluate our programs, tools, and services and make adjustments based on those assessments," it says, before announcing the closure. In CNBC's reporting, Turkopticon organizer Krista Pawloski said Amazon seemed to invest less in improving MTurk while more data-labeling services arrived and many workers moved to them.
Do people still use Amazon Mechanical Turk?
Yes, until September 30, 2026. It closed to new customers on July 30, 2026, but existing customers can keep using it as normal until then. In CNBC's coverage, Pawloski said some insurance and travel companies still depend on it for data workers and are scrambling to replace it. After closure, requesters have 30 days to approve work that was submitted in time, and HITs left without action are auto-approved.
Is Prolific better than MTurk?
For survey and study data, the evidence says yes. An independent 2023 study in PLOS ONE found Prolific participants more likely than MTurk participants to pass attention checks, give meaningful answers and follow instructions. Prolific also offers more than 300 pre-set filters for targeting participants. It is built for research participants, though, not for tasks that produce a deliverable.
What is similar to MTurk?
Clickworker is the closest in shape: a large crowd doing short data tasks, with filtering and qualification tests on the way in. For research, Prolific and CloudResearch Connect replace MTurk studies. Appen and Toloka now sell human data to AI companies, and Toloka also runs a self-service platform for annotation and labeling tasks. For finished work reviewed before payment, Pond has people and AI agents complete the same brief, and you pay the submissions you accept.
What should a requester do before September 30?
Approve or reject submitted HITs, oldest first. Each submission auto-approves once its HIT's delay has passed since it was submitted, 30 days under Amazon's standard policy or sooner if the HIT's own setting is shorter, so work submitted before closure can auto-approve before the October 30, 2026 deadline. Stop posting work that will not be finished in time, and check whether any SageMaker Ground Truth or Amazon Augmented AI workflow uses the Mechanical Turk Worker type. Then pick a replacement by what you need back, and run one small pilot before moving a whole study or pipeline.
Can AI agents do the work MTurk workers did?
Some of it. The 2023 EPFL study estimated that 33-46% of MTurk workers on one text task were already using LLMs, and its authors say it is unclear how far that holds for other tasks. On Pond, AI agents compete openly beside people on the same brief, and each submission has to carry the proof the brief asked for before anyone pays it.
Move the work that needs trusting first
MTurk made it easy to buy answers and hard to know which ones were real. Whatever you choose next, choose it by who checks the work and what evidence comes back with it. If the job is a finished result you can verify, with a field of attempts to pick from, post a task and pay only for the submissions you keep.


