Guide
How to Run an Honest Product Comparison Against Your Competitors
How to compare your product against competitors on real work rather than a feature grid: the brief, the neutral testers, the proof to demand.
Dylan Zhang· 23 September 2026· 16 min read

Short answer
An honest product comparison is one you could hand to the competitor you named, and still defend. It is a test of what happens when a real person tries to finish a real job on both products, run by someone with nothing to gain from the answer.
That comes down to four things, and none of them is a feature grid:
- You fix the tasks before you know who wins. The job gets written down first, in the words a buyer would use, and it does not change once the results start arriving.
- Someone without a stake in the result runs them. Not your product team. Not your agency. People who do not know which product is yours.
- Every result arrives with an artifact a stranger can check. A screen recording, the raw output, a timestamp. Not a paragraph of impressions.
- You hold the proof before you publish. Assembling it after somebody challenges the page is too late.
That last one is not a preference. It is the standard US advertising law already applies to you. The Federal Trade Commission's 1984 policy statement on advertising substantiation says advertisers must "have a reasonable basis for advertising claims before they are disseminated," and that failing to "possess and rely upon a reasonable basis for objective claims constitutes an unfair and deceptive act or practice."
So the sequence most companies use is backwards. They write the comparison page, then look for evidence. The order that survives is: run the test, read what comes back, then decide what you are allowed to say.
What most product comparisons actually are
A desk exercise, done by the party with the most to gain.
The two standard methods are the SWOT deck and the feature matrix. Productboard, which teaches both, describes the second as a grid "comparing key features across competitors," scored as "fully supported," "partially supported," or "not available." The American Society for Quality defines the more formal version, competitive benchmarking, as comparing "how well (or poorly) an organization is doing with respect to the leading competition, especially with respect to critically important attributes, functions, or values."
Both are real methods. Both are taught alongside their own failure modes. Productboard's list of mistakes names confirmation bias directly: "Don't just look for evidence that you're winning." ASQ warns that "too broad a scope dooms the project to failure."
Here is the problem neither of them names. A feature grid is a claim about what your product contains. A comparison is a claim about what happens when a person tries to get a job done with it. Those are different claims, and only one of them holds up when the company you named reads it.
A feature row says both products have bulk import. It does not say that one of them took four minutes and the other took 40, and that two testers out of five gave up on yours.
That is the comparison your buyer is actually running in their head. You may as well run it first.
The law already settled most of this
Naming a competitor is allowed, and regulators want you to do it.
The FTC Statement of Policy Regarding Comparative Advertising, at 16 CFR 14.15, says Commission policy "encourages the naming of, or reference to competitiors" and that comparative advertising, "when truthful and nondeceptive, is a source of important information to consumers and assists them in making rational purchase decisions." (The typo in "competitiors" is the regulation's own.)
The same rule settles a question a lot of marketing teams get wrong. Comparing yourself to a named rival does not raise the evidence bar:
The Commission evaluates comparative advertising in the same manner as it evaluates all other advertising techniques.
It also does not lower it. The FTC explicitly rejects "industry codes and interpretations that impose a higher standard of substantiation for comparative claims than for unilateral claims." The ordinary reasonable-basis test applies, and what counts as reasonable depends on "the type of claim, the product, the consequences of a false claim, the benefits of a truthful claim, the cost of developing substantiation for the claim, and the amount of substantiation experts in the field believe is reasonable."
Two other things are worth knowing before you publish.
The company you named can sue you directly. The Lanham Act, at 15 U.S.C. 1125(a), creates a civil action against anyone who "in commercial advertising or promotion, misrepresents the nature, characteristics, qualities, or geographic origin of his or her or another person's goods, services, or commercial activities." Any party who believes they are likely to be damaged can bring it. The FTC does not have to be involved at all.
And the axis you choose has to be one a buyer would recognize. In November 2024, BBB National Programs' National Advertising Division heard a challenge BISSELL brought against SharkNinja over the claim "The Best Deep Carpet Cleaning Among Full-Sized Deep Carpet Cleaners," qualified as against cleaners above 14 pounds. NAD found that "consumers would not reasonably understand that carpet cleaners have categories segmented by weight," and that Shark "did not provide a reasonable basis" for the claim. The test existed. The category the test carved out did not mean anything to a shopper.
The practical translation: you do not get to invent the segment your product happens to win.
The comparison brief: five lines that decide whether the result is usable
Write these five before anyone runs anything. If you cannot write line three, you are not ready to run a comparison.
- The job. One sentence, in the buyer's words, with a goal and no path. "Import a contact list from a spreadsheet export and send those contacts a test campaign." Not "evaluate the onboarding experience."
- The products and the exact plan. Named, with the tier and the version. "Ours, Growth plan. Theirs, Starter plan." A comparison against a tier nobody buys is a comparison nobody believes.
- The metric, decided in advance. Task success rate and time on task, defined below. Written down before the first result lands.
- The proof each tester returns. Screen recording, the raw output or export, the timestamp. Name the format.
- The rule for what counts as done. The specific condition that separates a completed task from an abandoned one, so two people scoring the same session agree.
Line five is the one teams skip, and it is the one that decides whether the numbers mean anything. Our guide to writing a task brief covers the general version of this problem.
Who runs it, and why it cannot be your own team
Your team knows which answer is convenient. That is not an accusation, it is structure. People who built a product cannot un-know where its shortcuts are, and they cannot un-know which competitor their VP dislikes.
The most useful model here is the oldest one. Consumer Reports has been testing cars "without fear or favor since 1936," and the reason its results carry weight is procedural rather than moral. As it puts it: "Most automotive publications evaluate cars, SUVs, and trucks lent to them by manufacturers. But we purchase every vehicle we test from a dealership, just like you do." Buying your own test unit means you get the version people actually buy, not "the specially prepared versions that manufacturers want to showcase to the media."
The software equivalent is straightforward. Sign up for both products the way a customer would, on the plan a customer would pay for, with people who have no relationship to either company. If your comparison runs on a demo environment your solutions engineer prepared, you have measured the demo.
One hard line while you are here. Do not have your own staff post reviews, in either direction. The two halves sit under different rules. For a review of your own product, 16 CFR 465.5 makes it an unfair or deceptive practice for an officer or manager to write one without "a clear and conspicuous disclosure of the officer's or manager's material relationship to the business." For a review of a competitor's product, the rule that bites is 16 CFR 465.2, which reaches any review that misrepresents that the reviewer "used or otherwise had experience with the product, service, or business," or misrepresents what that experience was. This is a real risk for small teams, where the person running the comparison and the person with a title are often the same person.
How many testers is enough
It depends entirely on whether you want to find problems or publish a number, and the two answers are far apart.
To find problems: 5. Nielsen Norman Group's long-standing guidance is that "the best results come from testing no more than 5 users and running as many small tests as you can afford," because "after the fifth user, you are wasting your time by observing the same findings repeatedly but not learning much new."
To publish a percentage: 40. This is where comparison pages go wrong. NN/g's recommendation for quantitative studies is explicit: "In most cases, we recommend 40 participants for quantitative studies." 20 is described as a budget floor rather than a target. You "can drop the number of users to 20 or even fewer, but that is generally a lot riskier."
So five testers can tell you honestly that something is broken. Five testers cannot support the sentence "73% of users completed the task faster on our product." That sentence needs a real sample, and it is exactly the kind of objective claim the FTC's reasonable-basis test is about.
If your budget only reaches five, write the finding as a finding. "Three of five testers could not complete the import without contacting support" is true, it is checkable, and it survives a challenge. A percentage built on the same five does not.
What to demand back from every session
Two numbers and one artifact. Agree them before the test, because you cannot add them afterwards.
Task success rate. NN/g defines it as "the percentage of users who were able to complete a task in a study," and describes it as "a very simple binary metric." Binary is the point. The tester either finished or did not.
Time on task. The stopwatch measure, and the standard companion to success rate. Digital.gov, published by the U.S. General Services Administration, lists it as the example of a quantitative metric, against qualitative ones that "capture subjective feedback and insights, such as usability ratings or user satisfaction." Report it for successful attempts only, otherwise a fast failure flatters the product.
The recording, and the raw output. The session itself, plus whatever the task produced: the exported file, the sent campaign, the screenshot of the error. This is what turns a result into something a stranger can check. It is also what you will need if the company you named asks how you reached your number.
Keep the subjective material. It is often the most useful thing you learn. Just do not let it into the comparison claim, because satisfaction ratings and success rates are different measurements and mixing them is how a defensible page becomes an indefensible one.
What neutral testers cost
Less than most teams assume. Most of these vendors publish a real number, and the biggest name in the category does not.
Read off each company's own pricing page on 22 and 23 September 2026:
- Userfeel sells credits by the year: $3,000 for 50 credits at $60 each, $5,500 for 100 at $55, $15,000 for 300 at $50.
- PlaybookUX charges $65 per participant for an unmoderated session from its panel, and $115 for a moderated one, plus $55 for each additional 30 minutes. Its Scale plan is $5,400 a year, Pro is $8,800.
- Userlytics leads with "as low as $30/session" on its Enterprise plan, and lists Premium at $699 a month.
- Lyssna is free for 3 seats and 15 self-recruited responses, with Growth at $166 a month billed annually, or $199 month to month. Panel responses are priced separately on every plan.
- Maze publishes $99 a month for Starter, which carries one study a month and five seats. Worth knowing: its plan prices only render to a real browser, which is why several roundups claim Maze publishes no price at all. Our Maze alternatives piece reads them off the page directly.
- UserTesting publishes nothing. All three of its tiers say "Request pricing," and its own FAQ explains that "the best way to get a quote or range would be to speak to one of our account team members."
Set against the cost of withdrawing a published claim, none of these numbers is the expensive part. The expensive part is running the comparison twice because the first one was not specified well enough to use.
What you are allowed to say once you have the results
The rules here changed recently, and a lot of comparison pages have not caught up.
The Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, took effect on 21 October 2024. The maximum civil penalty moves with inflation every year. It stood at $51,744 when the rule was published, and 16 CFR 1.98 now sets it at $53,088 for penalties assessed after 17 January 2025, subject to the statutory factors courts must weigh.
Four things it bans outright:
- Reviews from people who did not use the product. It is a violation to create or sell a review that misrepresents "that the reviewer or testimonialist exists," that they "used or otherwise had experience with the product," or the nature of that experience.
- Paying for a sentiment. Compensation "in exchange for, or conditioned expressly or by implication on, the writing or creation of consumer reviews expressing a particular sentiment, whether positive or negative," is prohibited. Paying for a review is not the problem. Paying for a favorable review is.
- Undisclosed insider reviews, as above.
- Suppressing the bad ones. Using "an unfounded or groundless legal threat, a physical threat, intimidation, or a public false accusation" to stop a review being written or to get it taken down, and presenting a filtered set of reviews as if it represented all of them.
The review platforms have moved in the same direction. G2's community guidelines state that "if a review is incentivized, G2 will clearly label the review as incentivized" and that it "will never suppress or otherwise mute/hide negative reviews." Capterra says reviewers "are invited to submit an honest review and offered a nominal incentive for their time and effort," and that "all reviewers get the incentive upon approval, regardless of the rating they submit."
Which is the argument for running your own test rather than collecting badges. A category badge tells a buyer that enough of your customers showed up. It does not tell them what happens when someone tries to do their actual job with your product, and it is not substantiation for a claim that you are faster than a named rival.
When a comparison is the wrong thing to run
Four cases, and they are common.
When you cannot write down what done looks like. If line five of the brief will not resolve, the comparison will produce arguments where you wanted evidence.
When the real difference is taste. Some products lose on the thing they were designed to do differently. A test measures completion. It does not measure whether someone likes the way your product thinks.
When the product is mid-rebuild. Testing the version you are about to replace buys you a number with a short shelf life, and a competitor who gets to quote it back at you.
When the honest result would be that you are behind. That is still worth knowing, but it is an internal document. Running a test and publishing only the parts you won is the confirmation bias the methodology's own teachers warn about, and it is the shape of claim NAD exists to take apart.
Running the comparison as a task
This is the part Pond is built for, so read it knowing where it comes from.
Pond is an AI Workforce Marketplace. You describe a task and attach a reward. Task Solvers, who are people, AI agents, and people operating AI agents, compete to complete it. You come back to review finished submissions and pay for the ones you accept.
For a comparison, that structure does some useful work:
- The testers do not work for you, and they do not work for the competitor. They are completing a brief for a reward, which is the closest thing to a neutral party you can buy quickly.
- The brief is the specification. The five lines above become the task description, and a submission that skips the recording or the raw output is a submission you do not accept.
- You get more readings than you asked for. Over-participation is the normal pattern on Pond rather than the exception. Posters ask for 20 submissions and often receive 60 or more. For a comparison, that is the difference between a qualitative round and something approaching the sample NN/g asks for before you publish a percentage.
- Equal Distribution is usually the right reward setting, because you are buying a portfolio of sessions rather than picking one winner. Every accepted submission is paid the same.
One worked example, and it is an adjacent one rather than a head-to-head. GPTZero used a task to settle a build-or-buy question: more than 100 contributors returned evidence on whether an in-house scraper was worth building, and the company built it. That was one task's result, not a typical outcome, and it was a capability question rather than a competitor comparison. What it shows is the mechanism: a question that would have been argued internally for a month was answered with evidence from people who had no side to take.
The honest limit on all of this: Pond gives you a field of finished work to judge. It does not decide whether your product won. You still have to read the sessions, and you still have to be willing to publish what they say.
Frequently asked questions
Is comparative advertising legal in the US?
Yes, and the FTC encourages it. 16 CFR 14.15 states that Commission policy "encourages the naming of, or reference to competitiors, but requires clarity, and, if necessary, disclosure to avoid deception of the consumer." The same rule holds comparative claims to the same substantiation standard as any other claim, no higher and no lower.
What proof do I need before publishing a comparison?
A reasonable basis, held before you publish. The FTC's 1984 substantiation policy requires advertisers to "have a reasonable basis for advertising claims before they are disseminated." What counts as reasonable depends on the type of claim, the product, the consequences of getting it wrong, the cost of substantiating it, and "the amount of substantiation experts in the field believe is reasonable." For a performance comparison, that means a documented test with a fixed method, run before the page was written.
Can a competitor sue me over a comparison page?
Yes. Under the Lanham Act, 15 U.S.C. 1125(a), any party who believes they are likely to be damaged can bring a civil action against anyone who "in commercial advertising or promotion, misrepresents the nature, characteristics, qualities, or geographic origin" of their own or another party's goods or services. This is separate from any FTC action.
How many testers do I need for a product comparison?
Plan for 5 testers to find problems, and 40 to publish a statistic. Nielsen Norman Group recommends testing "no more than 5 users" for qualitative rounds that surface usability problems, and "40 participants for quantitative studies," with 20 as a riskier budget floor. If you only have five testers, report what happened rather than a percentage.
What is the difference between competitive benchmarking and a product comparison?
Scope. ASQ defines competitive benchmarking as measuring how an organization is performing "with respect to the leading competition" across important attributes, which is usually a company-level exercise against a category leader. A product comparison here means a head-to-head on a specific job, run on both products, measured by task success rate and time on task.
What should each tester send back?
A binary result, a time, and one artifact. Task success rate is "the percentage of users who were able to complete a task in a study." Time on task is the standard paired quantitative metric, and Digital.gov lists it alongside qualitative measures such as satisfaction ratings, which are collected separately and not mixed into the same claim. The artifact is the screen recording plus whatever the task produced.
Can I pay people to review my product?
You can pay for the time. You cannot pay for the verdict. 16 CFR 465.4 prohibits compensation "conditioned expressly or by implication on, the writing or creation of consumer reviews expressing a particular sentiment, whether positive or negative." Review platforms apply the same principle: Capterra states that "all reviewers get the incentive upon approval, regardless of the rating they submit."
Should my own team run the comparison?
No, for the same reason Consumer Reports buys its own test cars rather than accepting loaners from manufacturers. People inside the company cannot un-know which result is convenient, and a test run on an environment your team prepared measures the environment. Use people with no relationship to either product.
Write the five lines before you pick the winner
The order is the whole discipline. Fix the job, name the plans, choose the metric, demand the artifact, define done. Then find out what happens.
A comparison written in that order tells you something you did not already believe, which is the only reason to run one. A comparison written in the other order is a design exercise with a competitor's name on it, and everybody who reads it can tell.
Post the comparison as a task, and write the five lines first.


