Questions to Ask Before Hiring a CRO Agency
The questions to ask a CRO agency before hiring, grouped by theme, with what good and bad answers sound like and a scorecard to compare suppliers side by side.
Hiring a conversion rate optimization agency is a bet on method more than on talent. Every agency's portfolio shows winning tests and rising charts; what separates a useful partner from an expensive one is how it decides what to test, whether its numbers survive scrutiny, and what you are left with when the engagement ends. The right questions to ask a CRO agency are the ones that expose that method before you sign, rather than six months in.
This guide is for marketing leads, e-commerce managers, founders and product owners who are comparing CRO suppliers, whether for a one-off research engagement or an ongoing testing program. It is organized as an interview script: the questions that matter most, grouped by theme, each with the reason to ask it, what a good answer sounds like and what should worry you. At the end there is a scorecard for comparing suppliers side by side and a full checklist you can take into the room.
If you have not yet settled whether an agency is the right model at all, start with our guide on getting the agency or in-house decision right. This article assumes you have decided to hire outside help and now need to choose well.
Why Choosing a CRO Agency Is Harder Than It Looks
CRO is unusual among marketing services because the output is mostly invisible until it is measured, and the measurement is easy to get wrong. A design agency's work can be judged by looking at it. A CRO agency's work is judged by statistics that the agency itself usually produces, from tracking it may have configured, on tests it chose. That makes the buyer unusually dependent on the supplier's honesty and competence, and it is why interview questions matter more here than in most hiring decisions.
Three constraints shape every honest answer you will hear. The first is traffic volume, which is the hard constraint: low traffic means tests take longer or cannot reach significance at all, and no amount of budget changes that. The second is development capacity: tests have to be built and winners have to be shipped, and without engineering time you get a document. The third is measurement: multiple platforms, offline conversions and consent requirements make attribution genuinely hard and genuinely expensive to get right. An agency that talks about all three unprompted is usually worth a second meeting.
If you only have time for a short interview, these are the questions that tell you the most:
- How do you decide whether our traffic can support testing at all?
- How do you verify our tracking before trusting any result?
- How do you set test duration, and when do you stop a test?
- Who will actually do the research, design, build and analysis?
- How do winning variants get into production?
- Who owns the test code, research and data when we part ways?
The rest of this guide expands those six into a fuller script, one theme at a time.
Fit and Process: Can They Work With Your Traffic and Your Team?
Fit questions establish whether the agency's method suits your site at all. Many poor engagements are not caused by bad work but by a testing program sold to a business that needed research, or research sold to a business that needed a paid media fix.
How do you decide whether our traffic can support testing at all?
Why ask it: Traffic volume sets the ceiling on what a testing program can do. A page with modest traffic and a low-frequency conversion may need months to detect anything but a large effect. An agency that sells a testing program without checking this is selling activity, not results.
A good answer sounds like: They ask for your traffic and conversion numbers by page and funnel step before quoting, run a sample size or minimum detectable effect calculation, and tell you plainly which pages can be tested, which can only be improved through research and best practice, and which are too small to bother with. They may recommend starting with a research engagement rather than a program.
A bad answer sounds like: Any variation on "we test everything" or "we'll run as many tests as possible." A promise of a fixed number of tests per month regardless of your traffic is a warning sign, because it means duration is being set by the contract rather than by the data.
What happens in the first weeks, before any test goes live?
Why ask it: The early weeks decide the quality of everything after. Agencies that rush to launch tests usually skip the audit and the research that would have told them what to test.
A good answer sounds like: A sequence with rough durations: a data audit of 1–2 weeks, research of 3–5 weeks, and a first test live around Week 5–7. They can describe what each phase produces, such as a tracking audit report, a ranked list of findings, and a test backlog with hypotheses tied to evidence.
A bad answer sounds like: "We can have a test live next week." It is sometimes possible, but a first test launched before anyone has checked the tracking or looked at user behavior is a guess with a statistics report attached.
How do you choose what to test first?
Why ask it: Prioritization is where research turns into money. With limited traffic, every test slot spent on a low-value idea is a slot not spent on a better one.
A good answer sounds like: A documented framework that weighs the strength of evidence, the size of the affected audience, the expected impact and the effort to build. They can show you an example backlog and explain why the top item is at the top. Our guide to experiment backlogs, done properly describes what a healthy one looks like.
A bad answer sounds like: Ideas that come from a generic list of best practices, a competitor's site, or the agency's favorite test from another client. "We always start with the button color" is a caricature, but versions of it still appear in pitches.
Craft and Quality: Will Their Numbers Hold Up?
Craft questions test whether the agency can run experiments that produce trustworthy results. This is the theme where polished pitches most often hide weak practice, so push for specifics and ask to see real artifacts.
How do you verify our tracking before trusting any result?
Why ask it: Every test result is only as good as the conversion data behind it. Duplicate purchase events, missing consent-mode configuration, broken cross-domain sessions and thresholded reports can all produce confident numbers that are wrong.
A good answer sounds like: They describe a concrete audit: reconciling analytics conversions against your own orders or CRM records, checking event definitions, checking how consent affects what is recorded, and checking sessions that cross domains such as a separate checkout or booking tool. They mention GA4 limits such as data thresholding and retention settings, which our guide on GA4 data retention and thresholding explains. For a fuller treatment of the audit itself, see auditing conversion tracking.
A bad answer sounds like: "We'll use whatever you have set up" or "the testing tool tracks conversions itself, so analytics doesn't matter." A testing tool's own counts still depend on correctly firing events, and nobody should analyze a test without knowing whether the numbers reconcile with real orders.
How do you set test duration, and when do you stop a test?
Why ask it: Stopping tests early when the numbers look good is the most common way CRO programs manufacture false winners. The duration question separates agencies that understand statistics from agencies that understand dashboards.
A good answer sounds like: Duration is calculated before launch from traffic, baseline conversion rate and the smallest effect worth detecting, and it runs in whole weeks to cover weekday and weekend behavior. Typical tests run 2–6 weeks each, set by traffic and not by preference. They stop early only under predefined rules, for example a sequential testing method chosen in advance, or when a variant is clearly broken.
A bad answer sounds like: "We stop when it hits significance" or "we call it as soon as there's a clear winner." Checking repeatedly and stopping at the first significant reading inflates the false positive rate. So does running a test for three days because the client wanted an answer before a meeting.
Red-flag answers on craft: any of these should lower an agency's score sharply.
- A guaranteed uplift or a promised conversion rate before they have seen your data.
- Test durations set by the contract or the calendar rather than by a sample size calculation.
- No mention of reconciling analytics against your own orders or leads.
- Case studies that report only winners and cannot say how many tests were run in total.
- Reluctance to share raw test data or let you check the analysis.
Can you show us a test that lost, and what you learned from it?
Why ask it: Most tests do not produce a winner. An agency that has run a real program has plenty of losing and flat tests, and how it talks about them shows whether it treats testing as learning or as sales material.
A good answer sounds like: A specific example: the hypothesis, the evidence it came from, what happened, and what the team changed in its understanding of the customer as a result. Good agencies treat a clear loss as useful, because it rules out an explanation and redirects the backlog.
A bad answer sounds like: "Our tests almost always win." Either the agency is testing trivially safe changes, stopping tests early, or not being candid. None of those is what you are paying for.
Team: Who Will Actually Do the Work?
The people in the pitch are not always the people on the account. CRO needs several distinct skills, including analytics, qualitative research, UX design, front-end development and statistics, and it is worth knowing exactly where each one sits.
Who will actually do the research, design, build and analysis?
Why ask it: A strong strategist backed by nobody who can build tests, or a strong developer with no one checking the statistics, produces lopsided work. Knowing the roles also tells you who to call when something breaks.
A good answer sounds like: Named roles, and ideally named people, for each discipline, with an indication of how much of their time your account gets. They are clear about what is done in-house and what is subcontracted, and they introduce the analyst and the developer, not just the account lead.
A bad answer sounds like: A single account manager who "coordinates the specialists" without being able to name them, or an evasive answer about offshore or freelance build resources. Subcontracting is not a problem in itself; hiding it is.
Do you build and ship tests, or do you need our developers?
Why ask it: Development capacity is one of the main drivers of cost and outcome. Tests have to be built and winners have to be shipped. Without engineering time you get a document.
A good answer sounds like: A clear division of labor: the agency builds test variants in the testing tool or in your codebase, follows your QA and deployment process, and tells you in advance how much of your developers' time will be needed to implement winners permanently. They ask about your release cycle, staging environments and any restrictions on client-side testing.
A bad answer sounds like: "We'll hand you the recommendations." If you have no development capacity, a recommendations deck may be all you ever get. Equally worrying is an agency that plans to inject heavy client-side scripts into your site without discussing performance or flicker.
Communication: What You Will See and When
Good communication in CRO is not frequent meetings; it is a predictable rhythm and honest reporting of uncertain results. Ask how the agency communicates when results are dull, because most of them will be.
What will we see each week and each month?
Why ask it: Reporting habits reveal how the agency thinks. A program that reports only test outcomes hides the research, the backlog and the reasons behind choices.
A good answer sounds like: A regular cadence with defined content: weekly status on live tests and upcoming launches, monthly review of results, learnings and the reprioritized backlog, and a periodic look at program-level impact. They explain that meaningful program results usually take 4–6 months, and they tell you what signals to expect before then. They also run structured retrospectives; see campaign retrospectives for the principles.
A bad answer sounds like: A dashboard login and a monthly slide deck of wins. Dashboards are useful, but a program without regular human explanation drifts toward whatever looks good on the dashboard.
How do you report a result that is inconclusive?
Why ask it: Many tests end without a clear answer. How an agency writes up those results shows whether you can trust the ones it calls winners.
A good answer sounds like: They report inconclusive as inconclusive, with the observed range of effects, what the test can and cannot rule out, and a recommendation: ship the simpler version, retest with a bolder change, or move on. They never round a flat result up to a win.
A bad answer sounds like: "It was trending positive," offered as if that were a result. Trends in underpowered tests are noise, and a supplier that reports them as wins will eventually lead you to ship changes that do nothing or worse.
Contracts: Terms, Exits and Accountability
CRO contracts deserve careful reading because the useful output accumulates slowly and the costs start immediately. The contract should match the timeline the agency itself described.
What is the minimum term, and how do we exit?
Why ask it: A testing program needs time: research, a first test, then several cycles of tests running 2–6 weeks each. A term that is too short guarantees disappointment; one that is too long with no exit locks you in if the relationship fails.
A good answer sounds like: A minimum term that matches the realistic timeline, a clear notice period, and a defined handover on exit. Some agencies suggest a research engagement first, which gives both sides a natural decision point before committing to a program.
A bad answer sounds like: A long lock-in with no performance review point, or automatic renewal buried in the terms. Also be wary of the opposite: a month-to-month arrangement where the agency has no reason to invest in research.
What are you accountable for, and what will you not guarantee?
Why ask it: Nobody can honestly guarantee a conversion uplift, because the result of a well-run test is unknown in advance. But an agency can and should be accountable for process: tests launched on schedule, QA standards, reporting accuracy and response times.
A good answer sounds like: A clear list of commitments on process and quality, and a plain statement that outcomes are not guaranteed. If they propose performance-based fees, they explain exactly how the baseline is measured and how external factors like seasonality and paid media changes are handled.
A bad answer sounds like: A guaranteed uplift figure. Performance fees tied to a baseline the agency defines and measures itself, with no reconciliation against your own data, are also a red flag.
Rights and Pricing: What You Pay For and What You Own
Pricing and ownership are where misunderstandings turn into disputes. Ask directly and get the answers in writing.
How is the work priced, and what is excluded?
Why ask it: CRO is priced in several ways, and the structures are not directly comparable. You need to know what each quote includes before you compare numbers.
A good answer sounds like: A clear structure and a clear statement of what drives it. For reference, reviewed US market ranges run roughly $4,000 – $15,000 for conversion research (analytics, session review, expert review and user testing, with prioritized findings) and $3,000 – $15,000 per month for an ongoing CRO program. Our own published starting rates are conversion research from $6,500 per engagement, with a turnaround of 3–5 weeks, and a testing program from $5,500 per month, sized before it runs. These are starting prices, not totals; the final figure depends on traffic volume, the number of funnels and markets, research depth, development capacity and measurement complexity. A good supplier also tells you about costs outside its fee, such as testing tool licenses and participant costs for moderated user testing with recruitment. For a deeper breakdown, see how much conversion rate optimization costs.
A bad answer sounds like: A single number with no explanation of scope, or a low monthly fee that turns out to exclude test development, analytics fixes or research. Another bad sign is a quote given before the agency has asked about your traffic.
| Service | Our starting rate | Typical US market range | Turnaround |
|---|---|---|---|
| Conversion research | From $6,500 per engagement | $4,000 – $15,000 | 3–5 weeks |
| Testing program | From $5,500 per month | $3,000 – $15,000 per month | Ongoing |
| Paid media management | From $2,400 per month | 10–20% of spend, or $1,500 – $10,000 per month | Ongoing |
If you manage our paid media too, how is that fee calculated?
Why ask it: Many CRO agencies also offer paid media management, and the fee structure affects incentives. Paid work is usually priced as a percentage of spend or as a flat retainer against it.
A good answer sounds like: A clear model and an honest word about its incentives. The typical US market range is 10–20% of spend, or $1,500 – $10,000 per month; our paid media management starts from $2,400 per month, with search, social and shopping managed against reconciled conversion data. A good agency explains how it avoids recommending more spend simply because its fee rises with it, and reports against your own orders rather than platform-reported conversions alone.
A bad answer sounds like: A percentage-of-spend fee with no discussion of the incentive, or reporting that relies entirely on the ad platforms' own attribution. Creative testing claims also deserve scrutiny; our guide to ad creative testing, done properly covers what rigorous looks like.
Who owns the test code, research, recordings and data?
Why ask it: Research findings, test code, design files and the experiment log are the durable value of a CRO engagement. If the agency owns them, you pay again to rebuild that knowledge after you leave.
A good answer sounds like: You own the deliverables, including research reports, user testing recordings (subject to participant consent terms), design files, test code and the full experiment log. Testing tool and analytics accounts sit in your name, with the agency given access. They also explain how any user-generated or customer content used in variants has been cleared for use.
A bad answer sounds like: Accounts held in the agency's name, test code that only runs inside the agency's proprietary tooling, or research you can see in meetings but not keep. Ask specifically, because default contract terms vary widely.
Red-flag answers on rights and pricing: walk away or renegotiate if you hear these.
- Your testing tool, analytics property or ad accounts will be created and owned by the agency.
- Research recordings or reports are not handed over at the end of the engagement.
- The quote excludes test development but the proposal promises a test cadence.
- A performance fee is based on numbers only the agency can see.
Rights questions extend to creative inputs too. Reviews, customer photos and social posts used in test variants need clearance just as they would in ads; our guide on rights for user-generated content in ads explains what to check.
Aftercare: What Happens After a Test Wins
A winning test running in a testing tool is not an improvement to your site; it is a temporary overlay. Aftercare questions reveal whether gains become permanent and whether knowledge survives the engagement.
How do winning variants get into production?
Why ask it: Programs often accumulate winners that are still served by the testing tool months later, adding page weight and fragility. Others never ship winners at all because nobody owns implementation.
A good answer sounds like: A defined path: a winner is documented, handed to your developers or built by the agency in your codebase, QA'd, and then the test is switched off. They track the queue of unshipped winners as a program metric. Some agencies also re-check a shipped winner's performance afterward to confirm it holds.
A bad answer sounds like: "You can leave the winner running in the tool." It is a stopgap at best. Equally concerning is an agency with no view on who implements, which usually means nobody does.
What do we keep if we part ways?
Why ask it: The best measure of an agency's value is what your team knows and owns when it leaves. A good engagement leaves behind a documented understanding of your customers, not just a list of changed buttons.
A good answer sounds like: A handover that includes the experiment log with hypotheses, results and learnings, the research repository, the prioritized backlog of untested ideas, all code and design files, and a walk-through session with your team. They are comfortable with this conversation because they expect you to judge them on it.
A bad answer sounds like: Vague reassurance, or a handover fee that was not mentioned at the start.
Scoring CRO Agencies Side by Side
Interviews blur together quickly. Score each supplier immediately after each meeting on the same themes, using the same scale, and weight the themes that matter most to your situation. The example below is illustrative: the agencies and scores are invented to show how the method works.
In this illustrative example, a business with moderate traffic and limited in-house developers weights craft and team most heavily, because statistical rigor and build capacity are its biggest risks. Each theme is scored from 1 (poor) to 5 (strong) and multiplied by a weight from 1 to 3.
| Theme (weight) | Agency A | Agency B | Agency C |
|---|---|---|---|
| Fit and process (2) | 4 | 3 | 5 |
| Craft and quality (3) | 3 | 5 | 4 |
| Team (3) | 2 | 4 | 4 |
| Communication (1) | 5 | 3 | 4 |
| Contracts (1) | 3 | 4 | 3 |
| Rights and pricing (2) | 3 | 4 | 5 |
| Aftercare (1) | 2 | 4 | 4 |
| Weighted total (max 65) | 39 | 52 | 55 |
Agency A gave the most polished pitch and scored highest on communication, but it could not name the people who would build tests and its answers on stopping rules were vague. On a weighted basis it falls well behind. Agencies B and C are close. B is the stronger on craft; C wins on fit and ownership terms because it recommended a research engagement first and offered to keep every account in the client's name. In a close result like this, the tie-breaker should be the question you care about most. Here, a reference call focused on how each agency handled inconclusive tests would be the sensible next step.
Two cautions. Weights should be set before the first interview, not adjusted afterward to favor the agency you liked. And any single red-flag answer on tracking or stopping rules should override a high total, because it undermines every number the agency will ever show you.
Before the interviews, it also helps to send each agency the same brief so that their proposals are comparable. Our guide to briefing an agency covers what to include.
The Full Interview Checklist
Take this list into each meeting and note the answers verbatim where you can. It follows the same themes as the sections above.
- How do you decide whether our traffic can support testing at all?
- What happens in the first weeks, before any test goes live?
- How do you choose what to test first?
- How do you verify our tracking before trusting any result?
- How do you set test duration, and when do you stop a test?
- Can you show us a test that lost, and what you learned from it?
- Who will actually do the research, design, build and analysis?
- Do you build and ship tests, or do you need our developers?
- What will we see each week and each month?
- How do you report a result that is inconclusive?
- What is the minimum term, and how do we exit?
- What are you accountable for, and what will you not guarantee?
- How is the work priced, and what is excluded?
- If you manage our paid media too, how is that fee calculated?
- Who owns the test code, research, recordings and data?
- How do winning variants get into production?
- What do we keep if we part ways?
After each meeting, ask for two artifacts: an anonymized example of a research report and an anonymized experiment log. What an agency produces for its existing clients is the most reliable preview of what it will produce for you. If it cannot share either in any form, that is itself an answer.
It also helps to know how a supplier describes its own process in public. Our page on how we run digital marketing and CRO work sets out the phases and pricing described above, so you can hold us to the same questions you ask anyone else.
Verdict Choose the agency whose answers about traffic limits, tracking verification, stopping rules and ownership are specific and uncomfortable in the right ways, not the one with the best case studies. A CRO partner that tells you what testing cannot do for your site, before you sign, is the one most likely to deliver what it can.
Spotted something wrong? Report an error on this page. We correct on the page and say what changed.