Why Creative Revisions Cost Agencies So Much

In agency work, there’s typically a standard two-revision clause. But find a project that sticks to it. By round four, the invoice doesn’t look anything like the original scope. Someone needs to bring it up, or nothing will get...

Why Creative Revisions Cost Agencies So Much

In agency work, there’s typically a standard two-revision clause. But find a project that sticks to it.

By round four, the invoice doesn’t look anything like the original scope. Someone needs to bring it up, or nothing will get done.

But you’ve been there. Nobody wants to be the one who risks the client relationship. Instead of risking it, the agency decides to swallow the gap to save the long-term relationship.

This is the visible cost of revisions. But there’s another cost that can cost an agency even more than the revenue. The work that survives those four rounds of feedback-by-committee isn’t always better, and it’s often far less objectionable. 

The very thing that made the work distinctive to the agency is typically the very first thing that’s traded away – because it’s that distinctiveness that started the disagreement in the first place.

The hardest part? Neither the visible nor the invisible costs are a client problem. It’s easy to feel that the client is the problem in the process. But the reality is that the core issue is that nobody in the room has a way to end an argument about “taste”. 

Avinash Kaushik, an early web-analytics pioneer, has given a name to what fills that vacuum: HiPPO, or the “Highest Paid Person’s Opinion”. In defining the HiPPO impact, Kaushik said,

“When a HiPPO is in the room and a difficult decision needs to be made but there’s no data or data analysis to determine the right course of action, the group will often defer to the judgment of the HiPPO.”

How does this apply here? Take that quote and replace “data analysis” with “audience evidence”. You’ll see the entire cycle in a single sentence.

Breaking Down the Cost of “Just One More Revision”

Agency accounting firm Alto has run the math on this concept. A project quoted at a 30% margin, hit with what could be considered moderate scope creep (about 25% more hours than the scope originally planned), drops to 12.5%.

On a $50,000 project engagement, that could mean a potential $15,000 in profit falls to $6,250. That’s a 58% hit from just a couple of “Let’s try one more pass at this” conversations. 

Agency consultancy Counta found the same shape, but from a different angle: a 40% margin campaign falling to 25% after just a few uncontrolled review cycles.

“Three Options, Two Rounds” Doesn’t Protect the Client Like You Think It Does

Many agencies will scope around the problem rather than solve it head-on. They’ll throw in three concepts, two rounds, and overage fees beyond that. But there’s a chance that that structure is focused more on protecting an agency’s time rather than the client’s outcome.

JD Graffam, who runs two digital agencies, says, “What’s really happening is the agency is struggling in a power dynamic.” He’s also pretty blunt about what the three “distinct” concepts end up looking like in reality: “…color and font changes.” 

What does this tell us? The standard fix for revision fatigue is typically a formalized version of the same issue: nobody involved can end the argument. So, the agency pre-negotiates how many times it’s willing to have the conversation.

That’s when audience testing – when done right – can break the dam wide open. Plus, it does so in a way that both sides can accept. Tools from platforms like goahead have made this kind of testing fast enough to run well before a project ever stalls, so what used to take a five-figure invoice and a six-week wait can now be turned around in about a day.

What Does Audience Testing Change?

Audience testing doesn’t always mean the right outcome. In fact, a badly built preference test can do the same thing as a committee revision round and sand down the distinctive option until it’s the safest one that wins. 

The only difference is that there are now five hundred strangers standing in for the five who made the call around the table.

Put a bold direction next to the safe one in front of a big enough room, and the familiar one is the one that’s usually picked. That’s because anything unfamiliar often reads as worse for someone seeing it in the moment, cold.

System1 and Kanter test advertising for a living. They’ve found that 40 to 45% of tested ads register no measurable effect at all. Worse yet, less than one in five clear a meaningful bar that’s worth using in a revision conversation.

What does audience testing change, then? Testing doesn’t settle what’s true as much as it redistributes who owns being wrong.

Before a test, a decision is defended by a gut feeling or by seniority. Whoever’s wrong ends up owning it personally. After a test, the decision is defended by something outside the room. It’s not personal.

Three Kinds of Brand and Advertising Research

“Brand and advertising research” is a pretty broad title that means more than you might expect. It’s worth knowing the difference between them. Picking the wrong type can take you the long way around to the same place you started in.

Brand Health Tracking

Brand health tracing is used to measure an established brand over time, looking at things like awareness, associations, competitor analysis, etc. It’s useful, but a single reading usually only provides a trend line, not something that can point to a solution.

Brand Perception Research

Brand perception research focuses on comparing intent with how customers experience the brand. 

This type of research data is usually qualitative and can be fairly rich, but it’s often better at raising a problem rather than helping pick between two fixes.

Brand Asset Testing

This is the testing that applies to revision disagreements because it’s typically focused on options that haven’t launched. There’s no market data to check because the thing hasn’t been launched into the world. 

You put this in front of people who look like the target audience, and you count what they pick. Now, a read that used to involve a five-figure invoice and a six-week wait can be turned around in about a day. 

How Agencies Put This to Work

The agencies that are doing this well are worth studying and emulating. There are some interesting patterns that show up among them:

Run the testing before the room forms an opinion.

Trying to run research after a project has stalled out is the least useful moment to do it. At this point, positions have hardened, and the data will look like an attempt to win the argument. 


Agencies that are seeing value here run the test before the client sees anything at all. You can then walk in with a few valuable, data-backed routes. And before they can argue, there’s an audience read attached. 

That reads as “Which do we build?” rather than “Which do you like?”

When the client’s gut feeling and the agency’s don’t match.

In some cases, clients are right (even if agencies don’t like to admit it!) But in just as many cases, the creative team is. But remember that an audience read changes whose opinion is being overruled. It’s no longer a colleague but a screened sample of the client’s own customers.

It is important to recognize that sometimes the verdict will favor the client. So, when you run it, it’s best to do so when you feel you are in the right – or clients may stop believing the results.

Flip the validation data into a deliverable, not a hidden cost.

Every extra round is already being paid for. It may just be in unbillable hours instead of a line item that was technically approved.

Scoping this into the SOW up front and giving it a name and price can help you turn this into an offered service within your deliverable.

There are Limits to this Method

Nobody who is running one of these platforms will tell you that this replaces judgment. In fact, if they do, that’s usually the moment to close the deck! 

System1 has been doing this for a long time, and they’re honest about the limits: “Ad testing isn’t there to tell brands what to make… The creative leap still belongs to you.” 

There are five limits that are worth considering:

A preference isn’t a purchase. Choosing a favorite in a study doesn’t automatically connect to spending money. It’s always important to treat preferences as input – not forecast data. This can only rank what’s on the table. If every single direction is weak, then the test will simply find the least-weak option. The final judgment still must come from the strategist. How you present the data can end up deciding who the winner is. When you can place a polished mockup on the table backed by audience research data, you’re showing ideas that have backing. The question decides the outcome just as much as the work itself. Asking “Which do you prefer?” or “Which would you buy?” are different, and more than you might expect. That’s why picking a different form of the same question in audience testing can lead to different winners within the same set of options A test can reproduce committee thinking, just under a different name. This is one that most pitches skip. A textbook case is the Pontiac Aztek – over-focus-grouped, every objection folded in, nobody left owning the final call on the decision. 

Evidence is great, but evidence still needs someone who is willing to make the call. Otherwise, a test simply translates consensus in a different way. 

How Many People Does This Take?

When running informal audience testing, the most common mistake is sample size. Eight coworkers splitting five-to-three typically produce a number that’s statistically indistinguishable from a coin flip.

The number you need depends on how close the real split is. Using the standard method for preference data (a one-sample binomial test against a confidence interval, per UX expert Jeff Sauro), you gain a closer, more useful result.

100 respondents confirms a split near 60/40 500 gets closer to 54/46 It takes ≈1000 to confirm something as tight as 53/47

Now, this isn’t an argument for always running a thousand-person study. That’s typically not a feasible option. But it does mean that when two options are close enough to cause disagreement, no realistic sample size separates them.

That’s useful in itself: the choice can now be made on cost or timeline instead of ego. Nobody has to win the argument now.

The Goal? Not “Right” Evidence. Aim for Neutral.

None of this conversation is really about taste, nor is it about people being difficult. Agencies have to determine who is going to end the revision conversation without losing something along the way. 

Outside evidence doesn’t necessarily fix this by offering more accurate information. With 80% of tested advertising still failing, accuracy is clearly not the mechanism. But it does offer a fix by being neutral enough that losing an argument to the data doesn’t cause anyone to admit defeat to the other party. 

Plus, this is a far less expensive option. When you can bring that external information into the room, you add one more tool that you can use to move out of endless revision arguments and into more, better-paying projects.