Four Rules for Marketing Reporting When Your Agent Responds Differently Every Week
When Opus 5 rolled out, our social posts got worse. We hadn’t changed the prompt. We hadn’t changed the context. We hadn’t changed the approval flow.
The first question in moments like that always points inward. Was the briefing too thin? Did we pile too many rules on top of each other? Did someone change the context without documenting it?
The cause was outside. The underlying model had been swapped. The agent and the instructions were the same — the engine was different.
The incident itself isn’t dramatic. What’s interesting is how long you search inside your own construct before you consider that the ground moved. Most marketing teams don’t have the language for that. Without language, decisions become random.
Why the same question yields two answers
If you ask an agent the same question twice, you will get two answers. Not completely different answers, but different. That’s not a malfunction. It’s how these systems work.
This breaks with a core assumption marketing has used for twenty years. A metric from the web analytics report sits still unless someone acts. If it moves, something happened. We’ve practiced that equivalence so long we forget it’s an assumption.
With an agent, it doesn’t hold. Its result is not a single value but a range where the result can fall. What you see is a draw from that range. Ask again and you draw again.
The problem is not that agents are unreliable. The problem is that we read a range as a point. Then we mistake noise for performance.
The principle we already followed — and the reason we now have
One of our operating rules is simple: be positive about the technology, but permanently critical of the output. Even when an agent performs well for months. Until now that was an attitude, and attitudes are contestable. Eventually someone rightly asks: when can we trust a system?
Since August we have a mechanical answer. The IAB published "Measuring Visibility in the AI Era," the first measurement standard for AI visibility. Part of the paper is a buying guide for measurement vendors — that belongs on the agency side. The other part describes something bigger: how to run an organization whose numbers move even when no one acted.
The decisive line is: a single measurement is not a measurement. If the result is a range, a long string of good results tells you statistically nothing about the next one. Skepticism is not mere caution; it is the correct reading of the data.
The paper frames its practices for AI visibility, but four of them apply to any agent a marketing team uses. Translating those practices for operation is our task, not the IAB’s.
Rule 1: Know the range of variability before you interpret
Before you judge an agent output, you must know how much it varies when nobody changes anything. Without that value every trend call is a guess. The IAB calls trend interpretation without a documented variability exactly that: guessing.
Concretely: run the same prompt ten times on the same day with the same context. Note the spread. That spread becomes your baseline.
A movement of four points is only a change if the baseline spread is smaller than four points. If the baseline is six, you observed nothing. This measurement takes about half an hour per agent. It is a prerequisite for any subsequent number to mean anything.
- 10 runs – Run the same prompt ten times on the same day
- 4 points – Minimum gap to treat a movement as a change (when the baseline spread is smaller)
- 30 minutes – Effort per agent to determine the variability
Rule 2: First ask whether it was the model
The IAB separates changes that come from the platform from those that come from the market. That distinction is practical. If a change occurs across multiple brands or across a category at the same time, it almost always comes from the model. If it is isolated, it’s likely true brand performance.
The paper warns: programs that don’t make this distinction report platform change as brand performance. The story that reaches leadership is then wrong.
Operationally: if your agent outputs shift from one week to the next, the first question is not “what did we do wrong.” The first question is whether the model was updated. These updates aren’t always announced and they shift results immediately and noticeably. We started searching inside our own construct before we asked the right question.
This rule saves time and prevents the worse outcome: a team tinkering with a working setup because something else changed.
Rule 3: Not every number should move budget
The IAB distinguishes directional data from decision-ready data. Both are legitimate. The error is treating the first as the second without seeing the difference.
An agent that scores leads provides a useful pre-sort. It doesn’t justify shifting budget. An agent that monitors competitors gives an early warning. It doesn’t deliver strategy.
An organization needs an explicit allocation: which class of decision each agent output can support. That allocation is a leadership decision, not a technical one. It costs an hour and stops the quiet upgrading of pre-screens into decision proposals.
Rule 4: Actively set the interpretive frame
The paper’s best section addresses reporting up. A leader who has read precise numbers for twenty years will read a range as imprecision. They do so correctly if the report offers no alternative. The task is not only to provide data. The task is to change the frame in which they are read.
A sentence pattern that works:
AI systems do not return the identical answer every time, even to the same question. We therefore report ranges rather than single-point values. The range shows what normal variability looks like. Changes inside the range mean nothing; changes outside it do.
And the rule of thumb behind it: “about 22 percent, plus or minus 4 points” is more useful for a decision than a bare “22 percent,” which pretends accuracy that does not exist. The range is not an admission of weakness. It is the information that was missing.
According to the IAB, only 16 percent of brands currently track their AI visibility systematically. That number should not cause fear; it highlights a gap. The paper’s data come mostly from the U.S., and it notes that the same query from Germany yields different values. Transfer the method, not the numbers.
- Know the variability Before you interpret metrics, be clear about how much an agent varies without intervention. That spread serves as a baseline and reveals what is true change and what is noise.
- Check for model changes first If results suddenly shift, the first hypothesis should be a platform or model update. Changes that appear across a category rarely reflect brand performance; they usually come from the model.
- Separate data classes Directional signals from agents are not a basis for budget or strategy decisions. A clear allocation of which output can support which decision class is leadership work.
- Set the interpretive frame Report ranges, not point estimates, to show normal variance. This clarifies when a change matters and enables sound decisions despite fluctuation.
What this means for building an agentic organization
Here the IAB paper ends and the real work begins. Building an agentic marketing organization means building an organization where an increasing share of the decision basis fluctuates. Three consequences.
First, ranges belong in the system, not in people’s heads. A value that’s measured once and stored applies to everyone. A value that lives only in one person’s mind will be estimated differently by each other person. This is the same logic we use to encode rules for agents instead of explaining them.
Second, a model change breaks the time series. The IAB prescribes a deliberate restart of the baseline with documentation. For an organization using agents productively, a model change is not an IT incident — it is a reporting event. If you don’t record it, you will compare numbers in six months that aren’t comparable.
Third — and here I need to be transparent about our own state — we measure something, but not the right thing. In our optimization loop the agent receives isolated performance numbers from each activity. That works and it improves the work. But it is not a measure of variability. We know whether an activity worked. We don’t yet know how much movement an agent produces without any activity. That’s the difference between optimization and measurement quality, and we are building the latter now.
This distinction matters more to me than any single rule above. Measuring is not the end. The question is whether you know your system’s movement or only its outputs.
What to ask differently at the next report
The four rules don’t require new technology. Variability won’t disappear with better models; it’s an inherent property of these systems. What must change is leadership practice.
At your next agent report one question makes the difference: how large is the variability when we do nothing? If you can’t answer that, you have a report without a benchmark. If you can, you decide from substance.
Frequently Asked Questions about Marketing Reporting with Agents (FAQ)
How do I measure an agent’s variability concretely?
Run the same prompt multiple times under identical conditions and document the spread of results. This baseline shows what variability is normal and when a change truly matters.
How often should I re-establish the baseline?
Re-establish whenever the conditions change — for example after a model update or a major setup change. Regular, scheduled re-measurement prevents you from extending unusable time series.
How do I detect a model change if it isn’t announced?
Watch for concurrent patterns across multiple brands or categories. If results shift in parallel across several places, that’s a strong sign of a platform or model update.
Which decisions can be based on agent outputs?
Separate orientation signals from decision-grade evidence. Pre-screens and early warnings are useful, but budget or strategy choices need separate validation.
How do I explain ranges in reports to leadership?
Introduce the interpretive frame proactively: report result ranges and explain their meaning. This makes clear that changes inside the range are normal and only outliers are actionable.
What in your reporting is moving right now without anyone doing anything?
Interested?
Let's find out together how we can implement these approaches in your organization.
Schedule a conversation now