Loop Engineering in Marketing: why your context has an expiration date
This spring we sat with a client on a campaign that had stopped moving. A surf and kiteschool in the Netherlands, modest budget, clear goal. The cost per lead was around €66 and stayed there. Not bad, not good, just stuck. We first did what everyone does: better images, sharper copy, new audiences. Nothing changed.
The error was in the question. We asked how to build a better ad. The right question was how to make the system learn faster. By August the cost per lead was €23. €800 in ad spend generated €5,600 in revenue, and 40% of the month’s revenue came through direct bookings.
- 66 € Lead-Preis – Ausgangssituation der Kampagne
- 23 € Lead-Preis – Ergebnis im August nach Loop-Anpassung
- 800 € Anzeigenbudget – Stand einem Umsatz von 5.600 € gegenüber
- 40 % Monatsumsatz – Anteil über Direktbuchungen
The difference has a name that has circulated since June 2026.
What is Loop Engineering?
Loop Engineering is the practice of designing not the single task you give an AI, but the cycle around it: what triggers a run, how the result is evaluated, what happens on failure, when something is good enough, and where the learning flows back.
The term comes from software development. It surfaced in June 2026 after essays by Addy Osmani, and IBM now maintains a definition page. The order is simple: Prompt Engineering shapes a single input. Context Engineering shapes the information environment the model sees. Loop Engineering shapes the cycle that repeats both until a goal is reached.
For us, Loop Engineering became its own discipline in the methodology this summer, on par with Process and Context Engineering. This is not a cosmetic rename. It shifts what gets decided in the first place.
The difference almost nobody names
Loop Engineering works so well in development because the evaluation is clear. The test runs or it doesn’t. The build compiles or it doesn’t. An agent can self-correct for hours because after each run it gets a hard, machine-readable signal whether it moved closer to the goal.
Marketing doesn’t give you that signal in the same form.
An ad with a high click rate can still attract the wrong people. Copy that harms the brand can convert in the short term. A channel that looks good can cannibalize revenue you would have acquired more cheaply elsewhere. And the metric you optimize shifts when you make it the target.
That’s why a naive transfer is dangerous. If you copy the development loop one-to-one into marketing, you build a system that efficiently optimizes in a direction nobody checked.
Concretely, this is what it looks like for us: with AI I can generate a hundred ad variants in twenty minutes. After several cycles the pattern is clear. A human still decides which three deserve real budget. Not because the model is weak, but because selection is an evaluation, not a calculation.
Example one: the campaign that improved itself
Back to the campaign. What we changed wasn’t the creative. It was the cadence.
The platform algorithm learns from signals. If you give it two new variants every three weeks, it learns slowly. If you give it several structurally different variants each week, it learns fast. The bottleneck was never the quality of a single ad. The bottleneck was how many meaningful tests we could actually launch per week, and that was limited by production time, not ideas.
The loop we built has three stations. The AI generates variants along clearly separated hypotheses — not a hundred versions of the same thought, but ten different ideas. A person selects with a stated rationale. The results flow back as notes into the context so the next round doesn’t start from zero.
The third station is the one that most often is missing. Without it you don’t have a loop; you have only a fast generator.
- Variants along separate hypotheses The AI generates not a hundred versions of the same idea, but clearly different approaches. That produces tests that actually tell you something about direction. The diversity is structural, not cosmetic.
- Human selection with a rationale A person selects the most promising ideas and documents the reasons. This evaluation is not a calculation; it’s a judgement with context. Budget goes only to what earns it.
- Feedback into the context Results return as notes so the next round doesn’t start at zero. Without this station you lose cumulative learning across cycles. That’s what separates a real loop from a rapid generator.
We learned one lesson the expensive way. When we first let agents edit campaigns directly, a run broke a conversion URL. The response was not to remove automation but to set a rule: agents read and prepare; humans write. We’ve enforced that boundary as a standard with clients since.
When an agent broke a conversion URL
When we first let agents edit campaigns directly, a run broke a conversion URL.
The response was not to remove automation but to set a rule: agents read and prepare; humans write. We’ve enforced that boundary as a standard with clients since.
Example two: the knowledge base that improves itself
The second loop is less visible and more important.
Most companies that built a knowledge base for their agents built it once. A document with positioning, tone, audiences, product knowledge. Cleaned, approved, stored. And then it sits.
The problem isn’t quality. The problem is the date.
Meta introduces new campaign types, Google shifts features, TikTok expands shopping. What was best practice four months ago is a footnote today. Knowledge about products, channels, and technology isn’t static; it has a half-life — and that half-life is shrinking. Previously, an experienced person kept this knowledge current without anyone calling it a process. They read, tested, and filed it mentally. Now a system can keep up, but only if someone builds the return path.
The second kind of obsolescence is harder to see because it has no date. An audience description doesn’t age; it was an assumption from day one. Who triggers purchase, which argument moves a conversation, where projects actually get decided — that’s not in the document; it’s in your campaign data and your calendar. If both sides don’t sit next to each other, you optimize against an image that was never tested.
And here it gets costly. A person working with a wrong audience assumption will notice it in conversation and correct it quietly. An agent won’t. It will play the assumption faster and more consistently than any team before. The technology amplifies your picture of reality. It doesn’t test it.
Who decides what a correction is worth
This brings us to the question that decides everything in marketing.
When knowledge continually flows in, a lot flows in. Not every platform announcement matters. Some items are marginal notes, some change your way of working, few change your planning. That classification is the real work, and it’s why these systems don’t function without experienced people.
A model can summarize a new feature. It can’t tell you whether it matters for your business. For that you need someone who has seen enough cycles to separate announcement from impact.
The expensive mistake here is rarely a wrong classification. It’s the absence of any instance that classifies. New knowledge lands in the archive, old knowledge sits next to it, and no one decides which one governs. The result isn’t worse output. The result is a system that answers the same question differently depending on the source, and from that point you can’t rely on either answer. A knowledge base forgives contradictions poorly. Redundancy it accepts.
What works for us is unspectacular: a weekly meeting whose sole task is to find contradictions and decide what holds. No new tool, a ritual.
How to start next week
Three steps small enough to do without a project.
Measure your test frequency, not your click rate. How many structurally different hypotheses has your most important campaign seen in the last four weeks? If the answer is under five, that’s your bottleneck, not the creative.
Hold a document up to reality. Put your audience description next to the list of the last twenty people you actually talked to. The gap is usually visible within an hour and then no longer ignorable.
Build the return path before you optimize the outbound path. When you discard an ad, write one sentence saying why and file it where your system will look next time. That one line is the difference between a loop and a carousel.
From the stuck €66 cost per lead, these loops produced €23, and from a channel that consumed budget we built one that returns seven times. It wasn’t the AI that did it. It was the decision that after every run someone looks and says what matters.
Häufige Fragen zu Loop Engineering im Marketing (FAQ)
Worin unterscheidet sich Loop Engineering von Prompt und Context Engineering?
Prompt Engineering shapes the single input. Context Engineering shapes the information environment for the model. Loop Engineering defines the repeated cycle of trigger, evaluation, and feedback until a goal is reached. It shifts the decision level from the individual artifact to the learning process.
Wie gehe ich im Marketing mit fehlenden harten Prüfsignalen um?
Instead of binary tests you need meaningful proxies and clear evaluation rules. Human judgement remains central because metrics like click rate or conversion can mislead in the short term. The important part is testing impact on the actual goal and resolving contradictions consistently.
Wie oft sollte ich neue Varianten in Kampagnen testen?
A higher cadence accelerates platform learning. Several structurally different variants each week, aligned to clear hypotheses, are more useful than rare, small changes. Crucially, each round must write insights back into the context.
Wie halte ich meine Wissensbasis aktuell, ohne Chaos zu erzeugen?
Build a clear return path for insights from campaign and customer data into the documentation. Establish a regular ritual that surfaces contradictions and decides what holds. That prevents different sources from permanently producing different answers.
Wo ziehe ich die Grenze zwischen Automatisierung und menschlicher Kontrolle?
Let agents prepare and let humans make the final decisions, especially where mistakes are costly. Automation speeds execution but does not replace judgement about relevance and risk. An explicit accountability rule prevents harm and ensures consistent quality.
The remaining question is not technical. If your system makes a hundred suggestions: who on your team has enough experience to pick the three right ones, and how much time does that person have?
Interested?
Let's find out together how we can implement these approaches in your organization.
Schedule a conversation now