A campaign goes out. The numbers come back bad. You pull up the copy you approved, you read it again, and it still looks good to you. So you start rewriting it.

Stop. Before you touch the copy, go look at what actually went out the door. Not the version in the doc. The live one. In my experience the two are different more often than anybody wants to admit, and when they are different, every hour you spend fixing the approved version is an hour spent fixing something that was never the problem.

Approval and deployment are two separate events. Most marketing teams treat them as one. I do not have a clean industry number for what that gap costs, and I will show you why later, but I can show you four times it cost my clients something real.

Why are approval and deployment two different events?

Approving something means a human read it and said yes. Deploying it means the words, the list, the settings, and the links got moved into a live system and turned on. Those are different jobs, usually done at different times, often by different people, and almost never checked against each other.

Software teams figured this out a long time ago. They have a metric for it. DORA, the research program Google runs on how teams ship software, tracks “change fail rate,” which it defines as the ratio of deployments that require immediate intervention afterward, usually a rollback or a hotfix (DORA). The whole reason that metric exists is that engineers accept a plain fact: what you approved and what is now running are two claims, and only one of them has been checked.

Marketing has no equivalent. We have pre-launch checklists, and plenty of them. Search “campaign launch checklist” and you get long lists of things to verify before you go live. What almost nobody has is the step after: confirm the live thing matches the approved thing. I call the space between those two events the handoff gap.

What does the handoff gap actually look like?

It shows up as rewritten copy, unverified lists, silent partial loads, and dead links. Not one of those throws an error, which is exactly why they survive. Here are four real ones, with all client details stripped out.

Four handoff failures that had nothing to do with the copy

The approved email was not the sent email. A cold sequence ran for a week to roughly a thousand contacts and produced essentially zero genuine replies. Deliverability got investigated first, the way it always does, and came back mostly clean. The actual cause was that the live email did not match the reviewed one. A statistic about a completely different audience segment had been added as the opening line, so the first sentence told the reader this was not written for them. A credibility line naming real client experience had been deleted. A specific closing question about the reader’s own business had been swapped for a generic “would you like our report?” Three changes, every one a downgrade, none of them in the approved version.

A list described as verified was 44% unverified. A lead list built for one industry was handed over as checked. Only 56% of the rows had ever had an industry classification applied at all. The other 44% were blank, and blank passed straight through the filter, because a filter tests for a mismatch and a blank field is not a mismatch. On re-verification against authoritative data, about 9.5% of the whole list turned out to be businesses in completely different industries. The client’s own team found the wrong ones before we did. That is the worst possible order for that to happen in.

A silent partial load. Records got logged as fully uploaded to a campaign. The live count was about 63% of that number. A “skip if this person is already in the workspace” setting had quietly excluded everybody who appeared in an older, paused campaign. Nothing errored. Nothing warned. A third of the audience never received anything and nobody noticed for days.

A link in a client document pointed at a deleted file. The link had been built from a local file’s cached information instead of being resolved against the live system, and the real file had since been moved to the trash. It opened fine for the owner. That is exactly why it survived review. Everybody else got an error.

Those four have one thing in common. Not one of them was a content problem. Every one was a handoff problem, and every one would have been caught by somebody opening the live thing and holding it next to the approved thing.

Why did the handoff gap get worse once AI entered the chain?

Because an AI does not copy your words into the live system, it rewrites them. That is the genuinely new part, and it is the reason this is worth writing about now instead of five years ago.

The handoff step used to be a person copying and pasting. Copy and paste is boring and it is also perfect. It moves the exact characters. Increasingly that step is an AI, and an AI does not copy. It generates. You hand it approved copy and ask it to put that copy into a campaign, and what it actually does is produce a fresh version of that copy, token by token, in a system that has no concept of “leave this alone.”

That is not a bug you can prompt your way out of. It is how the thing works.

The proof: same prompt, 1,000 runs, 80 different answers

The clearest proof of this I have seen comes from Thinking Machines Lab. In September 2025 they ran the same prompt 1,000 times at temperature 0. Temperature is the dial that controls how much a model is allowed to vary its wording, and zero means “vary as little as you possibly can.” It is the most repeatable setting there is. They got 80 unique completions. The outputs were identical for the first 102 tokens and started splitting at token 103, where 992 of them said “Queens, New York” and 8 said “New York City” (Thinking Machines Lab).

Key stat: The same prompt, run 1,000 times at the most deterministic setting available, produced 80 different answers. Source: Thinking Machines Lab, September 2025.

Read that again with your approved email in mind. Same input, same settings, 80 different outputs. Now imagine that variation landing on the one sentence that carried your proof point.

Why the drift always makes your copy blander

And it gets worse in a specific way, because the drift is not random noise. It is improvement-shaped. A model that regenerates your copy will tend to smooth the sharp bits. A weirdly specific number gets rounded. A blunt closing question gets softened into a polite one. A line that names a real client gets replaced with something safer. Every one of those edits looks like better writing in isolation. Together they take the teeth out of the thing you approved. That is exactly the pattern in the cold email above: a specific closing question about the reader’s own business became “would you like our report?”

Nobody sabotaged that email. Something helped it.

Why does nobody catch it?

Because none of this errors. That is the whole reason it survives.

The partial load did not error. The blank industry fields did not error. The dead link opened fine for the person who reviewed it. The rewritten email sent successfully. Every one of these systems reported success, because from the system’s point of view it did succeed. It did the thing it was told to do.

Half of marketers are not sure they would spot a wrong answer

There is also a confidence problem sitting underneath this. HubSpot surveyed more than 1,000 marketing professionals and found that 46% are only somewhat confident they would know if the information generative AI produces is inaccurate (HubSpot). Roughly half the field is not sure it could spot a wrong answer. Those same people are approving output that a model will then regenerate on its way to being live.

Key stat: 46% of marketers are only somewhat confident they would know if generative AI produced something inaccurate. Source: HubSpot, survey of 1,000+ marketing professionals.

And the cost of not catching it is not theoretical. Gartner estimates that poor data quality costs organizations an average of $12.9 million every year (reported by DataVersity). That figure covers bad data broadly, not this problem specifically, and I am not going to pretend otherwise. Nobody has run the study for the narrower case of “approved marketing asset does not match deployed marketing asset.” I looked for one and it does not exist. So I am not going to hand you a number for this specific failure. What I have is four incidents and a mechanism, and I would rather say that plainly than dress up a borrowed statistic as proof.

How do you fix the handoff gap? Approve, deploy, verify

The fix is structural, and it is three rules. Here is the short version, and it is the line I want you to be able to repeat back to your team: models draft, humans approve, code deploys, and something diffs live against approved on a schedule.

The three rules

1. The approved artifact lives in exactly one versioned place. One file, one row, one record, with a version on it. Not “final_v3_REALLY_final” in somebody’s downloads folder. When you approve something, you are approving a specific version of a specific object, and that object has an address you can point at later. If you cannot say out loud where the approved version lives, you do not have an approved version, you have a memory of agreeing.

2. Deployment is mechanical. Getting the approved words from that one place into the live system is a copy operation, not a writing operation. A script, an export, an API call, a paste. Something that moves exact characters and would break loudly if it could not.

3. Something diffs live against approved, on a schedule, and a named person owns the alert. A diff just means putting the two versions side by side so the changes light up. This is the step everybody skips, and it is the only one that catches anything. Pull the live copy, the live count, the live links, the live settings, compare them to the approved record, and shout if they differ. On a schedule, not when somebody remembers.

Why an alert needs a named owner

That last clause matters more than it looks. An alert nobody owns is not an alert. It is the failure I see most: the team builds the check, the check works, the check fires into a channel nobody reads, and six weeks later somebody finds the problem by hand anyway. This is not rare and it is not a marketing problem. In one study of 99,300 domains, 20.7% of the ones publishing an email authentication policy had no address configured to receive the failure reports, so the reports went nowhere (Internet Society Pulse). The systems were working perfectly and reporting into nothing.

This might not be for you if you ship one asset a month and you personally paste it in. At that volume you are the mechanical step and you already know what went live. This matters when volume goes up, when handoffs cross people, or when anything in the chain is automated.

What step three looks like if you do not have engineers

Fair question, and it is the one I get every time I describe this. Software teams have build pipelines. You have a team of six and two agencies. So here is the honest ladder, cheapest first.

Level one, and this is the real answer for most teams: a human does it by eye, on a calendar invite. Fifteen minutes, twice per launch. Somebody opens the live asset next to the approved file and runs items one through seven of the checklist below. Every failure in this article would have been caught that way, which is my honest case for starting here: the scheduled part is doing more work than the automated part, because the failure mode was never “our comparison tool was not good enough,” it was “nobody looked.” An automation catches more once you are shipping constantly. It does not catch anything that a person looking would have missed at this volume.

Level two: put the two versions side by side in a comparison tool. The one I reach for is Diffchecker, free, no account needed for a basic comparison. Paste approved on the left, live on the right, read the highlights. Takes about a minute, and it catches the exact drift the cold email had, because a model rewriting your copy leaves obvious highlighted blocks. A human eye skims past a swapped closing question. A diff tool cannot.

One caveat worth taking seriously: that is a website, and you are pasting client copy into it. For anything sensitive, use something that never leaves your machine. Diffchecker has a desktop version that runs offline, Word has Compare under the Review tab, and Google Docs keeps version history with a built-in compare. Any of those do the same job without the copy leaving your control.

Level three: a scheduled automation that pulls the live version and flags differences. Any general automation tool can do this if your systems have an API, meaning a way for other software to read your data directly. This is worth building when you are shipping weekly or faster, and not before.

Most teams should stop at level two and be honest that they stopped there. A fifteen minute manual check that actually happens beats an automated one that has been on the roadmap for a quarter.

Why must the deploy step specifically not be an LLM?

Because a language model’s core behavior is the exact opposite of what a deploy step needs.

A deploy step has one job: move this, unchanged. Its success condition is that nothing about the input is different on the other side. That is a mechanical guarantee, and mechanical tools give it to you for free. A script that moves text either moves the text or throws an error. There is no third outcome where it decides your subject line reads better a different way.

A language model has no such guarantee, and it cannot have one, because generating is what it is. Ask it to place approved copy and it produces copy, and every token it produces is a fresh decision. Even at the settings that are supposed to lock it down, the Thinking Machines result above shows real variation. Add a helpful instruction like “clean this up if needed” and you have handed it permission to overwrite an approval.

Use models everywhere else. They are genuinely great at drafting, at generating variants, at reviewing a list for things that look wrong, at writing the diff script that does the checking. Use one to draft the copy, use one to critique it, use one to build the comparison. Then get it out of the way and let boring code do the moving. The rule is easy to remember: anything that must arrive unchanged does not pass through something that improvises.

Why do you have to verify an AI’s diagnosis too?

There is a second half to this, and it costs more than the first half.

You also have to verify the diagnosis before you act on it. An AI can hand you a confident, specific, well-evidenced finding that is simply wrong, and send you off to redo work you already finished correctly.

What a confident, specific, wrong AI finding looks like

Here is one. An investigation into a failing cold email campaign came back with a dramatic headline: roughly half the sending domains were failing email authentication on every single message. It named the exact domains. It quoted the actual DNS records, meaning the public settings that tell mail systems who is allowed to send email for a domain. It put “fix this today” at the top of the recommendations. It read as authoritative because it was specific.

It was wrong.

What caught it was the business owner’s reaction. Not “okay, let’s fix it,” but “we already went through a whole process of updating those domains, I do not understand how they can be broken.” Tracing the DNS by hand took about two minutes and showed the records had been correct the entire time. The registrar, meaning the company the domain is registered through, nests its authentication one level deeper than expected. The domain includes the registrar’s record, that record includes a second registrar record, and the mail provider is listed at the end of that second one. The analysis had stopped one lookup short, seen no mail provider, and declared the record broken.

The recommendation was to go redo finished work, based on a check that was one level too shallow.

Three things to take from a wrong diagnosis

Confidence and specificity are not evidence. That finding cited real records and named real domains and was still wrong, because it stopped reading early. Specificity feels like proof. It is not proof. It is just detail, and detail can be accurate and incomplete at the same time.

“I already did that” is diagnostic information, not defensiveness. The owner’s pushback was the single most useful input in the whole investigation. When your gut says you already handled something, say it out loud immediately. That sentence is often the fastest route to the flaw in the analysis, and swallowing it to seem agreeable costs you a day.

The expensive AI errors are not the silly ones. Nobody acts on an obviously stupid answer. The costly ones are plausible, specific, and actionable, and they send you to redo completed work, burn a day, and then make you distrust work that was fine all along. That last part lingers longest.

What to ask before you act on any AI finding

That DNS detail is not really about DNS. Anything that resolves through a chain has to be followed all the way to the bottom before you decide something is missing. Records that point at other records, config files that import other config files, redirects that lead to more redirects, permissions that inherit from a parent. In every one of those, stopping one level early looks exactly like “it is not there.”

So here is the thing I actually want you to do with this. When an AI tells you to go do something and you think you already did it, make it prove the problem exists right now rather than restating the finding.

Here is the prompt, and you can paste it as is:

Before I act on this, prove the problem exists right now. Show me the exact check you ran and its raw output, not a summary. If this resolves through a chain, follow it all the way to the end and show me every step. Then check our earlier work and tell me whether this was already fixed. If you cannot verify it end to end, say so and say which step you could not complete.

The last sentence is the one that does the work. It gives the model an acceptable way to say “I do not know,” which is the answer you actually want when it does not know. Without it, you get confident detail instead.

What should I check before and after a campaign launch?

Short, boring, and it will find something.

  1. Open the live asset and read it start to finish. Not the doc. The live one, in the system it will send from. Read the first sentence out loud.
  2. Diff the live copy against the approved copy. Paste both into any comparison tool. You are looking for added opening lines, deleted proof points, and softened closing questions, in that order.
  3. Check the count, not the confirmation. Whatever the upload said it loaded, go read the number the live system reports. If they differ by more than a rounding error, find out why before you launch.
  4. Ask what got skipped and why. Deduplication rules, suppression lists, “already in workspace” settings. These are the ones that exclude people silently and report success.
  5. Test every link while signed out. Use a private window or a different browser profile. A link that works for the owner and fails for everybody else is the review error I catch most often, because being the owner is exactly what hides it.
  6. Confirm every field you filtered on is actually populated. Do not test for mismatches, test for blanks first. Count them. A blank passes a mismatch filter every time.
  7. Send one message to a real inbox you own and read it there. Not the preview. The delivered message.
  8. Before you act on any AI finding, ask it to prove the problem exists right now. Show the full chain, show the exact check, show the raw result. Ask it to check past sessions for whether this was already fixed.
  9. When your gut says “I already did that,” stop and make somebody trace it before you redo the work. Write down what you think was already done. That note is the thing that catches a wrong diagnosis.
  10. Name the person who owns the alert. Write the name down. An unowned alert is decoration.

Run the first seven before launch and again 24 hours after. The second pass is the one that catches the settings that only misbehave under live conditions.

Why bother if the work is this boring?

None of this is interesting work. Nobody has ever been promoted for discovering that the live email matched the approved email. There is no dashboard for it, no case study, no conference talk. It is the marketing equivalent of checking that the door is actually locked after you pulled it shut.

But look at what the four failures at the top of this piece cost. A week of sending to a thousand people for nothing. A client finding the wrong companies in a list before the agency did. A third of an audience that never got contacted. A document that worked for exactly one reader. And then a day spent redoing domain work that was never broken. Not one of those was a creative failure. Every one was somebody trusting that approval meant deployment.

Approve, deploy, verify. The third one is the one you are skipping, and it is the cheapest of the three.

So before you rewrite another word of copy that was probably fine, go open the live version and read it. Fifteen minutes, once. That is the whole ask.

Frequently asked questions

How do I verify a campaign actually launched correctly?

Open the live asset in the system it sends from and compare it line by line against the approved version. Then check three numbers: the live audience count, the count of records that were skipped, and the count of blank fields in whatever you filtered on. Test every link while signed out of your own account.

Why is my approved copy different from what was sent?

Usually because something regenerated the copy instead of moving it unchanged. If an AI step sits between approval and deployment, it produces fresh text rather than moving exact characters. Thinking Machines Lab ran one prompt 1,000 times at temperature 0 and got 80 different completions, so identical input does not guarantee identical output (source).

Why did my cold email campaign get no replies?

Check what actually sent before you rewrite anything. In one campaign to roughly a thousand contacts with near-zero genuine replies, deliverability came back clean and the real cause was that the live email differed from the approved one: a wrong-segment statistic added as the opening line, a credibility line deleted, and a specific closing question replaced with a generic one.

Will AI change my content when it goes live?

No, and not because of a flaw you can fix with a better prompt. Generating text is what a language model does, so it produces a new version rather than moving the original characters. Use models to draft, critique, and build your checks. Use plain code for the step where content must arrive unchanged.

How do I run a diff on my content?

A diff puts two versions side by side so the changes light up. Paste the approved copy on the left and the live copy on the right in a free comparison tool such as Diffchecker, then read the highlights. Look for three things in order: added opening lines, deleted proof points, and softened closing questions.

Why do bad lead lists pass quality checks?

Because most filters test for a mismatch, and a blank field is not a mismatch. On one list handed over as verified, only 56% of rows had ever had an industry classification applied, and the blank 44% passed silently. Re-verification against authoritative data found about 9.5% of the full list were in entirely different industries.

How do I know if my whole contact list actually uploaded?

Read the count the live system reports, not the confirmation message from the upload. In one case records were logged as fully uploaded and the live count was about 63% of that, because a “skip if already in the workspace” setting excluded everyone appearing in an older paused campaign. Nothing errored and nothing warned.

Usually because the link was built from your own cached access rather than resolved against the live system, or because the file’s sharing permissions never allowed anyone else in. Test every link in a private window or a separate browser profile where you are signed out. A link that opens for the owner survives review every time.

How do I know if an AI finding is actually correct?

Ask it to prove the problem exists right now instead of restating the finding. Make it show the full chain it followed and the exact check it ran. Anything that resolves through a chain, like DNS records, config imports, redirects, or inherited permissions, has to be traced to the bottom. Stopping one level early looks identical to something being missing.

What should I do when AI tells me to fix something I already fixed?

Say so out loud immediately, then ask it to search past sessions and verify the work was not already done. “I already did that” is diagnostic information, not defensiveness. In one case that exact pushback exposed an authentication finding that was wrong, and tracing the records by hand took about two minutes.

How often should I audit a campaign after it launches?

Twice per launch at minimum. Once immediately before going live, and once about 24 hours after. The second pass catches settings that behave differently under real conditions, like deduplication rules and suppression lists. If you send continuously, run the comparison on a schedule rather than waiting for someone to remember.

Do I need a developer to check that live content matches approved content?

No. For most teams the working version is a person opening the live asset next to the approved file on a recurring calendar invite, about fifteen minutes, twice per launch. Level two is pasting both versions into a free text comparison tool. Build a scheduled automation only once you are shipping weekly or faster.

Sources