Somewhere in the next three weeks, a marketing manager is going to open a performance review template and stall on the first section. The template asks for deliverables shipped, campaigns launched, content produced, and impact against goal. The marketer in question produced roughly four times the volume of last year. So did everyone else on the team. The manager’s honest read is that this person is one of the two or three best marketers in the org—and nothing in the template explains why.
That stall is the whole problem. H2 review and calibration cycles are opening now, on rubrics written for a world where marketing output was scarce and expensive. It isn’t. The scarce inputs moved somewhere else, and most performance systems have not followed them. What that produces is not just awkward paperwork; it produces calibration meetings where nobody can defend a rating, promotion cases that get denied for lack of evidence, and—worst—a reward signal pointing at behaviors that no longer create value.
This is a solvable problem, but not by adding an “uses AI effectively” competency to the existing form. That is the most common response and it is close to useless.
What Actually Broke
Three distinct things happened, and conflating them is why the fixes tend to miss.
Output volume stopped discriminating between people. For twenty years, throughput was a legitimate proxy for capability in marketing. Someone who shipped more good work was, on average, better. That correlation has substantially decoupled. When the marginal cost of a draft, a variant, a deck, or an analysis approaches zero, volume tells you about a person’s tool access and their willingness to click, not their skill. Two marketers producing identical volume can differ enormously in quality of judgment, and the rubric records them as equal.
The failure mode inverted. The classic marketing failure was not enough—campaigns that never shipped, content calendars with holes, analyses that arrived after the decision. The current failure is plausible and wrong: a competitive brief with a fabricated positioning claim, a segment analysis built on a metric the tool defined differently than the org does, a nurture sequence that reads well and contradicts what sales tells the same accounts. These are not volume failures. They are verification failures, and a rubric organized around production has no line for them.
Individual attribution got genuinely murky. When a marketer directs a tool that drafts, iterates, and publishes, the question “what did this person contribute?” has a real answer, but it is not visible in the artifact. The artifact looks similar whether the human supplied sharp strategic framing and caught three errors, or pasted a prompt and shipped what came back. Reviewing the artifact is now nearly uninformative about the reviewer’s actual question.
Notice that all three point the same direction. The work product has become weak evidence about the worker. Which means evidence has to come from somewhere else.
The Two Ways Teams Are Getting This Wrong
Before the fix, two failure patterns worth naming, because most organizations are currently in one of them.
The first is adoption theater: measuring AI usage as a performance dimension. Tool logins, prompts submitted, features adopted, percentage of work “AI-assisted.” This is appealing because the data is easy to get—every platform reports it—and it is exactly the wrong thing to reward. It measures compliance with a tooling initiative, not contribution to the business. Teams that install it get what they measure: high usage, indifferent output, and a quiet penalty on the marketer who correctly decided that a particular task was faster done by hand.
There is a sharper version of this hazard. As we discussed last week, consumption telemetry is about to become a serious operational concern for metered martech contracts—and that same telemetry is per-user. The temptation to feed it into performance conversations will be real, and it should be resisted. Consumption data is a cost-management input. Repurposing it as a productivity signal converts a budgeting tool into surveillance, invites gaming within days, and in several jurisdictions raises employment and works-council questions your legal team would rather hear about in advance.
The second failure is doing nothing and calling it judgment. Managers quietly override the rubric, rate on intuition, and reverse-engineer justifications. This produces the ratings most people would agree with—experienced managers are decent at recognizing good marketers—and it fails at everything else a performance system is supposed to do. It cannot be calibrated across managers, it cannot be defended when challenged, it cannot tell a mid-level marketer what to work on, and it distributes rewards in a pattern that is impossible to audit for bias.
What to Evaluate Instead
The useful reframe: stop asking what the marketer produced and start asking what they contributed that the tools could not. In practice that resolves into four things, all of which are observable if you collect the right evidence.
Problem selection and framing. Of everything this person could have worked on, did they work on what mattered? Did they define the problem well enough that execution was tractable? This has always been the difference between a senior and a mid-level marketer, and it is now most of the difference. It is also the dimension where a wrong answer is most expensive: perfectly executed work on the wrong problem is now cheap to produce in volume, which makes misdirection a much larger risk than underproduction.
Verification and error interception. Did this person catch what was wrong before it went out? This is the new craft skill, and it is genuinely distributed—some marketers reliably notice the fabricated stat, the stale pricing claim, the tonal miss, the number that contradicts the CRM; others reliably do not. It is also learnable and coachable, which makes it worth naming explicitly rather than folding into “attention to detail.”
Context and direction quality. AI output quality is dominated by the quality of the framing, constraints, and context supplied to it. A marketer who can articulate positioning, audience state, brand constraints, and success criteria precisely enough to get good output on the first or second pass is demonstrating a real capability, not a tool trick. The tell is efficiency of iteration and consistency of quality across different task types.
Durable assets and leverage. Did this person leave the organization better able to do the work than before? Reusable context, agreed metric definitions, prompt libraries the team actually uses, documented review standards, a workflow that removed a recurring bottleneck. This is where the definitional work we covered earlier this month shows up in a performance conversation—someone who nails down what “qualified pipeline” means for the whole org has produced more durable value than someone who shipped thirty assets, and only one of those shows up as a deliverable count.
Rewriting One Band Before the Cycle Closes
You cannot re-architect job levels in three weeks, and attempting it during an active cycle is how you end up with two half-finished systems. Do something smaller and finish it.
Pick the single level band where the mismatch is most acute—usually the mid-level rung where the difference between competent execution and real judgment determines promotion—and rewrite its expectations against the four dimensions above. One band, plain language, no new competency framework.
Then change what evidence you collect, because this is the part that actually determines whether the new rubric works. Instead of a deliverables list, ask each marketer for a short decision log: three to five consequential calls they made this half, what they chose not to do, what they caught before it shipped, and what they built that others now use. Managers should be prepared to add the cases the marketer will underreport—the near-miss they intercepted, the pushback that changed a campaign’s direction.
Bring that to calibration and use it to run the comparison that the old rubric could not: two marketers with similar output, different contribution. If the room can articulate the difference from the decision logs, the instrument is working. If it cannot, the logs were too thin and that is fixable next cycle.
Say plainly what is not being measured. Tell the team that AI usage is not a rating dimension, that consumption data is not performance data, and that hiding tool use is neither necessary nor rewarded. Absent an explicit statement, people assume the opposite and optimize accordingly.
The Harder Problem: The Rung That Disappeared
Underneath the rubric question sits a structural one that will not be solved this cycle and should not be ignored during it.
Marketing’s apprenticeship model ran on tactical work. You learned by writing the email nobody read, building the report nobody used, and getting your draft returned covered in edits. That work was the training data for judgment. It is precisely the work that has been automated first.
The consequence is straightforward and slow-moving: the pipeline that produces marketers with good judgment is being starved at the entry point, while the skills you now need most—verification, framing, taste—were historically acquired exactly there. Organizations optimizing headcount by removing the junior rung are borrowing capability from four years out, and 2027 headcount requisitions are being drafted right now on that basis.
Two things help, and neither requires new budget. First, put junior marketers in the verification seat deliberately rather than incidentally: reviewing AI output against sources, checking claims, comparing numbers to systems of record. This builds pattern recognition faster than producing first drafts ever did, because they see a high volume of near-miss work and have to say what is wrong with it. Second, restore structured critique. The edit loop was the mechanism, and it disappeared when the draft stopped being human. Replace it explicitly—reviews where work is examined in front of the person, with reasoning made audible.
Hiring should shift the same way. Portfolio review has lost most of its signal, because a strong portfolio is now easy to assemble. What still discriminates is a live exercise: hand a candidate AI-generated marketing work with three plausible errors embedded—a fabricated statistic, a positioning claim that contradicts the brief, a metric used incorrectly—and ask them what they would change before it ships. Strong candidates find the errors and explain why the piece was aimed at the wrong problem in the first place. Weak candidates polish the prose. That exercise takes twenty minutes and tells you more than any work sample.
The Honest Summary
Performance systems are a reward signal, and reward signals drift when the underlying work changes faster than the paperwork. Marketing’s work changed fast, and the paperwork still measures throughput—a metric that has quietly become a measure of tool access rather than talent.
The correction is not sophisticated. Evaluate what the tools cannot supply: choosing the right problem, catching what is wrong, directing the work precisely, and leaving the organization better equipped. Collect decision evidence instead of deliverable counts. Keep usage telemetry out of it. Fix one level band properly rather than all of them badly.
And then deal with the junior rung, because that is the decision with the longest shadow. Everything above depends on having marketers whose judgment is worth deferring to, and judgment has never come from anywhere except doing the work and being told what was wrong with it. If AI does the work and nobody is told anything, the marketers you will need in 2030 are not being made anywhere.