The situation

The client is a content studio producing short-form video and written social content for around thirty B2B and consumer accounts at once, spread across verticals with very little in common: dental practice software, a supplement brand, a freight broker, two fintechs, a handful of creators. They had already solved production. Editors, writers, and a stack of generative tooling meant they could ship a very large volume of finished assets in a week, and had the throughput numbers to prove it. What they could not scale was what happened before production: deciding what a given piece should actually say, to whom, and with what tension in the first two seconds.

That decision lived almost entirely with one person, their head of strategy. She could watch fifty videos in a vertical she had never touched and come out with a defensible point of view on what that audience responded to. She was the reason the good work was good, and she was also the reason nothing could grow: every brief either passed through her or was noticeably worse. The company was quoting new accounts it could not staff, because the angle work did not exist anywhere except in her head.

Why off-the-shelf didn't fit

They had tried three things before calling us, and each failed differently. The failures are what defined the scope.

Generic AI writing tools produced hooks that were grammatically fine and strategically empty. Asked for "ten hooks for a dental software brand," they returned ten variations on the same generic curiosity gap, because nothing in the prompt encoded which specific frustration a practice owner feels at 7pm doing insurance claims, or why that frustration converts where a broader "grow your practice" angle does not. The output was fluent, and fluency was never the bottleneck.

Social listening and competitive-intelligence platforms told them what performed (this video did 4 million views, this one did 12,000) but stopped precisely where the useful part started. Knowing an outlier exists is close to worthless; the value is entirely in why it worked, and in whether that reason transfers to a different audience. Every tool they evaluated treated the "why" as something the human would supply.

Hiring was the third. They had tried recruiting a second strategist twice. Both hires were competent and neither replicated the judgment, because there was nothing to onboard against: no taxonomy, no swipe library organized by anything other than the date it was saved, no written account of the frameworks she was applying. The knowledge existed only as an instinct one person exercised a hundred times a week.

Scoping the real workflow

The scope call ran across two sessions, and almost all of the useful material came from the same exercise we use whenever the workflow is a judgment call: we asked her to narrate. We pulled forty pieces of content (a mix of her clients' outliers, her clients' flops, and competitor outliers in verticals she had never worked) and had her talk through each one without letting her summarize. What was the tension. Who is this for. Why this audience and not the adjacent one. What would kill it.

Two hours of that surfaced something she had not articulated herself: she was not evaluating hooks. She was matching a specific pain to an angle shape to an audience's stage of awareness, and the hook was downstream output of that match. The frameworks were consistent enough to name (there were roughly eleven recurring angle shapes across everything we looked at) but she had never written them down, so every brief re-derived them from scratch.

That reframed the deliverable. This was never "generate hooks." It was: build the taxonomy she was already using, make it retrievable, generate briefs constrained by it, and measure which matches actually won, per vertical, so the taxonomy improved rather than calcifying around her existing assumptions.

We came out of scoping with four things to build against: an ingestion pipeline for outlier content, a structured angle taxonomy with a retrieval layer over an organized swipe library, a brief generator constrained by that taxonomy, and a measurement layer joining creative-level performance data back to the angle that produced it.

71%
Briefs shipped without strategist edits
2.4×
Median view-through vs. pre-system baseline
6 wks
Scope call to production

What we built

The ingestion side pulls outlier content per vertical on a schedule (anything performing above a rolling percentile threshold for its account size) and normalizes it into a common record: the transcript or copy, the first-two-second framing, the visual setup, the platform, the audience signals available, and the performance numbers in context rather than in absolute terms. Absolute view counts turned out to be actively misleading as a training signal, since a 200,000-view video on a 4,000-follower account is a much stronger signal than a 2 million-view video on an account that does that routinely.

On top of that sits the taxonomy and the retrieval layer. Every ingested piece is classified against the eleven angle shapes, the pain it addresses, and the awareness stage it assumes, with the classification itself producing a short structured rationale rather than just a label. The swipe library stopped being a folder of links and became something queryable: show me angle-shape matches for this pain in adjacent verticals where the audience is problem-aware but not solution-aware.

The brief generator sits on that retrieval. It does not write copy. It produces the thing the strategist used to produce: the pain, the angle shape, the audience and stage, the tension to open on, the specific reason this should work for this account, and two or three reference pieces from the library with an explanation of what to take from each and what not to. Writers and generative tooling downstream were already good; giving them a sharp brief was the whole intervention.

The measurement layer closes the loop. Every published asset carries the angle record that produced it, so creative-level performance data flows back keyed to pain, angle shape, vertical, platform, and awareness stage. Weekly, that becomes a legible view of which matches are winning where. It also feeds retrieval weighting, so the system's recommendations shift with evidence instead of drifting on the strategist's older assumptions.

Where it got hard

Our earliest version of the "why did this work" analysis reshaped the schedule. It produced explanations that read as genuinely insightful and were, on inspection, frequently unfalsifiable. Given any piece of content and its performance number, the model would construct a confident, plausible causal story. It constructed an equally confident story for the flops. Shown a video that did 12,000 views, it explained the failure just as fluently as it had explained a 4 million-view success. That is a system that cannot be trusted to teach anyone anything, and it is exactly the failure mode that makes this category of tool feel useful while being worthless.

So we stopped building the generator and built an evaluation harness first. We held out a set of forty pieces the strategist had never discussed with us, stripped the performance numbers, and had the system predict relative performance and state its reasoning. She scored the reasoning blind, without knowing which predictions were right. That gave us two separate measurements (was the call correct, and was the stated reason the actual reason) and they diverged more than either of us expected on the first run. Iterating against that harness, rather than against how good the output sounded, is the only reason the shipped version is worth anything. It cost about eight days that were not in the original quote; we flagged it the day we saw the problem and absorbed the time, because the estimation error was ours.

Distinguishing signal from coincidence in the measurement layer was the other hard part. Per-account sample sizes are small: a client publishing fifteen assets a month cannot support confident per-angle conclusions, and the first version of the weekly view happily reported that a given angle shape was "winning" on a sample of three. We ended up pooling within vertical and awareness stage across accounts, and gating the feedback loop behind a minimum sample before it moves any retrieval weight. The dashboard now says "not enough data yet" more often than it says anything else, which the team initially read as the system being broken and now reads as the system being honest.

Rollout & results

The system went live in week six, initially on four accounts across two verticals rather than the full roster, because we wanted the measurement loop running on real published work before it was informing briefs everywhere. Full rollout followed about three weeks later.

Three months in, 71% of generated briefs are going to production without strategist edits, against a pre-system baseline where effectively every brief either originated with her or came back marked up. Median view-through on assets produced from angle-engine briefs is running about 2.4× the studio's pre-system baseline for comparable accounts, though we would note that some of that is selection: the system is also better at declining to brief something it has no evidence for.

The studio has since onboarded two strategists who are productive against the taxonomy. Two previous hiring attempts had failed to do that. The head of strategy now spends her time on the accounts and verticals where the system has no coverage, and on adjudicating the cases it flags as uncertain.

"I was fairly sure my job wasn't encodable, and I was about eighty percent wrong. The part that isn't encodable turned out to be the part I actually want to be doing." Head of Strategy, client engagement

What we'd do differently

We would build the evaluation harness in week one as a planned deliverable, not reactively in week three once we caught the model explaining failures as confidently as successes. We had already learned this lesson on a lending engagement: any workflow with a judgment threshold needs a shadow-mode comparison built before the thing being judged. We did not carry it across because on the surface this looked like a content project rather than a calibration project. It was a calibration project. Anywhere a system produces explanations that a human will act on, the ability to check those explanations against held-out reality is the first deliverable, not the last.

The second thing: we would have pushed harder, earlier, on getting creative-level performance data out of the platforms rather than accepting account-level exports for the first two weeks. The measurement layer was the last piece to come online and it is what makes the system improve, so every week it was not running was a week the loop was open.

Have judgment sitting in one person's head?

If one person on your team is the reason the work is good, that's a constraint, not a compliment. Tell us what they actually do and we'll tell you how much of it is encodable.

Request Your Agent