Brand & Creative

OpenAI Automated the Editorial First Pass. Nobody Has Automated the Word No.

August 20, 2026

The team's judgment was already sitting in Slack threads and Google Docs comments, and mining it is the most copyable idea in B2B marketing right now. It also solves the easier half of the problem.

OpenAI Automated the Editorial First Pass. Nobody Has Automated the Word No.
Credit:
powered by

Make State of Brand one of your go-to sources on Google

Google Icon
Add State of Brand on Google

There is a document in every marketing organization that nobody has written down.

It's the list of things the senior editor always says. Cut the first paragraph, you're clearing your throat. This claim needs a number attached or it isn't a claim. Nobody has ever said this sentence out loud. The list lives in one or two people's heads and gets distributed one comment at a time, to whoever happened to send a draft that week.

Sandhya Simhan, who runs content marketing at OpenAI, has started writing it down. Not as a style guide. As a repo.

The setup takes a paragraph to explain. Nearly everything B2B marketing publishes at OpenAI gets an editorial pass from a human on her team, which makes that team a bottleneck by design. Product marketers, engineers and everyone else drafting upstream have no way to get useful feedback before the work lands. So the team is building a set of editorial agents, each with its own standards, examples and narrow job. They encode Dane Vahey's demo bar and Colin Fleming's narrative bar, along with strong examples of every format the team ships. The goal is a faster, better first pass: a blog post that sounds like the company, a report with an argument, a keynote outline that knows where it's going.

She started on day two of the job. She's still updating it. She thanks three colleagues for letting her steal their skills.

All of which deserves to be taken seriously, and it happens to be the most concrete answer anyone has published to a question this publication raised on Tuesday. It answers a slightly different question than the one that was asked.

The archaeology is the part worth stealing

Set the agents aside and look at step one of the recipe Simhan gives away at the end. Point Codex at the Slack discussions and Google Docs comments you're authorized to use, and go looking for editorial feedback that repeats.

That instruction is doing more work than everything around it.

Content teams in B2B have spent years producing a corpus and treating it as exhaust. Tracked changes. Suggestion-mode comments. The Slack thread where two people spent forty minutes arguing about whether a headline was a claim or a description. All of it is a record of one organization's taste, applied to one organization's work, by people who knew what they were talking about, and none of it has ever been searchable or transferable. When the person who wrote the comments leaves, it leaves.

It's also the one training corpus a competitor can't get hold of.

Which makes it the sharpest available rebuttal to what every marketing platform is currently selling. Every AI marketing tool now ships a brand voice feature, and they are all selling the same voice, because they are all built on the same public internet and configured from the same list of adjectives. Confident. Approachable. Clear. Four competitors fill in the same form and get back the same paragraph.

Nobody's comment history reads like anybody else's. If differentiation now depends on saying a narrow thing repeatedly, in language a model can't smooth into a category description, the raw material for it was never going to be in the brand guidelines PDF. It's in six years of arguments about sentences.

Mine it or don't. Just be clear about what it is, because for most content teams it's the only proprietary data they have.

A bar is a person before it's a document

Now the harder half.

The agents encode Vahey's demo bar and Fleming's narrative bar. Those are two named executives, and the naming isn't decoration. It's the mechanism. A bar means something inside that company because a specific person with specific standing enforces it, and everyone knows roughly what happens when work shows up underneath it.

An agent inherits the standard. It doesn't inherit the standing.

That gap is where Tuesday's piece landed. Once production cost stops filtering work, the scarce function isn't making things, it's refusing them, and refusing them without getting overruled by lunchtime. Fleming's own framing, that marketing shifts from controlling who participates to setting the standard for what ships, only holds if the standard has teeth in it.

A pre-draft agent gives a PMM the note. It can't give the PMM's VP a reason to accept the note. It can tell an engineer the narrative hasn't found its argument yet. It can't stop the deck going out on Thursday, because Thursday is when the customer is free. The comment that says this isn't there yet and the authority that says therefore it isn't shipping are two different objects, and only one of them fits in a repo.

That's not a flaw in what OpenAI built. It's a ceiling on the whole category, and it needs saying now, while the tooling is good enough that people will start mistaking one for the other. The editor-in-chief job appearing on marketing org charts exists to hold refusal authority. A well-built editorial agent means that person spends their political capital on drafts worth arguing about instead of on paragraph three. Real gain. Different gain.

The floor rises, and the volume rises with it

There's a knock-on effect the post implies without stating.

Better first passes mean more first passes. That's the intended outcome. PMMs and engineers who never had access to editorial feedback now have it, so more of them will draft, and the review queue fills with competent, on-brief, credible-sounding work produced by people who don't report to marketing.

Fleming's own warning was about perfectly competent garbage. A floor-raising agent is, mechanically, a competence-raising machine. It moves work from obviously not ready to arguably fine, and arguably fine is much harder to kill. The things that die easily are the things that are visibly bad. Nobody has ever shipped a tool that makes work easier to reject.

Brands don't get diluted by bad drafts. They get diluted by an unbroken supply of reasonable ones.

The corpus is a wasting asset

The last problem is structural, and it will show up in about eighteen months.

The agents work because senior editors spent years leaving detailed comments on unfinished writing. As soon as drafts start arriving pre-polished, that behavior changes. Nobody writes this claim needs a number when the claim already has one. Comments get shorter, sparser, more approving, and the corpus that made the system valuable stops being replenished by the system's own success.

The post says to keep updating the agents as the team reviews real work. Correct, and precisely the instruction that gets quietly dropped in month seven when everything seems to be working. Taste isn't a dataset you extract once. It's a byproduct of people continuing to argue about sentences in writing, and that argument now has to be protected on purpose, because it won't keep happening by accident.

Anthropic is paying around $300,000 for someone to argue about commas. OpenAI listed its B2B content role this spring at $266,000 to $295,000 plus equity, inside a team it calls content and narrative. Neither company is confused about what taste costs. The open question is whether an organization that successfully automates the expression of taste keeps paying for the production of it, once drafts stop coming back bleeding.

What most teams will build instead

Most teams who copy this will skip step one, because building agents is interesting and reading four years of Google Docs comments is not. They'll point Codex at the published archive instead: the blog, the reports, the decks that went out the door. Tidy, in one place, already approved by somebody.

That produces a different machine than the one Simhan is describing.

A published archive is a record of what cleared the bar. The comment history is a record of what didn't, and why, in the words of the person who caught it. Train on the first and you get a system that reproduces your house style, which is a formatting problem most teams had mostly solved. Train on the second and you get one that can tell you when something is off it. Only one of those is carrying any judgment, and it's the messier one.

The negative space is the asset. Everything a team decided not to run, plus the sentence someone left in the margin explaining why, is the only part of the record that holds a standard. The published version is the residue.

So take Simhan up on the offer she makes at the end of her post, which is to hand the whole thing to Codex and ask for your own version. Just be careful what you feed it. A system trained on your greatest hits will tell you that everything sounds like you. It will be right, and it will be useless.

And it still won't say no. It will tell a PMM the draft is weak, in your team's language, using your team's reasoning, which is already more than any brand voice slider has managed. Then it hands the draft back and goes quiet, and somebody in the building decides whether the thing ships. That job hasn't moved. It has only got harder to pretend it belongs to nobody.

Outlever Logo

If this caught your attention, that’s not accidental.


Decoration line

The best editorial systems don’t happen by accident. Outlever builds them.

Partial view of green concentric circles with a solid green dot on the outermost circle on a light background.Concentric green circles with a single solid green dot on a dashed circle on a light background.Minimalist design with faint curved lines and scattered small green dots on a white background.

Come back for the reason it lands.


Subscribe for the kind of thinking that makes people stop, read and come back.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.