Growth & Strategy

You Didn't Hire a Million Bad Employees. You Just Never Told Anyone What Good Looks Like.

August 2, 2026

A Hebbia founder's viral essay pins AI's failure on the workforce. He's right about the mechanism, and wrong about who it indicts.

You Didn't Hire a Million Bad Employees. You Just Never Told Anyone What Good Looks Like.
Credit:
powered by

Make State of Brand one of your go-to sources on Google

Google Icon
Add State of Brand on Google

"You Just Hired a Million Bad Employees." That's the headline George Sivulka ran on a16z's newsletter July 14, under a subhead announcing that humans are now cheaper than software for the first time in history.

Good headline. It was built to do a job and it did it. Sivulka founded and runs Hebbia, which sells AI into BlackRock, KKR, and the Air Force, and his argument to executives is that the agents they rolled out didn't replace anybody. They added a second workforce nobody knows how to manage. He calls the failure mode "looping," where agents call themselves over and over to make up for instructions a human never got right the first time. Fortune ran it ten days later. Hebbia cross-posted to its own blog after the fact. By now the vocabulary is turning up in meetings where nobody's read the essay.

Most of the pickup chased the cost line, which lands about where the tokenmaxxing scoreboard did in the spring. (The a16z URL still carries an older working title about goldrushes and loops, so the headline everyone quoted was the second attempt.) The paragraph that matters for brand and content leaders is buried near the end and has nothing to do with cost.

The coding exception

AI has thrown off real, defensible value in software engineering and almost nowhere else. Sivulka's explanation isn't that engineers are sharper or that the models were built for code. It's that code arrives with a verdict attached. Code runs or it doesn't, and every attempt gets scored instantly, by a machine, with nothing left to argue about.

He calls that the eval, and he thinks it's the whole reason coding slipped past the politics that stalled AI in every other department. Other use cases arrive when somebody builds the equivalent. Not when staff get better at prompting. Not when the chat window improves.

Evals are the new OKRs, he writes, a company's eval suite will end up being its most valuable asset, and no two companies will ever have the same one.

That's the operations-side version of what we've been arguing about brand since the spring.

Who the diagnosis lets off the hook

That's where the essay and its own headline stop agreeing.

If the eval is the mechanism, the employees were never the problem. Nobody ever wrote down what good looks like in a form anything could check against, and producing that document was never the workforce's job in the first place. It belongs to whoever sets the standard. Sivulka reckons one employee in a hundred can hand an agent usable context, and reads that as a shortage of clear thinkers. The same number measures something else entirely. Ninety-nine people out of a hundred have never seen a written specification of their own job, because their employer never made one.

"You just hired a million bad employees" puts the problem in staffing, a comfortable place for it if you signed off on the agent rollout, and a convenient one if you sell transformation work. The rubric went missing long before the agents showed up. Agents just made the gap expensive enough to see.

Marketing has the inverse problem

Engineering came with objective pass/fail and had to learn taste. Marketing has plenty of taste and has never written down a pass/fail.

Brand guidelines were always a specification of what good looks like. Voice, tone, what we say, what we'd never say. The document existed so a person could tell another person they'd missed. Nobody wrote it to be executed.

Then generation got cheap, and there was nothing to check output against except somebody's gut, and a gut doesn't scale to four hundred pieces a week. Vendor defaults filled the gap.

We went through those defaults in July. Jasper crawls your website, infers a voice out of whatever it finds, and applies it across every generation in the workspace by default. Its own documentation offers a voice that's helpful but not bossy. Typeface runs guidelines through a hub that validates copy against brand rules. Both are eval layers, shipping fast, and every company on the platform gets a version of the same one.

An eval that ships inside a product isn't an advantage. It's a floor, and it's the same trap we described when everything in marketing still worked and none of it differentiated anybody. Sivulka's own line is that generic evals leave an organization with no edge, and he's describing enterprise workflows rather than marketing copy. The logic carries over and bites harder here, because the output is what customers actually see.

What an actual marketing eval looks like

A description and a test aren't the same thing.

"Confident, human, direct" describes. It can't fail anything. Every brand in your category has some version of it, and a model handed those three words will regress to the identical mean for all of you. That's the flattening mechanism, and it's why a longer adjective list was never going to fix it.

A test looks different. There's a claim your brand won't make, and a draft that makes it gets killed. There's a competitor you name and three you don't. There's a structural rule, something like leading with the finding rather than the setup, that a reviewer can check in four seconds. Best of all, there's a library of rejected drafts with reasons attached, probably the most valuable and least maintained asset in any content operation.

Real evals encode a decision somebody made and could have made differently. Generic ones encode aspirations everybody shares. Only the first kind produces difference.

It's the operational form of the Great Flattening argument: using AI differently won't get you to distinctiveness, because the point of view has to exist before the model touches anything. Sivulka gets to the same place through cost accounting.

Where the argument gets expensive

Sivulka sells the fix. His essay closes by naming "AI transformation companies" as the next trillion-dollar category, and Hebbia sits inside the definition. The charts on token spend per employee and headcount growth after adoption run without sourcing. Ramp's study of 21,559 companies found the heaviest AI spenders hiring more rather than less, which supports his direction. He's also arguing the reverse of the case Alex Karp made on CNBC in July, using Palantir as his proof.

The second problem is harder, and he mostly walks past it. Code evals come cheap because correctness is objective. A marketing eval needs somebody with taste to convert judgment into rules and keep converting as the brand moves. That work is slow and senior and gets cut first whenever a company decides AI has made content cheap. Ford spent three years hiring back the engineers it let go before it topped J.D. Power for the first time in sixteen years.

The market has started repricing the judgment anyway. Anthropic is paying $300,000 for a standards editor. Mercury is paying $335,000 for a job that reads like an editor-in-chief's. Neither is a vanity hire.

The part nobody is calling a coincidence

Sivulka didn't put this on Hebbia's blog first. He ran it on a16z's newsletter, where the operators and investors already gather, with no product mention anywhere in it. Two weeks later his vocabulary was in Fortune and executives were using his framing to describe their own companies' failures.

"You Just Hired a Million Bad Employees" doesn't describe anything in the essay. It's a test, and a well-built one, failing on contact with any executive who's watched an agent burn a weekend of compute on nothing, and unfalsifiable for everyone else. Which is why it travels.

It also points the blame downward, at the hundred people who couldn't brief the machine, and away from the fact that nobody above them ever wrote the brief.

The argument underneath the packaging is better than the headline and less comfortable. The eval suite is the asset. Somebody has to sit down and write it. In most marketing organizations that person already works there, has the taste for it, and has never been asked.

Outlever Logo

If this caught your attention, that’s not accidental.


Decoration line

The best editorial systems don’t happen by accident. Outlever builds them.

Partial view of green concentric circles with a solid green dot on the outermost circle on a light background.Concentric green circles with a single solid green dot on a dashed circle on a light background.Minimalist design with faint curved lines and scattered small green dots on a white background.

Come back for the reason it lands.


Subscribe for the kind of thinking that makes people stop, read and come back.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.