How to audit your website copy for the AI era

TL;DR A copy audit for the AI era looks at three layers on every page. Vocabulary, structure, and voice. Pages that fail any one of the three start reading as generic to both human visitors and answer engines. The vocabulary…

A three-layer lens auditing website copy for vocabulary, structure, and voice in the AI era.
A three-layer lens auditing website copy for vocabulary, structure, and voice in the AI era.

TL;DR

  • A copy audit for the AI era looks at three layers on every page. Vocabulary, structure, and voice. Pages that fail any one of the three start reading as generic to both human visitors and answer engines.
  • The vocabulary layer is the easiest to score. A short banned-phrase list catches the AI-cliche set without forcing you to count words. Density is the signal, not single uses.
  • The structure layer asks whether each page carries a clean 40-60 word answer block, descriptive H2 questions, and a scannable shape an answer engine can lift cleanly.
  • The voice layer is the hardest to score and the most important. The test is whether the page sounds like a person who has done the work, or a person who has read about the work.
  • The audit is a weekly habit, not a one-time pass. One page a week, rewritten honestly, beats a six-month audit project that never ships its findings.

A small-business owner I know prints out her homepage and reads it aloud in her kitchen on a Saturday morning. She does this once a quarter. The kettle goes on, she reads the first paragraph, and somewhere around the third sentence she stops and frowns.

"Innovative solutions for forward-thinking businesses," her own homepage tells her.

She did not write that sentence. A drafting tool produced it nine months ago when she ran out of time before a launch. She approved it because it sounded professional and because she had a phone call in fifteen minutes.

The sentence has been on her homepage ever since. It is the first thing every visitor reads. It is the first thing the answer engines read when they crawl the page.

She closes the kettle, sits down, and starts a list of every page on the site that sounds the same way. The audit has begun.

Why does generic copy quietly become a liability in 2026?

The cost of generic copy is rising, not falling. In 2026, AI tools generate pages of fluent marketing text at zero marginal cost. Every competitor in your category can produce the same vocabulary. The pages that read as everyone-else copy stop earning the attention they used to earn from human visitors and answer engines alike.

The mechanism is the part that gets missed. AI tools reproduce the most prevalent patterns in their training data. In February 2026, a Brandfolio editorial traced the generic feel of AI content back to the training corpus, naming "pleased to announce" and "innovative solutions" as the kind of phrases that dominate the marketing-writing data the AI model learned from. Without explicit constraint, the AI reaches for what it has seen most often.

Answer engines synthesize from many pages at once. The page that restates the consensus has little citation value because the engine can already assemble the consensus on its own. The page that contributes a specific number, a named tool, a real testing result, or a counter-narrative is the page the engine cannot reproduce without quoting.

The 2026 read on this is consistent across the practitioner sources. Generic copy still ranks for some queries on traditional search. It struggles to win the answer-engine citation that drives a growing share of category-defining traffic.

The cost compounds. A homepage that read fine in 2024 may quietly drift into liability in 2026 even though nobody changed a word.

What does an AI-era copy audit actually look at?

A useful audit looks at three layers on every page. Vocabulary, structure, and voice. Each layer is scorable. Each layer maps to a fix the writer can ship the same week.

The vocabulary layer asks how many phrases on the page belong to the documented AI-cliche set. The 2026 inventory runs roughly twenty phrases.

"Leverage" and "synergy" at the abstract end, "robust" and "seamless" in the middle, "pleased to announce" and "in a world where" at the announcement end. The audit only needs the top ten as a starting screen.

The structure layer asks whether the page carries a clean 40-60 word answer block under the headline. Whether the H2 sections name real reader questions instead of stock tutorial labels. Whether scorecard tables and FAQ blocks sit where they earn the cite. A page that fails the structure layer often reads as competent prose to a human and as unliftable noise to an answer engine.

The voice layer asks whether the page sounds like a person who has done the work. A page that fails the voice layer reads as a person who has read about the work. The voice test is the hardest to score because it depends on judgement, not regex.

A page that passes all three layers is rare in a 2026 corpus. Most pages fail one. The audit is the discipline that names which one each page failed.

How do you score a page for the AI-cliche set without overcounting?

The trap with a banned-phrase list is treating it as a regex. A single use of "robust" in a 2,000-word piece is a style choice. Three uses per page is the AI-output pattern.

The Atlas knowledge base names the limit plainly. The cliche density is the signal, not the single use. A page that reaches for "leverage" once in a paragraph about leverage points has done nothing wrong. A page that uses "leverage" three times in 600 words has surfaced a vocabulary tell.

The practical score is a ratio. Count the cliche-set matches on the page, divide by the page’s word count, and multiply by 1,000.

A page with two matches per thousand words is clean. A page with five or more is in the AI-tell range and needs a vocabulary pass.

The other useful score is concentration. A page that piles three cliches into one paragraph reads worse than a page that spreads three cliches across the body. Concentration is what triggers the reader’s flinch.

The audit’s vocabulary verdict is one sentence. Either the page is in band, or it has a cliche cluster.

The fix is mechanical. Replace each cliche with a specific outcome. "Lightweight design" becomes "carry it all day without shoulder fatigue." Specificity is what evicts the cliche.

Which pages on a small site deserve the audit first?

Most small sites have somewhere between fifteen and sixty pages. Auditing every page in a week is not realistic. Triage matters.

The first audit candidate is the homepage. The homepage is the only page every visitor sees, including every answer engine that ever crawls the site. The cost of a generic homepage is the highest cost on the site.

The second candidate is the main service page. The page a buyer lands on after deciding the brand might be a fit. The page that has to convert the curiosity into a contact form.

The third candidate is whatever blog or comparison page draws the most search traffic today. The page already winning visits is also the page that pays back the audit fastest because the rewrite ships against existing demand.

The fourth candidate is any page that has not been touched in eighteen months or more. Time is the silent voice killer. A page that read fine in 2024 may have drifted because the vocabulary has shifted under it.

Everything else can wait. The audit’s job is not perfection. It is to surface the three or four pages where one Saturday of work moves the whole site.

How do you separate a voice problem from a structure problem?

Most audited pages fail one layer cleanly. A few fail more than one. The diagnosis matters because the two fixes ship in different timeframes.

A structure problem looks like this. The page has a clear voice, the opening reads as a person, and the body covers real points.

But the answer block is missing or buried in a callout box. The H2 sections name vague tutorial labels, and the FAQ block does not exist.

The fix is a half-day rewrite. The voice survives and the shape changes. The page now carries a 40-60 word answer block under the H1, question-shaped H2s, and a small FAQ block at the bottom.

The discipline is the 40-60 word answer block as the load-bearing AEO unit, and the structure layer of an audit is mostly that one shape applied across every page.

A voice problem looks different. The structure is fine, the page has an answer block, and the H2s read as questions. But the prose reads as a person who has read about the work, not a person who has done it.

The fix is harder. It takes a rewrite of the page’s first three paragraphs, anchored in a specific moment from your actual experience. A real client, a real number, a real before-and-after.

No fix-by-regex catches a voice problem. The voice has to be installed by a person who has the experience to install it.

A few pages fail both layers. The honest move on those pages is to leave them in draft and rewrite the page from a one-line brief, not from the broken version. Some pages are easier to rebuild than to fix.

What does an audit not measure, and where do you have to look elsewhere?

An audit measures the page as it sits on disk. It does not measure how the page performs in the wild. Three things sit outside the audit’s scope, and three other measurements catch what the audit cannot.

The first thing the audit misses is conversion. A page can pass vocabulary, structure, and voice and still convert at half the rate the previous version did. Conversion is a function of audience match and offer clarity, not just copy quality.

The audit’s job is the prose. The conversion test is a separate analysis.

The second thing the audit misses is citation. The audit can predict citation prospects. It cannot measure them. The only way to know whether the page earns answer-engine citation is to run the page’s core questions through ChatGPT, Perplexity, and Google’s AI Mode and check the citation set on each.

The third thing the audit misses is internal link health. A perfectly written page that no other page on the site links to is an orphan, regardless of how the prose reads. Link audits live in a different tool and a different discipline.

The audit’s honest scope is the prose itself. The other three measurements run alongside the audit in their own cadence. Conversion testing runs monthly, citation testing runs quarterly, and link audits run quarterly.

The four together cover the page’s health.

How do you turn the audit into a small, weekly rewrite habit?

A 40-page audit done in one sitting produces a 40-page punch list that never gets shipped. A one-page audit done every Saturday produces 50 audited and rewritten pages in a year, which is more than most small sites carry.

The weekly habit is the move. Pick one page, read it aloud, score it on the three layers, and note the failures.

Rewrite the failed layer the same morning. Push the rewrite the same afternoon.

The pace is the part that makes it durable. A page a week is small enough to ship through any week that life delivers. A page a week is also large enough that the corpus shifts visibly within three months.

The list itself lives in a single text file. Five columns: slug, last audited, vocabulary score, structure verdict, and voice verdict.

Updated each Saturday. Reviewed each quarter.

The habit also catches drift. A page audited clean in January may need a touch in October because the cliche set evolved or because a new pillar made the old one redundant. The weekly audit is the only forum where this kind of slow drift gets caught.

The discipline is small and repeatable. The compounding effect is real. A site that audited and rewrote one page a week for a year is a site that wrote roughly fifty rewrites of its weakest pages. That is the kind of work that moves a small-publisher domain into the citation set most owners assume only the big publishers ever reach.

Other questions worth answering

How often should the banned-phrase list itself be refreshed as language models evolve?

Three months is the practitioner cadence. AI models shift their defaults over time, and a list frozen in 2024 will miss roughly half the current tells. I add new phrases the week I notice them across three drafts.

Review the list each quarter and prune the phrases that no longer cluster. Brandfolio’s February 2026 editorial frames the pattern. Language models reach for the most common phrases in their training data.

What free tools help a writer spot weak phrasing in a draft?

Three tools earn their place. The browser find-and-replace catches the banned phrases at zero cost. Hemingway shows reading-level drift in real time. A read-aloud function in Word or Google Docs surfaces the rhythm tells that silent reading misses.

Reading the draft to yourself is still the best diagnostic. Brandfolio’s February 2026 editorial names the same root mechanism. The tool is your ear.

How do you teach a junior writer to spot these tells on their own drafts?

Two moves work in 2026. Sit next to them while you read one of their drafts aloud. Mark the phrases that make either of you flinch. The marking teaches more than a lecture ever does.

Give them the banned-phrase list as a single reference sheet. Ask them to mark three drafts on their own, and compare your marks to theirs.

What changes when the content was translated from another language?

Two patterns surface. The translated prose often carries the source language’s rhythm. A literal pass from English drops articles in Slavic readings. The tells double up.

Read the translated piece aloud in the target language. The phrases that snag are the ones to revise. Brandfolio’s February 2026 editorial traces the same root cause. Language models cluster on the phrases they have seen most often in their training corpus.

Where does keyword targeting fit alongside this kind of review?

Three layers come ahead of keywords. Keyword targeting fits below them, never above. A keyword-stuffed paragraph fails the vocabulary check. A keyword-anchored answer block in a real human cadence passes Google’s helpful-content read.

Pick the question your reader asks the engine. The target phrase follows from that question, in the reader’s own words.

Which page on your site should you audit this Saturday?

Open your analytics. Find the page that gets the most visits today, print it, and read it aloud in your kitchen with the kettle on.

The page you flinch at is the page the visitor flinches at, and the page the answer engine reads as generic too.

If you would like a second pair of eyes on the audit before you rewrite, you can contact me here. I will read the page on all three layers and tell you which one is the weakest and which fix to ship first. The first round is free and the rewrite is yours regardless of whether we work together after.

Similar Posts