Why AI summaries cannot replace customer interviews

TL;DR LLMs read review corpora the way someone reads sheet music — patterns, frequencies, clusters — and miss the one customer sentence an interview produces. AI is optimized to be neutral — neutral is a feature for topic tagging and…

AI theme map paired with three customer interviews, the two-track workflow this post argues for.
AI theme map paired with three customer interviews, the two-track workflow this post argues for.

TL;DR

  • LLMs read review corpora the way someone reads sheet music — patterns, frequencies, clusters — and miss the one customer sentence an interview produces.
  • AI is optimized to be neutral — neutral is a feature for topic tagging and a problem for writing headlines, because headlines are made of variance.
  • An interview is not a quote-gathering step. The quote is the byproduct of hearing how something was said.
  • The two-track workflow: use AI for theme mapping and pre-interview prep, use interviews to produce the verbatim sentences that become actual copy.
  • Pick one underconverting page this week, pair an AI summary of sixty reviews with three twenty-minute interviews, and rewrite the page using the transcripts, not the summary.

An LLM can map sixty customer reviews into six themes in under a minute. It cannot write the headline that belongs on the landing page.

Six weeks later, a competitor ships the same page. Same six themes, same generic voice. Both pages were written off the same pattern. The AI model saw the same shapes in both corpora because the shapes were the same.

The problem is not the AI model. The problem is what the AI model sees. An LLM reads reviews the way someone reads sheet music — chord progressions, keys, time signatures. What it cannot hear is what a guitar sounds like in somebody’s hands.

The weight on a note. The pause before a bar. The catch in a voice when the singer means it.

Patterns are not voice. That distinction is the whole argument.

Can I skip customer interviews now that an LLM can summarize reviews and support tickets?

No, but the question deserves a serious answer. AI-driven customer research tools have moved fast in the last year.

A koji.so analysis from April tracked product teams using AI interview platforms. Eighty-seven percent increased their research cadence by three times or more. Cost per qualitative insight dropped by over sixty percent. Turnaround fell from four to six weeks down to twenty-four to forty-eight hours.

Those numbers are real. Nobody is arguing interviews should stay slow or expensive. The argument is about what the cadence buys you.

A Harvard Business Review piece from the same week framed the state of things directly. "The process of collecting data from consumers is typically hard, slow, and costly," and generative AI "promises to improve this state of affairs."

That is true about collection. It is less true about the thing the collection is for.

In copywriting, the thing the collection is for is one sentence the customer said that you would not have written yourself. That sentence lives in the verbatim. AI summary smooths it out.

What does an LLM see in your review corpus, and what does it miss?

An LLM reads hundreds of reviews and returns clusters, frequencies, and sentiment averages. It will tell you that setup friction came up forty-seven times. It will tell you that price objections came up thirty-one times. It will give you the top three pains in a tidy bulleted list.

That is the chord chart. Topic, frequency, rough sentiment. Useful for a content calendar, a product roadmap, a prioritization meeting.

What the AI model does not return is the thing one customer said on a Tuesday night after a long week. Something like: "I need this thing to stop treating me like I’m about to quit."

Nobody types that in a review. Nobody writes it in a support ticket. It comes out once, in an interview, at minute twenty-two. The small talk is over and the person is tired enough to be honest.

That sentence is what a landing-page headline is actually made of. No summary produces it. The summary mode is designed to average.

Why does AI-summarized voice-of-customer produce generic headlines?

The same koji.so report surfaces the mechanism without quite meaning to. "AI maintains identical, neutral quality across every session — eliminating a major source of qualitative data variance."

Read that sentence twice. The word that matters is neutral.

AI is optimized to be neutral. That is a feature when you want reliable topic tagging across thousands of conversations. It is the problem when you want a headline.

Headlines are made of variance. Variance is the thing the AI model is specifically designed to smooth out. The writing honest headlines walkthrough shows what this variance sounds like once it survives the summary step and lands on the page.

Feed sixty verbatim customer quotes into an AI model. Ask for the top three pains. The three pains come back, phrased in the voice of a business-school case study.

The variance is gone. The case study is what you already had.

A cmswire piece from February put the consequence in plainer language. "The first letter in AI stands for ‘artificial.’ When life gets hard and customers need support, they value another ‘A’: ‘authenticity.’"

The reader hears artificial. Every time. That is why click-through holds and conversion does not.

What is the job of a customer interview the AI model cannot do?

Interviews were never about gathering quotes. The quote is the byproduct. The real job is to hear the exact words under the emotion.

That is three things the AI model cannot hear.

  • The word the customer reaches for after a pause.
  • The word they use when they are annoyed, not the word they use when they are composed.
  • The story they tell when you ask a second-level question the AI model would never ask.

A review says the onboarding was confusing. An interview subject, twelve minutes in, says: "I felt like I was taking a test I didn’t study for."

Same topic. Different sentence. One of those goes in the summary. The other goes on the hero block.

Phoebe Lown’s voice-of-customer work for long-form sales pages has been making this point for years. The interview is not a data-collection step. The interview is the thing the page is made out of. The find customer pain points walkthrough drills into the specific kind of sentence the interview produces and the summary does not.

When is AI summary the right tool, actually?

AI summary earns its seat on two jobs.

  • Topic mapping across a large corpus. What themes come up, how often, in what clusters. An AI model does that faster than a human, at higher recall.
  • First-pass pattern detection before the interview. Run the AI model across three hundred reviews. Pull the five loudest themes. Use those themes to structure the interview guide. That is not replacement. That is preparation.

The cmswire piece covers the shape in one line. "Automation can and, in many cases, should handle the simple. Humans, however, must own the complex, especially when emotions and trust are on the line."

Copywriting is the trust part. The summary prepares the interview. The interview writes the page.

What does a two-track VOC workflow look like?

The workflow splits cleanly once you stop asking AI to do both jobs.

Track one is AI-driven. Dump the review corpus, the support transcripts, and the sales-call recordings into an AI model. Ask for top themes, clusters, and sentiment. Use that output to write the interview guide and to decide who to interview.

Track two is human-to-human. Run five to eight interviews with the customers the summary pointed you toward.

Record them. Transcribe them. Do not summarize them.

When you draft, the summary lives on the outline page. The transcripts live in the sentences. That is the handoff.

Pattern informs structure. Voice writes copy.

The koji.so report is explicit about the scope boundary. The report’s own phrasing: strategic research planning, stakeholder communication, and synthesizing research into product direction still benefit from human judgment. In copywriting, the voice layer is where that judgment sits.

Other questions worth answering

How do I know when I have heard enough to stop scheduling more calls?

Roughly three to five conversations, usually. The signal is the same phrases repeating across different mouths. That repetition typically lands around the fifth call, sometimes the third. Eight is the upper end across most Zoom or Otter transcripts, and once you hear it twice, schedule one more and stop.

Where should I find people willing to sit for a 20-minute call?

Three pools work. Past buyers who replied to a rating prompt or a thank-you email rank first. Active users who answered in-app messaging come second. People who left long written feedback round out the set, and most of them say yes to a Calendly link and a small gift card.

How do I handle transcripts so the buyer’s exact wording survives to draft time?

Per koji.so’s April 2026 analysis, turnaround on AI-assisted qualitative work fell to 24-48 hours, so transcription is rarely the bottleneck. Open the transcript and copy the sentences that made you stop reading. Paste them into a single text file, labeled by speaker and date, untouched. The reworded sentence loses the catch you wanted to keep.

How often should I refresh buyer-language research as the market changes?

Annually for stable products, every six months in fast-moving categories. The koji.so April 2026 analysis found roughly 87% of product teams using AI interview platforms lifted their research pace three times or more. The wording your buyers use shifts as the category matures. Redo the work when the page stops sounding like the people it is written for.

What can you only get from a real customer interview?

Pick one active landing page that is not converting. Pull the last sixty reviews or support transcripts. Run a summary through an AI model — cluster by theme, count, and sentiment.

Then book three short interviews. Twenty minutes each. Bring two questions the AI summary pointed you toward and one question it did not.

Rewrite the page using the summary for structure and the transcript verbatims for every line of actual copy. Compare to the current version. If the new copy still reads like the old copy, you summarized instead of quoted.

If you want a second set of eyes on a page and the voice-of-customer work behind it, contact me here. I will tell you where the page is summarizing and where it is quoting. And what a reader hears when it summarizes.

The chord chart is useful. A guitar in somebody’s hands is the thing a page is actually made of.

Similar Posts