UserInsight
All posts
Feedback

How to Analyze Open-Ended Survey Responses (Without Coding Every Answer by Hand)

Open-ended survey responses hold the real answers and almost nobody reads them. Here is how to code, theme, and quantify free-text answers at scale, by hand and with AI.

By the UserInsight team

July 2026 · 9 min read

To analyze open-ended survey responses, you read the answers, assign each one a short label (a code), group those codes into themes, and count how often each theme appears so you can treat the text like data. That is thematic analysis, and done well it turns a pile of comments into a ranked list of what matters. The catch is scale: coding a few hundred answers by hand is slow and inconsistent, and coding a few thousand is a job nobody finishes. Here is the manual method in full, where it breaks, and how to do the same thing with AI without losing the rigor that makes it trustworthy.

Why open-ended responses are worth the trouble

Closed questions analyze themselves. You can chart a CSAT score or an NPS split in a spreadsheet in minutes. The open-ended box, the one where a customer explains in their own words what went wrong, is where the actual reason lives, and it is the part most teams skim once and file.

That is backwards. A score tells you something moved. The free-text answer tells you what to do about it. "NPS fell four points" is a fact you cannot act on. "NPS fell four points, and 214 detractors describe the same failed import step" is a roadmap item. The whole value of an open-ended question is in reading it, which is exactly the step that gets skipped.

The manual method: thematic coding, step by step

If you have under a hundred or so responses, doing this by hand is perfectly reasonable and worth understanding even if you later automate it. The process is the same one qualitative researchers have used for decades, and it is what dedicated research repositories automate; if that is the shape of your problem, our Dovetail alternative comparison covers where those tools help and where they stop.

1. Clean the data first

Export the responses and remove the noise: blank answers, "n/a", test submissions, duplicates, and obvious spam. Fix nothing about the wording itself, but get rid of anything that is not a real answer, because it will distort your counts later.

2. Read a sample before you code anything

Read fifty or so answers without labeling them. You are looking for the natural shape of the feedback: the recurring subjects, the vocabulary customers use, the range of sentiment. This stops you from imposing categories that do not fit the data.

3. Build a code frame

A code is a short label for what an answer is about: "pricing confusion", "slow support", "missing export", "loved the onboarding". Your code frame is the list of them. It can be flat (a simple list, faster to apply) or hierarchical (themes with sub-codes, more powerful but slower). Start flat unless you have a reason not to.

4. Code every response

Go through the answers and assign one or more codes to each. Some answers carry two ideas ("support was fast but the fix did not work") and should get both. Add new codes as genuinely new subjects appear, but resist letting the frame sprawl; if you end up with eighty codes, most are variations of a dozen real themes.

5. Group codes into themes and count

Roll related codes up into themes, then count how many responses fall under each. This is the payoff: instead of a wall of text you have "pricing confusion: 214 responses, mostly negative" and can rank themes by size and sentiment. Sort by the biggest negative clusters and you have your priority list.

6. Read the answers behind the top themes

Before you take anything to a decision, open the biggest themes and read a handful of the raw answers inside each one. This is the verification step, and it is where you catch a mislabeled cluster or a nuance the count hid. A theme you cannot trace back to real sentences is not one you should act on.

Where the manual method breaks

Thematic coding is rigorous, and at any real scale it falls apart for three reasons. It is slow: coding a thousand answers well is days of work. It is inconsistent: two people rarely code the same answers the same way, and the same person drifts over a long session. And it does not repeat: by the time you have finished last quarter's survey, this quarter's is already open, so it becomes a one-off exercise instead of a continuous signal. Teams that escape this usually stop running quarterly blasts altogether and switch to in-app surveys that trickle in a steady, analyzable stream of answers instead.

Word clouds and keyword counts are the usual shortcut, and they are worse than nothing, because they strip out meaning. A keyword counter reads "not easy to find" as being about ease, and "I would never recommend this" as an endorsement of recommending. Frequency is not meaning. You need something that reads the sentence in context, which is exactly the line modern text analytics tools cross when they cluster by meaning rather than by matching words.

Analyzing open-ended responses with AI

This is the part that changed. Language models are genuinely good at the coding step: they read each response, work out what it is about, group answers that mean the same thing, and score sentiment, including the cases keyword counts get wrong. What used to be days of manual tagging runs in minutes, and it does not drift the way a tired human coder does.

The workflow that holds up is a hybrid: let the tool do the first pass, then spend your time on judgment instead of tagging. Concretely:

  • Feed in all the responses and let the software cluster them into themes without you defining the code frame up front.
  • Sort the themes by size and by sentiment.
  • Open the largest negative clusters and read a sample of the underlying answers to confirm the label is right.
  • Take the verified themes forward, and ignore the long tail of tiny one-off clusters unless one is strategically important.

The one thing you must not give up when you automate is traceability. Accuracy you cannot check is not accuracy you can act on. Any theme the software shows you should link straight back to the exact responses it was built from, so you can spot a misgrouped answer and confirm a finding before it drives a decision. This is the line between an AI tool you can trust and a black box that hands you a chart you have to take on faith. Purpose-built survey analysis software treats that link as non-negotiable, so every clustered theme is one click from the raw answers behind it.

Turn themes into decisions

Analysis is not the goal; a decision is. Once you have verified themes with counts and sentiment, the work is to route each one to whoever can act on it. A theme about a confusing signup flow is a product problem. A theme about slow responses is a support-process problem. And when the recurring complaint is that people cannot tell what your product does, that is a conversion problem you can audit the page copy and layout for directly, rather than a survey problem at all.

The strongest version of this connects the answers back to behavior. A theme is far more actionable when you can see not just that 214 people mentioned a failed import, but that those same accounts stalled at the same step and half of them churned within a month. That join between what people said and what they did is what turns survey analysis from a reporting exercise into a prioritization engine.

Do it continuously, not once a quarter

The last shift worth making is from a periodic analysis push to a continuous one. When coding was manual, analysis had to be a project you scheduled. When it is automated, new responses can be themed as they arrive, so a problem shows up in the data within days of appearing rather than at the end of the quarter. That turns your open-ended answers from a retrospective report into an early-warning system, which is what they should have been all along.

The comments are usually the most valuable and least-read dataset a company owns. The only thing standing between you and them was the cost of reading, and that cost is gone.

Whether you code by hand or use a model, the principle does not change: read the answers, group what repeats, count it, and verify against the source before you act. The tools got faster. The discipline is the same.

See UserInsight surface the why

UserInsight unifies your usage, feedback, tickets, reviews and surveys, then surfaces why users churn and what to build next, each traced to the evidence. Aggregate and consented, with no PII.

Stop guessing why users churn

UserInsight unifies your usage, feedback, tickets, reviews and surveys, then AI surfaces why users churn and what to build next, on the sources where your user data already lives.

Usage, feedback, tickets, reviews & surveys · Traced to evidence · No PII

Aggregate and consented data · GDPR-friendly · for product, growth and CX teams.