- Contents
Open-ended survey questions, in-depth interviews, and focus group discussions produce the kind of data that numbers alone cannot capture. A respondent telling you why they do not want a vaccine is worth more than a hundred respondents ticking a box. But that richness comes at a cost, in money and time.
Before any of that text becomes an insight, someone has to code it.
Qualitative coding is the process of reading through unstructured responses and assigning labels, or codes, to segments of text so the data can be organized and analyzed. A response such as “the clinic was far and the queue took all morning” might be coded for distance, waiting time, and access to services. If you multiply that by 20,000 responses in five languages, you begin to see the problem.
We have covered the fundamentals of this process before in our guide to coding qualitative data. In this post, we look at what AI adds to the workflow.
How data coding works
There are two main approaches to coding qualitative data:
- Deductive coding begins with a codebook built in advance, usually from the research objectives or an earlier wave of the same study. Coders apply the existing labels to incoming responses. It is quicker and keeps findings comparable across waves, but it can miss themes nobody thought to look for.
- Inductive coding begins with the data. Coders read the responses, allow categories to emerge, and build the codebook as they go. It catches the unexpected, which is often where the real insight sits, but it takes longer and depends heavily on the coder.
In practice, most studies combine the two: a starting framework drawn from the objectives, then expanded as the data reveals themes the framework did not anticipate.
Why manual coding can be a bottleneck
The method itself is sound, but the constraint usually is human capacity.
An experienced researcher can code 500 responses in an afternoon. However, coding 50,000 responses across six countries is a different ballgame. And as volume grows, three problems compound.
The first is drift. The same coder applies a code slightly differently at response 4,000 than they did at response 40, usually without realizing it.
The second is disagreement. Two coders reading the same response will not always reach the same conclusion, which is why inter-coder reliability has to be tested, trained for, and tested again. That is more time on top of the coding itself.
The third is language. Taking the countries GeoPoll works in as an example, a single study might collect responses in Swahili, French, Spanish, Hausa, Arabic and English. Sometimes more. Each language needs coders fluent in it, and maintaining consistency across those teams is harder still.
Most researchers are familiar with the result – it takes a long, long time.
What AI brings to the process
As we have experienced at GeoPoll, large language models are good at the exact task coding requires: reading a passage of text and determining what it is about. But AI coding quality depends far less on the model itself than on how you set it up and manage it. Pointing a general-purpose chatbot at 20,000 responses and asking it to “find the themes” will rarely produce accurate output.

Used well, AI coding is a structured process. It starts with a codebook grounded in the research objectives and clear instructions that define each code, with examples of what belongs under it and what does not. Where the volume or specialization justifies it, models can be fine-tuned on previously coded data from similar studies, so they learn the categories and language patterns that matter in a given sector or market. Prompts are tested and refined against a hand-coded sample before the full dataset is processed. And throughout, quality control stays in human hands: researchers check agreement rates, review low-confidence and edge cases, and audit a portion of the output on every run. Generally, experienced experts have to properly create and curate the models to function as they would.
When those pieces are in place, AI transforms qualitative coding, with several benefits:
- Speed: Thousands of responses can be coded in the time a team would spend on a few hundred. Analysis begins while the topic is still current.
- Consistency: A model applies the same codebook logic to the first response and the hundred-thousandth. It does not tire or lose focus, which removes the drift that manual coding has always had to absorb.
- Multilingual coverage: AI can code responses across languages without assembling a separate coding team for each one. For multi-country studies, this is often the single biggest gain.
- Theme discovery: Models can group responses and surface recurring patterns that a predefined codebook would have missed, including minority themes that tend to get flattened when a human coder is working at speed.
- Depth: Sentiment, intensity, and the reasoning behind a response can be captured alongside the code itself, giving researchers the why and not only the what.
- Cost: The cost of coding at scale falls sharply. This changes what is worth asking, and teams stop rationing open-ended questions to keep analysis manageable.
The challenges researchers should plan for
Like we have been saying, AI-powered research is not a substitute for research rigor. AI-assisted coding is no exception, because there are weaknesses to address.
- Nuance can be missed: Sarcasm, local idiom, and indirect phrasing can be read literally – a response that means the opposite of what it says can be misleading.
- Models can be confidently wrong: An undertrained AI will assign a code it cannot justify as readily as one it can, and the output looks identical either way.
- Bias can be inherited: Models reflect the patterns in their training data. In studies covering underrepresented populations and languages, this can determine which themes surface and which do not.
- Themes can be over-flattened: Aggressive grouping can collapse genuinely distinct ideas into one tidy category, losing the specificity that made the open-ended question worth asking.
- Data protection matters: Open-ended responses may often contain personal detail. Any AI workflow handling them requires informed consent, proper anonymization, and compliance with applicable data protection law.
- Transparency is not optional: If you cannot explain how a code was assigned, you will struggle to defend the finding to a client, a donor, or an ethics committee.
Best practices for AI-assisted coding
Keep a researcher in the loop. The goal is AI-assisted coding, not AI-only coding. Researchers should set the research frame, decide what the codes mean, review the output, and own the interpretation of the findings. The model handles the volume, but the judgment calls about what a theme means for the client, and whether a finding holds up, stay with people who understand the study and its context.
Start from a defined codebook. Give the model a clear framework tied to the research objectives rather than asking it to invent categories on its own. Each code should have a plain definition, a description of what belongs under it and what does not, and a few real example responses, including borderline ones. Where codes sit close together, such as the cost of a service and the cost of getting to it, spell out the difference. Let the model propose new themes as they emerge, but have researchers review and approve each addition, and keep a version history so you know which codebook was applied to which data.
Validate against a manual sample. Before processing the full dataset, have experienced coders hand-code a sample that reflects the range of countries, languages, and respondent groups in the study. Run the model on the same sample, compare the results, and measure agreement for each code rather than relying on a single overall score, which can hide a code the model consistently gets wrong. Where the two disagree, find out why, refine the instructions, and test again. Agree on an acceptable level of agreement before you start, and treat the model exactly as you would a new coder joining the team.
Write explicit instructions. Vague prompts produce vague codes. A good coding prompt reads like the briefing you would give a trained coder: it explains the research objective, the question respondents were answering, the full codebook, and how to handle responses that fit more than one code or none at all. It should include worked examples of difficult cases such as sarcasm, negation, and passing mentions. Asking the model to return the exact phrase that supports each code, along with a confidence level, makes errors easier to spot and speeds up review considerably.
Code in the original language. Translating first and coding second loses nuance twice over, once in translation and again in coding. Code responses in the language they were given, then translate the output for reporting. Test the model’s performance in each language separately, since a model that codes English and French well may struggle with Hausa or Amharic, and account for code-switching and informal language such as Sheng or Pidgin, which are common in open-ended answers. Native-speaker researchers should be part of the review process.
Review the edges. Check low-confidence codes, rare themes, fallback categories, and outliers closely. This is where errors concentrate, and it is often where the most interesting findings sit. Audit a random sample of high-confidence codes as well, since a model can be wrong with full confidence. Compare results across subgroups and batches, and look inside large themes to make sure they have not merged distinct ideas, such as staff attitude, waiting times, and stock shortages, into one broad label.
Document the method. Record the model and version used, the codebook and prompt versions, the validation results, and the human review process followed. Clients, donors, and ethics committees will ask how findings were produced, and documentation also makes it possible to reproduce results or apply the same setup consistently in the next wave of a tracking study. Data handling belongs here too: note how personal information was removed and where responses were processed.
Keep the verbatims. Codes are an abstraction of what people said, not a replacement for it. Keep every code linked to the original response so any finding can be traced back to the words behind it. Reading the verbatims behind the major themes is also how researchers move from counting mentions to understanding what respondents actually mean.
How GeoPoll approaches AI-Powered Coding
GeoPoll has collected data in Africa, Asia, and Latin America for over a decade, in markets and languages where fieldwork is most demanding. Open-ended questions have always been part of that work, and so has the coding bottleneck that follows.
Over the last few years, we have been training models in the methods we use, and the languages we work with, no matter how underserved they are. One result is GeoPoll Senselytic, our AI-powered qualitative layer built to transcribe open responses, analyze them for sentiment, themes, and patterns, and deliver results without requiring a coding team in every market. Our experienced researchers still control the codebook and the interpretation, to ensure the results are as actionable as they should be.
If coding open-ended responses is the step holding your projects back, contact us to discuss how Senselytic can fit into your research.
