A step-by-step way to work
Review questions for bias first
Paste your draft survey and ask AI to flag leading, double-barrelled, loaded, or assumption-laden questions.
Check the scales and order
Ask whether your response scales are balanced and whether question order could bias answers.
Pilot the wording
Have it suggest clearer phrasings for any question a respondent might misread.
Code open-text at scale
After fielding, feed open-ended responses and ask for a thematic coding pass with example quotes — a major time-saver.
Keep significance to yourself
Use AI for question quality and qualitative coding, not for deciding whether a result is statistically meaningful.
Real prompts, and what good output looks like
Review these survey questions: 1) 'How satisfied are you with our fast, reliable service?' 2) 'Do you use our mobile and desktop apps?' (Yes/No)
1) Loaded — 'fast, reliable' primes a positive answer and assumes both are true. Fix: 'How satisfied or dissatisfied are you with our service?' with a balanced scale.
2) Double-barrelled — a Yes/No can't capture someone who uses only one. Split into two questions, or offer 'mobile only / desktop only / both / neither'.
Both fixes will materially change your results.
Here are 200 answers to 'What's the main reason you'd recommend us?' [pasted]. Code them into themes with counts and a representative quote each.
Ease of use – 71 responses – “everything is where I expect it to be.”
Customer support – 48 – “they actually answer fast.”
Price – 33 – “cheaper than what I had.”
Other/unclear – 48. Note: 'ease of use' and 'support' overlap in ~10 responses that mention both.
A prompt you can reuse
Fill in the highlighted parts and paste it into your AI tool of choice. Edit the output — it's a starting point, not a finished answer.
Help me design and analyse a survey. Don't make significance calls — just question quality and coding. [Design] Draft questions: [paste] Please flag leading, double-barrelled, loaded, or assumption-laden questions; check that scales are balanced; and note any order effects. Suggest cleaner rewrites. [Analysis] Open-text responses to '[question]': [paste responses] Please code into themes with counts and one representative verbatim quote per theme; flag overlaps and an 'other/unclear' bucket.
What AI gets wrong here
Treating coded counts as significance
Theme counts describe your respondents, not your whole user base, and small differences may be noise.
Missing sampling bias
AI can't see who didn't respond, and survey-takers differ from non-takers.
Over-trusting auto-coding
AI mis-files ambiguous responses into tidy buckets.
AI makes survey design tighter and open-text analysis dramatically faster, but a survey's validity rests on who you asked and whether a result is real — questions of sampling and statistics that AI's confident theme-counts quietly sidestep. Use it to improve questions and code responses; reserve the judgment about what the numbers mean for yourself.