
· 6 min read
Beyond the Likert Scale: Why Synthetic Research Needs Text-First Methods
Kira Warje
When researchers first tested LLMs as survey respondents, they followed the standard practice in consumer research: asking questions about products and eliciting direct numerical ratings on a Likert scale (ranging from 1 to 5).2 While real human responses naturally vary, LLM ratings clustered around the middle of the scale, showing less variation than real human surveys even though average LLM scores tend to correspond closely with human averages.3
LLM respondents may be suspiciously agreeable when asked for numerical ratings, but this doesn’t mean they’re fundamentally flawed.2 Where LLMs fall short is in distributional similarity; direct elicitation squeezes survey responses into middle-heavy distributions lacking realistic spread, flattening out the differences researchers rely on.
Why Direct Ratings Fail
There are two main reasons direct ratings produce overly narrow or systematically skewed distributions in synthetic research.
- Trained to agree: most models are specifically trained to be helpful and inoffensive, so they often hedge with noncommittal answers. When forced into a number on a Likert scale, this results in the numerical equivalent of “maybe.
- Skipped reasoning: Real survey respondents form language-based reactions to questions beforeassigning them a numerical rating. Asking a synthetic respondent for a number skips this important language-processing step, ironically an area in which LLMs are particularly adept.
Together, these contributing factors suggest that the problem with synthetic survey respondents lies in the questions we’re asking them. Forcing LLMs to pick numbers on a scale demands a single, discrete output where a genuine distribution of opinion should exist. Direct Likert ratings are simply the wrong format for generating human-like variation in synthetic research.
The Fix: Let LLMs Talk, Then Measure The Meaning
To address these challenges, Maier et al. propose an alternative called semantic-similarity rating (SSR): models answer questions in ordinary language first, then that free-text response is mapped onto a 5-point Likert scale.2 The mapping process works by comparing the response against a set of predefined reference statements for each scale point, such as “I would definitely buy this” for 5 and “I would never buy this” for 1. This is done through a machine-learning technique called embedding similarity, which mathematically measures how similar two pieces of text are in meaning.

Importantly, response similarities map to a probability distribution along the Likert scale rather than a single number. For example, an answer like “I’d probably try it because it’s easy to use and priced affordably” might map mostly to a 4, with some weight on 3 and 5 as well. Across a survey panel of synthetic respondents, this probability distribution allows SSR to closely replicate the shape and spread of real survey data.
SSR isn’t the only text-first method that’s been developed to address the failures of direct ratings; however, strategies like follow-up Likert rating fall short by ultimately collapsing free text back into a single integer. SSR, in comparison, mimics a real human moderator in an interview or focus group; it listens to an open answer from a synthetic respondent, then codes it into analyzable data.
While its use in synthetic survey data is relatively new, the concept of SSR has been around for decades. In the early 2000s, King et al. had survey respondents rate a set of hypothetical scenarios alongside their own subjective answers, then used differences in how people rated those same scenarios to statistically standardize responses onto a common scale.4 SSR methods like these remove the step where people (or LLMs) are forced to translate their own feelings to a single number themselves. Instead, external references become the measuring instrument, preserving the natural variation in responses.
The Evidence: Text-First Elicitation Maintains Realistic Response Distributions
Using existing consumer research data, Maier et al. put SSR to the test against 57 real personal care product surveys totaling 9,300 human responses.2 While direct rating LLMs already reached around 80% of human test-retest reliability on product rankings, SSR bumped that number up to 90%. More critically, SSR produced response distributions that more closely matched those of the real human surveys, quantified by high Kolmogorov–Smirnov (KS) similarity: a measure of per-survey similarity between synthetic and real distributions Additionally, because respondents provided rich language-based answers, every rating was accompanied by a readable rationale, producing both qualitative and quantitative data at the same time.
What This Means for Users
Although the use of SSR in synthetic research is in its infancy, studies like Maier et al.’s already show it to be a validated and practical complement to traditional research. Because SSR preserves the natural variation of open-ended responses, it’s well-suited to research questions where the shape of distributions matters as much as the average, such as identifying polarizing opinions or catching subtle shifts in consumer sentiment. Research methods that flatten this spread can make polarizing ideas look safe or hide changes in opinion that would otherwise prompt proactive decision-making.
With strong methods for mapping human distribution shapes, synthetic research even offers advantages over traditional research methods. Human surveys typically produce numerical ratings with no textual explanation, while interviews result in quotes with no standardized ratings. Text-first elicitation using SSR provides both the number and the free-text rationale, allowing results to be analyzed and interrogated. This feature makes synthetic research particularly useful early on, allowing for quick sanity checks before investing in full-scale human studies.
SSR in Artificial Populations
As synthetic respondents earn their place in research, the field is building a new generation of tools focused on asking the right questions and validating their answers. From The Decision Lab, Artificial Populations is one early example that allows researchers to run persona-accurate research panels in which synthetic participants interact and engage in natural conversation. Using SSR-style text-first elicitation, these panels are designed to better mirror the spread of real human data, avoiding clusters of uniform middle scores and instead surfacing the nuanced opinions that make survey data genuinely useful.
Sources
- Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J., Rytting, C., & Wingate, D. (2022). Out of one, many: Using language models to simulate human samples. ArXiv. https://doi.org/10.1017/pan.2023.2
- Maier, B. F., Aslak, U., Fiaschi, L., Rismal, N., Fletcher, K., Luhmann, C. C., Dow, R., Pappas, K., & Wiecki, T. V. (2025). LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings. ArXiv. https://arxiv.org/abs/2510.08338
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. Political Analysis, 32(4), 401-416. https://doi.org/10.1017/pan.2024.5
- King, G., Murray, C. J., Salomon, J. A., & Tandon, A. (2004). Enhancing the validity and cross-cultural comparability of measurement in survey research.American Political Science Review, 98(1), 191-207. https://ssrn.com/abstract=1082804
- Magnific. (n.d.). Digital mixer in a recording studio, close -up [Photograph]. https://www.magnific.com/free-photo/digital-mixer-recording-studio-close-up-concept-creativity-show-business-space-text_10108520.htm