· 6 min read

Three Synthetic Research Fixes That Look Right (But Fail)

Kira Warje

At first glance, output from synthetic research feels convincing. The numbers are plausible, the reasoning is logical, and the responses sound like things real people would say. Because synthetic data reads beautifully, it's easy to trust the output on sight. However, a research instrument can have face validity, meaning it looks right on the surface, but still be wrong. Recognizing the problem doesn’t automatically lead to good solutions, as seemingly sound fixes can still produce misleading output. Follow-up Likert rating (FLR), ipsatization, and ad hoc Likert scales each address a flaw in synthetic research, but end up flattening distributions, muddying comparisons, and generating output that looks right without actually measuring what it’s meant to.

Case 1: Follow-up Likert Rating Smooths Distributions

In synthetic research, direct Likert ratings produce unrealistically compressed distributions as synthetic responses cluster around noncommittal, middle-heavy values.1 This occurs because asking models for direct numerical answers sidesteps the thought processes that reflect accurate human opinions. Real survey respondents form language-based reactions to questions before assigning them a numerical rating. When someone is presented with a product and asked about their purchase intent, they might think “I’d try it if it were on sale” or “I’m really not a fan of this scent,” but they’re unlikely to pick a number without first processing this reaction. Asking a synthetic respondent for a number skips this important language-processing step, ironically an area in which LLMs are particularly adept. In FLR, the synthetic respondent first writes a free-text statement about its purchase intent. That text is then passed to a second instance of the same model, prompted to act as a “Likert rating expert,” which converts it into a single Likert score. While FLR restores the missing free-text step from direct Likert ratings, it reintroduces the original problem by reducing the text to a single integer. The second instance of the model, now taking on a scoring role, is still being asked for a direct number, so it flattens natural variation and gravitates toward "safe" middle values. A response like "I'd probably try it, but the price makes me hesitant" could easily be interpreted as a 3 or 4, and this spread carries important information about purchase intent.

To read more about how synthetic research can meaningfully preserve distributions with text-first methods, read our article on Semantic Similarity Rating (SSR).

On the surface, FLR can look like a better alternative to direct Likert rating because it restores the crucial form of the open-ended response, but it still loses crucial information to a flattened distribution.

Case 2: Ipsatization Strips Meaningful Differences from Data

Not every fix fails at the elicitation stage, as solutions to response biases like acquiescence can introduce new problems after responses are already collected. Acquiescence bias, the tendency to agree with survey statements regardless of content, is well-documented in humans and appears to affect synthetic survey respondents as well. Researchers have spent decades developing statistical methods to partial out acquiescence from surveys by estimating how much of each respondent's "agreeableness" reflects a general bias. One long-standing method, ipsatization, subtracts a respondent's overall mean score across all survey items from each answer, but this technique is known to overcorrect and distort comparisons, so it functions more as a blunt tool than a precise fix.2

Check out our article on acquiescence bias to learn about survey design solutions and statistical methods to correct the “yes-lean” in synthetic research.

Like FLR, ipsatization is an imperfect solution that appears to work because it addresses the target issue and produces cleaner output, but it also removes the real between-respondent differences that are necessary for meaningful comparisons.

Case 3: Custom Research Instruments Manufacture Validity

Just as statistical corrections can look good at face value, similar misdirections can show up earlier, at the instrument design stage. Fluent synthetic responses often create a false sense of construct validity. For example, an ad hoc 7-point Likert scale invented for a single study might appear fine, even if it has no established construct behind it. This is because language models are optimized to produce output that's coherent, confident, and internally consistent. More than a mere weakness, face validity in synthetic research can be actively misleading because the instruments are fine-tuned to look good regardless of what they're actually measuring. Coherence does not equal validity, meaning synthetic research might hold together internally but still fail to measure the specific traits it claims to measure.

Read about evidence-based methods to validate synthetic research instruments in our article on construct validity.

Ad hoc scales with face validity exemplify a common failure pattern in synthetic research: response generation methods, statistical corrections, and survey instruments can all appear rigorous while still falling short of accurate measurements.

The Common Thread: Form Without Function

The gap between looking right and being right can show up at every layer of a synthetic study. The above synthetic research solutions prioritize form over function at each level: FLR at elicitation, ipsatization at correction, and ad hoc scales at instrument design. All three are optimized to produce output that looks correct rather than a measurement that is correct. The good news is that researchers have already built methods to address these gaps. Solutions like semantic similarity rating (SSR), evidence-based survey design, and established construct validity measures give research platforms a solid foundation for generating valid synthetic data. This methodological rigor is what separates genuine AI-powered research instruments from chatbots dressed up as survey tools.

How Artificial Populations Closes the Gap

Artificial Populations was designed with the same scrutiny that exposed gaps in the imperfect fixes above. Where an ad hoc scale lacks a foundational construct, Artificial Populations anchors measures to previously validated, peer-reviewed instruments. Where a single model’s fluent answer creates the illusion of validity, Artificial Populations cross-checks results against independent AI providers, assigning each finding a confidence rating. Where a response strategy or statistical correction flattens data variance, Artificial Populations flags where fidelity is weaker so researchers can investigate further. As the field of synthetic research continues to evolve, Artificial Populations evolves alongside it, studying the gap that separates a real solution from a convincing but imperfect fix.

Sources

  1. Maier, B. F., Aslak, U., Fiaschi, L., Rismal, N., Fletcher, K., Luhmann, C. C., Dow, R., Pappas, K., & Wiecki, T. V. (2025). LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings.ArXiv. https://arxiv.org/abs/2510.08338
  2. Chan, W. (2003). Analyzing ipsative data in psychological research.Behaviormetrika, 30(1), 99-121. https://doi.org/10.2333/bhmk.30.99
  3. Magnific. (n.d.). Front view beautiful composition of different books [Photograph]. https://www.magnific.com/free-photo/front-view-beautiful-composition-different-books_12892768.htm

Test ideas with participant-
accurate Artificial Participants

Powered by behavioral science.

Decision-ready insights without recruitment delays or bias drift.

Three Synthetic Research Fixes That Look Right (But Fail) | Artificial Populations