Silicon sampling (RnD heading)
Silicon sampling (Rubric #RnD)
Today I spoke with a colleague who is engaged in a platform for conducting qualitative and quantitative research. (||Hey, Andrew.||). In the course of the conversation, I remembered about this interesting and rapidly developing approach to conducting them, where the LLM receives a sociodemographic “biography”. (backstory) She answers questions on her behalf. This allows you to reproduce the opinions of thousands of demographic groups without recruiting real participants. Then I became interested, and what are the key studies in this topic and so I got a selection below.
- Argyle et al., 2023 - "Out of One, Many" It is a classic foundational work that introduces the concept of "algorithmic accuracy." (algorithmic fidelity). GPT-3 It uses backstories from ANES and Pew. (later came out Political Analysis from Cambridge)
- Aher, Arriaga, Kalai (ICML 2023) - "Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies" Replication of classical sociological experiments (Milgram, Ultimatum Game and others.) A population of LLMs without a single live participant.
- Santurkar et al. (2023) - "Whose Opinions Do Language Models Reflect?" - Created a dataset. OpinionQA from 500 Controversial issues, compared 9 LLM with 60 American demographic groups. Conclusion: The discrepancy is comparable to the gulf between Democrats and Republicans on climate.
- Hu & Collier (2024) - "Quantifying the Persona Effect- the person explains less 10% variance, but gives a statistically significant increase. A linear relationship was found: the stronger the correlation of variables of a person with an opinion, the more accurate the simulation is.
- Tejaswani Dash et al (2025) - "Polypersona: Persona-Grounded LLM for Synthetic Survey Responses" open-source framework with dataset from 3 568 reply 433 unique persons, designed for instruction-tuning
- Taday Morocho et al. (2026) - "Assessing the Reliability" - tested 70 000+ pairs on World Values Survey data. Persona prompting does not give a stable improvement and in some cases worsens the result.
The approach looks like a wunder waffle, but there are some problems. - Ordering bias (Models systematically prefer the first answer
- Underrepresented groups - 65+, widowers, minorities reproduce worst Korrelation blindness Fine-tuned models are more marginal, but both techniques do not restore the latent correlation structure of the real population. In whitepaper Lin (2024) inSix Fallacies in Substituting Large Language Models for Human Participants“typical errors in the interpretation of such studies are systematically listed.
In general, if you make a platform for quantitative and qualitative research, then it makes sense to add it.silicon sampling" as a research approach:)
#AI #Engineering #RnD #Science #Research