Humans are, famously, the only animals that experience separately the concepts of what they want and what they want to want, a concept described and popularized by philosopher Harry Frankfurt in Freedom of the Will and the Concept of a Person as the difference between humans and animals or “wantons”, which do not have the capacity for second-order volitions. Such volitions manifest not only in what humans want, but often in how they behave. Examples of the chasm between these desires are ubiquitous:
A human may want to eat ice cream for dinner, but want to be the kind of person who wants a salad for dinner.
A human may want to buy a house in an area with little new construction or bans on further development, but want to want legislation that enables the construction of affordable housing
A human may want to bet on sporting events every weekend, but want to want to put $500 from every paycheck towards retirement
These sets of wants compete to determine actual behavior. In some cases, like the first, an individual who wants ice cream for dinner is motivated by a second-order volition to be the kind of person who eats a salad and buys Sweetgreen for dinner instead of a pint of ice cream. In other cases, like the second or third, an individual’s conscious or subconscious actions may actively betray their moral, political, or financial principles if they buy a house in a NIMBY-zoned neighborhood or bet away their savings.
From the outside, however, any distinction between a person’s first- and second-order desires is often not observable. The person who chooses a salad might love vegetables; the person who places large wagers on sporting events might not be worried about saving. What people do does not often reflect what they want, and, critically, what they say they want does not always reflect what they will actually do.
When asked to articulate their desires explicitly, studies show that humans consistently answer in accordance with their second-order volitions, expressing themselves as the people they aspire to be, not necessarily the people they are. In studies of survey design and psychology, this phenomenon is referred to as social desirability bias.
Polling and surveying human populations, therefore, has historically existed as a reflection of human higher-order desires, but not necessarily a strong predictor of current or future human behaviors. The advent of more advanced AI modeling techniques allows researchers to answer the same questions that traditional polls target, but exposes two different sets of answers: predictions of what humans might say they want, and simulations of how humans might actually behave, as reflected by their online behavior, preferences, and personalities. The difference between these sets of answers, and whether AI-polling companies can predict them accurately, lays the groundwork for meaningful changes in how companies, policymakers, and campaigns answer questions and make decisions.
The Rise of Synthetic Surveys
AI-generated poll results, also known as silicon sampling, synthetic audiences, or simply synthetic research, simulate the responses or reactions of particular population demographics to specific questions or events. The growing ecosystem of commercial synthetic survey firms allows companies or campaigns to not only conduct faster and cheaper market research, but also gives them the opportunity to forecast population behavior instead of simply survey responses.
Companies in the synthetic survey space identify this distinction explicitly, positing that “Humans are unreliable narrators of their own behavior. Memory fails, incentives distort answers, and social pressure warps what people say away from what they actually do.” These companies argue that measuring market outcomes and population behavior enables model predictions to be more accurate than traditional surveys. This distinction is similarly framed as offering tools that can answer “what will actually happen if we do this?” instead of simply “what do people think about x?”, i.e., modeling behavioral reactions instead of simply responses to questions; another company in the category, Simile, describes itself as the first platform built to simulate human behavior.
These companies point to real-world examples of the divergence as evidence; Aaru, a synthetic survey firm, partnered with client EY on a study aimed at replicating elements of EY’s 2025 Global Wealth Research Report, which included both surveyed preferences and real client behavior. The study notes that where real surveys and simulation diverged, the simulation was closer to real behavior: 82% of surveyed inheritance heirs claimed they would keep the same wealth advisor as their parents, compared to a simulated prediction that only 43% actually would, and industry studies showing the metric in reality is between 20% and 30%.
Similarly, surveys recorded that 69% of respondents wanted all their finances under one roof, whereas simulation put single-provider usage at 37%, and the real-world measurement at only 33%. The report also claims, without comparison to simulation, that survey respondents say they would pay three times more for hypothetical offerings than they will actually pay, and that 65% claim they prefer sustainable products while only 26% actually purchase them.
These differences could be attributed to surveys asking questions in a way that differs from asking individuals to predict their own behavior (i.e., many more people may hypothetically prefer buying sustainable products than can afford them), which further reinforces the thesis of synthetic survey companies that modeling behavior is more useful than surveying preferences.
While synthetic surveys are still a relatively young field in both academia and industry, they are already in use across a range of fields. In March 2026, the Public Sentiment Institute boosted a survey of 373 human respondents with 114 AI agent responses. The same month, Axios reported a finding from Aaru that “a majority of people trust their own doctors and nurses,” without noting anywhere that those respondents were models rather than patients.
Corporations are similarly integrating such technology; a paper from PyMC Labs and Colgate-Palmolive presented research showing that synthetic AI consumers tracked real purchase-intent patterns reasonably well across age and income brackets, though less consistently across gender, region, and ethnicity. The research also claimed that AI-simulated respondents gave qualitative feedback that was deeper and more critical than what actual human respondents wrote. While not confirmed, one former TikTok employee said the company employs “a fake AI viewer feed that tests your vid[eo] before humans ever see it… they literally run every post through an AI feed simulator first. If it flops there, it never reaches people.”
Efficiently modeling population behavior could have a meaningful impact on market research. Understanding how these predictions are generated, and what their limitations are, is critical to putting them in context as adoption grows.
How Synthetic Surveys Work
Synthetic surveys simulate the responses or reactions of individuals, rather than of populations at large, enabling the analysis of specific population demographics or even individual consumers. These systems rely on modeled behavior of individuals with particular demographic characteristics, including age, location, income, profession, education level, race, and religion, among others. Training these models revolves around comparing modeled behavior or responses to survey questions with real-world data.

Source: Stanford
The data used to build these models come from a variety of sources. Building a population usually begins with studies of real people, layered onto privately held records of how they behave. Simile says it retrains on a rolling basis, taking in fresh signals on how people are behaving, what the economy is doing, where prices sit, and how policy is shifting, and checking the result against live human responses at intervals.
Simile puts the volume of those checks at more than 7K a week, spread across demographic slices and real client problems, and scores them on how far the predicted spread sits from the observed one, how the magnitude of each effect compares, and how closely the predictions track what people actually did. This compounds the extent to which associated LLMs retain and train on user data, which is likely more representative of real human behaviors (what a person orders for dinner, where they search for a house, or how much money they save) than surveys or articulated preferences.

Source: Electric Twin
Advanced versions of these models continually update prediction context with streams of real-world media that personas would likely have consumed. Many companies offer clients the option to integrate proprietary customer data as well, including loyalty metrics, financial balances, or telemetry data. Synthetic surveys extend the universe of simulation-based AI research, which has gained traction within experimental physics, self-driving cars, and manufacturing line optimization.
The Synthetic Survey Landscape
Despite its infancy, there are already a variety of players in the synthetic survey landscape, each targeting slightly different approaches or swaths of customers. A review of these companies is not intended to be a competitive overview but rather to emphasize how quickly the industry has sprung up and how quickly these companies have achieved meaningful traction.
Newcomers
Aaru is an AI simulation company founded in March 2024 and valued at $1 billion as of September 2026, which is focused primarily on simulation of events and behaviors rather than survey responses. A white paper on the company’s site sets out its ambitions: “within two years, we will simulate the entire globe — from the way crops are grown in Ukraine to how that impacts production of oil in Iraq, trade through the Strait of Malacca, and elections for the mayor of Baltimore.” Aaru’s track record of political forecasts is not impressive (it gave Harris a 50.5% chance of winning the US presidential election on November 2nd 2024, a miss Fink has described in saying, “Statistically speaking, we’re within the margin of error — so we did well” and “A coin flip is a coin flip; 53-47 is not significantly different from 48-52.”
As of September 2026, the company’s published list of clients and partners includes businesses such as OpenAI, Accenture, Bayer, Comcast, EY, SharkNinja, Spindrift, and McDonald's. Dave Burwick, the CEO of Spindrift, has described the benefit of AI-assisted polling in saying, “The biggest challenge for consumer product companies is, how do you shorten that innovation cycle?” (Burwick is an investor in Aaru). In 2025, Burwick tasked Aaru with assessing soda, tea, energy drink, and smoothie products for a targeted demographic (25- to 35-year-old consumers with six-figure household incomes). The company’s simulations assessed fruit tea as the best opportunity, agreeing with Spindrift’s two-month real-world study of 500 human consumers. After the study, Spindrift launched an iced tea and squeezed fruit product based in part on this analysis.
Aaru has leaned explicitly into the distinction between human behaviors and reported preferences; the company’s website described case studies in which humans are unreliable narrators:

Source: Aaru
Simile is a synthetic survey company cofounded by Joon Sung Park in 2025, an AI communities researcher and the creator of Smallville, a 2023 virtual town of generative agents that first demonstrated the architecture required to simulate human individual and social behaviors in collectives. Unlike some synthetic survey companies, Simile agents map one-to-one to real-world individuals, grounding agent responses in answers from structured interviews with each person along with a real record of individual choices and behavioral data. This is distinct from synthetic-user products that start from a demographic profile played by a general-purpose model.
Simile raised $200 million at a $2 billion valuation in July 2026. Within five months of launching, Simile says it had run simulations numbering in the tens of millions for enterprises at the top of the Fortune list; the company’s website names CVS Health, Wealthfront, Suntory Beverages & Food, Gallup and Garnett Station Partners among its customers (CVS Health Ventures is also an investor in the company). The agents behind that work were built from 2.9 million consented responses contributed by more than 400K research participants, and CVS Health has used the platform to answer questions ranging from medication adherence to how a store should be laid out to whether a customer would give medicine to a pet.
Artificial Societies is a London-based synthetic survey company founded in 2024 to provide advice for strategic communications, including crisis and issue management, expansion to new markets, investor relations, and government messaging. The company was co-founded by James He, who led the first peer-reviewed paper on agent swarms at scale, observing how 33K chatbots behaved among themselves. His co-founder, former consultant Patrick Sharpe, agreed to form the company with him after seeing James simulate in 10 minutes the result of an experiment Sharpe had worked on for two months. The company uses both real-world surveys (mixed-method surveys, interactive focus groups, and interviews) in addition to multiverse experiments and scenario cascade simulation methods.
Other companies in this space include Electric Twin (founded in November 2023), Minds (founded in October 2025), and Evidenza (founded in January 2024). Companies with the explicit goal of testing against synthetic audiences include Replism (founded in May 2026) and Deepsona (founded in November 2025). The track record of these firms is also meaningful: Minds works with Accenture, Canva, Deloitte, and Stripe, among other customers, and Evidenza clients include ServiceNow, EY, SentinelOne, Mars, Salesforce, Toast, and SAS. Jim Lesser, the Chief Brand Officer of ServiceNow, said of the service, “I'd recommend NOT working with the Evidenza Team. The technology just might be game-changing, and we'd like to keep the competitive advantage for ourselves” (Evidenza's site carries a ServiceNow case study claiming that a year of conventional research was compressed into a single month, producing the persona work, the segmentation, and the creative brief at 12 times the usual pace).
Traditional Polling
In addition to dedicated AI polling companies, traditional polling institutes are experimenting with AI-generated data panels as a supplement to human respondents. Incumbents like Gallup, Qualtrics, Ipsos, and YouGov are all developing such functionality.
Gallup is independently validating Simile's simulated-response method. As of May 2026, its study of the method used agents generated based on interview data from 1K Gallup Panel members and benchmarked agent-based results against probability-based results. Gallup's early read was that the simulated distributions sat near enough to what people actually said for the work to be worth continuing, though it published no accuracy figure. Gallup has ruled out using simulated responses for published estimates or tracking, and has pointed to questionnaire pre-testing and hard-to-reach populations as areas of potential value.
Qualtrics launched synthetic survey panels for US consumers in March 2026, claiming to deliver results in hours at 50% the cost of human panels. The company claims that its models approximate human responses 12x better than general-purpose AI models. The synthetic survey platform offered by Qualtrics combines synthetic and human panels to enable teams to test early ideas fast and validate later-stage decisions with real test groups. Panels for the UK and Ireland, Canada, and Australia and New Zealand were slated for the first half of 2026 as of March.
As of September 2026, Ipsos focuses on synthetic data boosting, which enlarges real datasets when samples are small or uneven. The company favors tabular diffusion models, which it considers steadier than the older methods. Its SURE framework checks boosted data for statistical similarity, utility and fairness, rarity and novelty, and expert validation. Ipsos warns in the same breath that boosting done badly magnifies the static rather than the signal. The company is separately involved in a Stanford partnership studying simulated respondents.
YouGov Parallax builds AI twins of actual panelists, each anchored in several thousand recorded opinions and behaviors. Live panelists then check the findings within 30 minutes; YouGov keeps that step distinct from verification, which asks the narrower question of whether a twin answers as its real counterpart does. YouGov limits Parallax use to custom and ad hoc research, not tracking studies. It offers API and MCP access so clients' AI agents can query twins and trigger validation directly.
Comparing Approaches
Comparisons of various tools and approaches are still in their infancy. The Synthetic Research Index qualitatively compares simulation capabilities, grouped into population and behavioral simulation, synthetic respondents and personas, and synthetic research workflows (along with human-participant research systems and evaluation infrastructure).
The most important difference between systems for simulating population behavior is whether the agentic populations in each system are based on real individuals. Companies like Simile and Rehearsals base profiles on real individuals instead of synthetic combinations of demographic characteristics. These systems allow for the specification of real people, i.e., “Warren Buffett”, a given customer profile, or a selected real person based on a provided LinkedIn profile. Not all companies share exactly how the individuals in these populations are defined. The Rehearsals website clarifies, “Do people consent to being twinned? Yes. People consent directly through Rehearsals Research. Anyone can go through the participant experience themselves at rehearsalsresearch.com and review the privacy policy there.”
Other companies prioritize complete population demographic distributions over high-fidelity individual profiles. Artificial Societies, for example, aims to differentiate on network effects rather than individual accuracy, simulating how opinion moves through a social graph, surfacing second- and third-order impacts, beyond individual reactions in isolation.
Applications
Applications of synthetic surveys largely overlap with those of traditional surveys as of September 2026. Political operations stand up AI agents as stand-in voters, trying messages on them before anything reaches an actual electorate. Campaign organizations now expect pollsters and newsrooms to build synthetic cohorts for message testing among hard-to-reach voters. Civly, a synthetic survey data creator, describes one such exercise: a simulated electorate drawn from Johnson County voter-file records, polled again under each of five rival messages plus a no-message control, on a Kansas judicial amendment that reached the ballot in 2026. The company cautions on the same page that the movement it reports runs two to three times hotter than reality.
In marketing, synthetic panels have been used to test messaging and examine differences between audiences. In scenario planning, companies have used synthetic survey tools to plan product launches and market entry. For strategic communications, agents have been used to simulate earnings calls or press conferences, claiming 80% to 85% accuracy at predicting analyst questions.
Synthetic Survey Accuracy
Synthetic survey accuracy is difficult to assess for multiple reasons. First, it is not always clear, or obviously testable, whether simulations should aim to compare against real-world individual responses (which may be easier to obtain) or real-world behaviors (which may be more useful but harder to benchmark against). Second, synthetic surveys may accurately represent distribution responses while failing to accurately predict individual responses, undermining the supposed advantages of the technique.
Research from academic labs and AI polling companies documents both the achievements of synthetic surveys and their limitations. Work by researchers at Harvard, Stanford, and Transluce, published in Nature in 2026, showed GPT-4 predicting treatment effects across 70 preregistered, nationally representative US survey experiments, comprising 469 experimental effects and 119K participants, as well as human forecasters, including for unpublished studies the model could not have memorized. The paper flagged, however, that the model consistently overestimates how large effects are.
Statisticians in the field have attempted to answer the accuracy question in reverse, specifying the conditions under which inference from synthetic data is provably valid. Studies of this concept utilize task exchangeability, or the identification of past questions where real human data exists, and new questions that are similar enough to those past ones that the AI's track record carries over. One paper demonstrates this on public opinion surveys answered by AI-simulated respondents and on AI systems that grade other AI.
The survey test used Bisbee et al.'s dataset, in which GPT-3.5 imitated American National Election Studies respondents rating how warmly they feel toward groups and politicians on a 0-100 scale. With 11 thermometer questions crossed with three partisan subgroups, or 33 tasks per survey year, the authors measured the AI's errors on 2016 data and used them to predict 2020. Their ranges contained the true 2020 average for 97% of those tasks, versus only 3% for ranges built from synthetic data alone. The trade-off is precision: the typical range was 29.8 points wide. The authors concluded that synthetic data alone cannot justify a conclusion.
This view is shared by the report commissioned by the American Association for Public Opinion Research on AI use in surveys, a body of over 2K public opinion and survey research professionals. The report concludes that "a margin of error cannot be produced from synthetic responses," saying “At best, synthetic responses may function as proxies; at worst, they may systematically misrepresent the populations they purport to simulate. The field has not yet reached consensus on when, if ever, AI‑generated responses can stand in for human ones without fundamentally altering what is being measured. At a minimum, researchers should be transparent about and clearly distinguish between data derived from human respondents and data generated by AI systems.” The report names synthetic responses as the most risky use of AI in public opinion research.
Synthetic versus Real Individual Profiles
The efficacy of synthetic surveys to some applications may be meaningfully higher when simulated profiles are based on real individuals. This is particularly relevant when simulations are built from customer profiles or other sources of real-world named populations.
In the paper that formed the technical basis for Simile, first circulated as “Generative Agent Simulations of 1,000 People” and since retitled, Park’s team recruited over 1K Americans, interviewing each for two hours following the American Voices Project protocol; the interviews produced transcripts averaging 6K words per participant. The sample was stratified to approximate the adult US population on race, age, gender, education, region, and partisanship. The study used 150 core questions to benchmark simulation accuracy, each offering 3.31 response choices on average, giving a 30% chance of random accuracy.

Source: arXiv
The interview agents scored 65.7% before adjustment. Set against the 79.5% consistency participants managed with themselves two weeks later, that yields the normalized 83%. A demographics-only baseline reached 74%, so two hours of interview buys nine points over knowing a person’s demographics alone. This is the basis for modeling synthetic twins of real individuals instead of building synthetic profiles of non-existent individuals based on demographic data alone.
The paper's appendix separates two things these agents appear to be doing. Sometimes the answer is already sitting in the transcript, buried under a question about something else entirely, and the model only has to find it; the authors call this “direct retrieval.” Sometimes nothing in the interview speaks to the question at all, and the model reasons outward from what the person did say toward an answer they never gave, which the authors call “inference.” Another study found that embedding domain-specific contextual cues, such as 'Considering the environmental policies of the United States in 2022', improved prediction accuracy across four GSS government-spending domains by relative gains ranging from 6.7% to 23.6%. These findings form the basis for integrating real-world context and current events, as well as global data about individuals with certain demographic characteristics.
Other studies support this principle via different methods. A 2024 study by Moon et al. took a different route, having the model invent free-form life histories, gathered into what the authors call an anthology, and using those to anchor each persona. Personas built this way tracked the real distribution of answers more closely than personas specified directly, by as much as 18%, and did so across several open-weight models.

Source: arXiv
Zhao et al., separating the enrichment of a real respondent's partial profile (with interview data or environmental context, as documented above) from the wholesale manufacture of a dataset, set out four ethical worries about the practice: that skews already present in the data get magnified rather than corrected; that the underlying samples lean heavily Western, which limits how far any result travels; that generated answers can be passed off as real when nobody says otherwise; and that the way the data was produced has to be open enough for someone else to repeat it.
Distribution Matching
The most common assessments of synthetic survey accuracy measure distribution-wide accuracy. 1-MAE (1 minus Mean Absolute Error) measures the average size of the absolute differences between predicted values and actual values in a dataset. This metric is applied per-question or per-outcome across a survey, not on an individual basis. NDAM (Normalized Distribution Accuracy Measure) compares the same predicted and actual distributions but normalizes the result for how many answer options a question offers, so a three-way question cannot flatter the score the way it does under 1-MAE. Both still assess the distribution of answers as a whole rather than the fidelity of any individual agent. Scores above 90% are common against these metrics for commercial synthetic survey companies.
It is true that synthetic surveys, by simulating individual responses, outperform basic LLMs at predicting population responses to real-world questions. At the same time, however, these models can be unreliable at the level of individual respondents or specific subgroups, which is precisely where they purport to add value by simulating the responses of hard-to-reach subgroups of individuals.

Source: Electric Twin
Skeptics of synthetic surveys, however, have argued that most evaluations cheat by comparing distributions rather than predicting individuals, which conflates pattern matching with genuine respondent-level prediction, and identify a gap in how the field evaluates its own work. They propose a stricter alternative they call cross-survey transfer, in which a model is shown how one person answered a block of questions, then has to say how that same person answered a wholly unrelated block elsewhere in the questionnaire.
One experiment by survey group Verasight found that synthetic survey overall results came close to real poll averages, but subgroup and individual results did not. On Trump approval, the model returned 39% approval versus 42% in the real poll, and 60% versus 55% for disapproval. On a zoning policy question outside the model's training data, the LLM returned 59% support versus 28% actual, and 0% "don't know" versus 29% actual. For subgroups in the survey, errors averaged 8%, rising to 15% for Black respondents and 20% for respondents of other races. At the individual level, roughly 20% of Trump opponents were predicted to be supporters, with a similar error rate in the other direction.
Advantages
Proponents of synthetic surveys point to a variety of benefits compared to traditional surveys as advantages associated with the method. These benefits include the gap between real behavior and survey responses, sampling error associated with hard-to-reach populations, the cost and efficiency of simulated surveys, and the limitations to human-targeted surveys that AI is concurrently creating.
Synthetic surveys modeling behavioral reactions or preferences are based on data describing what humans with different demographic characteristics actually do, not how they might respond to a given survey. The reality of humans as unreliable narrators, or poor predictors, of their own behavior is avoided completely in synthetic surveys. One Aaru co-founder framed this risk as “there are massive issues when you’re using real people. You never know if someone is telling the truth.” While there may be inaccurate survey data incorporated into training, the focus of models on simulating behavior downweights the impact of such inaccuracies on predictions.
Synthetic sampling also, at least ostensibly, gives surveyors the power to access groups that would be otherwise hard to reach for surveys, like those that are less likely to respond to requests for interviews or the customers of a competitor’s brand, for example. Such benefits may be overstated, however, considering that such demographic minorities that are hard to reach for real-world surveys may also be underrepresented in the training data used to construct the models.
Such concerns are reinforced by academic research: Harvard researchers reported that ChatGPT could mimic Americans well enough that opinion splits cleanly along party lines, but lost the thread once it had to track how views diverge by age, race, or gender. Cheng et al. found that when a model is asked to speak as a politically or socially marginalized group, the portrayal slides into stereotype, and Durmus et al. built a benchmark spanning more than a hundred countries to show whose opinions models leave out.
Synthetic surveys are also theoretically faster, cheaper, and simpler than human polling. Real-world survey response rates have declined meaningfully over the last several decades, while costs associated with running these interviews have increased, making the appeal of synthetic surveys to corporations clear.
Evidenza, for example, advertises research timelines compressing from six months to six hours, and a move from a 2% response rate to 100% completion. The CMO of SentinelOne said in a testimonial, “In 24 hours, the team generated a comprehensive brand analysis that would have taken six months with traditional methods. Evidenza helps CMOs make smart decisions, quickly.” Similarly, one client of Electric Twin said: “It takes longer to write the question than it does to get the answer.” In a pharmaceutical-company case cited by Synthetic, another AI polling company, a round of expert interviews that had taken three months was reportedly compressed into hours. In the Aaru-EY study, the company generated 100K digital personas in hours, against a conventional survey that reached 3.6K wealthy individuals in over 30 markets and ran for six months.
Companies also attest to saving their clients money. Electric Twin says clients running more than 1.5K queries over six months have saved £300K compared to traditional research methods. The Director of Data Operations at the Times newspaper said of the efficiency advantage: “We went from rationing research to running it on demand. The same team, ten times the output - and we could prove it matched our real audience.” One Simile customer described the benefit as it pertains to product roll-out, saying, “We don’t need to fail fast in front of customers, we can fail safely in a controlled environment. Teams can test, learn, and refine until they’re confident it’s ready for the real world.”
Finally, synthetic surveys have been advertised as a means to combat AI bot responses to human-targeted polls. While there is conflicting evidence on how many bots already sit inside the online panels themselves, companies like Artificial Societies provide a quality check service auditing customer-collected real-world survey data for the fingerprints of bot responses.

Source: Artificial Societies
Limitations
Despite these benefits, there are real trade-offs and limitations associated with synthetic surveys. First, and most obviously, AI agents are not real people. Even the most complete and current training data may not reflect real-time changes in opinions or reactions. Further, AI-generated answers display systematic differences from the answers reported by real populations.
Academic research has shown that models can be poor replacements for human surveys in opinion or attitudinal polling because they tend to exhibit stronger biases and lower within-population variance compared to real humans (Bisbee et al. found the spread of LLM-generated responses consistently tighter than the spread of real ones, a phenomenon other researchers have replicated and named narrowing “variance collapse”). Further, models almost never admit uncertainty, offering far fewer non-answers than real respondents do, and are sometimes resistant to offering negative responses. John Hagner, a political pollster, has reported that “[in] the early experiments on this, they cannot get respondents to be as racist or sexist or, frankly, as negative as human respondents.”
The training process for these models may be partly responsible for such differences. Researchers Perez et al. showed that tuning a model on human preference data leaves its own mark on behavior, most visibly in greater sycophancy and firmer stated political views.
The training of models for synthetic surveys, and collection of data required to curate identities within the synthetic population, also introduces ethical concerns for the individuals being “twinned.” Akin to the predatory practice of surveillance pricing, the results of synthetic surveys could be used to target subgroups of customers (or at the extreme, distinct individuals) based on responses to surveys that those consumers did not endorse and may not even agree with.
Finally, there are some limits to official uses of synthetic survey data in the realm of traditional survey application, particularly as associated with claims that are regulated, studies that are clinical, and consumer testing that the law requires.
Critical Response
As noted above, the task force report issued by the American Association for Public Opinion Research deemed synthetic responses as “the most risky” of AI applications to public polling. Where synthetic responses are used, the report advises users to keep some real respondents in the study, either alongside the synthetic ones or through a mixed-subjects design. It explores two possible sequences: “validate-then-simulate”, where agreement with human data is established first, and generation is scaled only afterward, and, where circumstances force it, “simulate-then-validate”.
Professional political pollsters have expressed opposition to synthetic responses. Nate Silver, commenting in Eli McKown-Dawson’s Silver Bulletin piece, said:
“You might be able to train a model to make a reasonable estimate of what some hard-to-reach poll respondent would say — say, a young Black man who voted for Trump. (Such a person checks a number of boxes for a voter who is usually hard to reach in surveys.)… But you don’t actually know what these voters think unless you’re reaching them directly. If there’s a shift in opinion among this subgroup, you’re not going to detect it. So if I were running a campaign, I’d invest more in going the extra mile to find a representative sample of these voters. And then I’d hire some smart quants — perhaps with help from Claude et. al. — to figure out the implications for campaign strategy based on that proprietary data that my competitors didn’t have access to.”
These sentiments are echoed by pollsters like Natalie Jackson, a vice president at GQR Insights, who said: “I think politics should stay away from [synthetic sampling], because we’re trying to… represent the voice of the people… I would advise campaigns and media to stay away from AI polling.” Another pollster agreed, saying, “I think I’m just incredibly skeptical of this idea. I don’t think it’s research. At that point, you’re asking the machine to tell you what you already believe.”
Traditional polling agencies are approaching synthetic surveys with caution. In Gallup’s analysis of Simile’s method, Gallup published no accuracy figure and committed not to use synthetic surveys for its own population estimates or tracking measurements. Gallup plans to test how fast the method’s accuracy decays as questions move away from the interview topics and how fast it decays as the world changes. Gallup also conceded a reputational risk, acknowledging that the method could damage the public's confidence in polling.
Such agencies have already adopted AI in other forms, such as having AI sort open-ended answers into categories, letting a model put the questions to respondents itself, or in the construction of poll structures which pinpoint the locations and groups of people to target as the deciders of an election. AI is being used by dedicated companies like Listen Labs to make polling humans more efficient; Listen Labs reportedly runs hundreds of interviews at once, letting each participant answer on camera, by voice, in writing, or by sharing a screen.

