Artificial intelligence stand-ins for real survey respondents missed Americans' answers by an average of 12 percentage points across nearly 300 questions in a Pew Research Center experiment, raising a warning for anyone using synthetic polls to gauge public opinion.
The Pew study released Wednesday compared AI-generated answers with those of real participants in three waves of its American Trends Panel. The synthetic results were not merely a little off: The average error exceeded 15 points on about 28% of questions, and some individual response options missed by 20, 30 or more than 40 points. Pew concluded that, in this test, AI respondents were not a reliable replacement for surveys of people.
The experiment matters because synthetic polling promises rapid results without recruiting a new representative group of people. Pew did not ask a model for one national prediction. It created a digital counterpart for each actual panelist, supplying demographic information and prior political-typology answers, then administered the same questions in the same order and language the person had received. That design let researchers compare national estimates and demographic subgroups while retaining the structure of an ordinary survey.
Errors varied by subject and subgroup
The model overstated President Donald Trump's job approval in two 2026 surveys, estimating 46% in both January and April. Human respondents put approval at 37% and 34%, respectively. Pew said the AI missed a decline among Republican and Republican-leaning respondents even though it more closely matched Democrats' answers.
On a question about whether immigration officers should be allowed to cover their faces while working, 35% of synthetic Republicans approved, compared with 67% of real Republicans. The AI also overstated Republican support for U.S. strikes against Iran and missed awareness and views of data centers. It estimated that nearly all Americans who had heard of data centers viewed them as bad for home energy costs, the environment and neighbors' quality of life; actual human shares ranged from 41% to 53%.
Some errors made groups look far more uniform than they were. The synthetic survey estimated that 97% of Hispanic adults were at least somewhat likely to follow the World Cup, compared with 42% in the underlying human survey. It estimated that at least 80% of Democrats and Democratic leaners viewed billionaires as bad for the country, versus 45% of actual respondents. For Republicans and Republican leaners, it put favorable views of Israel above 90%, while the human result was 58%.
Pew found especially weak replication for respondents who had not attended college and those who used the internet infrequently or not at all. Across the demographic and behavioral groups it analyzed, none had an average error below 12 points. The researchers said they observed patterns consistent with models leaning on simplified assumptions about racial, political and other groups, though the experiment does not establish that every mismatch had the same cause.
AI answers flattened uncertainty and diversity
Almost half of the synthetic poll's questions had at least one answer option chosen by no AI respondent. None of the corresponding human-survey questions had a zero-selection option. The AI also tended to avoid saying it was unsure: On opinion questions offering that choice, human respondents used it 16% of the time, compared with 4% of the synthetic respondents.
For 13 factual-knowledge questions, real panelists answered correctly about half the time on average. The AI counterparts were correct about 80% of the time and in six cases predicted that at least 98% of Americans knew the right answer. A model's own knowledge of a subject, Pew found, can distort its estimate of what people know.
The distortion also appeared in everyday experiences. In the human poll, 30% of adults said they had trouble falling asleep every day, most days or never; the AI placed 98% in the two middle responses, some days or rarely. It estimated that just 1% of Americans had visited 10 or more foreign countries, compared with 13% of actual respondents.
Choice of model changed the apparent public mood
Pew ran a direct comparison on a January survey using Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.1. Neither matched the human panel especially well: Average absolute question-level error was 11.4 points for Opus and 13.3 for GPT. GPT tended toward more extreme response categories, while Opus favored middle options. Across questions with ordered answers, real people chose a middle category 45% of the time, compared with 31% for GPT and 56% for Opus.
That difference could change a news account even when the questionnaire and respondent profiles are identical. On abortion, Opus reproduced roughly the broad six-in-ten human majority favoring legality in all or most cases but understated strongly held positions. It estimated that nobody favored illegality in all cases, compared with 11% of real respondents. GPT suggested a much closer overall split. Neither depiction captured both the direction and intensity of human views.
The methodology says the three human survey waves were fielded Jan. 20-25, March 23-30 and April 20-26, 2026. The comparison used 6,700 panelists in the January wave, 3,398 in March and 4,981 in April, restricted to people who had also completed a 2025 political-typology survey. The January sample was drawn from a larger wave because of the cost of running the model, and weights were adjusted. Pew said the resulting figures can differ slightly from some previously published toplines. Most synthetic findings in the report come from Opus; GPT was used for a model-comparison subset.
The test focused on AI acting as the respondent, not on all uses of AI in survey work. Pew said models can help categorize genuine open-ended answers or write analysis code, and newer models or different synthetic-sample methods might perform differently. Its data do not prove that every future approach will fail. They do show that a plausible-looking synthetic percentage can conceal large, unpredictable errors, particularly for current events, minority views and specific groups. For now, the center says it has no plan to replace human respondents with AI-generated answers.



