---
title: "How synthetic polling results change based on the AI model"
description: "Survey experiments with Opus 4.6 and GPT-5.1 each paint a very different picture of U.S. public opinion, neither of which reflect reality."
date: "2026-09-30"
authors:
  - name: "Athena Chapekis"
    job_title: "Computational Social Scientist"
    link: "https://www.pewresearch.org/staff/athena-chapekis/"
  - name: "Arnold Lau"
    job_title: "Research Methodologist"
    link: "https://www.pewresearch.org/staff/arnold-lau/"
  - name: "Samuel Bestvater"
    job_title: "Associate Director, Data Labs"
    link: "https://www.pewresearch.org/staff/samuel-bestvater/"
  - name: "Sono Shah"
    job_title: "Former Associate Director, Research"
    link: "https://www.pewresearch.org/staff/sono-shah/"
  - name: "Andrew Mercer"
    job_title: "Principal Methodologist"
    link: "https://www.pewresearch.org/staff/andrew-mercer/"
  - name: "Aaron Smith"
    job_title: "Director, Data Labs"
    link: "https://www.pewresearch.org/staff/aaron-smith/"
url: "https://www.pewresearch.org/data-labs/2026/09/30/how-synthetic-polling-results-change-based-on-the-ai-model/"
categories:
  - "Artificial Intelligence"
  - "Methodological Research"
  - "Uncategorized"
---

# How synthetic polling results change based on the AI model

**About this research**

This Pew Research Center report examines whether AI-generated survey data can accurately duplicate the results of high-quality public opinion polls.

We conducted this methodological experiment to better understand an emerging method in survey research. The Center believes that speaking to the public is essential to measuring public opinion and has no current or future plans to use AI models to generate survey results. Read more about our [AI policy](https://www.pewresearch.org/decoded/2026/07/30/how-pew-research-center-is-and-is-not-using-ai-in-our-work-2/).

For this study, we had an AI model take on the role of real respondents from our [American Trends Panel](https://www.pewresearch.org/the-american-trends-panel/) (ATP) and answer the same surveys administered to human panelists. These surveys included many different question formats and asked about a wide range of topics, attitudes and behaviors.

#### Why did we do this?

The Center regularly studies and evaluates [advances and trends in polling research](https://www.pewresearch.org/topic/methodological-research/survey-methods/). As AI-based polling becomes [more widely used in the industry](https://www.nytimes.com/2026/04/06/opinion/ai-polling.html), we wanted to understand the data these “synthetic samples” produce and how it compares with high-quality surveys of real humans on topics of broad public interest. That question is the focus of this study; it does not address other uses of AI models in polling.

#### How did we do this?

We used a “digital twins” approach for this study, asking an AI model to adopt the personas of real humans who are members of the ATP. We gave the model a [wide range of information about each panelist](#_Conditioning_information), including their self-reported demographic information and their responses to questions on our [2025 political typology survey](https://www.pewresearch.org/politics/2026/06/10/beyond-red-vs-blue-the-political-typology/). We then showed the model questions from three different ATP surveys conducted in the first half of 2026 and asked it to respond to those same questions as the assigned persona. The model was given each question on the survey in order from start to finish, with exactly the same instructions and survey programming that the real humans who took it received.

To make a direct comparison between our AI and human polls, we only replicated surveys for the ATP panelists from each wave who had also completed the political typology survey. As a result, figures for U.S. adults in this report may differ slightly from those previously published by the Center.

We tested multiple models, but all synthetic results unless noted otherwise come from Anthropic’s Claude Opus 4.6. At the time we conducted this study, Opus 4.6 was the most recent Anthropic model available for commercial use and had the best performance of several we evaluated.

Here are the [survey questions from Wave 185](https://www.pewresearch.org/wp-content/uploads/sites/20/2026/01/PP_2026-01-29_views-of-trump_questionnaire.pdf), [Wave 190](https://www.pewresearch.org/wp-content/uploads/sites/20/2026/04/pg_2026.04.28_us-role-world_questionnaire.pdf) and [Wave 192](https://www.pewresearch.org/wp-content/uploads/sites/20/2026/05/PP_2026.5.11_national-problems_questionnaire.pdf) used for this analysis, the [detailed responses](https://www.pewresearch.org/wp-content/uploads/sites/20/2026/09/pl_2026.09.30_silicon-samples_topline.pdf) from the ATP and AI surveys, and the [methodologies](https://www.pewresearch.org/data-labs/2026/09/30/methodology-silicon-samples/) for the original ATP survey waves and our AI surveys.

A key feature of public opinion polling is its consistency and predictability. Although the individual participants may be different, two identical surveys fielded at the same time using the exact same methods should produce largely similar results within known boundaries.

But there is no guarantee of this when the “participants” taking the surveys are AI models. To explore how the choice of model can affect the results of the synthetic survey, we used both OpenAI’s GPT-5.1 and Anthropic’s Claude Opus 4.6 to simulate 6,700 real respondents from our [American Trends Panel](https://www.pewresearch.org/the-american-trends-panel/) (ATP) and generate synthetic survey data using a questionnaire we originally fielded in late January 2026. We then compared the models’ results against each other, as well as against the results from the human poll.

Looking at the survey average absolute error, neither of the models we tested replicated the human panel results especially well. Opus performed slightly better, with an average absolute error of 11.4 percentage points (compared with 13.3 percentage points for GPT).

To be sure, newer models and future innovations in synthetic sample construction could result in smaller overall error figures. The real story here is that **each synthetic sample painted a very different picture of the American public on multiple dimensions, and neither truly reflected actual public opinion.**

*This analysis is part of a larger evaluation of AI-generated synthetic samples in public opinion research. Read* *[a summary of the main findings](https://www.pewresearch.org/data-labs/2026/09/30/can-ai-stand-in-for-human-survey-takers-not-really/)* *and refer to the* *[methodology](https://www.pewresearch.org/data-labs/2026/09/30/methodology-silicon-samples/)* *for more details on how we conducted our synthetic poll and compared it with real survey results.*

The questions and topics on which GPT most closely mirrored actual public sentiment were almost entirely different from those on which Opus was the better performer. Still, there were [no topics or particular categories of question](#_Appendix_D:_Additional) that stood out as a strength of one model over the other.

Here are examples of ways in which the two models provided results that were different from each other – and also from our human poll.

### Example 1: Public views of politics and democracy

Our January ATP survey included a set of questions about broad political attitudes, like overall satisfaction in the way things are going in the United States, which ideological “side” is losing more often, the impact of voting, and whether there are clear solutions to the problems facing the country. These are all fairly long-standing trend questions with ample history to draw on, yet the two models produced wildly inconsistent results.

### GPT exaggerates dissatisfaction with the country’s direction, while Opus downplays clarity of solutions

*% who say …*

| Statement | U.S. adults | Opus synthetic respondents | GPT synthetic respondents |
| --- | --- | --- | --- |
| All in all, they are dissatisfied with the way things are going in the country today | 69 | 70 | 100 |
| There are clear solutions to most big issues facing the country today | 56 | 25 | 57 |
| Thinking about politics over the last few years, their side has been losing more often than winning | 63 | 63 | 100 |

Note: Figures for U.S. adults may differ slightly from those previously published by the Center because they are based on a subset of respondents to the original survey.

Source: Survey of U.S. adults and “digital twins” synthetic analysis using models set to low reasoning with extended profile information and expert reflection. Survey was conducted Jan. 20-26, 2026 (replicated March 2-3 with OpenAI GPT-5.1; and replicated March 9-12 and April 7-10 with Claude Opus 4.6).“Can AI Stand In for Human Survey-Takers? Not Really”

For instance, the Opus poll was almost exactly in line with actual public sentiment when it came to the share of Americans who are dissatisfied with the way things are going in the country today; the share who feel like their “side” has been losing more than winning in politics; and the share who say voting gives people like them some say in how the government runs things. Each of these opinions is held by a majority of Americans, but they are far from ubiquitous.

By contrast, our GPT poll depicts these views as quite literally universal – for instance, estimating that 100% of Americans are dissatisfied with the way things are going in the country.

On the share of Americans who agree that there are clear solutions to most big issues facing the country today, GPT more closely mirrors the views of the public, while Opus differs from true public opinion by roughly 30 percentage points.

### Example 2: Abortion attitudes

Abortion is an example of an issue on which neither model produced results that accurately reflect public sentiment.

### AI survey responses from 2 different models disagree on abortion legalization

*% who say abortion should be …*

| Mode | Legal in all cases | Legal in most cases | Illegal in most cases | Illegal in all cases | NET Legal | NET Illegal |
| --- | --- | --- | --- | --- | --- | --- |
| U.S. adults | 23 | 37 | 28 | 11 | 60 | 39 |
| GPT Synthetic respondents | 16 | 32 | 49 | 4 | 48 | 53 |
| Opus Synthetic respondents | 13 | 49 | 37 | 0 | 62 | 37 |

Note: Figures for U.S. adults may differ slightly from those previously published by the Center because they are based on a subset of respondents to the original survey.

Source: Survey of U.S. adults and “digital twins” synthetic analysis using models set to low reasoning with extended profile information and expert reflection. Survey was conducted Jan. 20-26, 2026 (replicated March 2-3 with OpenAI GPT-5.1; and replicated March 9-12 and April 7-10 with Claude Opus 4.6).“Can AI Stand In for Human Survey-Takers? Not Really”

In our human poll, around six-in-ten Americans said abortion should be legal in all or most cases, while around four-in-ten said it should not be legal. Opus produced the same general split, but greatly underestimated the share of Americans who hold very strong views on abortion:

- 23% of U.S. adults think abortion should be *legal in all cases*; our Opus poll estimated that share at just 13%.

- 11% think abortion should be *illegal in all cases*; Opus estimated that 0% of the public holds this view.

By contrast, GPT produced overall estimates that imply the country is about evenly split on the legality of abortion, with the view that it should be illegal slightly more common.

All told, neither synthetic poll illustrates where the country stands on this issue – and they fail to capture the full scope of public sentiment in different ways. Here, the choice of model alone has the potential to reshape the narrative around public opinion of abortion.

### Example 3: Partisan attitudes

The two models also painted very different pictures of the views held by Republicans and Republican leaners. Synthetic estimates for Republicans and those who voted for President Donald Trump in 2024 make for particularly striking examples of how using a different model can lead to very different conclusions about public opinion within a certain group.

### Model choice in creating AI-generated survey data leads to contradictory stories of attitudes among Republicans

| Topic | Response | U.S. adults | Opus synthetic respondents | GPT synthetic respondents |
| --- | --- | --- | --- | --- |
| Among 2024 Trump voters: strength of approval of Trump | Approve not so strongly | 19 | 53 | 9 |
| Among 2024 Trump voters: strength of approval of Trump | Approve very strongly | 65 | 46 | 88 |
| How to characterize China's relationship with the U.S. | Competitor | 60 | 85 | 56 |
| How to characterize China's relationship with the U.S. | Enemy | 28 | 15 | 44 |
| How to characterize China's relationship with the U.S. | Partner | 10 | 0 | 0 |
| Whether Republicans in Congress have an obligation to support Trump's policies and programs because he is a Republican president | Do not have an obligation to support Donald Trump's policies and programs if they disagree with him | 61 | 81 | 27 |
| Whether Republicans in Congress have an obligation to support Trump's policies and programs because he is a Republican president | Have an obligation to support Donald Trump's policies and programs because he is a Republican president | 38 | 19 | 73 |
| Whether the U.S. should send ground troops to Venezuela | Not sure | 17 | 0 | 0 |
| Whether the U.S. should send ground troops to Venezuela | Somewhat favor | 14 | 0 | 21 |
| Whether the U.S. should send ground troops to Venezuela | Somewhat oppose | 21 | 28 | 12 |
| Whether the U.S. should send ground troops to Venezuela | Strongly favor | 4 | 0 | 2 |
| Whether the U.S. should send ground troops to Venezuela | Strongly oppose | 43 | 72 | 66 |

Note: Figures for U.S. adults may differ slightly from those previously published by the Center because they are based on a subset of respondents to the original survey.

Source: Survey of U.S. adults and “digital twins” synthetic analysis using models set to low reasoning with extended profile information and expert reflection. Survey was conducted Jan. 20-26, 2026 (replicated March 2-3 with OpenAI GPT-5.1; and replicated March 9-12 and April 7-10 with Claude Opus 4.6).“Can AI Stand In for Human Survey-Takers? Not Really”

Our January 2026 ATP survey found that Trump’s overall approval rating among his 2024 voters was at 84%, with 65% approving very strongly and 19% approving not so strongly. Both of our synthetic polls overestimated his approval rating – Opus by 15 points and GPT by 13 points. But beyond the topline figures, each synthetic poll presented a different conclusion about this voter base:

- Trump voters in the GPT poll consisted almost entirely of those who approve very strongly.

- Trump voters in the Opus poll were evenly split between very strong and not so strong approval.

Across other questions in the survey, GPT estimates tended to support a view of Republicans as fervent Trump supporters with highly polarized partisan views. By contrast, the Opus poll would indicate that far more Republicans hold mixed or moderate views. Some questions where the two models present opposing views of Republican attitudes include:

- Whether Republicans in Congress have an obligation to support Trump’s policies and programs because he is a Republican president

- Whether or not Trump should work with Democratic congressional leaders

- Whether China is an enemy, competitor or partner to the United States.

- Whether or not the U.S. should send ground troops to Venezuela

They also present very different estimates of public sentiment toward prominent political figures. In our human poll, 74% of Republicans and Republican-leaning independents expressed a favorable view of Health and Human Services Secretary Robert F. Kennedy Jr. Opus estimated his favorability among Republicans at 85%, while the GPT poll estimated their opinion to be 63% *unfavorable*.

### Example 4: Magnitude of opinion and “extreme” answer options

The example of abortion notwithstanding, the GPT and Opus polls often agreed that the public generally leans in one direction or another on various issues. But often, the two models disagreed on the intensity of those attitudes.

For instance, 68% of U.S. adults in our ATP survey said they are concerned about the price of gasoline, split evenly between those who are very or only somewhat concerned. GPT and Opus both (incorrectly) estimated that nearly every American is concerned about this issue. But the GPT sample lumps most of the public into the “very concerned” category, while the Opus sample leans heavily toward “somewhat concerned.”

As it turns out, this kind of pattern appeared consistently throughout the survey. GPT tended to favor response options at the far ends of scales, such as “extremely” or “not at all.” By contrast, Opus tended to favor “middle” options like “somewhat,” “about right” or “neither.” This was broadly true regardless of what the scale was or what the question was about.

For all questions in the survey with response options arranged in a clear order:

- Human respondents chose a “middle” option 45% of the time, on average.

- GPT respondents did so 31% of the time.

- Opus respondents did so 56% of the time.

### GPT overestimates extreme response options like ‘always’ or ‘never,’ while Opus underestimates them

| Variable | Opus | GPT |
| --- | --- | --- |
| ABORTION3_W185 | -14.558333333333335 | 9.698010323010324 |
| ABORTION4_W185 | 0.06521043865442522 | 20.622530995974984 |
| ABRTLGL_W185 | -21.339632315242063 | -14.752845528455282 |
| ABRTVIEW_a_W185 | -23.88900528435412 | 29.911948516599693 |
| ABRTVIEW_c_W185 | -20.426720647773273 | 21.573279352226727 |
| BILLION_W185 | 16.354086781029274 | 31.63526796221044 |
| CHINA_US_ENEMY_W185 | -23.710827374872316 | 5.089172625127688 |
| CRYPT1_W185 | -31.66933867735471 | -9.069338677354708 |
| DATCEN_AWARE_W185 | -44.94635338856445 | -34.45175879396985 |
| DATCEN_SUPP_a_W185 | 26.913492486622033 | 12.333743661889258 |
| DATCEN_SUPP_b_W185 | 18.513323983169713 | 18.513323983169713 |
| DATCEN_SUPP_c_W185 | 23.227794347113317 | 22.303427653390045 |
| DATCEN_SUPP_d_W185 | 30.02504208242577 | 13.482757542163483 |
| DATCEN_SUPP_e_W185 | 25.74568288854003 | 16.443357307144694 |
| DEMCARE_W185 | -1.7346293662083134 | 6.936842105263164 |
| DTECON_W185 | -6.687563195146609 | 16.512436804853394 |
| DTSUCCESS_W185 | -3.7477341389728025 | 20.652265861027203 |
| ECON1B_W185 | 9.570352699740852 | 25.88666901605717 |
| ECON1_W185 | -13.033333333333337 | 19.766666666666666 |
| ECONCONC_ELEC_W185 | -31.268341708542714 | 23.631658291457285 |
| ECONCONC_ENG2_W185 | -17.400000000000002 | 38.21568431568431 |
| ECONCONC_GAS_W185 | -18.519157472417252 | 41.29982354656376 |
| ECONCONC_HC_W185 | 11.208835341365456 | 20.608835341365463 |
| ECONCONC_PRICE_W185 | 9.26533066132265 | 27.26533066132265 |
| ECONCONC_REAL_W185 | 3.4059177532597715 | 4.00591775325978 |
| ECONCONC_STCK_W185 | -30.918529707955692 | -31.11852970795569 |
| ECONCONC_UNEM_W185 | -22.400200400801605 | 9.659559358958155 |
| EXPAND2_GRNLND_W185 | -19.31334956274619 | -11.65449435647438 |
| EXPAND_GRNLND_MOD_W185 | 12.014302920054575 | 65.7279301745636 |
| FAVPOL_HEGSETH_W185 | -17.788642991936868 | -1.141743511740664 |
| FAVPOL_MUSK_W185 | 3.483407313464383 | -12.215789473684213 |
| FAVPOL_RFKJR_W185 | -1.35107871315126 | -14.626100479291525 |
| FAVPOL_RUBIO_W185 | -28.401536592016814 | -27.911497990114295 |
| FAVPOL_TRUMP_W185 | -11.785329349906355 | 28.312065439672807 |
| FAVPOL_VANCE_W185 | -16.12376278724316 | 11.850444946835523 |
| GUNSTRICT_W185 | -8.329238329238336 | 23.202293202293205 |
| IMMIGAGNT_ACCEPT_CITZ_W185 | -9.94096843537332 | 21.010783316378436 |
| IMMIGAGNT_ACCEPT_MASK_W185 | -20.514979757085 | 7.985020242914999 |
| IMMIGAGNT_ACCEPT_NGBH_W185 | -6.032171989618796 | 27.32168186423506 |
| IMMIGAGNT_ACCEPT_RLAN_W185 | -14.932388663967616 | 23.151795520216577 |
| IMMIGPPL_ACCEPT_REPT_W185 | -8.352845528455276 | 32.74715447154472 |
| IMMIGPPL_ACCEPT_TRCK_W185 | -14.923654619086086 | 22.63350253807107 |
| IMMIGPPL_ACCEPT_VDEO_W185 | -11.879896090422411 | 11.898627688101378 |
| IMMIG_FAVOPP_ASY_W185 | -5.336559464999837 | 26.7051987767584 |
| IMMIG_FAVOPP_DTN_W185 | -8.114720812182739 | 20.510454012992085 |
| IMMIG_FAVOPP_FEE_W185 | -21.2520325203252 | 30.54426377597109 |
| IMMIG_FAVOPP_MIL_W185 | -4.155983772819475 | 22.744016227180516 |
| IMMIG_FAVOPP_PAU75_W185 | -7.680286006128704 | 28.5197139938713 |
| IMMIG_FAVOPP_SCMED_W185 | -9.984755864715169 | 18.9768056968464 |
| LEADERCHAR_DSA_W185 | -11.309034907597535 | 13.64341754485492 |
| MARIJ_STRICT_W185 | -32.86678620398692 | 33.880722828091244 |
| MATTERSCONG3_W185 | -18.063974370735217 | 11.496165489404646 |
| MATTERSCONG_W185 | -6.18710976837864 | 26.00848582721696 |
| MEDESC_INT_W185 | -20.75236656596173 | -17.85236656596173 |
| MEDESC_OPEN_W185 | -16.187724011194423 | -17.777933801404217 |
| MEDESC_RESP_W185 | -33.370854271356784 | -19.055739156241664 |
| MEDESC_SKEP_W185 | -17.08796764408493 | -10.681561237678524 |
| MEDESC_TRAD_W185 | -10.063223067255327 | 16.982203978171718 |
| MEDESC_WRK_W185 | -26.770161290322584 | -21.758149278310572 |
| NEWPRES_ACTNS_W185 | -13.46693548387097 | 27.43306451612903 |
| NEWPRES_SUPP_W185 | -20.27671534713788 | 22.31227364185111 |
| QUALPRES_TRMP_ADV_W185 | -6.753768844221106 | 20.54623115577889 |
| QUALPRES_TRMP_ETH_W185 | -6.425352112676059 | 22.37464788732394 |
| QUALPRES_TRMP_FIRM_W185 | -10.53324681465385 | 31.52160804020101 |
| QUALPRES_TRMP_MF_W185 | -11.715721464465176 | 31.927135678391963 |
| QUALPRES_TRMP_PHY_W185 | -33.09170197387721 | 39.48087059869536 |
| QUALPRES_TRMP_RSP_W185 | -9.950736234724154 | 29.702416918429 |
| REPCARE_W185 | 8.568105322669219 | 16.311561866125764 |
| RUSSIA_US_ENEMY_W185 | -15.547365547365551 | 33.70188370188369 |
| TARIFFS_COUNTRY2_W185 | 1.741377373328831 | 21.160796792748272 |
| TARIFFS_INDIV2_W185 | -8.736931657824151 | 25.397202476309985 |
| TARIFF_APP_W185 | -10.06791406791406 | 32.0586814654951 |
| TAXBTHR_CMPLX_W185 | -15.078093306288032 | 25.321906693711973 |
| TAXBTHR_CORP_W185 | 9.106054490413712 | 17.30605449041373 |
| TAXBTHR_LWINC_W185 | -9.68829576296782 | 12.674066599394544 |
| TAXBTHR_WLTHY_W185 | 6.079764522636111 | 19.75369059656218 |
| TAXBTHR_YOU_W185 | -12.980761916932124 | 45.591996655826456 |
| TRUMPISSUE1_W185 | -12.153172205438054 | 22.646827794561936 |
| TRUMPISSUE2_W185 | -17.374420946626387 | 18.725579053373608 |
| USEXCEPT_W185 | -12.135923309788096 | -4.491979365844152 |
| VEN_ACTIONS_NATRSC_W185 | -38.81666878818737 | 10.609437562562562 |
| VEN_ACTIONS_TROOP_W185 | 14.865693430656924 | 10.465693430656934 |
| VEN_DTCONF_W185 | -8.024417426545085 | 34.27558257345491 |
| VEN_MADREMV_W185 | -10.740655283802496 | -4.168629298162969 |
| VEN_USROLE_W185 | -21.579291242212584 | -16.485597548518893 |
| VTPRIORITY_ATO_W185 | -7.137668848463143 | 28.21018329938899 |
| VTPRIORITY_GOVID_W185 | -13.603553299492404 | 6.569920173981075 |
| VTPRIORITY_ML_W185 | -16.302127659574474 | 23.39787234042552 |
| YOURTAXES_W185 | 3.3101522842639497 | 29.710152284263955 |

Note: Average percentage point error is calculated only for questions with ordered response categories (89 out of 119 questions tested). Extreme responses include options at the far ends of scales (such as “extremely” or “not at all”). Moderate responses include middle options (such as “somewhat,” “about right” or “neither”).

Source: Survey of U.S. adults and “digital twins” synthetic analysis using models set to low reasoning with extended profile information and expert reflection. Survey was conducted Jan. 20-26, 2026 (replicated March 2-3 with OpenAI GPT-5.1; and replicated March 9-12 and April 7-10 with Claude Opus 4.6).“Can AI Stand In for Human Survey-Takers? Not Really”

Put simply: GPT tended to depict Americans as having more extreme opinions than they actually do, while Opus tended to depict them as being more middle-of-the-road than they are in reality. It’s important to note that changes to how synthetic survey data is generated and future updates to the GPT and Opus models themselves could lead to different patterns. However, it is clear that the choice of model alone is enough to yield large, systematic differences between synthetic samples that were otherwise built in exactly the same way.[3. numoffset="3" Even models within the same family did not necessarily behave similarly to one another. GPT-5 nano (a “lighter weight” model within the GPT-5 family) produced very different estimates from both GPT-5.1 and Opus 4.6.]

When compared with the data from our real panelists, neither synthetic poll was able to provide a reasonably similar snapshot of the views of the American public. However, given that Claude Opus 4.6 had better performance *on average*, this is the model we selected to generate additional synthetic samples.

---

**Next:** [Acknowledgements](https://www.pewresearch.org/data-labs/2026/09/30/acknowledgements-silicon-samples.md)