Researching AI, Society & Democracy in Switzerland

Restricted (Over)view: How we collected and analyzed tens of thousands of Google’s and Bing’s AI summaries on the 10 million initiative

As part of the “RAISD” research project, AlgorithmWatch CH and the Department of Informatics at the University of Zurich examined how the AI-generated responses from Google and Bing portrayed the debate surrounding the 10-million initiative and which sources they used for this purpose. In this article, we explain our methodological approach.

Profile photo of Corinna Hertweck
Dr. Corinna Hertweck
Tech Researcher

Search engines increasingly push the integration of generative AI into their services, despite known risks, through so-called “AI summaries”. However, Switzerland’s current draft law, which is meant to regulate search engines and social media, might not account for the increasing significance of generative AI integrated into these services. We collected “AI summaries” on a high-stakes topic, the vote on the “10 million initiative”, to highlight the potential impact these changes by search engine providers might have on public discourse in Switzerland. Over the course of almost 9 weeks, we collected more than 45,000 AI summaries with more than 940,000 citations in five languages. We annotated a large share of the references for easier analysis and are now making this data available to journalists and the public in our newly released dashboard.

You can find the initial results of our analysis in our press release (in German or French). We will continue to publish further analyses based on the collected data in the coming weeks and months as we continue to analyze it.

Note that the numbers in the press release differ slightly from the numbers in this post, as we narrowed our data for the press release to a subset, while we report total numbers here. For more details, see the section on “How we analyzed the data”.


Content

  1. What are AI summaries?
  2. Why should we care about this development?
  3. How we collected the data
  4. How we labeled the data
  5. How we analyzed the data
  6. Our dashboard as an interactive analysis tool
  7. What we can learn from the data

What are AI summaries?

Search engines started out by showing their users a ranked list of links. They were not information providers, but transmitters of third-party content: they pointed users toward content produced by others. Over the years, that began to change. Features like featured snippets, knowledge panels, and direct answers started appearing on the results page itself, reducing the need to click through to any source at all. This phenomenon is sometimes called “zero-click internet.”

The integration of generative AI as “AI summaries” is the next and very consequential step in this trajectory, rolled out in 2024 for Bing and 2025 for Google. Rather than pointing users to a range of sources, AI summaries draw on multiple websites and their training data simultaneously, synthesize their content, and present a single authoritative-looking answer at the top of the page. This once again incentivizes users to stay on the search results page, further reinforcing “zero-click” search. Importantly, the answer also feels definitive, even when it is not. By design, it exposes users to less of the diversity and depth of perspectives that separate sources might offer. The search engine is no longer a transmitter; it has become an editor.

This shift is accelerating. At its I/O conference in 2026, Google announced plans to push AI even further into its core product, with traditional link-based results gradually receding into the background.

Why should we care about this development?

The concerns raised by this shift come from multiple directions.

Unaccounted-for power. Google has more than 80% of the search engine market in Switzerland and benefits from high levels of user trust. Research on earlier features like featured snippets has already shown that users place high levels of trust in such features. As AI summaries become the default response to many queries, the influence of a single company over what people read and, thus, what they might think grows considerably. Gradually and incrementally, this might also influence political choices and democratic deliberation.

Threatened access to reliable information. With AI summaries becoming so widespread, the public's exposure to reliable information depends to a large degree on those summaries. These then play an increasing role in democratic deliberation and decision-making.

Quality and accuracy. Concerns about the reliability of AI-generated answers have existed since the first chatbots launched. These range from misleading source selection and biased framing to outright factual errors, also called “hallucinations.” Researchers, civil society organizations, and media representatives have pointed out that making AI summaries the default response in widely used search engines exposes vastly more people to these risks. A recent ruling from a German regional court shows that search engines could be held legally liable for such errors.

A threat to media diversity. There are also serious concerns that AI summaries will significantly reduce traffic to news outlets and other public-interest websites. Fewer visits mean less advertising revenue and a weakened basis for quality journalism, with knock-on effects across the entire information landscape. AlgorithmWatch's Berlin office and representatives of German media filed a complaint against Google's AI Overviews with the German digital regulator, the Bundesnetzagentur. In the United Kingdom, the Competition and Markets Authority is imposing requirements that media companies should be able to opt out of having their content scraped for AI summaries.

Despite these concerns, there has been almost no systematic data on how AI summaries behave in Switzerland. Our research set out to address that gap. We chose a high-stakes topic: the Swiss People’s Party’s (SVP/UDC) 10 million initiative, which proposes to limit the resident population of Switzerland to 10 million people. Our aim was to document what AI summaries on politically sensitive content actually look like in Switzerland.

How we collected the data

Our data collection ran for approximately eight weeks, from April 13, 2026, to June 14, 2026. The first week, April 13-20, was a testing period during which the search queries still changed.

We built a tool to automatically ask questions to two search engines, Google and Bing, which are two of the largest search engines in Switzerland, from two locations, Zurich and Bern, twice each day. For each search query, the automated tool saved the search results page it saw, including the AI summary if there was one and the citations within the summary. Such tools are known as “scrapers.”

What’s the 10 million initiative?

The initiative "Keine 10-Millionen-Schweiz! (Nachhaltigkeitsinitiative)", officially translated as "No to Switzerland with 10 million! (Sustainability Initiative)", was submitted by the Swiss People's Party (SVP/UDC), the dominant party on the political right. The initiative would require the Swiss government to keep the country's total resident population below 10 million, primarily by restricting immigration. Switzerland currently has approximately 9 million residents.

The two sides of the campaign have chosen starkly different names for the initiative. Proponents call it the "sustainability initiative", framing their concerns around land use, infrastructure, and quality of life. Opponents call it the "chaos initiative", arguing that such drastic limits on immigration would devastate the economy and jeopardize Switzerland's bilateral treaties with the European Union. This naming contest is itself reflected in our data: we searched for both framings and tracked which sources each side cited.

The vote took place on June 14, 2026, and the initiative was rejected 55% to 45%.

To build our search queries, we started with the official name of the 10 million initiative listed on admin.ch and derived several variants (see search terms table below). We supplemented these with synonyms drawn from our team's knowledge of Swiss political discourse and cross-checked them against Google Trends data to make sure we were capturing the terms people actually search for. We also added language-specific suffixes (e.g., "summary," "pro," and "con,") to reflect the full range of how people might engage with ballot topics online. This process produced 456 distinct search queries spanning all four national languages of Switzerland (German, French, Italian, Romansh) and English. These search queries were stable after our testing period, so from April 21, 2026, onwards.

Search terms. The base terms fall into two groups: official names drawn from admin.ch (in full, long, and short forms) and synonyms. The synonyms were organized by framing, matching the categories used in the dashboard.

CategoryGermanFrenchItalianRomanshEnglish
Official nameKeine 10-Millionen-Schweiz! (Nachhaltigkeitsinitiative)Pas de Suisse à 10 millions ! (initiative pour la durabilité)No a una Svizzera da 10 milioni! (Iniziativa per la sostenibilità)Na ad ina Svizra da 10 milliuns! (Iniziativa per la persistenza)No to Switzerland of 10 million! (Sustainability Initiative)
Keine 10-Millionen-Schweiz!Pas de Suisse à 10 millions !No a una Svizzera da 10 milioni!Na ad ina Svizra da 10 milliuns!No to Switzerland of 10 million!
Nachhaltigkeitsinitiativeinitiative pour la durabilitéIniziativa per la sostenibilitàIniziativa per la persistenzaSustainability Initiative
10 million initiative10 millionen initiativeinitiative 10 millionsiniziativa 10 milioni10 million initiative
10 millionen schweizsuisse 10 millionssvizzera 10 milioni10 million switzerland
10 millionen schweiz initiativeinitiative suisse 10 millionsiniziativa svizzera 10 milioni10 million switzerland initiative
keine 10 millionen schweizpas de suisse 10 millionsno 10 milioni svizzerano to 10 million switzerland
10 millionen schweiz initiative svpinitiative udc suisse 10 millionsiniziativa svizzera 10 milioni udc10 million switzerland initiative svp
Sustainability initiativenachhaltigkeitsinitiativeinitiative durabilitéiniziativa sostenibilitàsustainability initiative
nachhaltigskeitsinitiative svpinitiative udc durabilitéiniziativa sostenibilità udcsustainability initiative svp
Chaos initiativechaos initiativeinitiative chaosiniziativa caoschaos initiative
chaos initiative svpudc initiative chaosiniziativa caos udcchaos initiative svp
Termination initiativekündigungsinitiative 2initiative résiliation 2iniziativa limitazione 2
kündigungsinitiative 2 svpinitiative udc résiliation 2iniziativa limitazione 2 udc

Suffixes. Each base search term was paired with eight suffix variants to capture the range of ways people search for ballot information, ranging from neutral to explicitly partisan. The "none" variant uses the bare term with no suffix.

SuffixGermanFrenchItalianRomanshEnglish
none
summaryzusammenfassungrésumésommarioresumaziunsummary
argumentsargumenteargumentsargomentiargumentsarguments
how_votewie abstimmen?comment voter ?come votare?co votar?how to vote?
propropourpropropro
yesjaouigeayes
conkontracontrecontrocuntercon
noneinnonnonano

German, French, and Italian search terms and suffixes were developed and reviewed by native speakers on our team. Our team also developed and reviewed the English ones. The Romansh translations were generated with the help of Claude (Anthropic's AI assistant) and were not independently verified by a native speaker due to a lack of resources. Feedback on these translations is welcome, even in hindsight. Due to this language barrier, we also chose not to translate the non-official search terms to Romansh. For English, we could not find official translations of the “termination initiative” synonyms and thus decided against using these terms in English.

Across the entire collection period, our scrapers executed a total of 192,976 searches. Of those, 47,728, so roughly one in four, returned an AI summary. That overall rate, however, masks a large difference between the two engines. Google served AI summaries in response to about half of all queries.

However, we faced issues with the collection of data from Bing: In the opening days of our data collection, Bing was returning AI summaries for around 80% of queries. Within days, that figure collapsed to the low single digits and at some point was almost stuck at zero. It is possible that Microsoft deployed bot-detection measures that recognized our automated queries and responded by serving irrelevant content, in some cases, results in different languages (and even different alphabets, e.g., in Chinese or Arabic) with no apparent connection to the search terms. We thus attempted several countermeasures to make our scraper less detectable, but none of them fixed the issue.

Throughout our collection period, Google repeatedly made small changes to the User Interface (UI) of its search results pages. This meant that our scrapers could sometimes not find the references and collect them, leading to many AI summaries for which we had zero citations saved. However, because we always saved the raw search results page for every query alongside the extracted structured data, we were able to detect these changes after the fact and apply fixes retrospectively. When looking for these UI changes, we simply looked for AI summaries with zero citations. Whenever we observed such patterns, we inspected the saved raw HTML files to understand what had changed in the page structure and then wrote targeted repair scripts that re-extracted the citation sources from the HTML. We identified about a handful of these structural changes over the course of our collection. In each case, having the raw HTML allowed us to recover the citation data that the original scraper had missed.

The technical challenges we faced in scraping data from the search engines, i.e., the bot detection of Bing and the UI changes of Google, are significant obstacles to independent research on AI summaries. Systematic investigation of these tools requires structured data access for researchers, which is currently a challenge even in the EU, which has stronger data access laws though many platforms are still resisting providing sufficient data for public interest research.

How we labeled the data

Collecting AI summaries was only the first step. The raw collected data, especially the dataset of links collected from the AI summaries, was not yet that helpful: Not every result was directly relevant to the 10 million initiative, as queries can surface summaries on loosely related topics, e.g., other initiatives related to the concept of “sustainability” in general. Moreover, raw citation counts tell you little on their own. To make sense of the data, we needed to understand who was being cited.

Hence, we annotated references along multiple dimensions, including:

  • Entity domains: connecting, for example, svp.ch and udc.ch to the SVP party or srf.ch and rtr.ch to SRG SSR (the Swiss Broadcasting Corporation).
  • Entity type: government, political party, media outlet, NGO, platform, or other
  • Link topic: Is this link about the “10 million initiative”, providing a general overview of all ballot items for June 14, 2026, is it about another (older) initiative or unrelated to any initiative?
  • Link position: Pro, leaning pro, overview, leaning con, con, unclear were the main positions we also show in the dashboard. Additionally, we also used markers for internal usage to flag long articles or articles in a language the annotator does not speak.
  • Link language: German, French, Italian, Romansh, English, or other.

This kind of annotation makes it possible to ask questions like: Which entities are cited most often for German Google search queries? How does the distribution of link positions shift depending on how a query is framed (e.g., suffixes “yes” and “pro” vs “no” and “con”)?

Using annotations in this way was feasible due to a structural feature of the data, which is that a large share of references in AI summaries go to a small set of links. The same set of sources tends to be cited again and again across thousands of queries. While we collected about 940,000 references across all AI summaries, these come from less than 14,000 unique URLs. Among these unique URLs, about 78 of them accounted for more than 50% of all references. This meant that annotating a modest number of high-frequency links and domains was enough to achieve coverage over a very large share of all citation instances. During two annotation sessions, seven people read through and annotated 638 of the most-cited links. This way, we covered approximately 82% of all references. All annotation work was done in a tool built for this annotation process to allow our team to classify links efficiently.

Coverage grows quickly as AI summaries tend to cite the same links over and over again. Even annotating a small set of links can lead to fairly high coverage of references.

To assess annotation consistency, 178 links were coded independently by at least a second team member. For those links, annotators agreed on which ballot initiative the link was about (so the link topic) in 87% of cases. Among those links with matching initiative labels, 83% also agreed on the exact position label, which indicates good agreement between annotators. The 47 links where annotations conflicted were manually reviewed and resolved by the lead investigator.

How we analyzed the data

In our analysis, which we reported on in our press release (in German and French), we narrowed our full dataset down to a subset using the following criteria:

Termination initiative. We included search terms from the category “termination initiative” in our data collection process. However, we filtered out data from these search queries in our analysis as we wanted the number of pro and con framings in our search queries to be roughly equal. We considered the “official name” and “10 million initiative” category in the search terms table as rather neutral and the “sustainability initiative” category as framed in a manner aligned with the “pro” side (arguably, it could also be seen as rather neutral as “sustainability initiative” was also the official name of the initiative). The “chaos initiative” and “termination initiative” categories are clearly terms used by the con-side, as they argue the initiative would lead to chaos and the termination of important bilateral agreements. This meant we had more con-framings than pro-framings. Given the choice between filtering out the “chaos initiative” or “termination initiative” data, we chose to keep the “chaos initiative” data as we found the chaos terminology to be more representative of the public debate.

The next two filters we applied only affected our results that involve references (so not the numbers of the AI summary frequency):

Links that we labeled. As mentioned, we achieved about 82% coverage in our data labeling step. When discussing the references in the AI summaries, we restricted our analysis to this subset of labeled links.

Links that focus on the 10 million initiative. Among all the labeled links, we further filtered our dataset to those links that we classified, after reading the linked material, as being about the 10 million initiative. We decided on this as this subset of the data only contains links directly relevant to the 10 million initiative, which is what is relevant for our research on the 10 million initiative. Note: Some links provide an overview of all ballot items that were voted on on June 14. While these might also take a position on the 10 million initiative (e.g., the voting recommendations by the parties), annotating them in a way that makes them useful for the analysis would have required a more complicated data structure for the annotation process. Seeing that only about 5% of our annotated links fell into this “ballot overview” category, we considered this share as negligible. Our analysis, thus, by default, filters for links that have been annotated as focusing on the 10 million initiative.

We notably did not exclude the testing period data from April 13-20, 2026, as this was the time when we collected the majority of our data on Bing. To keep things consistent and because it barely changed our results for Google, we also did not exclude this data from the Google analysis.

Note that we did not apply this filtering for this blog post and just reported total numbers from the raw dataset.

Our dashboard as an interactive analysis tool

The dashboard is designed for exploration. This means it has different views on the data and filters built in, so that users can explore the data from different angles.

The top of the interface shows summary statistics, such as how many queries returned AI summaries and how many returned the standard search results page without an AI summary. Below that on the right, users choose between different views of the data (e.g., most-cited domains, most-cited entities). In many views, users can choose between display styles, e.g., a bar or line chart. On the left, a filter panel lets users narrow the dataset to exactly the slice of the data that interests them, e.g., a certain time period, search engine, search query language, suffixes and more.

By default, the dashboard filters the data in the same way as we did for the analysis (see “How we analyzed the data”). These default filters can, however, be changed by the users in the filter panel.

What we can learn from the data

Our data confirms that generative AI is already deeply integrated into search engines in Switzerland. This is not just the case for the largest language groups, so German and French, but also for Italian, English, and even Romansh. At the scale at which search engines, and in particular Google, are used, this means that millions of AI summaries are seen by people in Switzerland every day. In Google’s recent I/O, it became clear that the importance of generative AI in search is going to grow in the near future.

Yet, our results underline that these AI summaries are not neutral. Rather, different AI summaries are imaginable, as shown by the observed differences between Google and Bing. They differ in which sources they cite and how often they cite these. We see the same for differences between languages. Differences are not random, but due to choices by the search engine providers. However, the criteria by which AI summaries are generated are not transparent at this point and difficult to investigate.

Therefore, it is clear that if we want to address the growing influence of AI in search in Switzerland, the draft law on communication platforms and search engines (KomPG/LPCom) must account for the increasing significance of generative AI within search engines. Without that, the arguably most consequential and increasingly central part of the search experience may fall entirely outside the law's reach.

AI transparency note

The team used Claude Code to write the scrapers, build the annotation tool, the dashboard and to support analysis. We deemed this proportionate due to the increased speed of programming and the cleaner code structure. We checked the plausibility of the results of the scrapers by manually searching some search terms and checking the frequency of AI summaries. To ensure the correctness of the analysis, we cross-checked the numbers from the dashboard with alternative calculations in Jupyter notebooks. The author also used Claude to write a first draft of the blog post based on detailed notes. She edited the text, along with multiple colleagues, to ensure correctness and clarity. Find out more about AlgorithmWatch’s guideline to use generative AI responsibly here.

AlgorithmNews CH – abonniere jetzt unseren Newsletter!

Ich bin mit der Verarbeitung meiner Daten einverstanden und weiss, dass ich den Newsletter jederzeit abbestellen kann.