Can Generative AI Make Culture More Homogeneous? Why Researchers Are Studying the “Averaging” Effect of AI

The most interesting problem with generative AI is not that it will start producing bad writing, weak images, or mediocre ideas. A more plausible scenario is almost the opposite: average quality rises while diversity falls. Each user gets a reasonably strong suggestion, a cleaner sentence structure, a more coherent story, or a more polished concept. The catch is that thousands of other people using similar models are receiving suggestions drawn from much the same statistical landscape.

Researchers describe this as homogenization, convergence, or a form of algorithmic monoculture. The “averaging effect” is a useful shorthand, but it does not mean that a model literally calculates an average of existing culture. The practical problem is subtler: outputs from different users begin to cluster around the same semantic regions, narrative structures, visual conventions, and associations.

That distinction matters. Two texts do not need to contain the same sentences to be creatively similar. They may have different characters, vocabulary, and length while relying on nearly identical premises, conflicts, and endings. In advertising, that can mean another campaign built around “authenticity.” In brand communication, another manifesto about changing the world. In visual generation, it can mean recurring compositions, faces, lighting styles, and stereotypical portrayals of particular groups.

That is what researchers are now trying to measure.

AI can improve individual work while making the group more alike

One of the most widely discussed experiments was published in *Science Advances* in 2024. It involved 293 writers who were asked to produce an eight-sentence story in English. They were assigned one of three themes: an adventure at sea, in the jungle, or on another planet.

Participants were randomly divided into three groups. The first wrote without AI assistance. The second could ask GPT-4 for one three-sentence story idea. The third could request up to five such ideas.

Participants used the option enthusiastically: 88.4% of those with access to AI requested at least one suggestion. The resulting 293 stories were then evaluated by an independent panel of 600 readers, producing a total of 3,519 evaluations.

At the level of an individual writer, AI performed well. Access to one AI-generated idea increased perceived novelty by 5.4%, while access to as many as five ideas increased it by 8.1%. A usefulness measure covering factors such as appropriateness, feasibility, and publishability increased by 3.7% and 9.0%, respectively.

The benefits were even larger for participants who had previously scored lower on the Divergent Association Task, or DAT, a test used as a proxy for creative potential. With access to five AI-generated ideas, their stories were rated as:

  • up to 10.7% more novel,
  • up to 11.5% more useful,
  • up to 26.6% better written,
  • up to 22.6% more enjoyable,
  • up to 15.2% less boring.

That is a genuine strength of generative AI. It can quickly raise the floor for someone who struggles with getting started, structuring a story, or extending an initial idea. In the experiment, the performance gap between less and more creative writers narrowed substantially.

The problem appeared when researchers stopped looking at each story in isolation and compared the entire set of outputs.

Stories written with AI assistance were more similar to other stories produced within the same condition. With one AI-generated suggestion, the increase in similarity amounted to 10.7% of the full range of variation observed in the no-AI group. With access to five suggestions, it was 8.9%.

The AI-assisted stories were also about 5% more semantically similar to the GPT-4-generated ideas themselves than stories written without AI support.

This is a practical example of anchoring. The first idea shapes the direction of later thinking. The writer still does most of the work, changes characters, adds details, and rewrites sentences, but the search has already begun inside a space defined by the model.

The same pattern can appear in a newsroom, design studio, or advertising agency. If five people receive five early ideas from the same system, each may produce competent work. The similarity only becomes obvious when all five outputs are placed next to one another: the differences are often in execution rather than in the underlying concept.

That creates the central paradox: AI can increase the creative performance of an individual while reducing the creative diversity of a group.

The averaging effect becomes more visible at scale

A more demanding test looked at 2,200 college admissions essays. Instead of asking only how similar two texts were, researchers developed a measure called the *diversity growth rate* — effectively, the rate at which diversity increases as more texts are added to a collection.

The question was straightforward: how much genuinely new semantic material does each additional text contribute to the whole set?

In the first study, researchers compared 100 real admissions essays written by applicants between 2018 and 2022 with 100 essays produced by GPT-4. The diversity growth rate was 0.58 for human-written essays and 0.18 for GPT-4. When the number of texts doubled, collective diversity increased by 0.40 and 0.13, respectively. The AI result was therefore only about 31% of the human level.

The researchers then tried to fix the problem.

A second experiment examined 400 essays divided into four groups: human-written essays, standard GPT-4 outputs, GPT-4 outputs generated with a simple instruction to be more creative, and GPT-4 outputs produced with modified generation parameters.

Simply telling the model to be more creative did not solve the issue. More aggressive changes to generation settings improved diversity more clearly, but the diversity growth rate of human essays was still about 67% higher than in the most diverse AI condition created through parameter changes.

A third experiment covering 1,600 essays produced an even sharper result. For standard GPT-4 outputs, the diversity growth rate was only 0.07, compared with 0.62 for human-written essays. The standard GPT-4 result did not even reach statistical significance. More elaborate prompting raised the model’s score to 0.31 — a substantial improvement, but still only around half the human level.

Across the three studies, the practical conclusion was striking: additional human-written essays increased collective semantic diversity roughly two to eight times more strongly than standard GPT-4 essays.

That does not mean every human essay was better. It measures something different. AI can generate an excellent single essay while still contributing much less novelty when the 50th, 100th, or 500th output is added to a large pool.

By August 2026, the effect had broader support. A meta-analysis covering 19 studies and 61 effect sizes found a statistically significant homogenization effect associated with generative AI use. The overall effect size was d = 0.334 — small, not evidence of cultural collapse, but consistent enough to remain meaningful after robustness checks.

Crucially, not every task behaved in the same way. Homogenization was stronger when participants worked inside a narrower idea space. In highly open-ended tasks requiring divergent thinking, the effect was weaker.

There are also results pointing in the opposite direction. One experiment involving more than 800 participants from over 40 countries found that extensive exposure to AI-generated ideas could increase the rate at which collective diversity changed. That matters because the simple claim that “AI always makes everything more similar” is wrong. The outcome depends on how suggestions are presented, how constrained the task is, and whether AI supplies one default direction or deliberately expands the search space.

A related problem appears in image generation, although it takes a different form. A 2025 study examined Stable Diffusion XL across six racial categories, two genders, 32 occupations, and eight traits. In the part of the experiment focused on facial homogenization, researchers generated 1,000 images for each analyzed group.

Representations of Middle Eastern people were among the most homogeneous. Male subjects were often compressed into a recurring set of visual features, including similar facial hair, skin tones, and traditional clothing cues. After the model was fine-tuned specifically to increase diversity, average similarity for that group fell from 0.61 to 0.41.

Here the issue is no longer just whether five images look alike. It is about compressing a diverse population into a small number of easily recognizable visual shortcuts.

For European markets, including Poland, the practical implication is straightforward. Writing a prompt in Polish does not automatically create cultural specificity. If the source of an idea for a campaign about a Polish family, a small town, a regional tradition, or a local holiday is mainly a global foundation model, the result can be polished and technically correct while remaining culturally interchangeable with material designed for several other countries.

There is still no basis for claiming that Polish, European, or global culture as a whole has already become measurably homogeneous because of AI. Most studies examine specific tasks, often in English, using particular model versions under controlled conditions. Models from 2023 or 2024 are also not identical to systems used several years later.

The evidence is strong enough to identify a mechanism and a scaling risk. It is not strong enough to declare that entire cultures have already converged.

How to use generative AI without producing the same culture in multiple versions

The weakest creative workflow looks like this: a team receives a brief, everyone opens a similar model, types some version of “give me 10 creative ideas,” and then chooses one of the first suggestions. It is efficient. That is precisely why it is risky.

If originality is the objective, the order should be reversed.

  1. Generate human ideas before opening the model. This does not require a three-hour workshop. Each person should write down several independent directions first. That prevents the first AI response from becoming the anchor for the entire process.
  1. Use AI later for expansion, criticism, and stress-testing. A model is often more useful as an adversarial editor than as the source of the original concept. It can identify predictable elements, challenge assumptions, suggest neglected audiences, or create alternative structures. That is generally safer than allowing it to define the entire creative space from the start.
  1. Do not treat “be more creative” as a safety mechanism. Research on essays shows that simple creativity prompting does not necessarily increase collective diversity. A better method is to deliberately separate the search space: one version can ban the most obvious motifs, another can be grounded in local source material, and a third can start from assumptions that directly oppose the original brief.
  1. Audit the batch, not just the best output. If a company produces dozens of ads, product descriptions, scripts, or concepts, those outputs need to be compared with one another. A basic audit can flag recurring story structures, metaphors, arguments, calls to action, character types, and cultural references. At larger scale, teams can use embeddings, semantic clustering, and cosine similarity.
  1. Supply local specificity from outside the model. Interviews with residents, customer language from real research, local archives, photographs, field reporting, and original primary material are more valuable than telling a model to “make it feel more Polish” or “more European.” Regional vocabulary, habits, and cultural detail should come from real evidence rather than from what the model statistically associates with a place.
  1. Using several models is a secondary safeguard, not the main one. Different model families may reduce dependence on one system’s default patterns, but they do not guarantee diversity. Models may still rely on overlapping data, similar conventions, and similar definitions of what counts as a strong answer.

The decision rule is fairly hard-edged. If the task is mainly about speed, consistency, standardization, and producing many acceptable variants, AI can enter the workflow early. Product descriptions, language cleanup, summarization, routine localization, and baseline copy variation fit that model well.

If the competitive advantage depends on an original idea, a distinctive authorial voice, local identity, or an unusual point of view, the process should not begin with the model’s answer. The team should build its own idea space first and introduce AI later.

A simple internal test can reveal whether averaging is already happening. Give the same brief to two groups. One produces a set of ideas without generative AI; the other uses it. Collect 15, 20, or 30 outputs from each group, then ask someone who does not know which process produced which material to cluster them by underlying concept.

There is no universal similarity threshold that automatically proves homogenization. The useful benchmark is your own control group. If 20 AI-assisted outputs consistently collapse into fewer genuinely independent concepts than 20 human-first outputs, the risk is real for that workflow — even if the single best AI-assisted result looks excellent.

This kind of audit also exposes the main inconvenience of anti-homogenization work: it slows down the very thing AI is best at accelerating. Teams must keep rejected variants, compare them, and sometimes discard a polished output simply because it resembles too many others. Increasing randomness can also produce stranger, less coherent material. Researchers adjusting generation parameters encountered exactly this trade-off: more variation does not automatically mean more meaningful creativity.

There is no single “anti-averaging prompt.” The strongest protection is a process in which the model does not get to define all of the first directions.

FAQ: Generative AI and cultural homogenization

Does generative AI actually reduce creativity?
Not in a simple sense. In several experiments, individual users produced work that was rated as more novel, better written, or more enjoyable when they had AI assistance. The problem appeared at group level: strong outputs from different people became more similar to one another.

Is there proof that AI is making entire cultures homogeneous?
No. Research has identified measurable homogenization in specific creative tasks, and meta-analytic evidence points to a small but statistically significant overall effect. That justifies monitoring the risk, but it does not prove that entire national or global cultures have already become homogeneous because of generative AI.

Is asking the model to “be more original” enough?
No. Experiments using different prompts and generation settings show that diversity can improve, but a simple creativity instruction does not reliably eliminate the gap between AI-generated collections and human-generated ones. The whole batch has to be evaluated, not just a single response.

Will using several different AI models solve the problem?
It can reduce dependence on one model’s defaults, but it is not a guarantee. A stronger setup combines independent human ideas, varied source material, and only then assistance from multiple models.

Does the averaging effect also apply to AI-generated images?
Yes, although it can appear differently. Image-generation research has found repeated visual features and stereotyped portrayals across social groups. Visual homogenization can therefore mean not only similar aesthetics but also a narrower range of ways in which people and cultures are represented.

How can a newsroom, agency, or company test for homogenization?
Run the same brief through two workflows: one without AI and one with AI. Produce at least a dozen outputs per group, then compare the number of genuinely distinct concepts, recurring motifs, structures, and phrasing patterns. At larger scale, add semantic similarity analysis. Your own human-generated control set is more useful than an arbitrary industry-wide similarity threshold.

What should a team do first when originality matters most?
Start with human ideas. Gather observations, source material, lived experience, and several independent concepts before asking a model for anything. Bring AI in later to expand, challenge, and test those directions rather than allowing it to decide where the creative search begins.

Leave a reply

Your email address will not be published. Required fields are marked *