Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
ThatPainter
AI

AI Beat Humans on a Creativity Test. What Does That Actually Mean?

A 2023 study found that chatbots beat average human performance on one divergent-thinking task—but the best humans still matched or exceeded AI.

By ThatPainter Team 6 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ThatPainter is reader-supported. When you buy through links on our site, we may earn an affiliate commission. Learn More

AI did not become universally more creative than humans. In a 2023 study, ChatGPT-3.5, ChatGPT-4, and Copy.ai scored higher on average than 256 human participants on a narrow divergent-thinking exercise: inventing unusual uses for ordinary objects. The strongest human responses still matched or exceeded the best AI responses.

That distinction matters. The result shows that language models can be remarkably effective at generating unusual associations. It does not prove that they have imagination, artistic intention, lived experience, taste, consciousness, or a humanlike desire to create.

What the study actually tested

The study used the Alternate Uses Task, a familiar psychology exercise for measuring one aspect of divergent thinking. Participants receive an ordinary object and must suggest uncommon uses for it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The objects were a rope, box, pencil, and candle. The researchers compared 256 human participants with ChatGPT-3.5, ChatGPT-4, and Copy.ai, which was based on GPT-3 technology. Each chatbot was tested 11 times for each object.

This is a useful test of how quickly and widely a system can move beyond an obvious association. It is not a complete test of creativity. A person can be excellent at generating possibilities yet poor at choosing, developing, or executing them.

What “beat humans” means

The headline refers mainly to average group performance. The chatbots produced higher average scores than the human group, particularly on measures of semantic distance. It does not mean that every AI answer was better than every human answer—or even that AI was better than the best human participants.

The paper’s central qualification is easy to lose in shortened headlines: the best human responses were at least competitive with, and in some comparisons better than, the best chatbot responses. Human performance was also more varied. Some people gave ordinary answers; others produced exceptionally strong ones.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A simple analogy is a calculator outperforming most people on arithmetic. That establishes impressive competence at the measured operation. It does not settle what mathematics means to a person, or whether the calculator understands a problem.

How the answers were scored

Semantic distance

One scoring method used an algorithm to estimate how far each answer was from an object’s conventional use. An ordinary use for a candle might involve light or heat. An answer conceptually remote from those associations would receive a higher originality-related score.

This captures one ingredient of novelty, but distance is not the same as quality. A bizarre answer may be far from the obvious use and still be useless, unsafe, or incoherent. Conversely, a practical idea may be valuable without being especially surprising.

Human ratings

Six evaluators also rated the answers for originality and creativity on a five-point scale. They were not told which responses came from AI. Human ratings make the comparison less dependent on an automated metric, although six raters remain a limited and potentially subjective judging pool. The full-text study describes both scoring approaches.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Neither method directly measures emotional impact, aesthetic power, cultural importance, feasibility, or whether an idea is genuinely new in human history.

Why language models can do well at this task

Language models have several advantages in rapid brainstorming:

  • They can generate many associations in seconds.
  • They have absorbed huge amounts of language containing metaphors, inventions, jokes, design concepts, and descriptions of unusual uses.
  • They do not experience embarrassment, fatigue, or social hesitation.
  • They can be explicitly instructed to prioritize novelty and unusualness.
  • They can produce a wide range of possibilities before settling on one.

Those properties can look like creativity in a benchmark. They do not, by themselves, show that the model is forming ideas from personal experience or pursuing a self-chosen creative purpose.

A creativity test is not creativity itself

Creativity is not one universally agreed ability. In ordinary creative work, people often care about several qualities at once:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Novelty: Is the idea original?
  • Usefulness: Does it solve a problem or serve a purpose?
  • Surprise: Does it move beyond an obvious continuation?
  • Fit: Does it work within a particular artistic, technical, or cultural context?
  • Development: Can the creator revise and improve it?
  • Meaning: Does it matter to someone or to a community?
  • Intention: Is someone trying to express, investigate, or accomplish something?

The Alternate Uses Task emphasizes divergent association. It says much less about the rest. A model can suggest an unusual use for a box without understanding why one idea would matter to a painter, a designer, a child, or a particular community.

Is AI “just remixing”?

That objection needs more care than a simple yes or no. Humans also create by combining memories, references, techniques, and existing ideas. Recombination alone does not disprove creativity.

The more useful questions are whether the combination is genuinely novel, useful or meaningful, appropriate to a goal, and capable of being evaluated and improved. We also need to ask who supplied the goal: the model, the user, or both.

It is possible for a model to produce a valuable new combination without having humanlike intentions. It is also possible for an answer to sound original while reflecting a familiar pattern or material encountered during training. Critics have raised the possibility that models may have seen examples of similar tasks or answers in their training data; that possibility is difficult to rule out completely, but it is not proof that any particular answer is copied. MIT Technology Review’s coverage discusses this broader interpretation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why real creative work is harder

Generating unusual possibilities is only one stage of making something worthwhile. Real creative work usually involves:

  • deciding which problem or subject deserves attention;
  • understanding an audience, place, history, or set of constraints;
  • selecting among many possibilities;
  • testing whether an idea works in practice;
  • revising after failure or criticism;
  • developing a coherent body of work over time; and
  • taking responsibility for the result and its consequences.

That difference is especially clear in painting and other visual arts. Producing many unexpected image prompts is not the same as developing a visual language, choosing what to leave out, responding to materials, or making a work that carries personal and cultural meaning. The benchmark does not test those things.

What later studies add

The 2023 finding was not an isolated result. Later research reported strong performance by GPT-4 on several divergent-thinking tasks, including alternative uses, consequences, and divergent associations. A 2025 Scientific Reports study compared GPT-4o, DeepSeek-V3, and Gemini 2.0 with a small human sample on divergent and convergent assessments, reporting that the tested systems outperformed that sample.

These studies broaden the evidence that generative AI can perform well on structured creativity tests. They do not turn those tests into a complete definition of creativity, and the systems tested in 2023 should not be treated as interchangeable with current products. See the later studies in 2024 and 2025 for their specific methods and limitations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the result is useful for

The finding is meaningful when the question is whether AI can:

  • produce many non-obvious associations;
  • help overcome a blank page;
  • suggest combinations outside a person’s usual habits;
  • generate rough concepts quickly; or
  • perform strongly on a defined verbal creativity task.

It is insufficient evidence for claims that AI can independently make meaningful art, originate a creative movement, understand why an idea matters, decide which human problems deserve attention, or create without prompts, training data, and evaluation.

The practical model is human plus AI

The most useful consequence is not “AI versus human.” It is a division of strengths.

AI can help with brainstorming, alternative compositions, rough drafts, counterexamples, variations in tone or structure, and deliberate escapes from habitual ideas. Humans remain responsible for framing the problem, selecting the worthwhile direction, checking originality and feasibility, supplying lived context, refining the work, and deciding whether the result deserves to exist.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a painter, that might mean asking an AI system for ten compositional departures from an initial concept, then rejecting most of them, testing one through sketches, and transforming the surviving idea through personal observation and material decisions. The value is in expanding the field of possibilities—not in treating every generated possibility as art.

How to interpret the headline precisely

The most accurate version is:

The tested chatbots scored higher than the average human group on one divergent-thinking benchmark, while exceptional humans still performed at least as well as the chatbots.

That is a significant result. It changes expectations about how capable machines can be at rapid idea generation. But it is a much narrower claim than “AI is more creative than humans.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Paint Desk

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.