October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
ThatPainter
archaeology

AI Hasn’t Solved a 60,000-Year-Old Cave-Art Mystery—but It May Offer a New Clue

Machine learning found possible sex-correlated patterns in modern experimental finger flutings, but it did not analyze or identify the makers of 60,000-year-old cave marks.

By ThatPainter Team 6 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ThatPainter is reader-supported. When you buy through links on our site, we may earn an affiliate commission. Learn More

No. Artificial intelligence did not identify the people who made 60,000-year-old cave marks. A Scientific Reports study published on October 16, 2025, trained image-classification models on finger flutings made by 96 modern adults. The experiment found potentially useful patterns in tactile marks, but performance on unseen data was unstable, the virtual-reality results were unreliable, and the researchers reported no independent archaeological validation. The work is a proof of concept, not a solved prehistoric mystery.

As an Amazon Associate I earn from qualifying purchases.

Read the primary study in Scientific Reports.

What the viral “60,000-year-old puzzle” claim gets wrong

Headlines described the project as AI solving the identity of prehistoric cave artists. That framing confuses the age of the archaeological phenomenon with the age of the data analyzed. The models never received photographs of 60,000-year-old grooves and did not name an ancient artist, determine a mark-maker’s species, or establish whether a particular person was male, female, or a child.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Instead, researchers led by Andrea Jalandoni at Griffith University collected new marks from living volunteers, then tested whether computer-vision models could classify the participants into two self-reported sex categories. The result suggests a research direction that could eventually be tested on archaeological material; it does not yet support an identification.

What are prehistoric finger flutings?

Finger flutings, also called digital tracings, are grooves made by dragging one or more fingers through a soft cave deposit. The material was often moonmilk, a calcium-carbonate-rich deposit that can be compacted by touch. Paleolithic examples are known from western Europe and Australia and span roughly 60,000 to 12,000 years before the present.

They are not the same as painted handprints or hand stencils. A stencil is made by blowing pigment around a hand; a fluting is a physical incision in a soft surface. Flutings may preserve clues about how many people participated, hand preference, movement, body position, and individual mark-making habits. Their social or ritual meaning remains uncertain.

Archaeological contexts associate finger flutings with both Homo sapiens and Neanderthals. That association does not assign every groove to one species, age group, sex, or individual.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the experiment worked

Participants and labels

The team recruited 96 adult volunteers in Australia during 2024 at the Australian Archaeological Association Conference, Griffith University, and SAE University College. Participants supplied information such as age, height, handedness, hand measurements, and self-reported sex. Children were excluded, and the sample was not designed to represent every population.

The prediction target was a binary survey label. It was not a direct measurement of biological sex, gender identity, artistic role, or cultural affiliation.

A tactile surface designed to resemble moonmilk

Each participant made nine marks: eight prescribed gestures and one freehand gesture. They worked on a purpose-built material intended to adhere to a vertical canvas, preserve grooves, and approximate the texture and behavior of moonmilk. Researchers photographed the marks under controlled conditions because obtaining enough genuine moonmilk for a large experiment would not be practical.

A virtual-reality comparison

Participants also made digital flutings with hand tracking in a virtual-reality environment using a Meta Quest 3 headset. VR offered repeatable recording, but it did not reproduce the resistance, moisture, grain, and tactile feedback of a physical cave surface. That difference is important because pressure, speed, finger angle, and wrist movement can change a groove.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

EurekAlert’s study summary describes both experimental settings.

The models and image sets

The researchers trained two convolutional neural networks, ResNet-18 and EfficientNet-V2-S, on images of the flutings rather than on a selected measurement such as finger-length ratios. Marks from a participant were kept within either the training or test split, preventing the same person’s examples from appearing in both.

Condition Training images Test images What it represents
Tactile material 573 126 Photographs of physical grooves
Virtual reality 666 152 Digitally generated flutings

The datasets were not perfectly balanced between the two labels, so accuracy alone cannot describe performance fairly.

What the AI was—and was not—asked to predict

  • It was asked to classify the modern participant’s binary, self-reported category.
  • It was not asked to identify a named person or distinguish one individual from another.
  • It was not trained to distinguish Neanderthals from Homo sapiens.
  • It was not trained to determine age, identify children, infer gender identity, or interpret cultural meaning.

The models detected statistical image patterns. The study does not show that those patterns correspond to anatomy, nor that they would survive changes in cave material, lighting, preservation, or movement.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the results actually show

Tactile marks contained a possible signal

For some tactile configurations, the paper reports area-under-the-curve (AUC) values above 0.85 during training. That indicates that the model found separable patterns in the experimental images. However, the authors also report a substantial and unstable gap between training and held-out test performance.

That pattern is consistent with overfitting: a model can memorize details specific to volunteers, the substitute material, camera setup, or laboratory conditions instead of learning a general feature of how people make flutings. A secondary article popularized an approximately 84% accuracy figure, but that number is meaningful only with its exact model, data split, class balance, and evaluation context. It is not a validated accuracy rate for ancient cave marks. See the secondary headline coverage for the source of that framing.

VR did not produce a dependable classifier

The virtual-reality results were inconsistent and did not provide sufficiently distinct, stable features for reliable classification. The absence of physical resistance is one plausible explanation, but the experiment does not prove that it is the only one.

Why this is not an ancient-artist identification system

The training world and the archaeological world are separated by several layers of uncertainty:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Different surfaces: a laboratory substitute is not every ancient moonmilk deposit. Resistance, moisture, grain size, and elasticity affect grooves.
  • Different images: camera, lighting, shadows, framing, and wall orientation can create signals a model detects more easily than the mark itself.
  • Different people and circumstances: ancient makers had unknown motor habits, body proportions, cultural practices, and working positions; the experiment included only modern adults.
  • Degradation: ancient grooves can erode, widen, overlap, or become partly obscured.
  • Out-of-distribution prediction: applying a model trained on modern experimental images to archaeological photographs would require independent testing against data from outside the original experiment.

The paper describes the method as requiring more samples and external validation before application to ancient sites. Its public code is available at FingerFluting-SexClassification.

Best Value
Sale
Modern Art. A History from Impressionism to Today (Bibliotheca Universalis)
  • Over 200 paintings, sculptures, photographs, and conceptual pieces
  • Length: 7.75in / 20cm, Depth: 2in / 5cm, Width: 6in / 15cm
  • By Hans Werner Holzwarth
  • Hardcover
  • 696 pages
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why researchers moved beyond finger-ratio measurements

Earlier attempts sometimes used the 2D:4D ratio—the relative lengths of the index and ring fingers—to infer the sex of a fluting’s maker. A groove is not a direct cast of a finger, however. Pressure, arm height, wrist and palm angle, humidity, surface properties, and later widening can all alter its apparent width and shape.

Machine learning offers a testable alternative: analyze the whole image rather than prespecifying one disputed measurement. But replacing a questionable measurement with a complex model does not remove the need for validation. A high-capacity network can also learn accidental features that researchers cannot easily interpret.

Why a cautious result still matters

A reliable method for comparing mark-making patterns could help archaeologists test assumptions about who participated in prehistoric art. Women’s contributions to ancient artistic activity have often been understudied, and a validated tool could examine sex- or age-related hypotheses instead of treating male authorship as the default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That goal requires careful language. A modern binary label is not the same as a complete account of biological sex or gender. Even a reproducible correlation would not show that sex determined an ancient person’s technique, social role, or artistic intention.

What would count as a genuine breakthrough?

  1. Larger, diverse samples: recruit participants across populations, ages, hand preferences, and physical characteristics, rather than relying mainly on conference and university groups.
  2. Realistic surfaces: test multiple cave-surface materials and document moisture, resistance, texture, and preservation effects.
  3. Independent external tests: evaluate locked models on images collected by different teams, cameras, lighting setups, and locations.
  4. Blind replication: preregister analyses and test whether new laboratories reproduce the result.
  5. Archaeological calibration: compare predictions with independently dated, well-preserved flutings while reporting uncertainty instead of forcing a categorical answer.
  6. Interpretability checks: determine whether predictions rely on groove geometry and movement patterns or on image artifacts and participant-specific context.

Bottom line

The 2025 study demonstrated that machine learning can find potentially informative patterns in finger flutings made by modern volunteers, especially on a tactile surface. It did not identify the makers of 60,000-year-old cave marks, distinguish Neanderthals from modern humans, identify children, or reveal artistic intent. Its real contribution is methodological: a promising experimental framework that remains preliminary until larger, independently validated studies show that the signal transfers to archaeological evidence.

Quick Recap

SaleBestseller No. 3
SaleBestseller No. 5
Modern Art. A History from Impressionism to Today (Bibliotheca Universalis)
Modern Art. A History from Impressionism to Today (Bibliotheca Universalis)
Over 200 paintings, sculptures, photographs, and conceptual pieces; Length: 7.75in / 20cm, Depth: 2in / 5cm, Width: 6in / 15cm
$23.30

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

More from the Paint Desk

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.