Seeing Van Gogh Through Data: The Hidden Limits of Scientific Methods & AI in Art Research

There is a particular way of speaking about Vincent van Gogh that has quietly changed in recent years.

Not in literary criticism.

Not in art theory.

But in the places where art becomes data.

In digital archives.

In image databases.

In algorithmic analysis models that process thousands of paintings simultaneously.

When navigating large digital art collections today, Van Gogh’s works no longer appear only as singular expressions of a human being who used color as existential condensation. They also appear as data points.

Color histograms.

Brushstroke densities.

Contrast distributions.

Compositional vectors.

Similarity clusters within large image datasets.

In some research projects, Van Gogh is embedded into broader stylistic networks: which painters of his time are statistically similar to him? Which visual patterns occur in his work with above-average frequency? How does his use of color change over time, measured in quantifiable color values?

The question is not whether this is interesting.

The question is what exactly we believe we are understanding through it.

Because at a certain point, something subtle but decisive happens:

Van Gogh becomes made explainable.

And that is precisely where the problem begins.


The Moment When Meaning Disappears from Data

If you observe such analyses long enough, a strange impression emerges.

Everything becomes precise.

Everything becomes comparable.

Everything becomes visualizable.

And yet something becomes quieter.

Not gone, but displaced.

Pic: Art Institute of Chicago on Unsplash

The experience of Van Gogh’s paintings—this restless, almost painfully vivid quality that cannot be reduced to mere color combinations—no longer appears in the models as a central element, but as a side effect.

The painting becomes an object with properties.

Not an event.

Here a fundamental tension becomes visible:

The more we measure, the more the focus shifts from meaning to structure.

This is not a flaw of the methods.

It is their consequence.


How Art History Learned to Count

For a long time, art history was an interpretive discipline.

It worked with:

- symbolic analysis

- style history

- iconographic interpretation

- historical contextualization

Paintings were read, not measured.

This has changed.

With the digitization of archives, something new emerged: large, searchable image corpora.

Suddenly, it became possible to compare thousands of works simultaneously.

This led to new methodological approaches now often grouped under terms such as “digital art history” or computational image analysis.

In these approaches, artworks are no longer only interpreted but systematically quantified:

- color distributions are statistically analyzed

- compositional structures are modeled geometrically

- motifs are classified in large databases

- similarities between artists are calculated algorithmically

Van Gogh then appears not only as an individual, but as a pattern within the data space of art history.

This is methodologically impressive.

And epistemologically ambivalent.


The Silent Shift in the Human Sciences

What becomes visible in art history is not an isolated case.

It is part of a broader movement.

Psychology took this path early.

Educational sciences followed.

Sociology as well.

Across disciplines, a similar trajectory appears:

Starting from qualitative, descriptive, or interpretive approaches, the focus increasingly shifts toward measurable variables.

Behavior is counted.

Learning outcomes are scaled.

Attitudes are transformed into questionnaires.

Even complex educational processes are translated into indicators.

This brings advantages.

Comparability.

Reproducibility.

Scalability.

But it also changes what counts as “knowledge” in the first place.


What Is Lost Without Disappearing

The problem is not that qualitative dimensions disappear entirely.

They do not disappear.

They are epistemically downgraded.

This means:

They are treated as secondary, interpretive, supplementary.

Not as central access points to the structure of reality.

In practice, this produces a quiet hierarchy:

- numbers are considered more stable than narratives

- statistics more objective than interpretation

- measurement closer to reality than meaning

But this hierarchy is not itself an empirical fact.

It is an epistemic decision.


Mixed Methods as a Repair Attempt

In response to this development, a counter-movement emerged in many disciplines: mixed methods.

The idea is simple:

If quantitative methods are too narrow and qualitative methods too subjective, combine both.

This produces triangulation:

a phenomenon is approached from multiple perspectives.

This sounds convincing.

And often it is.

But it does not necessarily solve a deeper problem.

Because even the combination of methods remains within the same framework:

it operates within what is already considered observable, measurable, or linguistically accessible.

The fundamental question remains:

Do we actually expand our understanding of reality—or only our descriptions of it?


A Brief Look at Metascience

In some fields, especially psychology and medicine, another movement has developed in parallel: metascience.

Here, the object of study is no longer primarily the world, but science itself:

  • reproducibility
  • publication bias
  • methodological standards
  • statistical robustness

This is an important development.

But here too, a certain boundary often remains:

critique is mostly directed at methodological processes within an established framework of scientific objectivity.

The question of the limits of that framework itself is rarely raised in a radical way.


Ontology, Epistemology, and the Blind Spot of Measurability

To clarify this, a simple distinction helps:

Ontology asks: What exists at all?

Epistemology asks: How can we know anything about it?

Methodology asks: Which procedures do we use for that?

In many empirical sciences, attention shifts strongly toward methodology.

Then toward epistemology.

The ontological question often remains implicit.

Here a central blind spot emerges:

Not everything that exists can be accessed, measured, or operationalized in the same way.

And not everything that can be measured exhausts what exists.


Back to Van Gogh

When we return to Van Gogh, this becomes tangible again.

Digital art history can say a great deal about him.

It can detect patterns, draw comparisons, reveal structures.

But it does not automatically answer the question many people actually care about:

Why does this painting affect us so directly?

Why does it have this emotional density?

Why does it seem to be more than the sum of its colors?

These questions do not disappear because we have better data.

They remain.

Only the language in which they are asked changes.


Conclusion: Between Measurability and Meaning

The development from classical interpretation to quantified analysis is not a distortion.

It is an expansion.

But it is not a completion.

What we observe is rather a shift:

  • from meaning to structure
  • from interpretation to comparability
  • from experience to data

And perhaps the real challenge is not to develop ever more refined methods.

But to remember that every method makes a certain kind of world visible—and renders others invisible.

Van Gogh remains a useful example in this regard.

Not because he can be measured.

But because he shows that what matters often begins precisely where measurability ends.

References 

Alvesson, M., & Sandberg, J. (2011). Generating research questions through problematization. Academy of Management Review, 36(2), 247–271.

Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533(7604), 452–454.

Bishop, D. V. M. (2019). Unreliable research: Trouble at the lab. Faber & Faber.

Borgman, C. L. (2015). Big data, little data, no data: Scholarship in the networked world. MIT Press.

Cobb, C. L., & Jockers, M. L. (2019). The digital humanities and literary studies. University of Illinois Press.

Ekbia, H., & Nardi, B. (2017). Heteromation, and other stories of computing and capitalism. MIT Press.

Feyerabend, P. (1975). Against method. Verso.

Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124.

Kitchin, R. (2014). The data revolution: Big data, open data, data infrastructures & their consequences. Sage.

Latour, B. (2005). Reassembling the social: An introduction to actor-network-theory. Oxford University Press.

McCarty, W. (2005). Humanities computing. Palgrave Macmillan.

Mitchell, T. M. (1997). Machine learning. McGraw-Hill.

Nosek, B. A., et al. (2015). Promoting an open research culture. Science, 348(6242), 1422–1425.

O’Neil, C. (2016). Weapons of math destruction. Crown.

Pinker, S. (2018). Enlightenment now. Viking.

Popper, K. (1959). The logic of scientific discovery. Routledge.

Ramsay, S. (2011). Reading machines: Toward an algorithmic criticism. University of Illinois Press.

Snow, C. P. (1959). The two cultures. Cambridge University Press.

Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley.

Wacquant, L. (2009). Punishing the poor. Duke University Press.


Inspired by HBS Puar 
Authored by Rebekka Brandt