A party game · Chapter 1
What will the machine say?
AlexNet is a 2012 neural network trained on ImageNet. It names any picture with one of exactly 1,000 nouns. We showed it 25 artworks. Guess which noun it picked.
The box is all AlexNet saw: the center square, shrunk to 224 × 224 pixels.
What did AlexNet call this?
What all four models said
Top 5 of 1,000 nouns · % = how sureGame over
Most confident misreadings
“We no longer look at images—images look at us.”
Chapter 1 doesn’t fault these networks for being wrong. A classifier can’t invent a category; it can only match a picture against the nouns it was trained on, and those nouns reflect the time, place and class of the people who chose them. The useful question is what the picture does to the machine: which of its 1,000 words it sets off.
Back to Chapter 1 · Invisible ImagesHow this works
Each artwork was run once, offline, through four classifiers: AlexNet (2012), ResNet-50 (2015) and ViT-B/16 (2020), all trained on ImageNet’s 1.28 million labelled photos, and CLIP ViT-B/32 (2021), trained on web image captions and here asked to choose among the same 1,000 ImageNet nouns. The three ImageNet models use their standard preprocessing: shrink the short side, cut out the center square, normalize the colors. The wrong answers in each round come from the other models’ guesses and AlexNet’s own runners-up.
Images are public domain (Met Open Access, Wikimedia Commons). Sources, code and known limits are in the README.