How to See Like a MachineCh. 1 · Invisible Images

A party game · Chapter 1

What will the machine say?

AlexNet is a 2012 neural network trained on ImageNet. It names any picture with one of exactly 1,000 nouns. We showed it 25 artworks. Guess which noun it picked.

Round
–
Score
0
Streak
0
Loading…

What did AlexNet call this?

Loading the first painting…

How this works

Each artwork was run once, offline, through four classifiers: AlexNet (2012), ResNet-50 (2015) and ViT-B/16 (2020), all trained on ImageNet’s 1.28 million labelled photos, and CLIP ViT-B/32 (2021), trained on web image captions and here asked to choose among the same 1,000 ImageNet nouns. The three ImageNet models use their standard preprocessing: shrink the short side, cut out the center square, normalize the colors. The wrong answers in each round come from the other models’ guesses and AlexNet’s own runners-up.

Images are public domain (Met Open Access, Wikimedia Commons). Sources, code and known limits are in the README.

The machine’s whole vocabulary

Every answer these classifiers can give is one of these 1,000 ImageNet classes. Nothing outside the list is possible.