# What will the machine say?

A party game for **Chapter 1, "Invisible Images"**, with a preview of Chapter 2. Players see a public-domain artwork and guess which of ImageNet's 1,000 nouns AlexNet gave it. The reveal then shows what four classifiers said.

Open `index.html` through a local server (see [Run it](#run-it)). It needs no build step and no server code. The page loads `predictions.json`, `artworks.json` and `images/` by relative URL.

## The concept

Chapter 1 (section iii) says a convolutional network trained on ImageNet was "quite sure" Manet's *Olympia* was a burrito, and that anyone holding something tends to be read as holding a cell phone. Paglen draws two points from this:

- A network can't invent a class. It can only relate a picture to the classes it was trained on.
- The training set records the historical, geographic and class position of the people who made it.

Fig. 1.2 (Paglen's *"Fire Boat"*, a synthetic high activation for AlexNet trained on ImageNet) shows the model this game uses. Fig. 1.5 (Rosler's *Semiotics of the Kitchen* as seen through an image classifier) shows the look the reveal copies: a box and a label stamped on an artwork. The vocabulary drawer previews Chapter 2: ImageNet's classes are WordNet nouns, and the 1,000-class subset the classifiers learned is the only vocabulary they have.

The game asks players to predict the machine's answers instead of reading the picture's meaning. To score, you have to stop asking what the painting is of and ask what white cloth, a gold border or a pointed arch will set off in the network.

## Models and why

| Model | Year | Weights | Trained on | ImageNet top-1 | What it sees |
|---|---|---|---|---|---|
| AlexNet | 2012 | torchvision `AlexNet_Weights.IMAGENET1K_V1` | ImageNet-1k photos | 56.5% | short side to 256, center 224×224 |
| ResNet-50 | 2015 | torchvision `ResNet50_Weights.IMAGENET1K_V2` | ImageNet-1k photos | 80.9% | short side to 232, center 224×224 |
| ViT-B/16 | 2020 | torchvision `ViT_B_16_Weights.IMAGENET1K_V1` | ImageNet-1k photos | 81.1% | short side to 256, center 224×224 |
| CLIP ViT-B/32 | 2021 | `openai/clip-vit-base-patch32` (Hugging Face) | 400M web image–caption pairs | n/a | short side to 224, center 224×224 |

- **AlexNet** is the guessing target because it is the network from the era Paglen writes about (Fig. 1.2). Its weights are kept on purpose rather than upgraded.
- **ResNet-50 and ViT-B/16** show that a network 25 points more accurate on ImageNet still answers with the same 1,000 nouns. They are often wrong just as confidently.
- **CLIP** never saw ImageNet's labels. It learned from web captions, and here it is scored zero-shot against `"a photo of a {class}."` for the same 1,000 class names. Forcing it onto ImageNet's list shows that the vocabulary limits it, not just the training photos.

Each ImageNet model uses its own standard `weights.transforms()` preprocessing. Every model runs once on MPS in eval mode, on the same resized JPEGs the page shows. The "% sure" in the page is the softmax probability. It is not a calibrated confidence, and CLIP's numbers are not on the same scale as the others'.

## Run it

From the repo root:

```sh
# play locally
python3 -m http.server 8103 --directory explorations
open http://localhost:8103/ch01-guess/

# regenerate images, artworks.json and predictions.json (about 3 minutes on an M5 Pro)
uv run python explorations/ch01-guess/precompute.py            # reuses images already in images/
uv run python explorations/ch01-guess/precompute.py --refetch  # downloads every image and record again

# check the page in headless Chromium (with the server above running)
uv run --with playwright python explorations/ch01-guess/verify_page.py
```

`precompute.py` has the curated list of artworks at the top, with Met object IDs and Commons file names. It only calls the Met's `/public/collection/v1/objects/{id}` endpoint and the Commons `imageinfo` API, with a pause and backoff, because the Met's CDN starts returning 403 after rapid requests. Without `--refetch` it skips downloads for images already in `images/` and re-applies hand-edited fields (titles, `depicted`) from the list. The first run downloads the three torchvision checkpoints (about 660 MB) into `~/.cache/torch`. CLIP loads from the local Hugging Face cache with `HF_HUB_OFFLINE=1`.

Outputs:

- `images/<id>.jpg`: 25 JPEGs, at most 900 px on the long side, quality 82.
- `artworks.json`: title, artist, date, institution, source URL, license, `theme`, `depicted` and `book_label` for each artwork.
- `predictions.json`: the 1,000 class names, model metadata and top-5 `[class_index, probability]` per model per image. Its `watch` block records each image's rank among all 1,000 classes for "burrito", "cellular telephone" and "Granny Smith". Its `diagnostics` block holds the full-frame AlexNet run and the Olympia crop probe described below.

## How the game works

- **Round.** Each game is 10 rounds. *Olympia* always comes first because the chapter names it; the rest are shuffled. There are four options: AlexNet's top-1 answer, one other model's top-1, one of AlexNet's own runners-up, and one more from the other models' top-5. Keys 1–4 pick an option and N goes to the next round.
- **Reveal.** The magenta box on the image marks the 224×224 center square AlexNet actually saw. Below it, four panels show each model's top 5, with AlexNet's answer highlighted wherever it appears. Up to two observations are generated from rules over the data: where "burrito" ranks (only for *Olympia*); "cellular telephone" if it ranks in the top 50 for a picture of someone holding something; where "Granny Smith" ranks on the apple pictures; whether all four models agree or all differ; whether AlexNet's answer appears in no other model's top 5; and whether CLIP makes a confident (≥ 40%) call that none of the ImageNet models made.
- **Score.** The scoreboard shows score and current streak and stays pinned to the top while scrolling. Options are at least 56 px tall so a phone can be passed around.
- **Vocabulary drawer.** "The machine's whole vocabulary" lists all 1,000 classes with their index numbers and can be searched. "burrito" and "cellular telephone" each match exactly one class. "person" and "painting" match nothing. "apple" matches only "pineapple" and "custard apple"; the drawer adds that "Granny Smith" is the only apple. "horse" matches only "horse cart"; the only horse class is "sorrel".
- **End screen.** For each artwork played, it takes the most confident top-1 answer from any of the four models and lists the five most confident. Answers that name something actually in the picture are left out. Those are listed by hand in each artwork's `depicted` field: for example, "volcano" for Hokusai (Mount Fuji), "brass" for the Benin plaque (ImageNet's "brass" is a memorial plaque) and "hoopskirt" for *Las Meninas*. The screen ends with one sentence from the chapter and a link back to it.

## What it shows

**Olympia and Las Meninas** (top-1, with probability):

| | AlexNet | ResNet-50 | ViT-B/16 | CLIP ViT-B/32 |
|---|---|---|---|---|
| *Olympia* | bassinet 38% | hoopskirt 11% | hoopskirt 43% | brassiere 41% |
| *Las Meninas* | barbershop 8% | hoopskirt 5% | vestment 22% | barbershop 60% |

**No model says "burrito" for Olympia.** This run doesn't reproduce the chapter's example:

- With standard preprocessing, AlexNet ranks "burrito" **39th of 1,000 (0.3%)**. ResNet-50 ranks it 296th, ViT-B/16 246th and CLIP 233rd.
- Squashing the whole frame to 224×224 instead of center-cropping changes AlexNet's answer to "hoopskirt" (53%), but "burrito" still isn't in its top five.
- The crop probe tried 147 square crops at five scales. "Burrito" was never AlexNet's top-1. Its best was 6.7%, on a crop of Olympia's torso and the paper-wrapped bouquet, where the top-1 was "Egyptian cat". ResNet-50 peaked at 2.8% and ViT-B/16 at 1.0%.

The book doesn't say which network, weights, reproduction or preprocessing produced "burrito", so this can't be checked directly. Several things differ from any 2012-era setup:

- torchvision's AlexNet follows Krizhevsky's 2014 "one weird trick" variant and was retrained by the PyTorch team. It is not the original 2012 release.
- Caffe-era pipelines used 227-pixel crops, BGR mean subtraction and often 10-crop averaging.
- The source image here is the Google Art Project scan, downsized to 900 px.

What the run does show is that "burrito" sits in AlexNet's neighbourhood for this painting (39th of 1,000) while being nowhere near the top for the newer models. The answer also depends on framing: the crop alone turns "bassinet" into "hoopskirt".

**Other results from this run:**

- **Everything becomes an ImageNet noun.** No answer names a painting or a person, because ImageNet-1k has no class for either.
- **Confident misreadings are common in modern models too.** ViT-B/16 calls *Krishna Spying on Radha* a "book jacket" (82%) and the Kongo *Mangaaka* figure a "mask" (80%). The Mughal album page of Shah Jahan and Dara Shikoh is a "tray" (63%).
- **Non-Western works are read as objects that look like them.** All four models call the Isfahan mihrab a "prayer rug", and AlexNet is 99% sure. Prayer rugs usually carry a niche-shaped design, so the network has a class for the rug but none for the architecture the rug copies. CLIP calls three Chinese and Japanese works "jinrikisha" (a rickshaw) and the Hokusai print "shoji".
- **Landscapes are read as geology.** Cole's *Oxbow* is a "volcano" for AlexNet (58%) and ViT-B/16 (64%). Hokusai's *Great Wave* is a "mixing bowl" for AlexNet (42%).
- **Holding things.** No picture of someone holding a letter, fan, book, pitcher or jewel gets "cellular telephone" in a top five. The closest are ViT-B/16 on Vermeer's *Woman in Blue Reading a Letter* (9th of 1,000) and AlexNet on the Shah Jahan page (29th). The chapter's cell-phone claim describes networks of its time on everyday photos; it doesn't hold up for these paintings.
- **There is no apple.** ImageNet-1k's only apple is "Granny Smith". On the three apple pictures its best rank is 19th (CLIP, Cézanne), and it falls to 346th or lower for the Maes (*Young Woman Peeling Apples*).

## What does this image do?

To a classifier, an artwork doesn't depict anything. It shifts probability across a fixed list of 1,000 nouns chosen from WordNet around 2010. *Olympia* sets off "bassinet", a baby's basket bed. The Isfahan mihrab sets off "prayer rug" at 99%. A Mandi miniature sets off "book jacket". These verdicts aren't interpretations that could be argued with. They are activations, and in a deployed system they would be passed downstream as facts.

For the players, the game turns this around. To score, you have to stop seeing the picture as a person sees it and predict what it triggers. That shift is what Chapter 1 asks for when it says images now look at us.

## Image sources and licenses

All images are public domain. Met images are Open Access (CC0) and were fetched as the object's full-resolution `primaryImage`, then downsized. Commons images are PD-Art reproductions of public-domain paintings, fetched as Commons thumbnails and downsized.

| Artwork | Artist | Date | Institution | License |
|---|---|---|---|---|
| [Olympia](https://commons.wikimedia.org/wiki/File:Edouard_Manet_-_Olympia_-_Google_Art_Project_3.jpg) | Édouard Manet | 1863 | Musée d'Orsay, Paris | Public domain (Wikimedia Commons) |
| [Las Meninas](https://commons.wikimedia.org/wiki/File:Las_Meninas_01.jpg) | Diego Velázquez | 1656 | Museo del Prado, Madrid | Public domain (Wikimedia Commons) |
| [Woman in Blue Reading a Letter](https://commons.wikimedia.org/wiki/File:Brieflezende_vrouw_Rijksmuseum_SK-C-251.jpeg) | Johannes Vermeer | ca. 1663 | Rijksmuseum, Amsterdam | Public domain (Wikimedia Commons) |
| [Young Woman with a Water Pitcher](https://www.metmuseum.org/art/collection/search/437881) | Johannes Vermeer | ca. 1662 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Portrait of a Young Woman with a Fan](https://www.metmuseum.org/art/collection/search/437391) | Rembrandt van Rijn | 1633 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Portrait of a Young Man](https://www.metmuseum.org/art/collection/search/435802) | Bronzino | 1530s | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [The Musicians](https://www.metmuseum.org/art/collection/search/435844) | Caravaggio | 1597 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Young Lady in 1866](https://www.metmuseum.org/art/collection/search/436964) | Édouard Manet | 1866 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Still Life with Apples and a Pot of Primroses](https://www.metmuseum.org/art/collection/search/435882) | Paul Cézanne | ca. 1890 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Young Woman Peeling Apples](https://www.metmuseum.org/art/collection/search/436934) | Nicolaes Maes | ca. 1655 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Adam and Eve](https://www.metmuseum.org/art/collection/search/336222) | Albrecht Dürer | 1504 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Madonna and Child](https://www.metmuseum.org/art/collection/search/438754) | Duccio di Buoninsegna | ca. 1290–1300 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [The Oxbow (View from Mount Holyoke, after a Thunderstorm)](https://www.metmuseum.org/art/collection/search/10497) | Thomas Cole | 1836 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Wheat Field with Cypresses](https://www.metmuseum.org/art/collection/search/436535) | Vincent van Gogh | 1889 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Composition VII](https://commons.wikimedia.org/wiki/File:Composition_VII_-_Wassily_Kandinsky,_GAC.jpg) | Vasily Kandinsky | 1913 | State Tretyakov Gallery, Moscow | Public domain (Wikimedia Commons) |
| [Under the Wave off Kanagawa (The Great Wave)](https://www.metmuseum.org/art/collection/search/45434) | Katsushika Hokusai | ca. 1830–32 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Irises at Yatsuhashi (Eight Bridges)](https://www.metmuseum.org/art/collection/search/39664) | Ogata Kōrin | after 1709 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Night-Shining White](https://www.metmuseum.org/art/collection/search/39901) | Han Gan | ca. 750 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Wang Xizhi Watching Geese](https://www.metmuseum.org/art/collection/search/40081) | Qian Xuan | ca. 1295 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [The Emperor Shah Jahan with His Son Dara Shikoh](https://www.metmuseum.org/art/collection/search/451283) | Nanha | ca. 1620 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Krishna Spying on Radha](https://www.metmuseum.org/art/collection/search/37990) | Unknown artist, Mandi (Punjab Hills) | ca. 1780–90 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [The Feast of Sada (Shahnama of Shah Tahmasp)](https://www.metmuseum.org/art/collection/search/452111) | Attributed to Sultan Muhammad | ca. 1525 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Mihrab (Prayer Niche), Isfahan](https://www.metmuseum.org/art/collection/search/449537) | Unknown artist, Iran | 1354–55 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Mangaaka Power Figure (Nkisi N'Kondi)](https://www.metmuseum.org/art/collection/search/320053) | Yombe-Kongo artist and nganga (ritual specialist) | ca. 1880–1900 | The Metropolitan Museum of Art | CC0 (Met Open Access) |
| [Plaque with Warrior and Attendants](https://www.metmuseum.org/art/collection/search/316393) | Edo artists, Benin City (Ìgùn Ẹ́rọ̀nmwọ̀n guild) | ca. 1540–70 | The Metropolitan Museum of Art | CC0 (Met Open Access) |

The book's text and figures are not reproduced here. The page quotes one sentence from Chapter 1, and this README cites figures by number.

## Known limitations

- **This isn't Paglen's setup.** "Burrito" isn't reproduced (see above). The torchvision AlexNet is a later re-implementation, and the book doesn't name its model, weights or preprocessing.
- **One crop, one reproduction.** Each verdict comes from a single center crop of one web reproduction. The crop discards a lot on wide images: AlexNet sees only the middle 42% of the width of Kōrin's screen and 60% of *Olympia*'s. Small changes in crop, color or JPEG quality can change the answer. The `diagnostics` block shows one such change.
- **The Met images are not always the whole work.** For the two Chinese handscrolls, the Met's primary image is a section of the scroll, and for Kōrin it is one of a pair of screens. The three objects (mihrab, *Mangaaka*, Benin plaque) are museum photographs against plain backgrounds, a different kind of image from the paintings.
- **"% sure" is a softmax output, not a calibrated probability.** CLIP's scale differs from the ImageNet models', and its zero-shot setup uses a single prompt template and torchvision's class names without any prompt tuning.
- **Hand judgments.** The `depicted` lists that keep "accurate" labels off the misreadings list are editorial. Decoys are drawn at random from the models' own guesses, so some rounds are much harder than others, and a decoy can be a near-synonym of the answer.
- **A small, curated set.** The 25 works were chosen to cover the brief: people holding things, apples, religious scenes, landscapes, abstraction, and 10 non-Western works. They are a provocation, not a sample. Except for four Commons images, all come from one museum's photography.
- **Fonts load from Google Fonts.** Without a network the page falls back to system fonts. Everything else is local.
- **The Met's search API changed.** `/public/collection/v1/search` was retired on 2026-10-01; its replacement is `/public/collection/v1.1/search` with `limit` and `offset`. `precompute.py` doesn't search, so it isn't affected, but older notes that use `/v1/search` will get HTTP 410.
