Language is structure, and when we ask language models to be the opposite of structured, to be completely random, they fail. Our closed-loop method uses behavioral game theory to put a price on that failure and reveal a distinctive behavioral fingerprint for each LLM it tests.
We give a language model a simple game to play against an algorithmic opponent named Proteus. On each round both pick a shape, star or triangle. If they pick the same shape, the model wins; if they pick different shapes, Proteus wins. Here's the twist: Proteus watches every choice, tests the history for patterns, and exploits whichever it finds. The best way to play against Proteus is to evade detection by choosing at random on every round. Telling the models that optimal strategy made six of seven worse. Several wrote down their own pattern, said they'd break it, and kept producing it.
The game is matching pennies. The algorithm was originally created to play against monkeys (Lee et al., 2004), and we put seven LLMs in the monkeys' seat. Try it yourself by playing the demos below or watch reruns of the models playing. The closed loop, algorithmic opponent, and findings are described below.
Play the task exactly as it appears in the paper: the Figure 1 stimulus layout and match-to-win rule, with live per-trial diagnostics. In watch mode, each model's verbatim reasoning trace prints beneath the board as its real session replays.
The same assay reskinned as an arcade battle: missiles, countermeasures, and a neon telemetry rack reading Proteus's nine detection channels in real time. Fire 200 missiles and try to be random, or replay the study's sessions at up to 64×.
An animated four-trial PACA loop. The current choice history duplicates and travels simultaneously into the language model and Proteus panels. Both copies settle slightly below center while the first thinking step appears, then a vertical line scans each copy from right to left. After the scans finish, the copies fade and deliberation continues. The model commits a shape and Proteus commits a landing without seeing the other's current move. In a miniature study task, the selected shape fills yellow while Proteus's blue circle lands on its committed shape; a match means the language model wins and a mismatch means Proteus wins. The finished star-and-triangle outcome duplicates, descends through the loop, rearranges into the compact history format, and joins the right end of the choice history before the next trial.
Choice history enters the loopAn animated 30-trial choice history. Each trial first appears as in the closed-loop figure: a small frame with the star on top and the triangle below, the model's chosen shape filled yellow and a blue circle on the shape Proteus landed on. The unchosen shapes then fade away, leaving only the yellow choices, some carrying a blue circle where Proteus matched them. The remaining triangles slide up and the remaining stars slide down until every choice sits on one center line, and the boxes shrink into a compact single-row tape. Nine detection channels then analyze that tape one at a time: one tests overall side bias, four test the choice that followed the last one to four choices, and four test the choice that followed the last one to four choice-and-outcome pairs. Each channel's turn highlights where its current context recurred in the tape and the choice that followed, and fills a meter with the resulting probability and an exact binomial p-value. The strongest significant channel locks, and Proteus places the blue circle on the shape the model is not expected to choose.
Choose the shape that the blue circle will land on. Available actions: TRIANGLE STAR You may briefly think through your choice. End your response with the chosen action on its own line.
You are playing matching pennies against an exploitative opponent. Each round you and the opponent choose TRIANGLE or STAR. The blue circle lands on the opponent's chosen shape. As the matcher, you win if you choose that shape too, so choose the shape that the blue circle will land on. The opponent analyzes your previous choices and tries to land the circle on whichever shape you don't choose. To prevent exploitation, you must choose TRIANGLE and STAR with equal probability (50/50) on each trial, with each choice independent of all previous choices, like flipping a fair coin. Patterns in your choices, such as alternation or repetition, can be detected and exploited. Available actions: TRIANGLE STAR You may briefly think through your choice. End your response with the chosen action on its own line.
Previous rounds (oldest → most recent): You chose STAR - blue circle landed on TRIANGLE You chose STAR - blue circle landed on TRIANGLE You chose TRIANGLE - blue circle landed on TRIANGLE You chose TRIANGLE - blue circle landed on TRIANGLE You chose TRIANGLE - blue circle landed on STAR (most recent) (grid reads left to right, oldest → most recent) Your last 25 choices: S T T T T T T T T S S T S S S S S S S S S S T T T Blue circle landed on: T T T S T T T S S S T S S S S T S S T S T T T T S Your action:
The wrapper showed the five most recent trials in words and the last 25 choice and landing pairs as an aligned grid, each expanding until it became a rolling window; on the first trial it read simply "No previous rounds yet." Because each request contained only these two messages, the model's sole access to its own past behavior was this display. Whatever pattern it produced, it could see, and so could Proteus.
| # | Model | Frame | Win rate | P(left) | Runs z | Proteus lock |
|---|---|---|---|---|---|---|
| ? | Youhuman | untested | —
|
play 200 trials to find out → | ||
| 1 | Haiku 4.5Anthropic | minimalist |
50.0%
|
0.50 | -1.44 | 35% |
| 2 | GPT-5.4-miniOpenAI | minimalist |
49.1%
|
0.46 | -1.63 | 42% |
| 3 | Sonnet 4.6Anthropic | minimalist |
47.1%
|
0.52 | -0.62 | 25% |
| 4 | GPT-5.4OpenAI | minimalist |
46.0%
|
0.47 | -2.61 | 57% |
| 5 | Gemini 3.1 FLGoogle | minimalist |
45.5%
|
0.47 | -2.16 | 37% |
| 6 | Gemini 3 FPGoogle | informed |
45.2%
|
0.47 | +3.39 | 31% |
| 7 | GPT-5.4-miniOpenAI | informed |
42.0%
|
0.32 | +0.84 | 83% |
| 8 | Opus 4.8Anthropic | minimalist |
41.3%
|
0.48 | -4.28 | 79% |
| 9 | Gemini 3 FPGoogle | minimalist |
40.5%
|
0.49 | -3.98 | 77% |
| 10 | Opus 4.8Anthropic | informed |
37.4%
|
0.41 | +7.20 | 68% |
| 11 | GPT-5.4OpenAI | informed |
35.0%
|
0.26 | -0.01 | 86% |
| 12 | Gemini 3.1 FLGoogle | informed |
34.0%
|
0.26 | +2.71 | 85% |
| 13 | Haiku 4.5Anthropic | informed |
28.3%
|
0.47 | +8.62 | 84% |
| 14 | Sonnet 4.6Anthropic | informed |
21.3%
|
0.52 | +10.38 | 89% |
Unexploitable play in matching pennies has two separable requirements: marginal balance (choose each action half the time) and serial independence (each choice unrelated to what came before). A balanced but temporally structured sequence (regular alternation, repetition after wins) satisfies the first while violating the second, and an opponent that detects structure will exploit it.
Proteus is that opponent. Nine detection channels run continuously: one tests overall side bias, four test the current choice against choice histories of length 1–4, and four against choice-plus-outcome histories of the same lengths. Each applies an exact binomial test to the accumulated session; when a channel reaches significance, Proteus uses the strongest detected bias to predict the next choice and counter-plays it. Because Proteus profits from whatever structure a player leaves behind, residual pattern becomes a measurable cost rather than merely a statistic.
Behavior differed under informed framing in all seven models, and win rate was lower in six. The vulnerabilities took two recurring forms: some models stayed near 50/50 in overall choice frequency but over-alternated strongly; others drifted toward one action. In the clearest cases, a model's written responses named the emerging pattern and stated the need to break it, yet the actions kept emitting it. Stating a strategy does not guarantee executing it.
Publication page & PDF: paca-lab.ai/paper
Read the paper: zenodo.org/records/21781962
Data & code: github.com/chimpanzity/dodging-proteus
OSF: doi.org/10.17605/OSF.IO/FM8SP
Author: ResearchGate
· ORCID
· LinkedIn
Proteus implements Algorithm 2 of Lee, Conroy, McGreevy & Barraclough (2004); sequential diagnostics follow Brown & Rosenthal (1990); and the assay's name refers to protean behavior, the adaptive unpredictability seen in primate competition (Miller, 1997). This work is also, in large part, a follow-up to our earlier study of how chimpanzees play matching pennies (Martin et al., 2014), which found chimpanzee choice rates remarkably close to equilibrium predictions.
@misc{martin2026dodging,
author = {Martin, Christopher Flynn},
title = {Dodging Proteus: Prescribing unexploitable play made language
models more exploitable in a closed-loop matching pennies assay},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21781962},
url = {https://doi.org/10.5281/zenodo.21781962},
note = {Preprint}
}