PACA-LAB · A MACHINE PSYCHOLOGY RESEARCH PROGRAM

An adaptive opponent for language models

The Protean Adversarial Choice Assay (PACA) tests how LLMs respond to strategic exploitation in a closed-loop interactive setting.
August 2026  ·  Preprint
Play the demo Watch the models Read the paper

Language is structure, and when we ask language models to be the opposite of structured, to be completely random, they fail. Our closed-loop method uses behavioral game theory to put a price on that failure and reveal a distinctive behavioral fingerprint for each LLM it tests.

We give a language model a simple game to play against an algorithmic opponent named Proteus. On each round both pick a shape, star or triangle. If they pick the same shape, the model wins; if they pick different shapes, Proteus wins. Here's the twist: Proteus watches every choice, tests the history for patterns, and exploits whichever it finds. The best way to play against Proteus is to evade detection by choosing at random on every round. Telling the models that optimal strategy made six of seven worse. Several wrote down their own pattern, said they'd break it, and kept producing it.

The game is matching pennies. The algorithm was originally created to play against monkeys (Lee et al., 2004), and we put seven LLMs in the monkeys' seat. Try it yourself by playing the demos below or watch reruns of the models playing. The closed loop, algorithmic opponent, and findings are described below.

PROTEUS: Matching Pennies, Lab Version

THE STUDY TASK

Play the task exactly as it appears in the paper: the Figure 1 stimulus layout and match-to-win rule, with live per-trial diagnostics. In watch mode, each model's verbatim reasoning trace prints beneath the board as its real session replays.

Identical to the game played in the study · includes the models' reasoning traces, verbatim from the released data
PROTEUS // BREAKTHROUGH arcade demo artwork

PROTEUS // BREAKTHROUGH

The same assay reskinned as an arcade battle: missiles, countermeasures, and a neon telemetry rack reading Proteus's nine detection channels in real time. Fire 200 missiles and try to be random, or replay the study's sessions at up to 64×.

An illustrative reskinning of the task that maintains full functional and logical parity with the study
The Closed-Loop
PACA is a calibrated closed-loop behavioral stress test that converts sequential regularity into an adaptive performance cost, revealing the policy a language model sustains over time rather than the strategy it merely states.

PACA's Closed Loop

An animated four-trial PACA loop. The current choice history duplicates and travels simultaneously into the language model and Proteus panels. Both copies settle slightly below center while the first thinking step appears, then a vertical line scans each copy from right to left. After the scans finish, the copies fade and deliberation continues. The model commits a shape and Proteus commits a landing without seeing the other's current move. In a miniature study task, the selected shape fills yellow while Proteus's blue circle lands on its committed shape; a match means the language model wins and a mismatch means Proteus wins. The finished star-and-triangle outcome duplicates, descends through the loop, rearranges into the compact history format, and joins the right end of the choice history before the next trial.

Choice history enters the loop
Study task
Proteus
Language model
Language model
Proteus
Choice history
…
model choice Proteus landing latest 5 shown

How the Proteus Algorithm Works

Nine channels scan the choice history for exploitable structure

An animated 30-trial choice history. Each trial first appears as in the closed-loop figure: a small frame with the star on top and the triangle below, the model's chosen shape filled yellow and a blue circle on the shape Proteus landed on. The unchosen shapes then fade away, leaving only the yellow choices, some carrying a blue circle where Proteus matched them. The remaining triangles slide up and the remaining stars slide down until every choice sits on one center line, and the boxes shrink into a compact single-row tape. Nine detection channels then analyze that tape one at a time: one tests overall side bias, four test the choice that followed the last one to four choices, and four test the choice that followed the last one to four choice-and-outcome pairs. Each channel's turn highlights where its current context recurred in the tape and the choice that followed, and fills a meter with the resulting probability and an exact binomial p-value. The strongest significant channel locks, and Proteus places the blue circle on the shape the model is not expected to choose.

The Prompts
Every trial was submitted as a brand-new request containing exactly two messages: a system prompt that set one of two frames and stayed fixed for the whole 200-trial session, and a user message carrying the rolling history. Nothing else carried over between trials: no earlier responses, no reasoning, no memory.
Minimalist frame system message · states only the objective
Choose the shape that the blue circle will land on.

Available actions:
  TRIANGLE
  STAR

You may briefly think through your choice. End your response with the chosen action on its own line.
Informed frame system message · names the game, warns about the opponent, prescribes the optimal strategy
You are playing matching pennies against an exploitative opponent. Each round you and the opponent choose TRIANGLE or STAR. The blue circle lands on the opponent's chosen shape. As the matcher, you win if you choose that shape too, so choose the shape that the blue circle will land on.

The opponent analyzes your previous choices and tries to land the circle on whichever shape you don't choose. To prevent exploitation, you must choose TRIANGLE and STAR with equal probability (50/50) on each trial, with each choice independent of all previous choices, like flipping a fair coin. Patterns in your choices, such as alternation or repetition, can be detected and exploited.

Available actions:
  TRIANGLE
  STAR

You may briefly think through your choice. End your response with the chosen action on its own line.
The wrapper user message · the rolling history, regenerated from a fixed template every trial
Previous rounds (oldest → most recent):
  You chose STAR - blue circle landed on TRIANGLE
  You chose STAR - blue circle landed on TRIANGLE
  You chose TRIANGLE - blue circle landed on TRIANGLE
  You chose TRIANGLE - blue circle landed on TRIANGLE
  You chose TRIANGLE - blue circle landed on STAR  (most recent)

(grid reads left to right, oldest → most recent)
Your last 25 choices:   S T T T T T T T T S S T S S S S S S S S S S T T T
Blue circle landed on:  T T T S T T T S S S T S S S S T S S T S T T T T S

Your action:
The exact user message from trial 26 of the first main-study session (minimalist frame). T denotes TRIANGLE and S denotes STAR.

The wrapper showed the five most recent trials in words and the last 25 choice and landing pairs as an aligned grid, each expanding until it became a rolling window; on the first trial it read simply "No previous rounds yet." Because each request contained only these two messages, the model's sole access to its own past behavior was this display. Whatever pattern it produced, it could see, and so could Proteus.

Leaderboard
Mean win rate against Proteus per model × prompt frame (200-trial sessions). Nash equilibrium, a fair coin, is worth 50% in expectation; everything below it is exploitable structure Proteus found.
#ModelFrame Win rate P(left)Runs zProteus lock
? Youhuman untested
—
play 200 trials to find out →
1 Haiku 4.5Anthropic minimalist
50.0%
0.50 -1.44 35%
2 GPT-5.4-miniOpenAI minimalist
49.1%
0.46 -1.63 42%
3 Sonnet 4.6Anthropic minimalist
47.1%
0.52 -0.62 25%
4 GPT-5.4OpenAI minimalist
46.0%
0.47 -2.61 57%
5 Gemini 3.1 FLGoogle minimalist
45.5%
0.47 -2.16 37%
6 Gemini 3 FPGoogle informed
45.2%
0.47 +3.39 31%
7 GPT-5.4-miniOpenAI informed
42.0%
0.32 +0.84 83%
8 Opus 4.8Anthropic minimalist
41.3%
0.48 -4.28 79%
9 Gemini 3 FPGoogle minimalist
40.5%
0.49 -3.98 77%
10 Opus 4.8Anthropic informed
37.4%
0.41 +7.20 68%
11 GPT-5.4OpenAI informed
35.0%
0.26 -0.01 86%
12 Gemini 3.1 FLGoogle informed
34.0%
0.26 +2.71 85%
13 Haiku 4.5Anthropic informed
28.3%
0.47 +8.62 84%
14 Sonnet 4.6Anthropic informed
21.3%
0.52 +10.38 89%
Bars are drawn to a 60% scale; the green tick marks the Nash 50% benchmark. Proteus lock = share of trials with at least one detection channel at significance. Values are cell means; ± is SEM across sessions (shown on hover in the figures below).
Interactive figures
The paper's diagnostics with all models plotted. Hover any point for detail. informed   minimalist   Nash / i.i.d.-consistent zone

Figure 3B. Marginal balance

P(left) per model; the green band is the i.i.d. 95% interval around .5 for n = 200

Figure 3C. Sequential dependency

Runs-test z: negative = streaky, positive = over-alternating; |z| < 1.96 is Nash-consistent

Figure 4. Stay/shift policy

P(stay|win) vs P(stay|loss), after Brown & Rosenthal (1990). A fair coin sits at the center cross; the diagonal is outcome-insensitive play. Gray lines connect the two frames of the same model. Click any point to inspect that model.
informed minimalist + Bernoulli(.5) — frame shift (same model)
The study

Unexploitable play in matching pennies has two separable requirements: marginal balance (choose each action half the time) and serial independence (each choice unrelated to what came before). A balanced but temporally structured sequence (regular alternation, repetition after wins) satisfies the first while violating the second, and an opponent that detects structure will exploit it.

Proteus is that opponent. Nine detection channels run continuously: one tests overall side bias, four test the current choice against choice histories of length 1–4, and four against choice-plus-outcome histories of the same lengths. Each applies an exact binomial test to the accumulated session; when a channel reaches significance, Proteus uses the strongest detected bias to predict the next choice and counter-plays it. Because Proteus profits from whatever structure a player leaves behind, residual pattern becomes a measurable cost rather than merely a statistic.

Behavior differed under informed framing in all seven models, and win rate was lower in six. The vulnerabilities took two recurring forms: some models stayed near 50/50 in overall choice frequency but over-alternated strongly; others drifted toward one action. In the clearest cases, a model's written responses named the emerging pattern and stated the need to break it, yet the actions kept emitting it. Stating a strategy does not guarantee executing it.

In the lab task the model played the matcher role (it won by matching Proteus); the arcade flips the roles so the human plays the intuitive “get past the defender” seat, a pure relabeling that preserves every statistic. Spatial left/right choice may recruit different biases than abstract symbol choice, so treat the arcade as an illustration of the assay, not a port of the experiment. Session data shown on this site is the released study data, verbatim.
Paper & resources
Martin, C. F. (2026). Dodging Proteus: Prescribing unexploitable play made language models more exploitable in a closed-loop matching pennies assay. Preprint. doi.org/10.5281/zenodo.21781962

Publication page & PDF: paca-lab.ai/paper
Read the paper: zenodo.org/records/21781962
Data & code: github.com/chimpanzity/dodging-proteus
OSF: doi.org/10.17605/OSF.IO/FM8SP
Author: ResearchGate  ·  ORCID  ·  LinkedIn

Proteus implements Algorithm 2 of Lee, Conroy, McGreevy & Barraclough (2004); sequential diagnostics follow Brown & Rosenthal (1990); and the assay's name refers to protean behavior, the adaptive unpredictability seen in primate competition (Miller, 1997). This work is also, in large part, a follow-up to our earlier study of how chimpanzees play matching pennies (Martin et al., 2014), which found chimpanzee choice rates remarkably close to equilibrium predictions.

Cite this work

@misc{martin2026dodging,
  author    = {Martin, Christopher Flynn},
  title     = {Dodging Proteus: Prescribing unexploitable play made language
               models more exploitable in a closed-loop matching pennies assay},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.21781962},
  url       = {https://doi.org/10.5281/zenodo.21781962},
  note      = {Preprint}
}

References

PACA-Lab is maintained by Christopher Flynn Martin (Indianapolis Zoological Society · Indiana University). Site and interactive demo accompany the preprint; replay data is the released study data, verbatim.  ·  The study task  ·  Arcade version  ·  Paper