PACA-LAB · PREPRINT · AUGUST 2026

Dodging Proteus: Prescribing unexploitable play made language models more exploitable in a closed-loop matching pennies assay

Christopher Flynn Martin
Department of Science and Research, Indianapolis Zoological Society · Department of Informatics, Luddy School of Informatics, Computing, and Engineering, Indiana University
ORCID 0000-0002-6510-4594 · August 2026 · Preprint — not yet peer reviewed
Download PDF DOI: 10.5281/zenodo.21781962 Play the assay

Abstract

In repeated play against an adaptive opponent, six of the seven language models tested won less often under a prompt that supplied the equilibrium strategy. We introduce the Protean Adversarial Choice Assay (PACA), which embeds a language model in repeated matching pennies against Proteus, our implementation of an adaptive opponent that Lee and colleagues introduced in 2004 to study decision making in macaques; it detects and exploits statistical regularities in the player’s choice history. Seven language models each played 200-trial sessions under two prompts: a minimalist prompt stating only the objective, and an informed prompt that named the game and the exploitative opponent, prescribed the equilibrium strategy, and warned against exhibiting patterns across choices. Behavioral policy differed under informed framing in all seven models, and win rate was lower in six. Joint analysis of overall choice frequency and sequential dependence, using the Brown–Rosenthal framework developed for mixed-strategy play, revealed signatures concealed by win rate alone: predictable over-alternation, sometimes despite near-balanced choice frequencies, and a persistent bias toward one action. In the informed Anthropic sessions, written responses often diagnosed emerging patterns without fully correcting them. Because Proteus profits from whatever structure a model leaves behind, residual pattern becomes a measurable cost rather than merely a statistic, and that cost exposes a dissociation: stating a strategy does not guarantee executing it.

Resources

Cite

@misc{martin2026dodging,
  author    = {Martin, Christopher Flynn},
  title     = {Dodging Proteus: Prescribing unexploitable play made language
               models more exploitable in a closed-loop matching pennies assay},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.21781962},
  url       = {https://doi.org/10.5281/zenodo.21781962},
  note      = {Preprint}
}
Part of PACA-Lab, a machine psychology research program. Contact: ResearchGate · LinkedIn