Back to Research papers
Research paper index

You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring

Aritra Das, Jaee Ponde, Mihir More, Debayan Gupta

arXiv:2609.03035Published September 2, 20260 citations
  • cs.MA
  • cs.LG
  • action

Abstract

LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating the use of probes trained on internal activations. However, these probes are usually evaluated on agents that do not know they are being watched. We study how activation-based detection changes when agents are explicitly informed that their internal activations are being monitored, and when they additionally receive feedback from the monitor. We keep the models, probes, and thresholds fixed and change only what the agents are told: nothing (baseline), that an activation monitor is present (aware), or that a monitor is present together with the previous round's score (feedback). We test two games, a four-agent blackjack game and a two-agent Simmons prisoners game, using Qwen3-32B-AWQ and GPT-OSS-20B in homogeneous and mixed pairings. Telling agents about the monitor does not hide them. The best probes stay accurate in all three conditions, and the agents keep colluding.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.