PAI-Bench is a new benchmark for persistent AI agents whose identity must survive restarts, updates and different roles. It separates factual recall from composition, behavioral enactment, resistance to conflicting prompts, persistence, lineage and governed identity changes.

Two fixed evaluation campaigns cover 16 synthetic profiles, 32 probes and three independently initialized target configurations, producing 1,536 retained responses. Scoring systems are kept outside the agent under test to reduce interference with its behavior.

A literal audit found direct-parent identifiers in all 48 answers to atomic questions, but in only one of 48 implicit self-portraits. On eight profiles, explicitly naming the relevant fields raised the joint appearance of three identifiers from zero of eight responses to seven of eight, despite the same four-sentence limit. That shows information can be available to a model without being expressed as part of its working identity.

Evaluator choice also mattered: replaying identical responses produced a Claude headline score 12.5 percentage points below Astra’s. The study uses one target sample per condition, which limits statistical claims. As a new arXiv preprint, its main contribution is a reproducible protocol for distinguishing memory from consistent identity behavior rather than a verdict on any one deployed agent.