While AI control research focuses on the safe deployment of potentially misaligned models, a new study identifies a practical challenge: most control protocols require the ability to instrument the model and its pipeline. This requirement is frequently unmet by regulated organizations that access frontier models through APIs or managed endpoints, resulting in what the authors call bounded sovereignty.
The Four-Layer Access Typology
The research presents a four-layer typology of access—covering data, model, infrastructure, and interaction—and introduces the sovereignty discount cost. This term refers to the portion of the "control tax" dedicated to compensating for restricted access via:
- Contractual agreements and vendor assurance.
- Specialized architectures and third-party audits.
- Residual risk acceptance or the reduction of system scope.
Simulated Findings and Implications
A synthetic access-ablation experiment involving 1.35 million simulations, analyzed through an anonymized national payments infrastructure scenario, demonstrated how different access levels impact safety outcomes. The findings indicate that:
- Full interaction logs are necessary for effective diagnosis.
- Pre-execution gateways are required to enable real-time intervention.
- Access to internal traces and model-version control are vital for explaining incidents after they occur.
- Restricting the scope of a system can increase safety, though it may decrease overall utility.
The authors conclude that developers of safety protocols should be transparent about the specific access conditions their methods require, as many organizations must operate within significant technical constraints.