Efficient Evaluation for AI Control via Subsampling

- 337706
Abstract
Favorite this paper
How to cite this paper?
Abstract

Evaluating AI control protocols is the problem of determining whether a given safety pipeline can prevent a potentially misaligned AI model from causing harmful outcomes, even under intentional subversion. A central challenge in this field is that a full evaluation requires executing the protocol across a large grid: given P problems, C protocol configurations (e.g., suspicion threshold values), and K red-team attack strategies, a complete assessment demands PCK executions, which is computationally prohibitive when any of these dimensions is large.

We propose a methodology to estimate Safety and Usefulness metrics for AI control protocols using only a fraction of the full execution budget. The approach adapts the PromptEval framework for efficient LLM benchmark evaluation to the control setting, combining a structured probabilistic model for outcome imputation with a balanced subsampling strategy for variance reduction.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Escola de Matemática Aplicada, Fundação Getulio Vargas
Track
  • ST09 - Computational Modeling
Keywords
LLM
AI Control
Efficient evaluation