To cite this paper use one of the standards below:
Evaluating AI control protocols is the problem of determining whether a given safety pipeline can prevent a potentially misaligned AI model from causing harmful outcomes, even under intentional subversion. A central challenge in this field is that a full evaluation requires executing the protocol across a large grid: given P problems, C protocol configurations (e.g., suspicion threshold values), and K red-team attack strategies, a complete assessment demands PCK executions, which is computationally prohibitive when any of these dimensions is large.
We propose a methodology to estimate Safety and Usefulness metrics for AI control protocols using only a fraction of the full execution budget. The approach adapts the PromptEval framework for efficient LLM benchmark evaluation to the control setting, combining a structured probabilistic model for outcome imputation with a balanced subsampling strategy for variance reduction.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper