Efficient Evaluation for AI Control via Subsampling

- 337706
Resumo
Favoritar este trabalho
Como citar esse trabalho?
Resumo

Evaluating AI control protocols is the problem of determining whether a given safety pipeline can prevent a potentially misaligned AI model from causing harmful outcomes, even under intentional subversion. A central challenge in this field is that a full evaluation requires executing the protocol across a large grid: given P problems, C protocol configurations (e.g., suspicion threshold values), and K red-team attack strategies, a complete assessment demands PCK executions, which is computationally prohibitive when any of these dimensions is large.

We propose a methodology to estimate Safety and Usefulness metrics for AI control protocols using only a fraction of the full execution budget. The approach adapts the PromptEval framework for efficient LLM benchmark evaluation to the control setting, combining a structured probabilistic model for outcome imputation with a balanced subsampling strategy for variance reduction.

Compartilhe suas ideias ou dúvidas com os autores!

Sabia que o maior estímulo no desenvolvimento científico e cultural é a curiosidade? Deixe seus questionamentos ou sugestões para o autor!

Faça login para interagir

Tem uma dúvida ou sugestão? Compartilhe seu feedback com os autores!

Instituições
  • 1 Escola de Matemática Aplicada, Fundação Getulio Vargas
Eixo Temático
  • ST08 - Modelagem Computacional
Palavras-chave
LLM
AI Control
Efficient evaluation