Para citar este trabalho use um dos padrões abaixo:
This paper addresses the optimization of responses generated by Large Language Models (LLMs) at inference time. While current state-of-the-art approaches often rely on optimal stopping frameworks or combinatorial methods that operate over the prompt space, these techniques are frequently constrained by the high cost of multiple LLM calls. We propose a novel heuristic framework based on a Biased Random-Key Genetic Algorithm (BRKGA) capable of performing active search within the response vector space. Our method optimizes candidate responses at the lexical level, guided by a semantic reward model, to identify high-quality neighborhoods of an initial generation. Experimental results demonstrate that the proposed pipeline outperforms the Best-of-5 sampling baseline in 70% of the evaluated instances. Furthermore, our approach achieves a 60% reduction in computational costs, requiring significantly fewer LLM inference calls while maintaining superior response quality and grammatical coherence.
Com ~200 mil publicações revisadas por pesquisadores do mundo todo, o Galoá impulsiona cientistas na descoberta de pesquisas de ponta por meio de nossa plataforma indexada.
Confira nossos produtos e como podemos ajudá-lo a dar mais alcance para sua pesquisa:
Esse proceedings é identificado por um DOI , para usar em citações ou referências bibliográficas. Atenção: este não é um DOI para o jornal e, como tal, não pode ser usado em Lattes para identificar um trabalho específico.
Verifique o link "Como citar" na página do trabalho, para ver como citar corretamente o artigo