To cite this paper use one of the standards below:
This paper addresses the optimization of responses generated by Large Language Models (LLMs) at inference time. While current state-of-the-art approaches often rely on optimal stopping frameworks or combinatorial methods that operate over the prompt space, these techniques are frequently constrained by the high cost of multiple LLM calls. We propose a novel heuristic framework based on a Biased Random-Key Genetic Algorithm (BRKGA) capable of performing active search within the response vector space. Our method optimizes candidate responses at the lexical level, guided by a semantic reward model, to identify high-quality neighborhoods of an initial generation. Experimental results demonstrate that the proposed pipeline outperforms the Best-of-5 sampling baseline in 70% of the evaluated instances. Furthermore, our approach achieves a 60% reduction in computational costs, requiring significantly fewer LLM inference calls while maintaining superior response quality and grammatical coherence.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper