Para citar este trabalho use um dos padrões abaixo:
Combinatorial optimization problems have increasingly been addressed using Deep Reinforcement Learning (DRL) due to its ability to learn decision policies directly from experience. However, most DRL models applied to optimization are structured as construction heuristics, where agents build solutions sequentially and often become specialized to the distribution of a specific problem. In contrast, traditional metaheuristics are typically problem-independent but rely on stochastic search mechanisms and predefined hyperparameters without adapting their search behavior from previous exploration. To address this gap, this paper proposes a DRL-based architecture structured upon the continuous hypercube of Random-Key Optimizers (RKOs). By separating the continuous search process from the deterministic decoding procedure, the model allows the agent to perform iterative search in the continuous space while remaining agnostic to the underlying binary optimization problem. Experimental results show that the learned search policy transfers in a zero-shot setting across unseen binary NP-hard problems without additional training.
Com ~200 mil publicações revisadas por pesquisadores do mundo todo, o Galoá impulsiona cientistas na descoberta de pesquisas de ponta por meio de nossa plataforma indexada.
Confira nossos produtos e como podemos ajudá-lo a dar mais alcance para sua pesquisa:
Esse proceedings é identificado por um DOI , para usar em citações ou referências bibliográficas. Atenção: este não é um DOI para o jornal e, como tal, não pode ser usado em Lattes para identificar um trabalho específico.
Verifique o link "Como citar" na página do trabalho, para ver como citar corretamente o artigo