To cite this paper use one of the standards below:
Combinatorial optimization problems have increasingly been addressed using Deep Reinforcement Learning (DRL) due to its ability to learn decision policies directly from experience. However, most DRL models applied to optimization are structured as construction heuristics, where agents build solutions sequentially and often become specialized to the distribution of a specific problem. In contrast, traditional metaheuristics are typically problem-independent but rely on stochastic search mechanisms and predefined hyperparameters without adapting their search behavior from previous exploration. To address this gap, this paper proposes a DRL-based architecture structured upon the continuous hypercube of Random-Key Optimizers (RKOs). By separating the continuous search process from the deterministic decoding procedure, the model allows the agent to perform iterative search in the continuous space while remaining agnostic to the underlying binary optimization problem. Experimental results show that the learned search policy transfers in a zero-shot setting across unseen binary NP-hard problems without additional training.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper