To cite this paper use one of the standards below:
Recent developments in generative AI models are transforming the way humans work. One of the key features that enabled such transformations is that anyone can easily call AI models via API (Application Programming Interface). However, drug discovery tasks usually involve considerations about data confidentiality and, thus, sending the SMILES specification of small molecules that may have novel biological activities to third-party servers raises concerns about data security and intellectual property protection. One way to solve this problem is by running open-source large language models (LLMs) locally. However, LLMs can be computationally expensive, and they can have slow performance even on high-end hardware. One way to reduce the hardware requirement of LLMs and make their performance faster is to use quantization, a process through which the numerical precision of the neural network weights is reduced. Quantized models are expected to be less resource-intensive, making it possible to run LLMs on conventional hardware, but they may show reduced reasoning capability. In this work, we assess the impact of quantization on the performance of different LLMs to generate valid small molecules. Specifically, we tested the GPT-OSS-120B, Gemma-4-31B, Mistral-Small-4-119B, and Qwen-3.6-35B-A3B models. We evaluated the original models (precision of 16 bits), and models quantized to 8, 6 and 4 bits in two tasks. Interestingly, all the models generated drug-like molecules in one of the tasks, even without having been prompted to generate molecules with such properties. In the task where the models were required to generate small molecules similar to a reference set, Gemma 4 was the model that had the best performance in following the chemical space of the reference set. Our results indicate that quantization has only a marginal impact on model reasoning for these tasks, while it can lead to dramatic increases in the speed of molecule generation. Our results have the potential to speed up workflows in drug discovery that rely on LLMs run locally without compromising the generative ability of LLMs.
This work was supported by the Brazilian Biosciences National Laboratory (LNBio), part of the Brazilian Center for Research in Energy and Materials (CNPEM).
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper