EFFECTS OF QUANTIZATION ON THE PERFORMANCE OF LARGE LANGUAGE MODELS TO GENERATE SMALL MOLECULES

Vol 4, 2026 - 345835
Abstract
Favorite this paper
How to cite this paper?
Abstract

Recent developments in generative AI models are transforming the way humans work. One of the key features that enabled such transformations is that anyone can easily call AI models via API (Application Programming Interface). However, drug discovery tasks usually involve considerations about data confidentiality and, thus, sending the SMILES specification of small molecules that may have novel biological activities to third-party servers raises concerns about data security and intellectual property protection. One way to solve this problem is by running open-source large language models (LLMs) locally. However, LLMs can be computationally expensive, and they can have slow performance even on high-end hardware. One way to reduce the hardware requirement of LLMs and make their performance faster is to use quantization, a process through which the numerical precision of the neural network weights is reduced. Quantized models are expected to be less resource-intensive, making it possible to run LLMs on conventional hardware, but they may show reduced reasoning capability. In this work, we assess the impact of quantization on the performance of different LLMs to generate valid small molecules. Specifically, we tested the GPT-OSS-120B, Gemma-4-31B, Mistral-Small-4-119B, and Qwen-3.6-35B-A3B models. We evaluated the original models (precision of 16 bits), and models quantized to 8, 6 and 4 bits in two tasks. Interestingly, all the models generated drug-like molecules in one of the tasks, even without having been prompted to generate molecules with such properties. In the task where the models were required to generate small molecules similar to a reference set, Gemma 4 was the model that had the best performance in following the chemical space of the reference set. Our results indicate that quantization has only a marginal impact on model reasoning for these tasks, while it can lead to dramatic increases in the speed of molecule generation. Our results have the potential to speed up workflows in drug discovery that rely on LLMs run locally without compromising the generative ability of LLMs.

This work was supported by the Brazilian Biosciences National Laboratory (LNBio), part of the Brazilian Center for Research in Energy and Materials (CNPEM). 

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Brazilian Biosciences National Laboratory (LNBio), Brazilian Biosciences National Laboratory (LNBio), Brazilian Center for Research in Energy and Materials (CNPEM), Campinas, SP 13083-970, Brazil
  • 2 Brazilian Biosciences National Laboratory (LNBio), Brazilian Center for Research in Energy and Materials (CNPEM), Campinas, SP 13083-970, Brazil
Track
  • 3. Drug design and delivery
Keywords
DRUG DESIGN
LARGE LANGUAGE MODELS (LLM)
QUANTIZATION
GENERATIVE AI
SMALL MOLECULES