Chemometric Methods for Classification and Discrimination of Chinese Patent Medicines Using Near Infrared Spectra
Near-infrared (NIR) spectroscopy is a fast and nondestructive analytical technique. It has attracted considerable attention in pharmaceutical industry for quantitative analysis, qualitative analysis and on-line control of pharmaceutical products. For examples, the technique has been extensively studied for quantifying active principal ingredients (API), excipients and water content in pharmaceutical products, and to monitor the production process of pharmaceutical products, such as assessing tableting process, monitoring blend uniformity of solid dosage forms or API concentration in powder mixing process. On the other hand, discrimination of geographic origins or manufacturer and identification of counterfeit drugs have been an important task in pharmaceutical industry. However, in some cases, e.g., the same pharmaceutical product from different manufacturers, there is no significant difference. Therefore, efficient methods are needed to classify the similar samples by exploring the tiny difference between the products. With the aim to establish approaches for rapid identification of medicines, chemometric methods for classification and discrimination were studied in our works.
Two datasets were prepared for the studies. The first dataset includes NIR spectra of five Chinese patent medicines (CPM) produced by different manufacturers in 12 classes. 22 spectra of each class were used for calibration and more than 30 spectra for each class were used for an independent validation. The second dataset includes 192 NIR spectra of Chinese patent medicines in seven classes. All the spectra were recorded on an MPA FT-NIR spectrometer (Bruker, Germany) in the wavenumber range 3999.7-11995.3 cm-1 with the digitization interval 3.857 cm-1. In the calculations, the variables from 4246.6 to 8913.7 cm-1 were used.
In order to use more information in principal component analysis (PCA), principal component accumulation (PCAcc) method was proposed in our previous works. In this study, PCAcc was used in discrimination of Chinese patent medicines. In the method, an accumulation strategy is utilized to combine the classification information contained in multiple PC subspaces by using a rotation, a projection and a summation operation. The results show that, among the 12 classes of Chinese patent medicines, 8 classes are correctly classified, and a total of ten samples are misclassified for the other four classes. Compared with the results obtained by principal component analysis (PCA), radial basis function artificial neural network (RBF-ANN) and partial least squares discriminant analysis (PLSDA), PCAcc produces the best classification. Another method for discrimination of medicines was proposed. The method performs principal component analysis (PCA), at first, on the spectra of multi-class samples, and then constructs the optimal set of orthogonal discriminant vectors using the PCA scores by maximizing Fisher’s discriminant function. Therefore, the filters for discriminating the samples can be obtained by transforming the loadings with the discriminant vectors. Applying the filters onto the spectra of new samples, the difference between samples of different class can be obtained and the difference can be used for discrimination of these samples. With the second dataset of 192 CPMs, all the samples of the seven classes are correctly identified.