A novel, simple, fast and effective way for data pre-processing selection based on experimental design
NIR data is often used for chemometrical data analysis. The selection of optimal pre-processing for such data is one of the bottlenecks in chemometrics. Many different pre-processing methods are available for e.g. baseline correction or scatter correction, but it is not clear on beforehand which method—or combination of methods—should be used for which data set. The process of pre-processing selection is currently often based on “trial-and-error”: a subset of pre-processing methods is selected, their performance on a chemometric model is assessed and the method with the best model performance is selected as most optimal. However, it has already been shown that this approach often does not lead to an optimal pre-processing.
In this study, we have developed a novel, simple and effective approach for pre-processing selection based on experimental design. Partial Least Squares (PLS) model performance of a few different pre-processing methods and combinations thereof (called ‘strategies’) is evaluated according to the design. Interpretation of the main effects and interactions subsequently enables the selection of an optimal pre-processing strategy.
The approach has been validated on a selection of different spectroscopic data sets. The main data set used deals with the prediction of concentration NaOH and NaOCl from NIR spectra obtained on different mixtures of these compounds. The full data set consists of a training data set of 65 spectra and a validation set of 6 spectra; each spectrum contains 1102 data points.
The first step of the approach consists of a full factorial experimental design, in which each pre-processing step is evaluated as separate factor. The design assesses the influence of 4 different factors—baseline correction, scatter correction, smoothing and scaling—on the Root Mean Square Error of Prediction (RMSEP) of a PLS model. The low level setting for each factor always indicates “do nothing”, while the high level implies a specific setting, e.g. Standard Normal Variate (SNV) for scatter correction. Based on interpretation of the outcome of the design (main effects and interactions), some factors are deemed relevant and some irrelevant.
In the second step, the optimal method is obtained for each relevant factor by evaluating model performance of a larger selection of methods corresponding to that factor. This ultimately leads to an optimal pre-processing strategy for the data under study.
Evaluation and interpretation of the results from the experimental design shows that a pre-processing strategy is obtained for both compounds that is very close to the true optimal strategy. The latter is obtained by evaluating RMSEP for all available pre-processing strategies (almost 5,000!), which is a very time-consuming process. Our approach thus identifies pre-processing strategies with a large increase in model performance within only a fraction of the time that would be required to evaluate all possible pre-processing strategies.
The presented approach is generic and can easily be applied to new spectroscopic data sets. Results on other data sets confirm the ability of this experimental design approach to provide an optimal pre-processing strategy within reasonable time (e.g. 15-30 minutes calculation time).