Validation: still an unexplored land?
In multivariate data analysis, descriptive or qualitative/quantitative models are often built for the sake of interpreting the experimental results or to be able to make predictions on future samples. In this context, whatever the method concerned, no modelling approach should prescind from an estimate of the reliability or the validity of the proposed interpretation or of the resulting predictions; in a single word, from a proper validation. But the concept of validation is even more general, encompassing questions such as whether an appropriate model was chosen, or if outliers and/or highly influential points are present in the data set, or again not only whether the optimal dimensionality was chosen but also if the selected subspace remains stable over different samplings of the same population.
Accordingly, it is of utmost importance that the validation schemes adopted reflect the questions answers are sought for: an improper validation can be even more dangerous than performing no validation at all, if one is deluded to have behaved correctly.
However, despite this key role, still many papers are published and presented, which seem to ignore these fundamental issues, lacking a proper validation strategy or even not considering validation at all. Extreme cases of such a behavior can include situations (reported in the literature) where, for instance, even replicate measurements taken on the same samples are split between training and test sets, not to mention the validation of underlying hypotheses which is often neglected. Any estimates of the prediction error or classification accuracy should be validated across subgroups of objects due to replicates, batches of raw material, sampling site, instrument, season, etc. This can be achieved by systematic cross validation or cross model validation (CVM). If the aim is to find the best subset of variables or optimize the model on other criteria, CMV is the recommended approach.
In the present communication, these general concepts and the risks associated with an incorrect validation will be illustrated with some real world examples, mostly involving spectroscopic data sets. The interpretational ascpects of correct validation will also be discussed in terms of the true underlying model, pure spectra, and estimates of the Net Analyte Signal.