EVALUATION OF MULTIVARIATE CALIBRATION MODELS TRANSFERRED BETWEEN NEAR INFRARED INSTRUMENTS
In an industrial setting where multiple NIR instruments are used for the same measurements it is important that the instruments are providing the same results for identical samples. Nevertheless, the responses from two similar instruments for the same sample measured under the same conditions will be different. The calibration models constructed on one instrument will thus not necessarily be valid for another instrument. Constructing high quality multivariate calibration models using e.g. partial least squares (PLS) regression may require up to hundreds or thousands of samples – with reference values – collected over a long period. Consequently, constructing high quality calibration models is expensive and time consuming. Therefore, models are often constructed on one instrument and then transferred to other instruments.
Traditionally, models transferred from one instrument to other instruments are evaluated in a so-called master/slave setting. The model is developed on the master instrument and is then evaluated on the slave instrument. The error, e.g. the Root Mean Squared Error of Prediction (RMSEP), obtained for the slave instrument is used for evaluating how similar the instruments are. However, this RMSEP includes not only the uncertainty between instruments responses but also the uncertainty in the reference measurements. Methods eliminating the uncertainty originating from reference methods are necessary, in order to easier evaluate how similar the results from different instruments are. In this study approaches for eliminating the uncertainty in the reference method are tested and compared to the master/slave approach.
A total number of 84 flour samples were included in the study. All 84 samples were measured on 10 different NIR instruments, five NIR instruments were DS2500 (FOSS Analytical A/S, Hillerød, Denmark) and five NIR instruments were InfraXact (FOSS Analytical A/S, Hillerød, Denmark). Ash and protein content were quantified for all 84 samples and used as reference values for PLS modeling. Ash was quantified by calcination and protein was quantified by Kjeldahl digestion.
Three different metrics for comparing results from different instruments were developed and applied to the five InfraXact instruments and the five DS2500 instruments, respectively. In a cross-validation loop each instrument was “left out”, subsequently. The calibration model was constructed on the remaining four instruments. The calibration model was applied to the “left out” instrument and the predictions were collected. In the first approach (similar to the traditional master/slave setting) the collected predictions were compared to the reference values. In the second approach, the predictions were compared to the median of the predictions. In the third approach, calibration models were constructed on each individual instrument and the predictions were calculated for both the model instrument as well as the four other instruments. A model was calculated for each instrument and the calculated predictions compared for each instrument model avoiding the use of reference values.
When comparing the first (traditional master/slave) approach to the second and third approach the systematic differences between instruments were clearer in the second and third approaches. The five DS2500 instruments provided very similar results, whereas the five InfraXact instruments provided slightly different results.