Global* vs Regional Calibrations: A Case Study
Global calibrations encompassing many sources of variation (e.g. seasonal, geographical, varietal or recipe differences) are often showed to have superior robustness and performance accuracy over regional calibrations that are based on only a limited number of samples.
To illustrate this concept in a large scale context we have made use of three independent NIRS feed calibration datasets used for protein measurement: An INGOT (AuNIR) dataset comprising more than 12,000 samples from around the world and two regional feed datasets sourced from China and Brazil with ca. 1200 and 400 samples respectively. All spectra were measured in reflection mode covering wavelength range 1100-2500 nm. WinISI software (FOSS Analytical) was employed for the modelling work.
Predictive performance of the global calibration was tested on the regional datasets as independent validation sets (RMSEPs=0.49 and 0.75 for China and Brazil datasets respectively), and the results were compared with the case where the regional calibrations were used on the large dataset (RMSEPs=1.90 and 2.23).
Attempts were also made to study the significance of the size of this global calibration dataset by monitoring its performance as it was subjected to different systematic sample removal regimes (random selection, selection based on reference values and or selection based on spectral variability).
*Global in this context refers to large datasets derived from many sources of variation (not to be confused with Foss’ proprietary global PLS modelling)