34371

Optimizing PLS predicting models with indepent validation set rather than cross validation

Favorite this paper

Two forage calibration database (1100-2498, 2nm) for alfalfa hay and corn silage that were developed over a period of 15 years, were used to develop and optimize variance scale PLS predicting models for crude protein (CP), ADF, NDF, Ca, K and P. Spectral mathematical pretreatments like scatter correction (SNV and Detrend) and derivatives (1st derivative, gap and smoothing over 4 datapoints) were kept constant for all of the models, without optimization. The evaluation compared the number of principal components (PC), which were optimized either with cross validation or by the use of independent validation sets, made with samples and reference values from the certification program run by the National Forage Testing Association (USA).

Using cross validation tended to use a greater number of PC for ADF and NDF than necessary to minimize SEP corrected for bias (SEPc). On average the cross validation used a couple of extra PC not necessary for optimal prediction as evaluated on independent validation sets. For CP the behavior was not consistent with the fiber fractions. For alfalfa hay the cross validation yielded a model (SEPc= 0.50%DM) using the maximum number of PC allowed which was set at 16. Using 15 or 14 PC had similar performance (SEPc= 0.49 %DM) to what was obtained with 16. With corn silage, CP model developed with cross validation had 12 PC (SEPc=0.40 %DM), but lowest SEPc on independent validation sets was obtained with 15 PC (0.32 %DM); an effect opposite to what was seen for fiber fractions. For minerals the relationship between spectra and composition was very weak and the optimization didn’t have much effect.

Using fewer PC used in the optimized models has the effect to reduce the magnitude of beta coefficient of the final calibrations, increasing the robustness of prediction. A data set of alfalfa hay with samples ground with different mills, was predicted for ADF and NDF using the optimized or the cross validated models. Using the equations developed with fewer PC resulted in an improvement of precision of about 40% for ADF (repeatability 0.75 vs 1.31 %DM) and about 10% for NDF (0.84 vs 0.97 %DM). In conclusion, for optimal performance of NIR calibrations, fine tuning of predicting model may require the use of an independent validation set that may reduce the number of PC used, maintaining accuracy, but increasing robustness of calibrations