34388

New tools to solve large complexity in large scale soil NIR libraries thought the use of memory-based learning

Favorite this paper

1. Introduction
One of the majors constrains for producing relevant soil information relies on the fact that the conventional methods of soil analysis are expensive and time-consuming. Visible (vis) and near infrared (NIR) spectroscopy as a complement to conventional soil analysis methods, can be used to optimize the efficiency of the soil analyses. However, in contrast to small scale soil vis–NIR models (e.g. field), large scale (e.g. regional, continental) models usually present a much lower prediction performance due to high variation of soil constituents which produces complex non–linear relationships between soil vis–NIR spectra and soil attributes. In this respect, large scale soil characterization by means of vis–NIR requires large amounts of samples in order to cover the soil variability adequately. This implies that the construction of any reliable large scale soil vis–NIR library usually derives in a large and complex dataset. Therefore, the use of soil spectral libraries demands reliable algorithms for accurate soil characterizations able to deal with massive and non–linear data. In this work, we present a detail analysis of memory–based learning memory–based learning (MBL, a.k.a local modeling) algorithms and its potential to produce reliable information from large and complex soil vis–NIR libraries. Furthermore we present a new open source software package for performing such kind of MBL analyses.

2. Experimental
In MBL, for each sample for which a given attribute (yui) needs to be predicted from its spectral data (xui), its neighbors (most spectrally similar samples) are searched in a spectral library (Xr). Then a (local) model is calibrated with these (reference) neighbors and the prediction of yui is finally carried out. In this work we investigated 10 different similarity metric algorithms used to select the most spectrally similar samples, 3 strategies for using the similarity information and 4 regression algorithms for fitting the local models. We also investigated on the interactions between these different aspects.
For this investigation we used two vis–NIR soil libraries: a continental one (C–SSL, n = 19.000) containing samples from all the EU countries and a soil spectral library of the world (G–SSL, n = 3.643). We evaluated the performance of the algorithms for predicting clay content. All the analyses were carried out in R using the “resemble” package which we developed for performing MBL analyses.

3. Results and discussion
We obtained a wide range of results in terms of prediction performance for the different configurations of MBL algorithms investigated. For example, for the prediction of clay content in the C–SSL, the RMSE varied between 8.61% and 4.81% and the R2 varied between 0.59 and 0.87. The results indicate that the similarity/dissimilarity metric used is the most important aspect of a MBL algorithm since it largely impacts its prediction performance. The new strategies proposed here for assessing the spectral similarity between samples outperformed the conventional ones. The results also indicate that the information on the response variables should be always taken into account for building models of spectral similarity/dissimilarity.