Abstract
The terroir concept relates to the specific growing conditions of landscape, geography, climate, and, importantly, the interactions of people through farming practices that impact the composition of a primary product and its gustatory characteristics. Defining compositional measures of terroir is challenging because many attributes combine to create a wine’s terroir. In the present study, a temporal and spatial investigation comprising targeted measures of viticultural practice, grape and wine composition, and sensory ratings was undertaken across subregions in the Barossa Geographical Indication. Data were arranged into multiple blocks associated with viticultural variables, grape or wine composition, and sensory ratings. Statistical approaches to analysing the data included k-means clustering, ANOVA Multiblock Orthogonal PLS (AMOPLS), random forest ensemble (RF), boosted classification decision trees (BCT) and artificial neural networks (NN). Two or three clusters of samples were evident in vintage k-means models, and the clusters were correlated with vineyard elevation and temperature accumulation. AMOPLS consistently extracted predictive latent variables associated with five levels for the subregion explanatory factor for vintage models, but predictive scores for subregion were not evident when the entire data set was decomposed in a single model, suggesting that subregions were inconsistent for the expression of vine, grape, and wine composition over the duration of the study. Optimised RF or BCT models performed on par with similar overall classification errors of between 11 and 14 % using a reduced feature set derived through recursive elimination of less important features. Feature importance in the two decision tree ensembles varied slightly, with RF models selecting features from most data blocks and BCT employing more features related to wine composition. The NN model was the least accurate model for sample classification assignment. k-means clusters were heavily dependent on measures of grape and wine volatiles with contributions from grape amino acid data. AMOPLS features of importance were distributed throughout the data blocks with viticultural characteristics, grape carotenoids, phenolics, and wine polysaccharides making the largest contributions to model outcomes. These results provide insight into different data modelling approaches, which may provide similar outcomes for classification or clustering of samples, but the influential features within specific models may be different; this will impact the interpretation of the importance of these measures of composition in terroir models.
| Original language | English |
|---|---|
| Article number | 8381 |
| Number of pages | 19 |
| Journal | Oeno One |
| Volume | 60 |
| Issue number | 2 |
| DOIs | |
| Publication status | Published - 13 May 2026 |
| Externally published | Yes |
Keywords
- chemometry
- IVAS 2024
- machine learning
- oenology
- viticulture
Fingerprint
Dive into the research topics of 'On terroir – The choice of model emphasises different measured attributes in data sets'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver