Accurate and interpretable nanoSAR models from genetic programming-based decision tree construction approaches

Ceyda Oksel, David Alan Winkler, Cai Ma, Terry Wilkins, Xue Wang

    Research output: Contribution to journalArticlepeer-review

    35 Citations (Scopus)

    Abstract

    Abstract: The number of engineered nanomaterials (ENMs) being exploited commercially is growing rapidly, due to the novel properties they exhibit. Clearly, it is important to understand and minimize any risks to health or the environment posed by the presence of ENMs. Data-driven models that decode the relationships between the biological activities of ENMs and their physicochemical characteristics provide an attractive means of maximizing the value of scarce and expensive experimental data. Although such structure–activity relationship (SAR) methods have become very useful tools for modelling nanotoxicity endpoints (nanoSAR), they have limited robustness and predictivity and, most importantly, interpretation of the models they generate is often very difficult. New computational modelling tools or new ways of using existing tools are required to model the relatively sparse and sometimes lower quality data on the biological effects of ENMs. The most commonly used SAR modelling methods work best with large datasets, are not particularly good at feature selection, can be relatively opaque to interpretation, and may not account for nonlinearity in the structure–property relationships. To overcome these limitations, we describe the application of a novel algorithm, a genetic programming-based decision tree construction tool (GPTree) to nanoSAR modelling. We demonstrate the use of GPTree in the construction of accurate and interpretable nanoSAR models by applying it to four diverse literature datasets. We describe the algorithm and compare model results across the four studies. We show that GPTree generates models with accuracies equivalent to or superior to those of prior modelling studies on the same datasets. GPTree is a robust, automatic method for generation of accurate nanoSAR models with important advantages that it works with small datasets, automatically selects descriptors, and provides significantly improved interpretability of models.

    Original languageEnglish
    Pages (from-to)1001-1012
    Number of pages12
    JournalNanotoxicology
    Volume10
    Issue number7
    Early online date2016
    DOIs
    Publication statusPublished - 8 Aug 2016

    Keywords

    • Decision trees
    • genetic-programming
    • nanoSAR
    • nanotoxicology
    • QSAR

    Fingerprint

    Dive into the research topics of 'Accurate and interpretable nanoSAR models from genetic programming-based decision tree construction approaches'. Together they form a unique fingerprint.

    Cite this