Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Applications of machine learning in the genetics of brain disorders

Bracher‐Smith, Matthew and Escott‐Price, Valentina ORCID: https://orcid.org/0000-0003-1784-5483 2026. Applications of machine learning in the genetics of brain disorders. Su, Li, ed. Artificial Intelligence in Neuroscience, Wiley, pp. 205-230. (10.1002/9781394278886.ch08)

Full text not available from this repository.

Abstract

Prediction is an important part of achieving improved outcomes in neurology and psychiatry. Early diagnosis with genetics may assist interventions to delay onset and reduce the progression rate of the disease. However, genetic prediction of brain disorders only became feasible relatively recently through polygenic risk scores (PRSs) following the discovery of robust risk loci in association studies. This approach relies on univariable tests of association and typically assumes additivity within and between loci. Presently, PRSs are the most effective method for genetic prediction of brain disorders, yet they still only explain a small fraction of liability. A contrasting approach is machine learning (ML), which has evolved as a subset of artificial intelligence for learning complex patterns from labelled data. ML methods are an enticing option in genetics, as they allow for multivariable predictive modelling, complex predictor relationships including interactions and can learn from datasets where the number of predictors exceeds observations. They therefore represent a distinct option from PRS, which may address several of its limitations. However, their application to genetics in psychiatric and neurodegenerative disorders and diseases is relatively recent and has not been widely adopted. Initial challenges are expected when any novel methodology is introduced to a field and recent evidence has helped clarify predictive performance and methodological concerns in psychiatric and neurodegenerative genetics. The aim of this chapter is first to present recent systematic evaluations of the predictive performance of a range of ML models for prediction of brain disorders, including popular approaches such as neural networks (deep learning), tree-based ensemble models (gradient boosting and random forests) and support vector machines. We will highlight common practices and pitfalls in current applications of ML in the field using evidence from systematic reviews and simulated data. We show that poor reporting and widespread inadequate modelling approaches that introduce data leakage are common causes of high risk of bias in analysis, with key steps in model development and validation often being under-reported or absent from studies. Furthermore, we illustrate with simulations how small choices in common pre-processing approaches to genetic data can influence inflation of prediction estimates. Second, we emphasise best practices in methodology and reporting for improving future studies, highlighting exemplary studies and signposting to efforts, which call for researchers to adhere to more rigorous ML workflows and reporting procedures. Given widespread high risk of bias and the small sample sizes typically used in the literature, it is important to ensure robust analysis methods are adopted. ML approaches have made genuine progress in several areas in the biosciences, such as protein structure prediction, but realisation of their potential in disease genetics will require improvements in study design, implementation and reporting. Finally, we consider the future of ML in the genetics of complex diseases, which will continue to be used more extensively in both academia and the industry due to its ability to analyse complex patterns in datasets. Genetic data is classed as sensitive data under General Data Protection Regulation (GDPR), and most large genetic datasets require strict permissions and exact descriptions of usage. We discuss advances in federated learning (FL), its potential to address concerns for data privacy and learning from non-identically distributed data and recent applications of FL to omics analysis in brain disorders.

Item Type: Book Section
Date Type: Publication
Status: Published
Schools: Schools > Medicine
Publisher: Wiley
ISBN: 9781394278855
Last Modified: 07 May 2026 10:45
URI: https://orca.cardiff.ac.uk/id/eprint/186862

Actions (repository staff only)

Edit Item Edit Item