MELLODDY: cross pharma federated learning at unprecedented scale unlocks benefits in QSAR without compromising proprietary information

Wouter Heyndrickx; Lewis Mervin; Tobias Morawietz; Noé  Sturm; Lukas  Friedrich; Adam Zalewski; Anastasia  Pentina; Lina Humbeck; Martijn Oldenhof; Ritsuya Niwayama; Peter  Schmidtke; Nikolas  Fechner; Jaak Simm; Adam Arany; Nicolas Drizard; Rama Jabal; Arina  Afanasyeva; Regis  Loeb; Shlok  Verma; Simon Harnqvist; Matthew Holmes; Balasz Pejo; Maria  Telenczuk; Nicholas  Holway; Arne  Dieckmann; Nicola  Rieke; Friederike  Zumsande; Djork-Arné  Clevert; Michael  Krug; Christopher  Luscombe; Darren  Green; Peter  Ertl; Peter  Antal; David  Marcus; Nicolas  Do Huu; Hideyoshi  Fuji; Stephen  Pickett; Gergely  Acs; Eric  Boniface; Bernd  Beck; Yax  Sun; Arnaud  Gohier; Friedrich  Rippmann; Ola  Engkvist; Andreas H. Göller; Yves  Moreau; Mathieu N. Galtier; Ansgar  Schuffenhauer; Hugo  Ceulemans

doi:10.26434/chemrxiv-2022-ntd3r

Biological and Medicinal Chemistry

Search within Biological and Medicinal Chemistry

MELLODDY: cross pharma federated learning at unprecedented scale unlocks benefits in QSAR without compromising proprietary information

13 October 2022, Version 1

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Federated multi-partner machine learning can be an appealing and efficient method to increase the effective training data volume and thereby the predictivity of models, particularly when the generation of training data is resource intensive. In the landmark MELLODDY project, each of ten pharmaceutical companies realized aggregated improvements on its own classification and/or regression models through federated learning. To this end, they leveraged a novel implementation extending multi-task learning across partners, on a platform audited for privacy and security. The experiments involved an unprecedented cross-pharma dataset of 2.6+ billion confidential experimental activity data points, documenting 21+ million physical small molecules and 40+ thousand assays in on-target and secondary pharmacodynamics and pharmacokinetics. Appropriate complementary metrics were developed to evaluate predictive performance in the federated setting. In addition to predictive performance increases in labeled space, the results point towards an extended applicability domain in federated learning. Increases in collective training data volume, including by means of auxiliary data resulting from single concentration high-throughput and imaging assays, continued to boost predictive performances, albeit with saturating return. Markedly higher improvements were observed for pharmacokinetics and safety panel assay-based task subsets.

Keywords

federated learning

multitask learning

small molecule drug discovery

MELLODDY

QSAR

Supplementary materials

Title

Description

Actions

Title

Complementary figures and details

Description

Complementary figures and additional details, including on the hyperparameter search.

Actions

Title

Data preparation manual

Description

Full details on the data preparation procedure

Actions

Supplementary weblinks

Title

Description

Actions

Title

MELLODDY-TUNER

Description

MELLODDY-TUNER, data preparation package

Actions

View

Title

Model predictive performance evaluation

Description

Model predictive performance evaluation

Actions

View

Title

Pseudolabel data preparation

Description

Pseudolabel data preparation

Actions

View

Title

Scaffold-based multi-party simulation split

Description

Scaffold-based multi-party simulation split

Actions

View

Title

Public data preparation

Description

Public data preparation

Actions

View

Title

Model manipulation software

Description

Model manipulation software

Actions

View

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Version History

Oct 13, 2022 Version 1

Metrics

4,860

2,163

Views

Downloads

License

The content is available under CC BY NC ND 4.0

DOI

10.26434/chemrxiv-2022-ntd3r

Funding

Innovative Medicines Initiative

831472

Author’s competing interest statement

M.N.G. and M.T. are employed and own stocks in the company Owkin commercializing the underlying Federated Learning Platform based on the open source Substra software. The remaining authors have no conflicts of interest to declare.

Ethics

The author(s) declare that they have sought and gained approval from the relevant ethics committee/IRB for this research and its publication.

MELLODDY: cross pharma federated learning at unprecedented scale unlocks benefits in QSAR without compromising proprietary information

Authors

Abstract

Keywords

Supplementary materials

Supplementary weblinks

Comments

Version History

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share