Back

Molecular Graph Contrastive Learning with Parameterized Explainable Augmentations

Wang, Y.; Min, Y.; Shao, E.; Wu, J.

2021-12-06 bioinformatics
10.1101/2021.12.03.471150 bioRxiv
Show abstract

Learning generalizable, transferable, and robust representations for molecule data has always been a challenge. The recent success of contrastive learning (CL) for self-supervised graph representation learning provides a novel perspective to learn molecule representations. The most prevailing graph CL framework is to maximize the agreement of representations in different augmented graph views. However, existing graph CL frameworks usually adopt stochastic augmentations or schemes according to pre-defined rules on the input graph to obtain different graph views in various scales (e.g. node, edge, and subgraph), which may destroy topological semantemes and domain prior in molecule data, leading to suboptimal performance. Therefore, designing parameterized, learnable, and explainable augmentation is quite necessary for molecular graph contrastive learning. A well-designed parameterized augmentation scheme can preserve chemically meaningful structural information and intrinsically essential attributes for molecule graphs, which helps to learn representations that are insensitive to perturbation on unimportant atoms and bonds. In this paper, we propose a novel Molecular Graph Contrastive Learning with Parameterized Explainable Augmentations, MolCLE for brevity, that self-adaptively incorporates chemically significative information from both topological and semantic aspects of molecular graphs. Specifically, we apply deep neural networks to parameterize the augmentation process for both the molecular graph topology and atom attributes, to highlight contributive molecular substructures and recognize underlying chemical semantemes. Comprehensive experiments on a variety of real-world datasets demonstrate that our proposed method consistently outperforms compared baselines, which verifies the effectiveness of the proposed framework. Detailedly, our self-supervised MolCLE model surpasses many supervised counterparts, and meanwhile only uses hundreds of thousands of parameters to achieve comparative results against the state-of-the-art baseline, which has tens of millions of parameters. We also provide detailed case studies to validate the explainability of augmented graph views. CCS CONCEPTS* Mathematics of computing [->] Graph algorithms; * Applied computing [->] Bioinformatics; * Computing methodologies [->] Neural networks; Unsupervised learning.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Briefings in Bioinformatics
354 papers in training set
Top 0.2%
18.1%
2
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.1%
12.2%
3
Bioinformatics
1204 papers in training set
Top 3%
9.5%
4
Journal of Computational Biology
48 papers in training set
Top 0.2%
4.8%
5
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.2%
4.0%
6
Nature Communications
5641 papers in training set
Top 36%
3.2%
50% of probability mass above
7
PLOS Computational Biology
1863 papers in training set
Top 10%
3.2%
8
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
9
Patterns
78 papers in training set
Top 0.7%
3.1%
10
PLOS ONE
5266 papers in training set
Top 41%
2.7%
11
Scientific Reports
3612 papers in training set
Top 45%
2.3%
12
Journal of Cheminformatics
29 papers in training set
Top 0.3%
2.3%
13
BMC Bioinformatics
457 papers in training set
Top 4%
2.1%
14
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 0.6%
2.0%
15
Nature Computational Science
55 papers in training set
Top 0.7%
1.7%
16
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.7%
17
Nature Machine Intelligence
70 papers in training set
Top 2%
1.4%
18
Advanced Science
286 papers in training set
Top 6%
1.3%
19
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.3%
20
iScience
1154 papers in training set
Top 27%
1.1%
21
PNAS Nexus
159 papers in training set
Top 2%
1.1%
22
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 36%
1.1%
23
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.1%
24
IEEE Access
35 papers in training set
Top 1%
1.0%
25
Expert Systems with Applications
11 papers in training set
Top 0.3%
1.0%
26
GigaScience
212 papers in training set
Top 5%
0.8%
27
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
28
Journal of Molecular Biology
232 papers in training set
Top 5%
0.6%
29
PRX Life
42 papers in training set
Top 1%
0.6%