首页|HyFactor: A Novel Open-Source, Graph-Based Architecture for Chemical Structure Generation

HyFactor: A Novel Open-Source, Graph-Based Architecture for Chemical Structure Generation

扫码查看
Graph-based architectures are becoming increasingly popular as a tool for structure generation. Here, we introduce novel open-source architecture HyFactor in which, similar to the InChI linear notation, the number of hydrogens attached to the heavy atoms was considered instead of the bond types. HyFactor was benchmarked on the ZINC 250K, MOSES, and ChEMBL data sets against conventional graph-based architecture ReFactor, representing our implementation of the reported DEFactor architecture in the literature. On average, HyFactor models contain some 20% less fitting parameters than those of ReFactor. The two architectures display similar validity, uniqueness, and reconstruction rates. Compared to the training set compounds, HyFactor generates more similar structures than ReFactor. This could be explained by the fact that the latter generates many open-chain analogues of cyclic structures in the training set. It has been demonstrated that the reconstruction error of heavy molecules can be significantly reduced using the data augmentation technique. The codes of HyFactor and ReFactor as well as all models obtained in this study are publicly available from our GitHub repository: https://github.com/Laboratoire-de-Chemoinformatique/HyFactor.

Akhmetshin Tagir、Lin Arkadii、Mazitov Daniyar、Zabolotna Yuliana、Ziaikin Evgenii、Madzhidov Timur、Varnek Alexandre

展开 >

University of Strasbourg

Kazan Federal University

2022

Journal of chemical information and modeling

Journal of chemical information and modeling

EISCI
ISSN:1549-9596
年,卷(期):2022.62(15)
  • 35