Dictionary of 140k GDB and ZINC derived AMONs

We present all AMONs for GDB and Zinc data-bases using no more than 7 non-hydrogen atoms (AGZ7)---a calculated organic chemistry building-block dictionary based on the AMON approach [Huang and von Lilienfeld, Nature Chemistry (2020)]. AGZ7 records Cartesian coordinates of compositional and constitutional isomers, as well as properties for ∼140k small organic molecules obtained by systematically fragmenting all molecules of Zinc and the majority of GDB17 into smaller entities, saturating with hydrogens, and containing no more than 7 heavy atoms (excluding hydrogen atoms). AGZ7 cover the elements H, B, C, N, O, F, Si, P, S, Cl, Br, Sn and I and includes optimized geometries, total energy and its decomposition, Mulliken atomic charges, dipole moment vectors, quadrupole tensors, electronic spatial extent, eigenvalues of all occupied orbitals, LUMO, gap, isotropic polarizability, harmonic frequencies, reduced masses, force constants, IR intensity, normal coordinates, rotational constants, zero-point energy, internal energy, enthalpy, entropy, free energy, and heat capacity (all at ambient conditions) using B3LYP/cc-pVTZ (pseudopotentials were used for Sn and I) level of theory. We exemplify the usefulness of this data set with AMON based machine learning models of total potential energy predictions of seven of the most rigid GDB-17 molecules.

Identifier
Source https://archive.materialscloud.org/record/2021.61
Metadata Access https://archive.materialscloud.org/xml?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:materialscloud.org:819
Provenance
Creator Huang, Bing; von Lilienfeld, Anatole
Publisher Materials Cloud
Publication Year 2021
Rights info:eu-repo/semantics/openAccess; Creative Commons Attribution 4.0 International https://creativecommons.org/licenses/by/4.0/legalcode
OpenAccess true
Contact archive(at)materialscloud.org
Representation
Language English
Resource Type Dataset
Discipline Materials Science and Engineering