Mechanistically transparent models for predicting aqueous solubility of rigid, slightly flexible, and very flexible drugs (MW<2000) Accuracy near that of random forest regression.
{"title":"Mechanistically transparent models for predicting aqueous solubility of rigid, slightly flexible, and very flexible drugs (MW<2000) Accuracy near that of random forest regression.","authors":"Alex Avdeef","doi":"10.5599/admet.1879","DOIUrl":null,"url":null,"abstract":"<p><p>Yalkowsky's General Solubility Equation (GSE), with its three fixed constants, is popular and easy to apply, but is not very accurate for polar, zwitterionic, or flexible molecules. This review examines the findings of a series of studies, where we have sought to come up with a better prediction model, by comparing the performances of the GSE to Abraham's Solvation Equation (ABSOLV), and Random Forest regression (RFR) machine-learning (ML) method. Large, well-curated aqueous intrinsic solubility databases are available. However, drugs may be sparsely distributed in chemical space, concentrated in clusters. Even a large database might overlook some regions. Test compounds from under-represented portions of space may be poorly predicted, as might be the case with the 'loose' set of 32 drugs in the Second Solubility Challenge (2020). There appears to be still a need for better coverage of drug space. Increasingly, current trends in predictions of solubility use calculated input descriptors, which may be an advantage for exploring properties of molecules yet to be synthesized. The risk may be that overall prediction approaches might be based on accumulated uncertainty. The increasing use of ML/AI methods can lead to accurate predictions, but such predictions may not readily suggest the strategies to pursue in selecting yet-to-be-synthesized compounds. Based on our latest findings, we recommend predictions based on both 'grouped' ABSOLV(GRP) and 'Flexible Acceptor' GSE(<i>Φ</i>,<i>B</i>) models with the provided best-fit parameters, where <i>Φ</i> is the Kier molecular flexibility index and <i>B</i> is the Abraham H-bond acceptor strength. For molecules with <i>Φ</i> < 11, the prudent choice is to pick the Consensus Model, the average of ABSOLV(GRP) and GSE(Φ,B). For more flexible molecules, GSE(Φ,B) is recommended.</p>","PeriodicalId":7259,"journal":{"name":"ADMET and DMPK","volume":"11 3","pages":"317-330"},"PeriodicalIF":3.4000,"publicationDate":"2023-08-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10567068/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ADMET and DMPK","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.5599/admet.1879","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2023/1/1 0:00:00","PubModel":"eCollection","JCR":"Q2","JCRName":"CHEMISTRY, MEDICINAL","Score":null,"Total":0}
引用次数: 0
Abstract
Yalkowsky's General Solubility Equation (GSE), with its three fixed constants, is popular and easy to apply, but is not very accurate for polar, zwitterionic, or flexible molecules. This review examines the findings of a series of studies, where we have sought to come up with a better prediction model, by comparing the performances of the GSE to Abraham's Solvation Equation (ABSOLV), and Random Forest regression (RFR) machine-learning (ML) method. Large, well-curated aqueous intrinsic solubility databases are available. However, drugs may be sparsely distributed in chemical space, concentrated in clusters. Even a large database might overlook some regions. Test compounds from under-represented portions of space may be poorly predicted, as might be the case with the 'loose' set of 32 drugs in the Second Solubility Challenge (2020). There appears to be still a need for better coverage of drug space. Increasingly, current trends in predictions of solubility use calculated input descriptors, which may be an advantage for exploring properties of molecules yet to be synthesized. The risk may be that overall prediction approaches might be based on accumulated uncertainty. The increasing use of ML/AI methods can lead to accurate predictions, but such predictions may not readily suggest the strategies to pursue in selecting yet-to-be-synthesized compounds. Based on our latest findings, we recommend predictions based on both 'grouped' ABSOLV(GRP) and 'Flexible Acceptor' GSE(Φ,B) models with the provided best-fit parameters, where Φ is the Kier molecular flexibility index and B is the Abraham H-bond acceptor strength. For molecules with Φ < 11, the prudent choice is to pick the Consensus Model, the average of ABSOLV(GRP) and GSE(Φ,B). For more flexible molecules, GSE(Φ,B) is recommended.
期刊介绍:
ADMET and DMPK is an open access journal devoted to the rapid dissemination of new and original scientific results in all areas of absorption, distribution, metabolism, excretion, toxicology and pharmacokinetics of drugs. ADMET and DMPK publishes the following types of contributions: - Original research papers - Feature articles - Review articles - Short communications and Notes - Letters to Editors - Book reviews The scope of the Journal involves, but is not limited to, the following areas: - physico-chemical properties of drugs and methods of their determination - drug permeabilities - drug absorption - drug-drug, drug-protein, drug-membrane and drug-DNA interactions - chemical stability and degradations of drugs - instrumental methods in ADMET - drug metablic processes - routes of administration and excretion of drug - pharmacokinetic/pharmacodynamic study - quantitative structure activity/property relationship - ADME/PK modelling - Toxicology screening - Transporter identification and study