Posts

Getting Real with Molecular Property Prediction

Image
Introduction If you believe everything you read in the popular press, this AI business is easy. Just ask ChatGPT, and the perfect solution magically appears. Unfortunately, that's not the reality. In this post, I'll walk through a predictive modeling example and demonstrate that there are still a lot of subtleties to consider. In addition, I want to show that data is critical to building good machine learning (ML) models. If you don't have the appropriate data, a simple empirical approach may be better than an ML model.  A recent paper from Cheng Fang and coworkers at Biogen presents prospective evaluations of machine learning models on several ADME endpoints.  As part of this evaluation, the Biogen team released a large dataset of measured in vitro assay values for several thousand commercially available compounds.  One component of this dataset is 2,173 solubility values measured at pH 6.8 using chemiluminescent nitrogen detection (CLND), a technique currently consid...

Using Counterfactuals to Understand Machine Learning Models

Image
While machine learning (ML) models have become integral to many drug discovery efforts, most of these models are "black boxes" that don't explain their predictions.  There are several reasons we would like to be able to explain a prediction.  Provide scientific insights that will guide the design of new compounds.  Instill confidence among team members.  As I've said before, a computational chemist only has two jobs; to convince someone to do an experiment and to convince someone not to do an experiment.  These jobs are much easier when you can explain the "why" behind a prediction.  Debugging and improving models.  Improving a model is easier if you can understand the rationale behind a prediction.   As I wrote in a previous post ,  identifying and highlighting the molecular features that drive an ML prediction can be difficult.  One recent promising approach is the counterfactuals method published by Andrew White's group  at ...

Build a QSAR Model in 8 Lines of Python

Image
This post is just a pointer to a Jupyter notebook.  https://colab.research.google.com/github/PatWalters/practical_cheminformatics_tutorials/blob/main/ml_models/QSAR_in_8_lines.ipynb The code is in this git repo.  https://github.com/PatWalters/practical_cheminformatics_tutorials

Getting Inside the Mind of the Medicinal Chemist with Machine Learning

Image
For over two decades, many people, including me , have been writing programs attempting to replicate "medicinal chemistry intuition".   This ability to identify molecules that would be considered drug-like can be valuable in various areas, including purchasing compounds for screening collections and prioritizing molecules output by generative algorithms or other denovo design methods.  Currently, the most widely used approach for evaluating drug-likeness is the QED method , published by Andrew Hopkins and coworkers at Pfizer in 2012.  QED uses a weighted combination of calculated properties and structural alerts to generate a drug-likeness score for a molecule, with a higher score indicating a more drug-like molecule.  Most recently,  many generative molecular design methods have used QED as part of their objective function.  A new paper from scientists at Novartis and Microsoft presents an alternate approach, called MolSkill, for quantifying dr...

Working With Drug Data from the ChEMBL Database

Image
When working on drug discovery projects, it's handy to have access to a set of chemical structures and associated data for marketed drugs. If you're considering introducing new functionality, someone invariably asks whether that functionality has been used in a marketed drug. It's also helpful to compare the properties of a new compound or compounds to those of marketed drugs. Early in my career, I remember a new medicinal chemist asking Josh Boger, the founder of Vertex Pharmaceuticals, what they should do on their first day of work. Boger responded, "read the Merck Index so you can see what a drug is supposed to look like". Recently a few papers have been published showing how the properties of drugs have changed over time. I thought it might be helpful to create a notebook showing how to extract and clean drug data from ChEMBL and use it for subsequent analysis. The Jupyter notebook is available here on GitHub and can also be run here on Google Colab . In th...

Generative Molecular Design - We Need to Raise the Bar

While it's great that we're now seeing papers describing the experimental validation of generative algorithms for molecular design, we need to consider the significance of these findings and put them into the appropriate context.  Over the last five years, we've seen an explosion in the number of papers describing methods for generative molecular design. The 2018 paper by Gómez-Bombarelli, which launched the field, has already been cited more than 2,100 times. For those unfamiliar with the area, generative molecular design algorithms learn the distributions and associations of chemical functionality from a training set, then sample these distributions to generate new molecules. This molecule generation task can be coupled with one or more scoring functions to generate molecules meeting a specific objective, such as a predicted binding affinity. These methods can be considered similar in spirit to techniques for generating photorealistic images , art , or text that have be...