Posts

The Solubility Forecast Index

Image
Introduction Recently, I've seen a number of deep learning models designed to predict the aqueous solubility of drug-like molecules.  Despite the advantages brought about by techniques like graph neural networks, I have yet to see a commercial or open-source method that outperforms the venerable Solubility Forecast Index (SFI).  I've written about the challenges associated with predicting aqueous solubility before , so I won't revisit that discussion.  Needless to say, this is a difficult problem.   The SFI, published in 2010 by Alan Hill and Robert Young at GSK, provides a simple, elegant equation for estimating aqueous solubility.   SFI   =   c L og D pH7.4   +   #Ar Where  c L og D pH7.4   is the calculated partition coefficient of all neutral and ionic species of a molecule between pH 7.4 buffer and an organic phase, and #Ar is the number of aromatic rings.  This seems pretty simple and should be easy to calcula...

Useful RDKit Utilities

There's a lot of useful functionality in the RDKit .  My problem is remembering where all of the most useful bits are, and how to use them.  In order to make my life, and perhaps yours, a little easier, I put together a Python package called " useful_rdkit_utils ".  Some of what's in there is simply a repackaging of existing functionality to make it easier to use (at least for me).  In other cases, there are functions I borrowed from elsewhere, and there are a few new ideas introduced.  One interesting component in the library is a REOS class that encapsulates the functionality in the rd_filters package I released a few years ago.   I made the package easy to install.  All you have to do is " pip install useful_rdkit_utils ".  The GitHub repo also has Jupyter notebooks that demonstrate some of the functions in the package.  I'm planning to continue to add to the package, and I'm very open to pull requests with corrections and additions...

Picking the Highest Scoring Molecule(s) From Each Cluster

 Here's a quick post based on a conversation with a friend who wanted to be able to cluster a set of docked molecules based on fingerprints and select the highest scoring molecule(s) from each cluster.  As usual, Pandas made this super easy.    The code for this example can be run on Colab  and is also available as a Gist . 

Exploratory Data Analysis With mols2grid and Bemis-Murcko Frameworks

Image
One of the most common tasks in Cheminformatics is exploratory data analysis (EDA).  Given a new dataset, we often need to rapidly explore the chemistry in a set containing hundreds, or even thousands, of molecules.  One useful technique for EDA is the Bemis-Murcko framework .  This technique, originally published by Guy Bemis and Mark Murcko, provides a simple but elegant means of grouping molecules.  Bemis-Murcko frameworks (also known as scaffolds) are created by successively removing monovalent atoms until only ring atoms and linker atoms remain.  There are a few nuances having to do with the removal of exocyclic double bonds and the maintenance of aromaticity, but the method itself is very easy to understand.  There are two versions of the Bemis-Murcko framework, which are sometimes confused.  In the first version, illustrated in the top row of the figure below, monovalent atoms are removed until only ring atoms and linker atoms remain. ...

Similarity Search and Some Cool Pandas Tricks

Image
In this post, we're going to take a look at molecular similarity searches.  Molecular similarity is central to a lot of what we do in Cheminformatics.  It's important for identifying analogs and understanding SAR.  Molecular similarity is also at the core of many clustering methods that we use to understand datasets or design screening libraries.   In this example, we'll be using the chemfp package by Andrew Dalke.  Chemfp has both free and paid tiers.  With the free tier, you can perform similarity searches on smaller datasets, like the one we're using here.  For larger datasets, you need to purchase the paid version.  Chemfp is a great package. If you're using it for production drug discovery, you should buy a license.   In addition to performing searches with chemfp, we'll also go over a few Pandas tricks that will enable us to rapidly process the output from chemfp.  Here's a link to the tutorial notebook on  Google C...

Building a multiclass classification model

 A pointer to the fastpages site. 

Practical Cheminformatics - The Directory

In no particular order, here's a hopefully useful, topical organization of the posts I've written over the past few years. Resources and Reviews A Highly Opinionated List of Open Source Cheminformatics Resources AI in Drug Discovery 2020 - A Highly Opinionated Literature Review Clustering Viewing Clustered Chemical Structures in a Jupyter Notebook Clustering 2.1 Million Compounds for $5 With a Little Help From Amazon & Facebook Self-Organizing Maps - 90s Fad or Useful Tool? (Part 1) Self-Organizing Maps - The Code (Part 2) Molecule Generation Automatic Analog Generation With Common R-group Replacements Predictive Models Predicting Aqueous Solubility - It's Harder Than It Looks Assessing Interpretable Models High-Performance Computing Fast Parallel Cheminformatics Workflows With Dask Wicked Fast Cheminformatics with NVIDIA RAPIDS Databases What Do Molecules That Look LIke This Tend To Do? Adding Chemical Structures to a Recent COVID-19 Drug Repurposing Dataset Filtering ...