Posts

Scaffold Hopping? It's Complicated

Image
As seems to be the case these days, this post was motivated by a comment I saw in the blogosphere.  In one of the myriad discussions on applications of AI in drug discovery, someone wrote:  "I have yet to encounter a machine learning algorithm which predicts a true scaffold hop (say from Viagra to Cialis). From that standpoint, a tool like ROCS which looks at abstract but general features like shape and electrostatics is better than a lot of ML." The comment got me thinking about something that has bugged me for a long time.   What exactly constitutes a scaffold hop?  Should we consider Viagra to Cialis a scaffold hop? (hint, I don't think so, stick with me and I’ll explain)  What is a scaffold hop? Let’s start by taking a step back and looking at some of the classic work of Hans-Joachim Böhm and Martin Stahl.  In their 2004 paper , Böehm and Stahl highlighted three different three-dimensional approaches to scaffold hopping.  ...

Filtering Chemical Libraries

Image
This post was partially motivated by a recent post from Karl Leswing describing how to use the DeepChem package to do virtual screening on a large database.  As part of the tutorial, Karl used the HIV sample file that is part of the DeepChem distribution to build a model.  This model was then used to select compounds from the ZINC database .  The tutorial is nice and the methodology is explained in a manner that is easy to follow.  The problem is that the molecules selected by the model are not what I would consider "drug-like". The molecules reported in Karl's post have aryl sulfonic acid groups.  In fact, molecule D has 2 sulfonic acids AND a thiol.  To someone who has spent a bit of time looking through HTS data, these molecules scream false positive .  If an enterprising young computational chemist were to take this set of hits to an experienced medicinal chemistry colleague, my guess is that he/she would simply shake his/her head...

Cheating at Word Cookies with Python

Image
Ok, so this isn't really a Cheminformatics post, but it's kind of about combinatorial optimization, which is the sort of thing we tend to do a lot of.  Lately, my family has been obsessed with an iOS game called "Word Cookies".  The game is simple (an example is shown below), given a set of 5 letters, you have to construct a set of 3, 4 and 5 letter words.  My wife and daughters are really good at this game, I'm not.  It's funny that the hard part, at least for me, is coming up with the shorter words.   Anyway, being a data geek, I immediately thought about how I could write a Python script to solve this puzzle.  If you think about it, it's pretty easy to just brute force a solution.   Generate all 3, 4 and 5 letter permutations Check each permutation to see if it's a word  Fortunately, there are lots of useful Python libraries and we can do this in just a few lines of code.  Let's take a look.  First off, w...

Free Wilson Analysis

Image
How often have you been in this situation?  You're working on a drug discovery project, you're in mid to late lead optimization, and you're wondering "what have we missed".  "Are there promising combinations of substituents that we didn't synthesize"? Wouldn't it be great if you could look at all the substituent combinations that you didn't make, and identify the combinations that appear promising? Wow, you ask, what is this snazzy new technique?  Actually, it's not a new technique, it's Free-Wilson analysis, which was originally published in 1964 .  The method is pretty simple and can be illustrated through an example like the one in a 1998 review by Hugo Kubinyi .  In this blog post, I'll walk through the method using a Python implementation in GitHub .   1. R-group Decomposition Let's assume we have a set of molecules built around this Markush structure .  In addition, let's assume that we have this set of m...