Benchmarking "One Molecular Fingerprint to Rule Them All"
In a preprint that is currently on Chemrxiv, Capecchi and coworkers describe a new molecular fingerprint that can be used for similarity searching and machine learning. In the paper, the authors demonstrate how their new fingerprint, called MAP4, performs well when used to build classification models. I wanted to see how well the MAP4 fingerprints performed when used to build regression models. Please note that my intent here is not to be critical of the paper or the authors. I was curious to see how the method would perform with the sorts of models I typically build and thought it might be useful to share my results. I commend the authors for releasing their work as a preprint and for making their code available. In my opinion, this is how science should be done. As usual, all of the code and data that I used to do this analysis is available on GitHub . Rather than detail every step in my analysis here, I'll point the interested reader to the corresponding Jupyter note...