back
86 comments
> “…the focus of AI in drug discovery must shift from doing what can be done - such as modelling data that is readily available, but that is unlikely to move the needle - to doing what should be done, even if this requires, for example, substantial data generation…” It’s a worthy goal, but I think that many involved in this work might be thinking, even unconsciously, “You first”.

This is the problem with AI for all of science - not just drug discovery. Applied ML has spread like wildfire through academia over the past decade - this started well before the LLM hype. It’s the perfect honey trap: research is painstaking and slow, ML offered a shortcut, and best of all, it just needs data. Research produces lots and lots of data! Surely this will be a match made in heaven.

I’ve watched the same pattern play out at least four or five times now in various roles.

(1) Propose an ML-guided approach to material/chemistry discovery/optimization.

(2) Gather existing data (real, experimental data).

(3) Realize there’s less than about 50 true rows of data on the outputs of interest.

At this point, you either: (4a) revert to traditional methods but keep the veneer of using ML to save face, or (4b) pivot to computational/simulation work or a high-throughput system that’s very far removed from your original problem, but allows you to keep playing with ML toys

It’s really bad. I left the industry. I don’t know how long it will take for people doing real science to take back the reins (and the funding).

The real value right now is in figuring out how to generate robust data cheaply and quickly. I'd wager that the effect of a good model on marginal data is small, but the effect of a marginal model on great data is probably quite large.
Here is an article by Pat Walters on the usefulness of ML in drug discovery. This article is a response to another one making the case that utility of ML models are very limited in drug discovery

https://patwalters.github.io/Response-to-Peter-Kenny/

> (4a) revert to traditional methods but keep the veneer of using ML to save face

I haven't worked in the industry side of things but in academia everyone kind of agrees that gradient boosting trees are some of the best models to do these things.

The obvious question is what limits getting more data? Astronomy (especially in Australia) has been quite good at designing surveys to answer multiple scientific questions with reasonable amounts of data (and then fed into ML systems like the cannon). Sadly one of the consequences of the LLM hype is the increasing cost of doing this, so "AI" is actually making things worse not better.
Yes indeed, I call it the revenge of statistics! Pseudoreplication rears its ugly head once again. Nonparametrics are certainly useful, but once you know things about the DGP, it’s often possible to do something “traditional” that’s better. Probabilistic programming, for life!
(3) seems like a problem in its own right? Basing science, traditional or newfangled ML, on such small amounts of data looks pretty weak.
It was always easy to come up with new molecules/drugs/materials. The thing that changed is the scale that it can happen with the new ml-based approaches.

What hasn’t changed is finding ones that are manufacturable/synthesizable.

Even if you find 1 million new stable molecules, there no guarantee that even one of them is manufacturable.

I'm a structural biologist at a mid-sized biotech. I use AI tools daily. They make accomplishing the same things I was able to accomplish before quite a lot faster and easier. They don't help me magically accomplish new things that I couldn't previously.

For example, it helps me install academic software, debug things. It helps me take a large dataset and write scripts to ask questions. It helps me go through experiment drafts to see if I'm missing things. It helps me remember obscure formulas I use every 6 months. It has not, at least in my experience, come up with anything truly novel.

A concrete example: AlphaFold is great...to come up with a starting model for a chimeric fusion or something. What would have taken me 1-2 hours fumbling around in PDB or CIF files is now a quick prompt.

Serious question - have you tried applying AI tools to more of your job, and in a goal-seeking fashion? Have you hit roadblocks?
> For example, it helps... debug things... and write scripts to ask questions.

Im guessing this isnt code that needs to "scale", that needs to "be elegant", that you arent focused on maintainability for the next decade. That its built for purpose and left behind.

Its all the code that for a programer would normally be in this matrix https://xkcd.com/1205/ (is it worth your time) -

do you feel this is the same trade off of UI builders like android studio (or msvb6). you do in minutes what you previously did in 2, 3 hours.

is it all the work? no, but it's a part that's early on and have high perceived impact.

then, as you progress, that tool actually gets in the way and a new feature that would take 2 hours, now is around 2 days.

I think the real win here is for idiots like me:

A) no education

B) no resources

C) not smart enough to be a self-taught bio-hacker

Everyone hears "AI is going to cure disease" and pictures some cure-all pill from a bio lab which is what I feel this paper is hinting at is missingb but that's the top of the funnel; I'm at the bottom where patients live and that is where AI is already quietly working. Its just not being benchmarked.

I built https://crohns.ai. I set out to make an AI-native clinical-trial manager with a feedback loop (DDP) and ended up somewhere completely different: instead of chasing a new "drug" which is totally out of my grasp; financially, intellectually etc... I used it to codify a care protocol that helped me avoid a flare after I got laid off, lost my insurance, and lost access to Skyrizi.

cool. i like what you've done here with a virtual panel that can answer questions ... i have similar ideas for kabuki syndrome (currently: https://www.thekabukipapers.org). talk? marstall at gmail.
> Skyrizi

How are those biologics? Did you have to visit the doctor to get injections frequently?

Derek Lowe discusion of the paper https://www.science.org/content/blog-post/so-how-ai-drug-dis...

I think that was originally linked but got changed to the £30 to Elsevier version for some reason.

Derek Lowe as a science communicator, and others like him, is sorely needed to understand the real meaning and significance of the study and others. I say that as someone with a PhD in chemistry who's been to plenty of presentations on drug discovery topics.

It's difficult to calibrate statements made by other scientists unless you're well embedded within a field: Is this someone whose opinions matter? Are they the subject matter expert they make themselves out to be? Is this research itself truly impactful? Is it really 5 years until it will be realized outside of academic labs? Etc...

It's difficult to decipher questions around credibility because they rely on real-world interactions and associations that extend beyond the physical tokens of paper counts, publication venues, citations, and author lists that typically lag behind the front of human knowledge which is generated from real-world interactions. It can be simple things, like the insightful question a grad student, with minimal publication history, asks in a seminar.

Of course, the paywall is also unhelpful too, but a good, brief commentary by an appropriate commentator is a better link for 99% of prospective readers compared to most "peer reviewed" (scare quotes because that's a real question nowadays) articles.

OP here: the title of the thread still links to Derek Lowe's blog post, but the article discussed in the blog post was added to the body of the original post (not by me).
As to

>clinically relevant impact is, so far, disappointingly limited

it could be that the AI tools have to get to some threshold before they are very useful? Like with the Economist talking to Hassabis:

>AlphaFold itself took six years of work to predict its first protein structure, and then one year to follow up with what he describes as the structures of “all 200m proteins known to science”. He hopes a similar speedup will happen inside Isomorphic.

Need one for hair loss ASAP.
Indeed. My plans for a youthful Mohawk are being stymied by the lack of AI promised medical breakthroughs.

Wheres my follicles dammit?

there is a AI designed drug for hair loss that i know of.

its slow-release oral minoxidil formulation called MINX. AI helped with the formulation [1].

its in in similar category as VDPHL01. Hundreds of millions if not a billion dollars has been invested into Veradermics, and their main product is VDPHL01 (also an extended-release oral formulation).

[1] https://x.com/anagenxyz/status/2071601868841595082

Track KX-826, clascoterone, and VDPHL01
Already exists (finasteride), only problem is that it castrates you chemicallym
A good set of clippers.
Absci has a good asset coming out pretty soon that you could try to get on the trial.
We're all about to come face to face with this reality. This dance can only last so long.
>We're all about to come face to face with this reality. This dance can only last so long.

Only for values of 'all' that exclude well-connected members of the billionaire class and their select associates.

if you think the state of the art in this area is something you'll hear about from a guy that looks like santa in an academic journal, your investments deserve what's about to happen to them.
Obviously the missing part (which we already have for software and math) is that we need agents to be able to run automated loops in the real world. That basically requires robots. I think we'll be there in less than 5 years.
[AI drug discovery] was never the hard part.
This made me laugh, more than expected, but I did visit LinkedIn just before so that could explain it. Thanks!
I hope to be lucky enough to never take a drug that ai had any impact on the discovery or study
How refreshing was this article vs. all the slop?

The lack of comparable data and testability really does seem to be a challenge. I wonder if people would be more willing to collect and share lots of health data if the collecting company was a non-profit dedicated to anonymizing it.

need one for brain plasicity. it would be nice to be able to easily learn a foreign language or musical instrument naturally.
There is one!

Ketamin should have huge impacts on neuro/brain plasticity when used properly (i.e. in therapy)

Easily? Have you ever actually learned a second language or an instrument as a child?
Psilocybin has this effect. Source: I can't remember where I read it, so low confidence.
the comments in this thread is what makes me love hn
> "The paper goes on to make recommendations for AI companies and investigators, and these are well worth reading. The common theme is that people need to think more about why they’re doing certain techniques or using certain technologies, rather than just using them because they’re newly available."

Please. Please let some people with power and influence understand this lesson sooner rather than later. I understand the reasons that's unlikely to occur, but usually the impact isn't quite so drastic and expensive as this is. Just because something is new and shiny doesn't mean that it'll produce the outcomes you need at the other end, and until it's shown that capability your approach to it should be MODERATE.

Drug discovery scientists think about what they're doing and why ALL THE TIME. AI stuff is just another tool.

It's also worth mentioning that drug development timelines typically exceed the interval in which these technologies have been available (or at least effective). Measuring impact will take a long time.

nerd-fanboi proposal: Articles by national treasures (like Derek Lowe, Raymond Chen) should be highlighted with specific identifiers on HN - like a distinctive title font or an ascii diamond ◊.
No thanks. Social media needs less hero worship. Just RSS whoever you like.
Has this yet produced a treatment for AI psychosis?

No? Well fancy that! :)