It is March 11, 2026 when I write these lines. A bit late to crown my favorite 2025 read. But then I read the latest blog post from Dap where he highlights the value of benchmarking, so I second that here. This was my favorite 2025 read:
Have protein-ligand cofolding methods moved beyond memorisation?
Peter Škrinjar, Jérôme Eberhardt, Gerardo Tauriello, Torsten Schwede, Janani Durairaj
doi: https://doi.org/10.1101/2025.02.03.636309
It is a benchmarking paper of AI-based cofolding methods. Who is interested in that? Turns out – many people in my field. Cofolding was hot in 2025: at the CADD GRC 2025, it felt like every second talk showed figure 1A of this preprint:

Figure 1A from Škrinjar et al. (2025), CC-BY 4.0.
Briefly on the preprint: the benchmarking is about AlphaFold3 and a selection of YAAFCs (Yet Another Alpha Fold Clone). And the answer to the question of the title is: No. Cofolding methods are wonderful in memorizing known systems, but performance drops when no similar protein-ligand systems are available.
Going by the headlines at the breakfast table, this was not obvious, and only in hindsight not surprising. What is surprising though: I cannot see this work in press although the preprint appeared first online in February 2025 (13 months ago!). The preprint has already many citations and has become a cornerstone of how cofolding methods must be evaluated today, as shown by OpenFold3, Pearl, and IsoDDE (“AlphaFold4”).
I want to end this post by thanking the authors of this study. Figure 1A was my favorite plot in 2025.
Disclosure: I know several of the authors personally.
From preprint to NSMB — 15 months
Great news: my favorite 2025 plot has finally been published in a peer-reviewed journal on May 8, 2026. In a good journal, namely in Nat Struct Mol Biol:
doi: https://doi.org/10.1038/s41594-026-01797-5
The paper now has a new title:
Evaluating generalization in protein–ligand cofolding methods
A cool add-on in the publication is a comparison to physics-based docking, as shown in Figure 7 of the paper. It shows that redocking with Glide is insensitive to similarity in the training set – but only in the twice over-idealized scenario where the ground truth receptor and ground truth pocket are used. Glide docking into AF3 or SWISS-MODEL generated models is equally unsuccessful at low similarity as co-folding. The authors speculate that induced fit docking could improve docking into models, and I heartily agree with this based on my own experience.
