A new competition benchmark published in Nature Biotechnology puts AI antibody design tools through experimental validation, with mixed results.
The AIntibody challenge, modeled on the CASP structure-prediction contests, gathered 511 AI-designed or predicted antibodies from 29 organizations. Teams faced three tasks: in silico affinity maturation, ranking high-affinity clones within clusters of similar antibodies, and designing complementarity-determining regions for proteins outside the training selection.
Several groups produced developable antibodies with binding affinities below 100 picomolar, evidence that AI can genuinely optimize antibodies in well-defined biological settings. But the wins did not transfer across tasks.
Affinity prediction proved especially weak. Except for a single model, ranking high-affinity clones from clustered datasets performed worse than picking clones at random. Out-of-library design also varied sharply, with many submissions failing to beat standard experimental selections.
The authors, led by researchers at Los Alamos National Laboratory and collaborators across industry and academia, say the results separate durable advances from hype while exposing where the field still falls short. Affinity prediction and cross-task generalization remain the biggest gaps for AI in antibody engineering.
