A benchmark test found leading AI agents completed only 28% of healthcare administrative tasks on a first attempt.