Image-guided surgery depends on a hard computational problem called 2D/3D registration, working out exactly where a patient’s anatomy sits relative to a live X-ray beam. Get it wrong, and a screw, needle or catheter travels to the wrong place.
The two established approaches both stall. Intensity-based optimization needs per-case tuning and resists generalizing. Deep learning runs fast but wants enormous sets of paired X-rays and ground-truth poses, and a network trained on pelvises cannot read a spine.
A team spanning MIT, Harvard Medical School, Massachusetts General Hospital, Saint Luke’s Marion Bloch Neuroscience Institute and Boston Children’s Hospital describes a third path in Nature. The framework, xvr, uses physics-based differentiable X-ray rendering to synthesize training images from a patient’s own preoperative CT or MR volume at random virtual camera poses. No manual annotation of real X-rays is required.
One model, many anatomies
A foundation network pretrained on thousands of whole-body volumes from public datasets is then fine-tuned to the individual patient in about five minutes, with augmentation that mimics real fluoroscopy. A canonical resampling step lets the same network handle machines with different detector sizes and source-to-detector distances.
The authors call it the largest assessment of 2D/3D registration on real fluoroscopy to date, spanning public femur and cerebral angiography benchmarks plus hospital data. Errors fell by an order of magnitude, with submillimeter success rates ranking first across benchmarks. Code and datasets are open.
