Stowers Institute scientists built a tool that reveals what a DNA model has learned, position by position. Called PISA, the method doubles as a bias detector. It exposed an artifact baked into genomic data and gave the team a way to scrub it out.
Attribution sits at the heart of the technique. PISA walks a prediction at one genomic position back to every other base that shaped it, then renders a base-by-base map of the model’s knowledge. The work landed in Nature Communications and went public August 25. Julia Zeitlinger led the study with Stanford’s Anshul Kundaje; Stowers AI Fellow Charles McAnany is the first author.
The lab folded PISA into BPReveal, its latest BPNet-based deep-learning layer. First up was MNase-seq, the standard technique for locating nucleosomes, where DNA coils around histone proteins to form spool-like structures. The assay’s enzyme trims exposed DNA but prefers certain sequences, so the raw data carried two overlapping signals and the model picked up both.
The problem with older attribution tools was compression. They squeezed each base’s influence into a single score, and opposite signals could erase each other. PISA never compresses. The enzyme’s sequence preferences left a telltale pattern stamped across the full-resolution maps. The team isolated that pattern, built a separate network that reproduced the artifact, and subtracted its output. What remained was a network shaped by biology alone.
With the artifact gone, PISA found sequences that help place nucleosomes, their influence stretching hundreds of base pairs on either side, often unevenly. The lopsided influence flagged chromatin domain boundaries, regions normally charted with expensive 3D sequencing. Zeitlinger likened the added visibility to super-resolution microscopy, where more pixels reveal structure that was always there.
