Researchers at the Stowers Institute for Medical Research have developed an innovative interpretation method that reveals, base by base, the insights gained from DNA by a deep-learning model. This method, known as PISA, effectively eliminates hidden experimental biases from genomic data. The details of this groundbreaking research were published in a paper in Nature Communications in August 2026 and were officially announced by the institute on August 25, 2026.

Understanding Sequence-to-Function Neural Networks
Sequence-to-function neural networks are designed to process raw DNA input and forecast the outcomes of various genomic experiments, including transcription factor binding and nucleosome organization. However, these models often lack transparency regarding the reasoning behind their predictions. PISA, which stands for pairwise influence by sequence attribution, addresses this gap by tracing a model’s prediction at a specific genomic location back to every other base that contributed to it. This results in a two-dimensional map that showcases the learned information at single-base resolution rather than just the model’s predictions.
The research was spearheaded by Julia Zeitlinger at Stowers, in collaboration with Anshul Kundaje from Stanford University. Charles McAnany, a Stowers AI Fellow, served as the first author. PISA operates within BPReveal, the latest extension of BPNet, a deep-learning framework developed by the team in 2021.
Disentangling Biological Signals from Experimental Bias
The researchers applied PISA to MNase-seq, a widely used assay for mapping nucleosomes—structures formed when DNA wraps around histone proteins. This assay utilizes an enzyme that selectively cuts exposed DNA while preserving nucleosome-protected DNA. The selectivity of the enzyme introduces two overlapping signals into the data, complicating the model’s learning process.
Traditional interpretation tools tend to summarize each base’s influence into a single value, which can obscure essential details as positive and negative effects may cancel each other out. In contrast, PISA retains full-resolution information, allowing the enzyme’s sequence preference to emerge as a distinct fingerprint on the maps. By mathematically extracting this signature, the team trained a separate model solely on the bias, which they then subtracted from the original, resulting in a refined model focused exclusively on biological signals.
Insights from Bias-Corrected Models
Within the bias-corrected model, PISA unveiled DNA sequences that play a crucial role in positioning nucleosomes, with effects that extend hundreds of base pairs in both directions. Many of these effects were asymmetric, influencing one side of the nucleosome differently from the other. Following this asymmetry led the researchers to chromatin domain boundaries, which delineate the interactions between regulatory sequences and genes. Notably, the model identified thousands of boundaries from nucleosome data alone, often with greater precision than traditional 3D chromatin mapping methods.
The team utilized the biology-focused models to design synthetic DNA sequences that were predicted to arrange nucleosomes in specific configurations. They experimentally tested a subset of these designs, confirming that the model’s learned rules could generate testable hypotheses rather than merely describing existing data.
Complementary Advances in Genomic Modeling
This work arrives at a time when significant investments are being made in increasingly sophisticated sequence models, such as Google DeepMind’s AlphaGenome, which predicts the impacts of regulatory variants across the genome. While PISA complements these efforts, it specifically focuses on understanding the sequence features that inform a model’s predictions after they have been made.
The method’s applicability has already expanded beyond the Zeitlinger lab, with collaborators implementing PISA in separate software packages. Notably, Stowers neuroscientist Neşet Özel has adopted the method for tackling different biological questions.
Bridging the Gap in Genomic Research
This research contributes to a growing understanding of the relationship between genetic variation and disease, although it does not directly produce drug candidates. Most disease-associated genetic variations occur within regulatory DNA rather than coding regions. Zeitlinger emphasizes the importance of placing variants in binding sites or at domain boundaries to propose potential mechanisms, highlighting the need for expertise in both deep learning and experimental biology—a persistent gap in the field.
While the validation of designed DNA sequences was limited to a subset of experimental tests, the demonstration of bias correction in MNase-seq data is a significant advancement. The authors also indicate that PISA can be applied to various data types, enhancing its versatility.
Conclusion
What this research establishes is a novel approach for auditing genomic models, enabling the correction of biases stemming from experimental conditions and allowing for the extraction of precise sequence rules. PISA not only enhances the interpretability of deep-learning models but also opens avenues for designing and testing synthetic biological systems. As this method continues to evolve, it promises to advance our understanding of genomics and its applications in biotechnology.
- Key Takeaways:
- PISA reveals how genomic models interpret DNA data with high precision.
- The method effectively isolates biological signals from experimental biases.
- Insights gained can inform the design of synthetic DNA sequences for experimental validation.
Read more → www.unite.ai
