Markus Hoffmann, Tiago Vaz, Shreeti Chhatrala, Lothar Hennighausen
{"title":"Data-driven projections of candidate enhancer-activating SNPs in immune regulation.","authors":"Markus Hoffmann, Tiago Vaz, Shreeti Chhatrala, Lothar Hennighausen","doi":"10.1186/s12864-025-11374-7","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Millions of single nucleotide polymorphisms (SNPs) have been identified in humans, but the functionality of almost all SNPs remains unclear. While current research focuses primarily on SNPs altering one amino acid to another one, the majority of SNPs are located in intergenic spaces. Some of these SNPs can be found in candidate cis-regulatory elements (CREs) such as promoters and enhancers, potentially destroying or creating DNA-binding motifs for transcription factors (TFs) and, hence, deregulating the expression of nearby genes. These aspects are understudied due to the sheer number of SNPs and TF binding motifs, making it challenging to identify SNPs that yield phenotypic changes or altered gene expression.</p><p><strong>Results: </strong>We developed a data-driven computational protocol to prioritize high-potential SNPs informed from former knowledge for experimental validation. We evaluated the protocol by investigating SNPs in CREs in the Janus kinase (JAK) - Signal Transducer and Activator of Transcription (-STAT) signaling pathway, which is activated by a plethora of cytokines and crucial in controlling immune responses and has been implicated in diseases like cancer, autoimmune disorders, and responses to viral infections. The protocol involves scanning the entire human genome (hg38) to pinpoint DNA sequences that deviate by only one nucleotide from the canonical binding sites (TTCnnnGAA) for STAT TFs. We narrowed down from an initial pool of 3,301,512 SNPs across 17,039,967 nearly complete STAT motifs and identified six potential gain-of-function SNPs in regions likely to influence regulation within the JAK-STAT pathway. This selection was guided by publicly available open chromatin and gene expression data and further refined by filtering for proximity to immune response genes and conservation between the mouse and human genomes.</p><p><strong>Conclusion: </strong>Our findings highlight the value of combining genomic, epigenomic, and cross-species conservation data to effectively narrow down millions of SNPs to a smaller number with a high potential to induce interferon regulation of nearby genes. These SNPs can finally be reviewed manually, laying the groundwork for a more focused and efficient exploration of regulatory SNPs in an experimental setting.</p>","PeriodicalId":9030,"journal":{"name":"BMC Genomics","volume":"26 1","pages":"197"},"PeriodicalIF":3.5000,"publicationDate":"2025-02-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMC Genomics","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1186/s12864-025-11374-7","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"BIOTECHNOLOGY & APPLIED MICROBIOLOGY","Score":null,"Total":0}
引用次数: 0
Abstract
Background: Millions of single nucleotide polymorphisms (SNPs) have been identified in humans, but the functionality of almost all SNPs remains unclear. While current research focuses primarily on SNPs altering one amino acid to another one, the majority of SNPs are located in intergenic spaces. Some of these SNPs can be found in candidate cis-regulatory elements (CREs) such as promoters and enhancers, potentially destroying or creating DNA-binding motifs for transcription factors (TFs) and, hence, deregulating the expression of nearby genes. These aspects are understudied due to the sheer number of SNPs and TF binding motifs, making it challenging to identify SNPs that yield phenotypic changes or altered gene expression.
Results: We developed a data-driven computational protocol to prioritize high-potential SNPs informed from former knowledge for experimental validation. We evaluated the protocol by investigating SNPs in CREs in the Janus kinase (JAK) - Signal Transducer and Activator of Transcription (-STAT) signaling pathway, which is activated by a plethora of cytokines and crucial in controlling immune responses and has been implicated in diseases like cancer, autoimmune disorders, and responses to viral infections. The protocol involves scanning the entire human genome (hg38) to pinpoint DNA sequences that deviate by only one nucleotide from the canonical binding sites (TTCnnnGAA) for STAT TFs. We narrowed down from an initial pool of 3,301,512 SNPs across 17,039,967 nearly complete STAT motifs and identified six potential gain-of-function SNPs in regions likely to influence regulation within the JAK-STAT pathway. This selection was guided by publicly available open chromatin and gene expression data and further refined by filtering for proximity to immune response genes and conservation between the mouse and human genomes.
Conclusion: Our findings highlight the value of combining genomic, epigenomic, and cross-species conservation data to effectively narrow down millions of SNPs to a smaller number with a high potential to induce interferon regulation of nearby genes. These SNPs can finally be reviewed manually, laying the groundwork for a more focused and efficient exploration of regulatory SNPs in an experimental setting.
期刊介绍:
BMC Genomics is an open access, peer-reviewed journal that considers articles on all aspects of genome-scale analysis, functional genomics, and proteomics.
BMC Genomics is part of the BMC series which publishes subject-specific journals focused on the needs of individual research communities across all areas of biology and medicine. We offer an efficient, fair and friendly peer review service, and are committed to publishing all sound science, provided that there is some advance in knowledge presented by the work.