Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Targeting Outer Membrane β-barrel Proteins of Burkholderia mallei to design a Multi-epitope Vaccine Against Human Glanders
Kapoor, J.; Panda, A.; Kumar, S.; Bandyopadhyay, A.Abstract
Burkholderia mallei, a facultative intracellular Gram-negative pathogen, is the causative agent of glanders that primarily affects solipeds and is sporadically transmitted to humans. Current interventions mainly rely on antibiotics; however, increasing antimicrobial resistance and the lack of a licensed vaccine further complicate disease management. Surface exposed outer membrane {beta}-barrel (OMBB) proteins serve as excellent targets for vaccine development. In the present study, a consensus-based computational framework was employed on the B. mallei turkey2 proteome that identified 59 OMBB proteins - including porins, TonB receptors, autotransporters, and efflux components. These OMBB proteins were leveraged to predict B- and T-cell epitopes which were manually curated, and mapped onto the corresponding protein models to identify surface-exposed epitopes with direct accessibility to the host immune cells. These epitopes were linked together to construct a multi-epitope vaccine (MEV) that was predicted to be antigenic, and soluble upon overexpression. The tertiary structure of the MEV was generated which was used for molecular docking with TLR4 and TLR2. Molecular dynamics simulation and flexibility analysis confirmed the structural stability of the MEV-TLR4/TLR2 complexes. In-silico immune simulation showed the capability of MEV to induce a strong immune response. Codon optimization and in-silico cloning were performed to evaluate its efficient expression in the E. coli host. The findings suggest that surface exposed OMBB proteins can serve as promising antigenic candidates for designing an MEV construct.
bioinformatics2026-09-15v3TCRdenoise - an unsupervised similarity-based approach for denoising of TCR-pMHC specificity data
Lund, J. M.; Deleuran, S. N.; Nielsen, M.Abstract
Public repositories of T cell receptor (TCR)-peptide-MHC (pMHC) interactions constitute a critical resource for studying adaptive immunity and developing predictive models of TCR specificity. However, recent evidence suggests that a substantial fraction of reported TCR-pMHC interactions may be incorrectly annotated, limiting the quality of downstream analyses and machine learning applications. Here, we present an unsupervised sequence similarity-based framework for denoising peptide-specific TCR repertoires. The method combines pairwise TCR similarity metrics derived from TCRbase and TCRdist3 with hierarchical clustering and a novel adaptation of the silhouette score designed to address the prevalence of singleton clusters and highly imbalanced cluster structures. By incorporating a pseudo-cluster containing singleton and background TCRs, and by optimising both clustering distance thresholds and minimum cluster-size criteria, the proposed approach identifies TCRs likely to represent true antigen-specific binders while filtering putative noise. Using experimentally validated repertoires from TCRvdb, we demonstrate that the modified silhouette score closely tracks clustering solutions that maximise separation between binding and non-binding TCRs, achieving strong agreement with independent validation based on the Matthews correlation coefficient. Extension to a large collection of peptide-specific TCR data revealed a strong negative correlation between the percentage of TCRs classified as noise and the predictive performance of peptide-specific binding models. Further, the denoising classification labels on this data set were corroborated using structural modeling confidence scores of the peptide-TCR interface extracted from a refined AlphaFold 3 modeling pipeline. Additionally, retraining NetTCR on denoised data improved internal cross-validated performance compared with models trained on the full data set, whereas models trained exclusively on TCRs classified as noise performed close to random. Together, these results demonstrate that sequence similarity-based denoising can effectively enrich for biologically meaningful TCR-pMHC interactions and improve the quality of training data for predictive immunological models. The proposed framework provides a scalable strategy for improving the reliability of public TCR databases and facilitating the development of more accurate TCR specificity prediction methods.
bioinformatics2026-09-15v2Effects of UniProtKB restructuring and taxonomic database restrictions on downstream Unipept peptide-centric profiling
Vande Moortele, T.; Van de Vyver, S.; Binke, B.-B.; Van Den Bossche, T.; Dawyndt, P.; Martens, L.; Verschaffelt, P.; Mesuere, B.Abstract
Metaproteomics identifies the proteins present in a microbial community and, from them, which organisms are active and what functions they carry out. Tools such as Unipept assign peptides to taxa and functions by matching them to UniProtKB proteins and taking the lowest common ancestor of the matching taxa, deriving functional annotations from the same matches. This analysis therefore depends directly on the content of UniProtKB, which was substantially restructured in 2025-2026, reducing it by more than 100 million protein entries. We investigated how these changes affect taxonomic interpretation and how restricting the database to sample-specific taxa identified by rRNA profiling modifies that effect. Fixed, non-redundant peptide lists from human-gut and marine-hatchery studies were reanalysed with Unipept against UniProtKB releases 2025_03, 2025_04, and 2026_02. Peptide mapping coverage fell from 85.9% to 73.3% (gut) and from 82.3% to 68.7% (marine) between release 2025_03 and 2026_02, yet dominant taxonomic profiles remained robust: family- and genus-level distributions changed little and all 15 dominant gut species were retained. Root-level assignments dropped from 18.7% to 7.0% (gut) and 25.8% to 14.3% (marine). Restricting UniProtKB 2026_02 to taxa detected by large-subunit ribosomal RNA (LSU rRNA) profiling reduced coverage to 64.4% (gut) and 41.8% (marine), while genus- and species-level proportions changed by less than one percentage point. The major UniProtKB restructuring of 2025-2026 therefore narrowed mapping breadth without destabilizing the dominant taxonomic profiles in these two datasets.
bioinformatics2026-09-15v2SEAHORSE: A Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments
Quackenbush, A.; Kolluri, J.; Biju, R.; Nhong, S.; DeConti, D.; Wu, H.; Quackenbush, J.; Saha, E.; Eicher, T. D.Abstract
Large public molecular atlases such as the Genotype-Tissue Expression (GTEx) project and The Cancer Genome Atlas (TCGA) invite systematic discovery, yet most analyses remain hypothesis-driven and interrogate a tiny fraction of possible relationships among phenotypic, clinical, and molecular variables. We developed SEAHORSE (Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments), a discovery engine and accompanying R package that exhaustively precomputes all pairwise associations across heterogeneous data types and presents them as a searchable association landscape. Using GTEx (948 donors, 43 tissues, 154 phenotypes), SEAHORSE generated 341,008 phenotype-phenotype associations, 125,246,938 phenotype-gene associations, and 10,269,910,410 gene-gene correlations. In parallel analyses spanning 33 tumor types in TCGA, SEAHORSE generated 625,042 phenotype-phenotype associations, 183,080,369 phenotype-gene associations, and 12,096,948,950 gene-gene correlations. Across GTEx, height was repeatedly associated with enrichment of transcriptional programs, most strikingly the Kyoto Encyclopedia of Genes and Genomes (KEGG) term "Pathways in Cancer," significant in 17 tissues, offering a molecular entry point into long-reported links between stature and cancer risk. Height was also associated with immune, cardiovascular, and neurologic programs. In TCGA, age was consistently associated with WNT signaling, translation, cell differentiation, and cell cycle programs across tumors. These findings illustrate a new paradigm: large cohorts should be treated not merely as repositories for testing preconceived hypotheses but as association landscapes that can generate unexpected biological hypotheses.
bioinformatics2026-09-15v2Ryder: Epigenome normalization using a two-tier model and internal reference regions
Cao, Y.; Ge, G.; Zhao, K.Abstract
Motivation: Sequencing-based epigenomic profiling methods are powerful but suffer from technical variability that complicates cross-sample comparisons and can obscure true biological signals. While existing normalization methods using spike-in controls or computational approaches have been proposed, they often rely on assumptions that may not hold across diverse experimental conditions or require additional data types. Results: We present Ryder, a flexible and robust Python package for the normalization of epigenomic signal tracks. Ryder leverages stable internal reference regions, such as invariant CTCF binding sites, to correct for technical artifacts genome-wide. Our results show that it effectively adjusts both background noise and signal intensity, ensuring accurate signal alignment across samples while preserving genuine biological differences. We demonstrate that Ryder performs robustly across diverse assays including DNase-seq, CUT&RUN, ATAC-seq, MNase-seq, and ChIP-seq, with or without spike-in controls. By reducing technical noise, Ryder improves the detection of genuine biological changes, such as quantitative reduction of chromatin accessibility at key enhancer elements by depletion of BRG1, a key subunit of the chromatin remodeling BAF complexes. Availability and Implementation: The Ryder source code, documentation and test data are freely available at: https://github.com/YaqiangCao/ryder . The software version used in this study is archived at Zenodo: https://zenodo.org/records/21267457 .
bioinformatics2026-09-15v2Characterization of Outer Membrane β-barrel Proteins of Pasteurella multocida for Multi-Epitope Vaccine Design against Human Pasteurellosis
Panda, A.; Kapoor, J.; Kumar, S.; Bandyopadhyay, A.Abstract
Pasteurella multocida is a facultative anaerobic, Gram-negative coccobacillus that causes pasteurellosis in companion animals, livestock, and poultry and poses a significant zoonotic risk to humans through bite wounds, scratches, licking, and transfer of bodily fluids. Although vaccines are available for livestock and poultry, no vaccine is currently licensed for human use. In this study, we systematically identified and characterized 29 outer membrane {beta}-barrel (OMBB) proteins in P. multocida Past9 proteome and classified them into functional categories, including TonB-dependent receptors, porins, autotransporters, adhesins, and efflux pumps. B-cell, cytotoxic T-lymphocyte (CTL), and helper T-lymphocyte (HTL) epitopes were predicted from the identified proteins and screened based on antigenicity, non-allergenicity, and non-toxicity. Epitopes conserved across eight human-infecting P. multocida strains and located within the extracellular loop (ECL) region were incorporated into a multi-epitope vaccine (MEV) construct. The designed MEV was predicted to be antigenic, non-allergenic, and soluble. Its tertiary structural model was iteratively refined and validated. Molecular docking with human toll-like receptors 4/2 (TLR4/TLR2) predicted stable interactions, further supported by 100 ns molecular dynamics simulations. Immune simulation of the MEV construct predicted a strong simulated immune response. Furthermore, codon optimization and in silico cloning supported the feasibility of recombinant MEV expression in E. coli. The construct was benchmarked against OmpH, a well-known antigenic protein and exhibited broadly comparable predicted immune responses and receptor-binding energetics. This study proposes a designed MEV candidate against human pasteurellosis and highlights OMBB proteins as potential immunogenic targets for vaccine development.
bioinformatics2026-09-15v2Interpretable Machine Learning Reveals Complementary Age-Related Signatures in the Oral and Gut Microbiome
Ruthbah, C. A.; Sadi, T. H.; Jahan, N. E. S.; Adib, A. N. M. T.Abstract
Whether combining microbiome data from multiple body sites improves prediction, and whether different sites carry complementary or redundant information, are distinct questions that most studies conflate into a single accuracy metric. This work makes two contributions, one methodological and one biological, using paired stool and oral cavity microbiome samples from 44 subjects across two age groups, healthy adults and newborns (Ferretti et al., 2018). Methodologically, we show that a subject-matched fusion design combined with SHAP-based (SHapley Additive exPlanations) site attribution can detect complementary information between body sites even when no measurable accuracy gain results. This is a pattern that conventional model comparison would misread as a null result. Gut (stool) composition alone achieved near-perfect classification (area under the receiver operating characteristic curve, AUC = 1.00), and combined stool-oral models never exceeded this ceiling. A null baseline, bootstrap confidence intervals, and preprocessing sensitivity checks confirmed that this ceiling reflects genuine biological signal rather than an artifact. Despite the flat accuracy curve, SHAP analysis of the fused model showed that oral cavity features carried more total feature importance than stool features (58.1% versus 41.9%), indicating that the model draws on real, non-redundant information from both sites. Biologically, the taxa driving this pattern include Malassezia restricta, Staphylococcus epidermidis, and Prevotella melaninogenica. These taxa behave in a manner consistent with their established roles as early colonizers of the neonatal gut, skin, and oral cavity, once their model-specific behavior is verified directly against abundance data rather than inferred from the literature alone. An independent, substantially larger paired-cohort study using a different analytical method reports a compatible pattern. Together, these results support a model of oral-gut microbiome maturation as two distinct, complementary processes, and demonstrate that detecting this kind of relationship requires examining a model's internal reasoning rather than its accuracy alone.
bioinformatics2026-09-15v1Accounting for pseudo-replication of Linkage Disequilibrium for contemporary Ne estimation
Zhou, L.; Hui, T.-Y. J.; Burt, A.Abstract
The Linkage Disequilibrium (LD) of unlinked loci can be used to estimate contemporary effective population size (Ne) of one to a few generations ago. In genomic datasets loci on different chromosomes are considered unlinked, but there are many more pairs of unlinked loci than there are independent pairs of chromosomes, resulting to confidence intervals (C.I.) being too narrow if the non-independence is not taken into account. Simulations were run to investigate the correlation structure among LD of unlinked loci, which can be expressed by the LD of loci along the same chromosomes, based on a discovery of a novel Random Probe LD estimator. We classify the correlation into two categories: overlapping of loci and disjoint pairs. The former is induced from the same locus being considered twice and is the stronger form of correlation. These correlations feed into {rho}, a parameter to quantify the degree of pseudo-replication in a dataset, and further a correction formula from which C.I. can be properly inferred. We demonstrate the use of our method via an analysis of genomic data from the malaria-transmitting Anopheles gambiae s.s mosquitoes. Apart from the point and C.I. estimates, we find that Var((r^2 ) ) is inflated by about 550 times due to pseudo-replication, highlighting the danger of not handling genetic correlation properly.
bioinformatics2026-09-15v1Cell-level random splits leak group-owned answers in single-cell benchmarks
Sun, S.; Cang, H.Abstract
Machine learning models in single-cell biology increasingly forecast differentiation, reprogramming and therapeutic response from early transcriptomic profiles. Testing whether a model has learned real biology requires held-out cells. Single-cell data, however, are grouped: cells from the same clone, patient or batch share the same label. A random split therefore places relatives of each test cell, carrying its label, in the training set, and a model can score well by memorizing a relative instead of learning a transferable rule. Grouped validation removes this leakage but leaves far fewer independent units behind each error bar. Here we show how to estimate this leakage before training any model, from two properties of the data: exposure, the fraction of test cells with relatives in training, and retrievability, how often a nearest-neighbor search returns such a relative rather than an unrelated cell. Across lineage-barcoded and patient data, exposure determines whether a random split opens a leakage channel, and retrievability determines how much it can inflate the score. The inflation is negligible where cell state has decoupled from ancestry, much larger where clonal sisters remain close in expression space, and in a patient cohort large enough to overturn a clinical conclusion. We also provide eakcheck, which computes both properties in seconds, before the outcome model is fitted.
bioinformatics2026-09-15v1Hierarchical temporal transformer for cancer grade prediction and cross cancer transfer learning from pathology reports
Brimo, N.; Anand, R.; Harb, H.; Serdaroglu, D. C.Abstract
Language models for cancer clinical reports carry two blind spots. They read each report in isolation, ignoring how a patients disease changes across visits, and they are evaluated only on cancer types present in their training data. We present the Hierarchical Temporal Transformer (HTT), a two-level architecture that addresses both. Level 1 encodes each report with BiomedBERT adapted by low-rank adaptation (LoRA). Level 2 is a temporal transformer that reads a patients full report sequence using a continuous-time positional encoding built from the measured number of days between visits, with learnable cancer-type embeddings supplying per-family conditioning. Two experiments test the two capabilities separately, since no fully open corpus contains longitudinal reports for many cancer types. On a controlled synthetic corpus of sequential radiology reports, in which progression phrases are inserted from templated trajectories, HTT reaches a validation AUROC of 0.942 against 0.881 for a single-report baseline and transfers to held-out pancreatic cancer at 0.995 against 0.949 while the two models are indistinguishable on a 60-patient test set. On 4,786 real pathology reports from the TCGA-Reports corpus spanning 14 cancer types HTT predicts tumor grade for three types withheld entirely from training, reaching AUROC 1.000 on thyroid carcinoma, 0.960 on sarcoma and 0.808 on lung squamous cell carcinoma. The mean held-out AUROC of 0.923 equals the in-distribution test AUROC of 0.923, so transfer to unseen cancer families incurred no measurable penalty. Ablation on the real corpus shows that the transfer is carried by the pre-trained encoder rather than by the temporal components, which, with one report per patient, contribute 0.39 AUROC points. Grade-related pathological language therefore appears to be learnable in a cancer-type agnostic way, which points toward unified cancer NLP systems that require no per-type retraining.
bioinformatics2026-09-15v1Accurate and scalable decontamination of imaging-based spatial transcriptomics via optimal transport
Chen, Y.; Liu, Y.; Chao, Z.; Han, S.; Zeng, Y.; Yu, B.; Zhang, F.; Wu, A.; Wang, J.; Chen, H.; Xiao, J.; Yang, C.Abstract
Imaging-based spatial transcriptomics enables molecule-resolved profiling of gene expression and tissue organization in situ. However, segmentation errors, transcript spillover and three-dimensional cell overlap can introduce misassigned transcripts into cell-level expression profiles, compromising biological interpretation and obscuring genuine signals. Existing methods either remove suspect expression at the cost of signal loss or lack a biologically grounded criterion for transcript assignment. Here we present CellDot, an optimal-transport framework that determines the fate of each transcript by retaining it in its host cell, reassigning it to a plausible neighboring cell or removing it as background. By integrating reference-guided expression compatibility with spatial information and data-adaptive constraints, CellDot enables accurate and traceable molecule-level correction while preserving biologically meaningful variation. In evaluations across multiple human tumor datasets, CellDot exhibited superior performance compared to existing decontamination methods, successfully restoring spatial expression patterns that matched independent cross-platform measurements. Moreover, it significantly enhanced the recovery of cellular states, intercellular communication, and spatial niche programs. Our experiments using real data demonstrated CellDot's scalability and established it as the only method applicable to a whole-transcriptome Atera dataset, underscoring its distinct advantages in the field of spatial transcriptomics.
bioinformatics2026-09-15v1Codon language model scores provide information beyond protein language models for missense variant interpretation
Chen, R.; Palpant, N.; Foley, G.; Boden, M.Abstract
Predicting variant effects remains a central challenge in genomics. Protein language models (PLMs) capture amino-acid-level sequence constraints, whereas codon language models operate on coding sequences and may retain information that is lost upon translation. Here, we tested whether scores from the codon language model CaLM provide predictive information beyond protein-level representations for missense-variant interpretation. Across 71,436 ClinVar missense variants from 11,554 genes, adding CaLM to PLM baselines produced modest but reproducible improvements under gene-held-out cross-validation. PLM-only ensemble controls and explicit mutational-context analyses indicated that this improvement could not be explained solely by generic ensembling or simple codon-substitution features. Aggregating CaLM probabilities across synonymous codons attenuated codon-degeneracy-associated discordance while preserving most of the broader differences between CaLM and PLM scores. Gene-level analyses further showed that CaLM contribution varied continuously across genes and depended partly on the protein-model background. Across ClinMAVE functional assays, however, improvements were less consistent, indicating that codon-protein complementarity is context dependent rather than universal. Together, these results identify a modest but reproducible component of variant-effect information in CaLM-derived codon-level scores that is not fully captured by protein-level language-model representations.
bioinformatics2026-09-14v4Differential co-localisation analysis of multi-sample and multi-condition experiments with spatialFDA
Emons, M.; Scheipl, F.; Gunz, S.; Purdom, E.; Robinson, M. D.Abstract
Advances in spatial omics data generation have led to an explosion in new datasets that record the spatial location of transcripts and proteins. However, challenges remain in the analysis of spatial omics data. One important analysis is differential cellular co-localisation (CCoL): the quantification of the clustering, or spacing, of one or more cell types across multiple conditions. Our framework spatialFDA combines methodology from spatial statistics with functional data analysis to accurately quantify and test for differences between conditions in CCoL across spatial scales. Using two simulation studies, we show that spatialFDA performs well in controlled settings. Furthermore, spatialFDA recovers known biological processes in type-1 diabetes and adds insights about the CCoL strength in space. spatialFDA is readily available as an open-source Bioconductor R package.
bioinformatics2026-09-14v2Benchmarking niche identification via domain segmentation for spatial transcriptomics data
Wang, Y.; Chen, Y.; Yang, L.; Wang, C.; Cai, J.; Xin, H.Abstract
Tissue niches are spatially organized microenvironments in which coordinated multicellular interactions shape cellular states and biological functions. Currently, niche identification is routinely performed using domain segmentation frameworks. While interrelated, spatial domains and niches are not fundamentally equivalent. The former emphasizes intra-domain compositional consistency and transcriptomic homogeneity, whereas the latter is defined by the emergent properties of localized signaling gradients and the functional reciprocity between key cell lineages. Here, we present a high-resolution reference by thoroughly annotating single-cell resolution CosMx ST data of a human follicular lymphoid hyperplasia lymph node, a dynamic, non-compartmentalized tissue containing several critical immune niches defined by specific lineage architectures. We systematically benchmarked 16 contemporary domain segmentation algorithms, demonstrating that most methods in their default configurations fail to recapitulate biologically defined niche boundaries. Our analysis reveals that the definitive, disjoint spatial distributions of key functional lineages are frequently obscured by the stochastic infiltration of peripheral cell types. Such reduction in the spatial signal-to-noise ratio represents a primary bottleneck for existing algorithms, which prioritize local transcriptomic variance over global architectural logic. Following this observation, we demonstrate that strategic weighting of core functional lineages can restore the resolution of spatial niches in select domain segmentation frameworks. Cross-comparison against compartmentalized tissues further underscores the unique challenges of niche identification in non-mechanically separated environments and clarifies the fundamental divergence between structural domain segmentation and functional niche discovery. Our work delineates the limitations of current paradigms and advocates for the development of specialized computational approaches tailored specifically to the complexity of functional microenvironments.
bioinformatics2026-09-14v2SpaCoEx: Sparse Gene Selection for Spatially Varying Co-expression in Spatial Transcriptomics
Jian, M.; Pei, S.; Alterovitz, G.Abstract
Spatial transcriptomics enables gene expression to be measured while preserving tissue location, but most existing analyses focus on spatial variation in individual genes or expression-defined domains. Here, we introduce SpaCoEx, a sparse spatial representation framework that integrates gene-expression levels with spatially varying gene-gene co-expression. SpaCoEx first estimates local co-expression matrices from neighboring spatial spots, maps them into a log-Euclidean representation, and performs structured gene selection by retaining or removing the full row and column associated with each gene. The selected genes are then used to construct both expression-level features and local co-expression features, which are combined through an -weighted joint representation for downstream spatial analysis. We applied SpaCoEx to human cutaneous squamous cell carcinoma and annotated human breast cancer spatial transcriptomics datasets. In the cutaneous squamous cell carcinoma dataset, SpaCoEx selected 29 of 45 keratinocyte-related genes while preserving 96.74% of the spatial co-expression variation. In the breast cancer dataset, SpaCoEx identified spatially varying co-expression between B2M and HLA-C, a biologically meaningful major histocompatibility complex (MHC) class I antigen-presentation gene pair. Their local correlation was significantly higher in cancer-associated regions than in non-cancer regions (mean difference = 0.30, spatially adjusted SE = 0.045, P<1.0x10^(-10)), whereas B2M and HLA-C expression individually did not differ significantly between cancer and non-cancer. In benchmarking against manual tissue annotations, co-expression-only SpaCoEx achieved the strongest spatial coherence (percentage of abnormal spots [PAS] = 0.076), while the joint expression/co-expression representation achieved the highest annotation agreement, with an adjusted Rand Index (ARI) of 0.584 at = 0.60 and and normalized mutual information (NMI) of 0.663 at = 0.40. By integrating marginal gene-expression information with local gene-gene co-expression structure, SpaCoEx provides a sparse, low-dimensional, and interpretable representation of spatial transcriptomics data that captures complementary aspects of tissue organization beyond expression-based variation alone.
bioinformatics2026-09-14v2ForceFlowAb: physics-aware mixture-of-experts flow matching model for antibody CDRs design
Li, Z.; Lv, Z.; Zhang, G.Abstract
Abstract Motivation: Antibodies are a major class of therapeutic molecules, and their recognition of target antigens is largely mediated by complementarity-determining regions (CDRs), making antigen-conditioned CDR design a central problem in antibody engineering. Recent generative methods have enabled antigen-conditioned co-design of CDR sequences and structures, but their limited capacity to capture local interface heterogeneity and lack explicit energy-based guidance during sampling, which may result in unfavorable antibody-antigen interaction energies. Overcoming these limitations requires methods that better represent diverse interface environments via adaptive routing and incorporate physical guidance to steer sampling toward energetically favorable conformations. Results: We present ForceFlowAb, a physics-aware mixture-of-experts flow-matching framework for antigen-conditioned CDR sequence-structure co-design. The framework models heterogeneous interface environments through specialized expert routing and applies differentiable force-field guidance during sampling to guide CDR generation toward energetically favorable conformations. For CDR-H3 design, ForceFlowAb achieved more favorable antibody-antigen interaction energies than FlowDesign and Diffab, with improvement rates (IMP) of 46.5% versus 35.0% and 35.5%, respectively. For simultaneous six-CDR design, ForceFlowAb also outperformed Diffab, with IMP values of 16% versus 9%. These results suggest complementary roles for interface-adaptive modeling and energy-based guidance, with the former capturing binding-mode diversity and the latter leveraging physical constraints to ensure biophysical feasibility. Availability and implementation: The web server is freely available at http://zhanglab-bioinf.com/ForceFlowAb. The source code and implementation are available at https://github.com/iobio-zjut/ForceFlowAb. Contact: zgj@zjut.edu.cn Supplementary information: Supplementary data are available at Bioinformatics online.
bioinformatics2026-09-14v1PHACTn enables training-free, context-independent inference of nucleotide variant tolerance across the genome
Yildirim, C.; Kuru, N.; Adebali, O.Abstract
Accurate prediction of single-nucleotide variant (SNV) tolerability across the entire human genome remains a fundamental challenge in computational genomics, particularly for non-coding regions where the regulatory landscape is vast and poorly understood. Machine learning classifiers suffer from data circularity and demographic bias, while genomic language models demand massive computational resources and offer little biological interpretability. Here, we present PHACTn (Phylogeny-Aware Computing of Tolerance for nucleotide variants), a training-free, parameter-minimal method that infers nucleotide variant tolerability by traversing the mammalian phylogenetic tree and explicitly modeling the evolutionary independence of observed substitutions and their distance from the query species. With only 4 interpretable parameters, no training and no GPU requirement, PHACTn outperforms all evaluated tools on non-coding variants curated from both the ClinVar, and on non-coding variants potentially responsible for selected Mendelian diseases curated from OMIM. Additionally, it achieves state-of-the-art performance on variants within the informative range of alignment-based inference. These results establish that principled probabilistic phylogenetic modeling captures evolutionary constraint signals that large-scale sequence models fail to recover, offering a powerful, accessible, and mechanistically transparent alternative for genome-wide variant effect prediction.
bioinformatics2026-09-14v1TransBind2: Improving Transcription Factor-DNA Binding Prediction with Multimodal Data and Bidirectional Cross Attention
Basnet, S.; Cheng, J.Abstract
Accurate genome-wide prediction of transcription factor (TF)-DNA binding remains challenging because many models focus mainly on DNA sequence and overlook chromatin context and TF structure. We previously developed TransBind, a protein-aware model that combines TF and DNA representations through cross-attention. Here, we introduce TransBind2, which improves on TransBind in several ways. It incorporates DNase-seq accessibility and genome mappability tracks as additional input, uses a biomodal protein language model (ProstT5) to capture both TF sequence and structure, and applies bidirectional cross-attention so DNA and protein features can refine each other. We also frame prediction as binary classification of individual <DNA bin, TF, cell type> triplets, allowing the model to generalize to new TFs and cell types. Across 690 human ChIP-seq experiments covering 161 TFs and 91 cell types, TransBind2 achieves a macro AUROC of 0.9648 and AUPR of 0.4215, outperforming TransBind and other baselines, with a [≥]12.67% relative AUPR gain. The model trained on human data also performs well in cross-species zero-shot prediction on mouse data. Saliency analysis shows that it can identify TF-binding peaks with a median error of 12-38 base pairs (bps) despite being trained on window-level labels. Ablation studies further show that TF structure, chromatin accessibility, and bidirectional attention each improve performance. Overall, these results show that combining TF structure with chromatin context leads to more accurate and generalizable TF-DNA binding predictions.
bioinformatics2026-09-14v1LIGER2: Scalable Single-Cell Integration with On-Disk Datasets
Wang, Y.; Robbins, A.; Gadhvi, G.; Welch, J. D.Abstract
Correcting batch effects and integrating single-cell sequencing datasets has been a crucial step in large-scale biological studies. Many methods have been published for this task, and excel in various scenarios. Our previous work, LIGER, leveraging integrative non-negative matrix factorization (iNMF), stands out in providing an interpretable low-dimensional representation. To adapt to the modern need for integrating millions of cells, we developed a highly-optimized parallel factorization solution with on-demand loading from disk. The upgraded LIGER algorithm shows significant improvements in time and memory efficiency for single-cell data integration. We also developed a new downstream embedding alignment method significantly improved performance in conserving biological variation while still aligning corresponding cell types across datasets.
bioinformatics2026-09-14v1ImmuneLens: linking transcriptional states and TCR clonotypes through disentangled multimodal learning
Duan, Z.; Wang, Y.; Li, C.; Li, G.; Cao, Y.; Bai, X.; Yang, F.; Song, S.Abstract
Single-cell multi-omics technologies simultaneously capture the transcriptome and TCR sequence of T cells, providing an opportunity to study the relationship between transcriptional states and clonal architectures. However, jointly modeling the relationships between transcriptional states and TCR sequences while preserving modality-specific information remains challenging. Here, we present ImmuneLens, an interpretable multimodal representation learning framework designed for paired single-cell transcriptome and TCR sequence data. ImmuneLens supports the construction of a transferable multi-cohort immune reference atlas and enables unsupervised mapping of external query data. The complementarity between GEX and TCR information improves the stability of antigen-specificity prediction. In neoadjuvant immunotherapy cohorts, ImmuneLens resolves response-associated T cell heterogeneity and reveals links between clonal expansion and CD8 T cell functional states. Overall, ImmuneLens provides a
bioinformatics2026-09-14v1Heterogeneous graph neural networks with biological prior knowledge for interpretable drug repurposing in triple-negative breast cancer
Fernandez-Lozano, C.; Ferreiro, D.; V.-del-Rio, P.Abstract
Drug repurposing offers a cost-effective path to new therapies for triple-negative breast cancer (TNBC), a subtype with limited targeted treatment options. We present PRECISION, a framework integrating transcription factor (TF) regulatory networks, protein-protein interactions, and drug-target edges into a heterogeneous graph neural network (GNN) to identify TFs mediating drug sensitivity and prioritize repurposing candidates. The knowledge graph has 23,498 nodes and approximately 830,000 edges from CollecTRI, OmniPath and the PRISM screen. In a fair cell-line hold-out evaluation, the GNN matches ML baselines in global prediction (Pearson r = 0.76). Per-drug mechanistic attributions via Integrated Gradients on the trained GNN, replicated across three independent training seeds and robust to baseline choice (Spearman's rho = 0.94 between the mean and the random Gaussian baselines), highlight stress-response (CREB3L1), epithelial-mesenchymal transition (EMT; KLF8, ZEB1), stromal/TNBC-specific (AEBP1, MZF1), and epithelial (SPDEF) regulators as stable mediators of drug response. Multi-cohort validation in SCAN-B (n = 7,397), METABRIC (n = 1,979), and TCGA-BRCA (n = 1,072) shows that predicted drug sensitivity is associated with overall survival in 623 drugs in SCAN-B and 74 in METABRIC (FDR < 0.05). Fisher's meta-analysis identifies 551 drugs at FDR < 0.05, validated by positive controls paclitaxel (p_adj = 8.6x10-3), docetaxel (3.1x10-2), epirubicin (3.2x10-6), and talazoparib (1.4x10-4). Paired Wilcoxon tests in 39 AURORA-US patients with matched primary and metastatic samples confirm that 4 of the 10 IG-identified TFs (CREB3L1, KLF8, AEBP1, SPDEF) are significantly altered during metastatic progression after Bonferroni correction. An explicit rule applied to the PAM50-adjusted Cox multivariate results (penalizer = 0.1) selects seven candidates (osimertinib, saracatinib, erlotinib, brigatinib, pelitinib, entinostat, trametinib), revealing pharmacological convergence on the EGFR signaling axis.
bioinformatics2026-09-14v1Heterogeneous epigenetic regulatory patterns link mammalian aging, development, and mortality
Tikhonov, S.; Dmitriev, S. E.Abstract
Aging is often described as a monotonic accumulation of cellular damage, yet all-cause mortality follows a U-shaped trajectory with age, suggesting non-monotonic molecular changes. We investigated links between early childhood development, aging, and chronic diseases by analyzing DNA methylation in mammalian blood. A meta-analysis of 16 human chronic diseases revealed heterogeneous methylation signatures that formed 2 major disease clusters distinguished by their associations with development and sex-related methylation changes. Although epigenetic entropy increased monotonically across the lifespan, several diseases reduced blood DNA methylation entropy independently of blood cell composition. Across mammals, many CpG sites, particularly in intergenic regions, followed U-shaped age-related methylation changes that paralleled mortality curves. Based on these patterns, we developed epigenetic clocks that predict expected mortality across species and tissues and are effective in detecting a range of disease models. Overall, our findings reveal fundamental links between epigenetic regulation during development, aging, and chronic diseases.
bioinformatics2026-09-14v1OmniTCR: a foundation model unifying T cell receptor recognition prediction and conditional sequence generation
Zeng, F.; Feng, D.; Song, D.; Ding, L.; Tan, Z.; Lei, Q.; Lei, W.; Guo, A.-Y.Abstract
T cell receptor (TCR) recognition prediction and receptor generation are traditionally modelled separately, leaving vast TCR sequence collections disconnected from smaller TCR-peptide-MHC datasets. Here we present OmniTCR, a 113-million-parameter autoregressive foundation model pretrained on 328 million formatted human immune-sequence records. Sequence-type tokens and complementary component orders enable joint learning from individual TCR chains and partial or complete TCR-pMHC associations. On unseen epitopes, OmniTCR achieved AUPRCs of 0.7009 for peptide-TCR{beta}; recognition and 0.8235 for TCR-pMHC interaction prediction, exceeding the strongest evaluated comparators by 0.3396 and 0.3451, respectively. It distinguishes cancer from healthy repertoires across 11 independent pan-cancer cohorts (mean AUROC, 0.9436). The model achieved the highest sequence recovery on internal and external generation benchmarks. Structural modelling supported the plausibility of selected pMHC-conditioned CDR3{beta}; candidates. OmniTCR bridges heterogeneous immune sequence data, providing a foundation for computational immunology and receptor design.
bioinformatics2026-09-13v1pydreg: a fast Python package for identifying active cis-regulatory elements from nascent transcription
He, A. Y.; Danko, C. G.Abstract
Background: Active promoters and enhancers generate characteristic patterns of RNA transcription that can be measured through nascent RNA sequencing. dREG is a leading method that uses these patterns to identify active cis-regulatory elements across the genome, allowing regulatory activity and gene transcription to be profiled in the same experiment. However, its reference implementation was developed around an R-based workflow and a legacy GPU-accelerated support vector machine library that have become increasingly difficult to maintain and deploy. Findings: To improve future usability of dREG, we developed pydreg, a Python port of dREG. pydreg preserves the original pretrained models and peak calling procedure from dREG while using contemporary numerical libraries for CPU and GPU computation. pydreg achieves 4.5 and 5.4-fold reductions in runtime and peak host memory, respectively, compared to dREG while producing near identical peak calls. Conclusions: pydreg reduces practical barriers to running dREG locally, improves runtime and memory usage, integrates readily with Python-based genomics workflows, and provides a maintainable foundation on modern computing infrastructure. Availability and Implementation: pydreg is implemented in Python 3.11+ and is freely available under the GPL-3 license at https://github.com/adamyhe/pydreg and from PyPI via pip install pydreg[gpu] (for CUDA acceleration) or pip install pydreg[mlx] (for Apple Metal acceleration).
bioinformatics2026-09-13v1Detection-Guided Beamforming for Efficient Bat Localisation
Alessandri, R.; Gilmour, L.; Tan, J. J.; Osses, A.; Stowell, D.Abstract
Passive acoustic monitoring is widely used to study wildlife, but current approaches provide limited insight into the spatial behaviour of animals. In bat ecology, reconstructing flight trajectories is essential for studying habitat use, movement patterns, and interactions, yet it remains difficult to achieve under field conditions. Acoustic cameras offer a potential solution by enabling sound source localisation, but their practical application is limited by the computational cost of beamforming and by the non-stationary, broadband, and transient nature of echolocation calls. In particular, exhaustive beamforming over wide ultrasonic bandwidths and dense spatial grids becomes infeasible for continuous monitoring. In this work, we propose a detection-guided beamforming framework for efficient localisation of free-flying bats. The method exploits the sparsity of echolocation signals by restricting beamforming to detector-identified time-frequency regions and combines this with physically consistent short-time analysis parameters and dense spatial sampling. The framework is evaluated on field recordings acquired with a Sorama CAM iV64s acoustic camera. Results show that detection-guided processing reduces the number of beamformer evaluations by approximately 77 times, corresponding to a reduction of about 98.7 % in computational effort, while preserving spatial resolution. At the same time, the proposed parameter configuration improves the stability and sharpness of reconstructed trajectories by avoiding artefacts associated with temporal averaging. These findings demonstrate that high-resolution acoustic localisation can be achieved under realistic computational constraints, supporting the integration of acoustic cameras into ecological monitoring workflows. The proposed framework enables the extraction of spatial trajectories from passive acoustic monitoring data, facilitating spatially resolved analyses of bat behaviour in field conditions.
bioinformatics2026-09-13v1Towards reconstruction of the human interactome from positive and negative experimental evidence
As, J.; Pelz, K.; Bernett, J.; Battini, F.; List, M.; Blumenthal, D. B.; Schaefer, M. H.Abstract
Protein-protein interactions (PPIs) have been detected and reported in the millions, but while they are used in many different contexts for better understanding cellular processes in health and disease, the knowledge of the human PPI network is far from complete, containing many false positive measurements and being highly biased. Both to chart the extent of those problems and to solve them requires not just a knowledge of high-confidence positive interactions, but also likely non-interacting protein pairs. However, this information is typically not reported in PPI studies. We developed a methodology to reconstruct this knowledge from existing PPI data. We reconstruct the experimental search space in which PPI screens have been performed and then create a model that informs how likely a PPI is real given its testing and observation frequency. We argue that negative protein pairs allow us to estimate the error rates of experimental and computational screens. We show how this knowledge could be incorporated for calibration. Finally, we evaluate a simple machine learning approach to PPI prediction and propose how such negative data can be used for training instead of random protein pairs. Together, our results show that reconstructing the experimental search space recovers a largely overlooked layer of information from existing PPI data that can help guide a more accurate and complete mapping of the human interactome.
bioinformatics2026-09-13v1Benchmarking Twist Genotyping-by-Sequencing Against Whole-Genome Sequencing in Nuclear Families
Klugerman, J.; Iossifov, I.; Ye, K.Abstract
Genome-wide genotyping is widely used in human genetics research, and targeted sequencing-based approaches such as the Twist Bioscience genome-wide SNP capture platform (GxS) have emerged as alternatives to conventional SNP arrays. Here, we evaluated GxS genotype calls from 555 individuals in 184 nuclear families against matched whole-genome sequencing (WGS) calls and compared platform performance with that of the Illumina Infinium Global Screening Array-24 (GSA), which was evaluated in 987 individuals from 279 nuclear families. Genotype data were harmonized across platforms, and analyses were restricted to overlapping SNP loci. Across all callable positions, mean per-SNP call rates were 98.26% for GxS and 98.67% for GSA. Overall SNP concordance with WGS was 99.79% for GxS and 99.87% for GSA, and mean per-individual concordance was also 99.79% and 99.87%, respectively. Per-trio Mendelian violation rates of GxS are about 10 times those of WGS, while those of GSA are about 4 times those of WGS on average. These results indicate that GxS performs slightly worse than GSA by key concordance and inheritance metrics, while still showing strong overall agreement with WGS.
bioinformatics2026-09-12v3Leakage-Safe Cross-Cohort Transcriptomic Survival Analysis in Pancreatic Ductal Adenocarcinoma: A Corrected Reanalysis
Markarian, M. B.; Houssigian, G. V.Abstract
Background: Transcriptomic prognostic signatures in pancreatic ductal adenocarcinoma (PDAC) are vulnerable to endpoint misspecification, heterogeneous specimen inclusion, platform effects and information leakage. We re-audited and replaced an earlier batch-harmonized alive/dead classifier. Methods: GSE71729 was audited for assay and composition but excluded from prognosis because its GEO Series Matrix does not provide an auditable patient-level survival join. Development used 176 TCGA-PAAD primary tumors with 92 deaths. A penalized Cox model was evaluated by repeated nested 5-fold cross-validation with filtering, scaling, outcome-guided screening and tuning learned inside each training partition. Cross-platform transport was tested without joint ComBat by holding out 79 survival-labelled primary tumors from GSE85916 (57 deaths), using a per-sample percentile-rank representation. Discrimination was summarized by Harrell's concordance index (C-index) with patient-bootstrap 95% confidence intervals (CIs). The five genes highlighted in version 1 were re-examined by cohort-specific Cox models and random-effects meta-analysis. Results: Internal TCGA C-index was 0.630 (95% CI 0.567-0.694). TCGA-to-GSE85916 C-index was 0.527 (0.435-0.607); reverse-direction transport was 0.576 (0.508-0.647). A fixed five-gene Cox model trained in TCGA gave external C-index 0.569 (0.476-0.657). None of the five random-effects associations survived Benjamini-Hochberg correction across the panel; heterogeneity was high (I2 74%-93%). Conclusions: The corrected data support exploratory transcriptomic signal within TCGA but not a validated transportable prognostic model or validated five-gene biomarker panel. Cohort mixing after batch correction is not evidence of biological preservation or external validity. The revised open workflow provides a reproducible basis for larger, prespecified multi-cohort studies.
bioinformatics2026-09-12v2CRISPR-enhanced assessment of variants of unknown significance nominates oncology therapeutic targets and drug repositioning opportunities
Savino, A.; Oikonomou, A.; Perrone, F.; De Lucia, R. R.; De Pietri, L.; Belattar, Y.; Brown, L.; Grau, M. L.; McCarten, K.; Najgebauer, H.; Perron, U.; Azzolin, L.; Livanova, A.; Cremaschi, P.; Lopez-Bigas, N.; Sottoriva, A.; Coelho, M. A.; IORIO, F.Abstract
Interpreting infrequent somatic variants remains a challenge in cancer genomics. We developed CRISPR-VUS, a framework that uses public Cancer Dependency Map data to identify Dependency-Associated Mutations (DAMs) - variants linked to increased host-gene dependency - with resolution extending to singleton events. Analysis of 977 cell lines across 36 cancer types identified 2,376 DAMs in 1,383 genes, including 1,260 not established as cancer drivers. DAM-bearing genes converge on oncogenic networks, while recurrence in histology-matched tumours, functional-impact predictions, tractability and pharmacological associations enable systematic prioritisation. Prime editing showed that the prioritised NSCLC-specific RTN4IP1-A80T DAM conferred a significant competitive growth advantage in a lung epithelial model, nominating a candidate driver allele. Exploratory pharmacological testing showed a greater maximal response to istaroxime in ATP1B3-I189M-bearing than in ATP1B3-wild-type cells. CRISPR-VUS combines dependency-based rare-variant discovery with evidence-guided prioritisation to nominate candidate drivers, therapeutic targets and drug-repositioning hypotheses. Interactive results are available at https://vus-portal.fht.org/.
bioinformatics2026-09-12v2DNAharvester: A Nextflow Pipeline for Analysing Highly Degraded DNA from Ancient and Historical Specimens
Sharif, B.; Kutschera, V. E.; Oskolkov, N.; Guinet, B.; Lord, E.; Chacon-Duque, J. C.; Oppenheimer, J.; van der Valk, T.; Diez-del-Molino, D.; D. Heintzman, P.; Dalen, L.Abstract
Ancient DNA (aDNA) research has advanced rapidly with the development of high-throughput sequencing, enabling genome-wide analyses of large collections of prehistoric specimens. However, analysing palaeontological and archaeological material with highly degraded DNA constitutes a major bioinformatic challenge. DNA from such samples is characterised by short fragment lengths, low endogenous content, post-mortem damage, and cross-species contamination, which can increase spurious mapping and reference bias, affecting downstream population genetic inferences. We present DNAharvester, a modular and reproducible pipeline designed specifically for processing highly degraded DNA from ancient and historical specimens. DNAharvester integrates metagenomic filtering, competitive mapping, adaptive aligner selection (incorporating BWA-aln, BWA-mem, and Bowtie2), and systematic evaluation of reference bias and spurious mapping. By incorporating flexible mapping and filtering strategies, the pipeline can be adapted to varying sample preservation, focusing on maximising authentic data recovery. DNAharvester features subworkflows for iterative assembly of mitogenomes, identification of genomic repeats and CpG sites, taxonomic classification, microbial/pathogen screening, genetic sex determination, and variant calling. To accommodate varying sequencing depths, the pipeline supports diploid variant calling, genotype likelihood estimation, and pseudo-haploid random allele calling. Implemented in Nextflow, DNAharvester provides a highly scalable, containerised framework that enhances reproducibility, portability, and robustness in aDNA analyses. We validated the pipeline using simulated and empirical datasets, demonstrating its ability to systematically mitigate complex background contamination while preserving authentic genomic signals. By streamlining complex bioinformatic tasks through simple configuration files, DNAharvester establishes a standardised approach for analysing aDNA datasets and makes genomic analyses of ancient remains accessible to the broader research community.
bioinformatics2026-09-12v2CasanovoGUI: a cross-platform desktop application for deep learning-based de novo peptide sequencing with Casanovo
Wen, B.; Li, K.; Riffle, M.; MacCoss, M. J.; Bittremieux, W.; Noble, W. S.Abstract
De novo peptide sequencing detects peptides directly from tandem mass spectra without a protein sequence database, and deep learning has substantially advanced its performance. Casanovo, one such widely used model, is distributed as a Python command-line program. Consequently, installation, GPU and dependency configuration, and manual parameterization can be challenging for many bench scientists and are a recurring source of errors. Interpreting and validating the resulting predictions poses a further challenge. We present CasanovoGUI, an open-source Java-based desktop application that makes all of Casanovo's main analysis functions available through a point-and-click interface on Windows, macOS, and Linux. On first use, CasanovoGUI automatically installs a private Python environment and Casanovo with a GPU-matched build, requiring no prior software setup. The GUI provides access to Casanovo's analysis functions and configuration parameters, streams live progress, and integrates results interpretation: annotated spectra with per-residue confidence scores in the PDV viewer, and mismatch-tolerant mapping of de novo peptides back to a reference proteome. CasanovoGUI is available at https://github.com/Noble-Lab/CasanovoGUI.
bioinformatics2026-09-12v2VARION: A Network Propagation Framework for Individual Patient Somatic Mutation Interpretation in Cancer Molecular Subtyping
Kwon, T.; Park, Y.-G.; Choi, J.-G.Abstract
Accurate molecular subtyping of individual cancer patients from somatic mutation data remains a challenge in precision oncology research. Existing network-based stratification (NBS) methods treat all mutations equivalently, require full-cohort batch processing, and do not demonstrate generalization to independent datasets without retraining. To address this, we present variant interpretation via the adaptive network pRopagatION (VARION), which integrates population-level variant constraint scoring with protein-protein interaction (PPI) network topology. The Adaptive Topology-aware Random Walk with Restart (ATR-RWR) algorithm weights each mutated gene by {varphi}g = {surd}(GIS(g) x {rho}topo(g)), where GIS (Gene Intolerance Score) reflects population-level functional constraint, propagated across a shared PPI network; subtype assignment then uses cosine similarity to TCGA-derived reference centroids, enabling real-time single-patient classification. Across ten TCGA cancer cohorts (n = 2,417), VARION achieved 77.7% accuracy for ovarian cancer (OV), 69.5% for glioblastoma (GBM), 90.2% for cholangiocarcinoma (CHOL), and 75.4% for gastric cancer (STAD). A controlled benchmark applying two alternative clustering methods (PyNBS; a dense autoencoder) to identical ATR-RWR propagation matrices recovered no significant driver enrichment (OR = 1.79 and 1.52, n.s.), versus OR = 144.29 (p = 1.77x10^-12) for VARION, confirming that the GIS-weighted centroid architecture, not propagation alone, drives performance; generalization without retraining was further confirmed in two independent cohorts (ICGC CCA, n = 396; PCAWG, n = 110; OR = {infty}, p < 5x10^-9). Together, these results indicate that VARION's GIS-weighted centroid architecture enables individual-patient molecular subtyping that outperforms existing NBS and graph-learning clustering approaches, with high sensitivity for clinically actionable rare subtypes and robust cross-platform generalization.
bioinformatics2026-09-12v1FlashDeconv reveals resolution horizons in atlas-scale spatial transcriptomics
Yang, C.; Chen, J.; Zhang, X.Abstract
Coarsening Visium HD resolution from 8 to 64 m can flip cell-type co-localization from negative to positive (r = -0.12 [->] +0.80), yet many widely used compositional deconvolution workflows require coarsening or subsampling at million-bin scale. Here we introduce FlashDeconv, which combines leverage-score importance sampling with sparse spatial regularization to achieve competitive benchmark accuracy while processing 1.6 million bins in 153 seconds on commodity hardware. Systematic multi-resolution analysis of Visium HD mouse intestine reveals a tissue-specific resolution horizon (8-16 m), the scale at which this sign inversion occurs, validated by Xenium ground truth. Below this horizon, FlashDeconv provides, to our knowledge, the first sequencing-based quantification of Tuft cell chemosensory niches (15.3-fold stem cell enrichment). In a 1.6-million-bin human colorectal cancer cohort, FlashDeconv uncovers neutrophil inflammatory microdomains co-localized with immunoregulatory dendritic cells (mRegDC) at the tumor-stroma interface, spatial niches largely missed by discrete-label summaries, with RCTD doublet mode labeling only 2.3% of hotspot bins as neutrophil singlets.
bioinformatics2026-09-11v5Fast and accurate taxonomic domain assignment of short metagenomic reads using BBERT
Alekhin, D.; Alon, M.; Sidi, T.; Perez Mazeh, S.; Carmi, G.; Finkel, O. M.; Erez, A.Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses.
bioinformatics2026-09-11v4scPyviewer: a Python-native interactive viewer from AnnData single-cell data
Xuan, H.; Huang, Y.; Bian, J.; Liu, X.Abstract
Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.
bioinformatics2026-09-11v2SenSASP: A Unified, Multi-Layer Database of Senescence and SASP Genes
Xuan, H.; Huang, Y.; Bian, J.Abstract
Research on cellular senescence and the senescence-associated secretory phenotype (SASP) draws on independently curated gene resources that differ in scope, identifiers, and update cycles, making cross-resource integration error-prone. We unified four widely used resources, CellAge, GenAge, the SenMayo signature, and the Reactome Cellular Senescence pathway, onto a single canonical identifier (the Ensembl gene ID) and enriched every gene with three annotation layers absent from all four inputs: cross-species conservation, tissue and cell-type expression, and high-confidence protein-protein interactions. Unification collapsed 1,460 summed source entries into 1,250 unique genes (210 redundant entries removed, 14.4%) while preserving full source provenance: 173 genes are corroborated by two or more resources and two (IL6, JUN) by all four. The three annotation layers reach 95.8%, 97.9%, and 93.0% of genes, with 89.4% annotated across all three. A 500-gene random sample of identifier mappings was validated against HGNC and Ensembl (98.0% exact match). The result, SenSASP, is a single, machine-readable, provenance-tracked database of harmonized identifiers and net-new functional context, illustrated here with a gene-prioritization score and a tissue-expression atlas. SenSASP is freely available at https://xuan13hao.github.io/sensasp/
bioinformatics2026-09-11v2Kintsugi decides, gene by gene, where spatial transcriptomics borrows information
Yang, C.; Zhang, X.; Chen, J.Abstract
Subcellular spatial transcriptomics captures where RNA is in tissue, but a single location holds too few molecules of any one gene to estimate composition alone. Every current method fixes in advance where to borrow -- a smoothing scale, a cell outline or a factor model -- and the fixed choice shapes what is visible. Kintsugi removes the fixed choice and lets held-out molecules decide, gene by gene, how much to borrow from spatial neighbours and from other genes at the same location. On a lung section measured by both Xenium and Visium HD, the data-chosen allocation placed an epithelial programme where the Xenium molecules were, ahead of smoothing, cell segmentation and a factor model; the result replicated across tissues and against protein. Across a 45-core pulmonary fibrosis cohort, separating composition from captured amount shows that a fibroblastic focus is not a place with more RNA but a place with different RNA: 2.8-fold higher in activated-fibroblast composition while segmented nuclear density is at most 1.08-fold higher.
bioinformatics2026-09-11v2Histology-Aware Graph for Modeling Intercellular Communication in Spatial Transcriptomics
Wang, X.; Tao, C.; Jiang, Y.; Jiang, Y.; Liu, H.; Jiang, Z.; Zhu, P.; Que, N.; Xi, J.; Price, S.; Mou, Y.; Xu, J.; Li, C.Abstract
Cell-cell communication (CCC) is essential to how life forms and functions. Recent tools achieve single-cell-resolved CCC inference utilizing spatial transcriptomics (ST). However, most ignore the modeling of tissue contexts surrounding cells, causing high false-positive/negative rates. Here, we propose HARMONIC, a CCC inference method integrating multimodal ST and hematoxylin and eosin (H&E)-stained images. HARMONIC causally modeling the transcriptomic-to-contextual relationships for CCC inference. The state-of-the-art performance was verified across ST platforms, species and healthy/diseased status, on both synthetic and biological samples. HARMONIC was applied in various real-world scenarios, especially on tissues with clear morphological boundaries, including cortical layers in mouse brain, medullary-cortex structures in mouse kidney, as well as tumor-stromal/immune interface. Significant refinement of false-positive/negative predictions was observed compared to ST-only CCC tools.
bioinformatics2026-09-11v2Distinct geometries, comparable interfaces: binding modes and thermodynamic implications in conventional and single domain antibodies
Hauser, A.; Dangla-Pelissier, G.; Cazals, F.Abstract
Heavy-chain only antibodies, produced by the adaptive immune systems of camelids and cartilaginous fish, complement canonical antibodies comprising both heavy and light chain variable domains. Using an integrated interface model that unifies interfacial atoms--including solvent molecules, contacts, buried surface areas, and interface curvature measures, we shed light on two aspects of antibody binding that have so far remained elusive when comparing single domain (SdAb) and double domain (DdAb) antibodies. First, contrary to previous reports of smaller SdAb interfaces, we show that SdAb achieve an interface size comparable to that of DdAb despite using a single variable domain and fewer interface residues, a consequence of a geometric pattern driven by convexity and curvature effects. Second, we show that SdAb exhibit a broader diversity of binding modes than previously reported, with a prominent role played by FR regions. Finally, we discuss the thermodynamic implications of these findings for the design of high-affinity single domain antibodies, with particular relevance to protein engineering and design.
bioinformatics2026-09-11v2Geomosaic: a flexible bioinformatics platform integrating complementary metagenomic analyses from sequencing reads to genomes
Corso, D.; Taccaliti, E.; Barosa, B.; Giovannelli, D.Abstract
Metagenomic analyses can be performed at multiple analytical levels, including read-based, assembly-based, and genome-resolved approaches, each capturing complementary biological information while introducing distinct analytical biases and trade-offs. However, existing workflows are commonly optimized for a single analytical strategy, making it difficult to integrate these complementary representations within a unified, reproducible framework. Here we present Geomosaic, a modular framework that integrates complementary analytical representations of metagenomic data, from reads to genomes, within a single scalable, customizable, and reproducible workflow. Built on a graph-based architecture implemented in Snakemake, Geomosaic enables users to construct complete end-to-end workflows or execute individual analytical modules while selecting among interchangeable software packages. The framework supports read preprocessing, quality control, taxonomic and functional profiling, assembly, genome reconstruction, genome-resolved annotation, custom HMM-based analyses, and automated downstream result aggregation. Automatic generation of execution scripts, modular workflows, and multiple analysis entry points make Geomosaic accessible to researchers approaching metagenomic analyses for the first time, while providing the flexibility and control required by expert users. Native support for HPC environments enables efficient analysis of datasets ranging from individual projects to large-scale metagenomic surveys. Rather than treating read-, assembly-, and genome-resolved metagenomics as alternative analytical strategies, Geomosaic integrates them as complementary representations of the same biological system, allowing users to move seamlessly between community-wide patterns and organism-resolved functional interpretation. By combining workflow flexibility, computational reproducibility, standardized analysis-ready outputs, and extensive documentation, Geomosaic provides a unified platform for environmental metagenomic analyses and facilitates reproducible downstream ecological and evolutionary investigations.
bioinformatics2026-09-11v1CIDER: detecting changes in gene regulatory networks that are associated with changes in phenotype
Jung, W. J.; Ding, M.; Liao, S.; Erdenebaatar, Z.; Brent, M.Abstract
Changes in gene regulatory networks may drive quantitative traits, or may transmit the effects of one trait, such as blood lipid level, on another, such as cardiovascular health. Yet the standard tools, differential correlation and differential network analysis, compare two discrete groups, while the contexts of interest - circulating lipids, inflammation, and blood glucose - vary continuously; applying them forces dichotomization, discarding within-trait variation. We introduce Continuous Interaction-based Differential Edge Regulation (CIDER), which tests whether a gene regulatory network edge, the relationship between a transcription factor and its target gene, varies with a continuous trait: the target gene's expression is modeled as a function of the TF's expression level, the trait, and their interaction, with the interaction coefficient measuring the trait dependence. To limit multiple testing, CIDER tests only the edges of a reference regulatory network. A generalized additive extension detects interactions that change the shape of the relationship, not only its slope, including forms that cannot be expressed as a difference between two correlations. In simulations it outperformed four two-group methods across sample sizes, effect sizes, and noise levels, with most of its advantage from keeping the trait continuous. In whole-blood transcriptomes from four independent human cohorts across ten quantitative health traits, CIDER identified 63 replicated cases in which a TF's regulation of its target varies with the trait, including coupling of the glucocorticoid-receptor (NR3C1) to the granulocyte colony-stimulating-factor receptor (CSF3R) that strengthens as triglycerides rise, and a pair whose regulation reverses direction across the observed range of C-reactive protein.
bioinformatics2026-09-11v1scOLAR: Ontology-Anchored Open-Set Annotation of Single-Cell RNA-seq Data
Liu, Y.; Yi, S.; Yin, H.; Ju, W.Abstract
Single-cell RNA sequencing profiles cellular heterogeneity at atlas scale, making automated annotation essential. However, target datasets often contain novel cell types missing from incomplete references. We present scOLAR, an ontology-guided open-set framework that learns prototypes over the Cell Ontology and uses both reference and target expression to annotate known classes while detecting unfamiliar populations. Guided by ontology hierarchies and decision-boundary regularization, scOLAR penalizes coarse-lineage misclassification and groups novel cells without requiring predefined cluster counts. Across benchmarks, scOLAR achieves a novelty-detection AUROC of 0.9726 and an average precision of 0.9871, enabling structured post-hoc lineage-level interpretation of populations absent from the reference.
bioinformatics2026-09-11v1SpectroVQ: Noise-Aware Compression of Proteomics Data via Vector-Quantized Deep Learning improves MS/MS data storage and Peptide Identification
Lam, H.; Li, J. H. W.; Hoque, A.Abstract
The amount of proteomics data generated has dramatically grown for the past decade due to the wider accessibility to mass spectrometers and technological advances. Current data storage and compression techniques largely treat mass spectra as meaningless series of numbers, wasting storage on useless noise and limiting the compression ratio. Here, we present SpectroVQ, a noise-aware vector-quantized autoencoder to compress and denoise peptide tandem mass spectra without any prior annotation by exploiting peptide fragmentation pattern using deep-learning Evaluation results showed that SpectroVQ can preferentially retain useful signals in spectra from diverse peptide ions, including those in unseen datasets. SpectroVQ achieved over 3-fold increase in compression ratio over mzMLb while maintaining over 0.9 in average cosine similarity and 90% agreement in peptide identifications. In addition, we develop a novel strategy to increase peptide identifications by ~15% via ordinary library searching, by leveraging the tuneable denoising capability of SpectroVQ.
bioinformatics2026-09-11v1Early terminated transcripts and missing proteins reflect artifacts in bacterial proteomes
Insana, G.; Martin, M. J.; Pearson, W. R.Abstract
The high redundancy of many bacterial proteomes can be used to evaluate proteome quality and identify sequence errors. We have used MMseqs2 clustering with subsequent filtering to identify clusters that contain sequences from at least 50% of the clustered proteomes to build sets of core proteins that include proteins from 95% of the clustered bacteria. These clusters typically capture more than 80% of proteins in the bacteria. Because these clusters have highly uniform length (the median cluster has more than 99% of its proteins at the mode length), short (<75% of mode length) or long (>133%) proteins are likely artifacts. Most "outlier" proteins are found in fewer than 10% of clusters, and "high-outlier" clusters are over-represented in a small fraction of proteomes, which often have poor proteome BUSCO fragment scores. Short-outlier proteins are artifacts; at least 80% of short-outlier genomes contain mode-length copies of the protein, which were missed because of frame-shifts, termination codons, or initiation codon choice. MMseqs2 clustering with 50% participation provides robust sets of core bacterial proteins and can be used to identify lower-quality proteomes and proteins.
bioinformatics2026-09-09v5Vizitig a pangenome and pantranscriptome explorer
Degardins, B.; Paperman, C.; MARCHET, C.Abstract
Vizitig is the first platform for real-time exploration and querying of DNA and RNA sequence de Bruijn graphs across many samples, unifying visualization, metadata, and flexible search. It constructs compacted colored de Bruijn graphs from raw sequencing reads and reference sequences, then provides an interactive web interface for graph exploration. It integrates raw and reference-based data, handles complex variation, and provides a human-readable feature-based (also referred to as metadata) query language with scalable graph loading. Its domain-specific query language supports composable searches combining sequences of arbitrary size, genomic features such as gene or exon identifiers, and experimental factors such as sample identifier or abundance thresholds. On-demand subgraph loading retrieves only regions of interest, enabling interactive exploration of large datasets in the graphical user interface without loading the entire graph into memory. We demonstrate Vizitig's capabilities through case studies in pantranscriptomics and pangenomics. In pantranscriptomics, we recover fusion transcript breakpoints on long and short reads. In pangenomics, we explore sequence variations across yeast, rice, nematode, and human pangenomes. Vizitig scales from small virus genomes to human-scale pangenomes (our largest experiments comprises up to 20 assembled human haplotypes, on a laptop). Vizitig enables fast, reproducible analysis in both pangenomics and pantranscriptomics while providing a deployable and user-friendly working environment.
bioinformatics2026-09-09v3Pareto optimization of masked superstrings improves compression of pan-genome k-mer sets
Plachy, J.; Sladky, O.; Brinda, K.; Vesely, P.Abstract
The growing interest in k-mer-based methods across bioinformatics calls for compact k-mer set representations that can be optimized for specific downstream applications. Recently, masked superstrings have provided such flexibility by moving beyond de Bruijn graph paths to general k-mer superstrings equipped with a binary mask, thereby subsuming Spectrum-Preserving String Sets and achieving compactness on arbitrary k-mer sets. However, existing methods optimize superstring length and mask properties in two separate steps, possibly missing solutions where a small increase in superstring length yields a substantial reduction in mask complexity. Here, we introduce the first method for Pareto optimization of k-mer superstrings and masks, and apply it to the problem of compressing pan-genome k-mer sets. We model the compressibility of masked superstrings using an objective that combines superstring length and the number of runs in the mask. We prove that the resulting optimization problem is NP-hard and develop a heuristic based on iterative deepening search in the Aho-Corasick automaton. Using microbial pan-genome datasets, we characterize the Pareto front in the superstring-length/mask-run space and show that the front contains points that Pareto-dominate simplitigs and matchtigs. Finally, we demonstrate that Pareto-optimized masked superstrings improve pan-genome k-mer set compressibility by 12-19% when combined with neural-network compressors, achieving less than 1.2 bits per k-mer in common scenarios.
bioinformatics2026-09-09v3Sampling in structure-token space enables accurate prediction of multiple protein conformations
Wang, Z.; Yu, Y.; Zheng, W.-M.; Yu, C.; Bu, D.Abstract
Protein function is fundamentally mediated by ensembles of distinct metastable states. However, existing methods, such as AlphaFold 3, typically exhibit a bias toward predicting a single dominant state, failing to capture alternative conformations or provide robust metrics for identifying high-quality multi-state conformations. Here, we present MultiStateFold (MSFold), a framework that integrates Parallel Tempering into the discrete structure token space of the ESM3 protein language model. By conceptualizing the model's latent space as an implicit energy landscape, MSFold enables global exploration and barrier crossing, thereby overcoming the local sampling limitations inherent in base generative models. Across a benchmark of 313 multi-conformation pairs, MSFold sets a new performance standard: it achieves the highest success rate in modeling native states and substantially outperforms leading methods, including AlphaFold 3, on challenging alternative conformations, while maintaining competitive accuracy for primary structures. Furthermore, we propose Sequence Log-Likelihood (SLL), a novel confidence metric derived from sequence-structure consistency. Our results demonstrate that SLL offers a modest improvement over standard metrics such as pTM and pLDDT. This work establishes a new paradigm for conformational sampling, bridging classical statistical physics with protein language models.
bioinformatics2026-09-09v3Evaluating Few-Shot Meta-Learning using STUNT for Microbiome-Based Disease Classification
Peng, C.; Abeel, T.Abstract
The human gut microbiome is increasingly explored as a diagnostic indicator for disease, yet machine learning models trained on metagenomic data are often constrained by limited sample sizes and poor cross-cohort generalizability. Meta-learning, a machine learning paradigm that optimizes models for rapid adaptation to new tasks with limited examples, offers a promising strategy to address this by leveraging the potential shared microbial structure across publicly available metagenomic datasets. Here, we evaluated STUNT, a framework combining self-supervised pretraining with metric-based meta-learning (Prototypical Networks), for few-shot microbiome-based disease classification. Using over 5,000 species-level gut metagenomic profiles from 57 cohorts in GMrepo v2, we meta-trained STUNT on 52 cohorts and evaluated the pretrained embedding on five held-out disease cohorts covering rheumatoid arthritis (RA), gestational diabetes mellitus during pregnancy (GDM), non-alcoholic fatty liver disease (NAFLD), diabetes mellitus, type 1 (T1D), and inflammatory bowel disease (IBD). We compared Prototypical Networks, Logistic Regression, and Random Forest with and without STUNT-derived embeddings across shot sizes of 1 to 10 samples per class. We found that STUNT-derived embeddings provided a modest benefit only under extreme data scarcity (one labeled sample per class) and this advantage rapidly diminished and reversed with additional samples, indicating that the meta-learned representations impose an information bottleneck limiting access to task-specific signals. Classification performance varied substantially across cohorts, consistent with PERMANOVA-estimated microbiome-disease separability. These results highlight the need for representation learning approaches that preserve disease- and cohort-specific variation and suggest that intrinsic biological signal strength is the primary determinant of classification success.
bioinformatics2026-09-09v2Hub Facade: View Track Hubs in Integrated Genome Browser
Freese, N. H.; Raveendran, K.; Sirigineedi, J. S.; Chinta, U. L.; Badzuh, P.; Marne, O.; Shetty, C.; Naylor, I.; Jagarapu, S.; Loraine, A. E.Abstract
Summary: Genome browsers are essential for understanding genomic data in detail but use incompatible data repository formats and access protocols, limiting their use. We present a new web application Hub Facade that translates the Track Hub format used by the UCSC Genome Browser into the Quickload format used by the Integrated Genome Browser (IGB), and vice versa. We used the Facade's translation capability to write a single-page web application for IGB users to search and add UCSC-curated Hubs to IGB for exploration and visual analysis. We created a new IGB App (GenArk Genomes) that uses the Facade to automate importing genome assemblies from GenArk, a Hub-based repository with nearly 50,000 genomes hosted by the UCSC Genome Browser team. An example use case investigating alternative splicing of human gene MEOX1 shows how using both browsers to view the same data via the Hub Facade promotes understanding and discovery. Availability and Implementation: Hub Facade is free, open-source software deployed at translate.bioviz.org. Code is available from git repositories at [bitbucket.org|github.com]/lorainelab/hub-facade. The use case is available as Supplemental File 1.
bioinformatics2026-09-09v2FUSED: A Functional Representation for Joint Structural and Elemental Analysis of Protein Ligand Binding Sites
Priyankara, T. M. S.; Ellingson, L.Abstract
Ligand binding site representations are central to the analysis of protein-ligand interactions, with applications in functional characterization, binding-site comparison, and ligand recognition. Many descriptor-based approaches characterize ligand binding sites using a fixed distance threshold from the ligand, despite substantial variability in how such thresholds are defined and the possibility that relevant structural and compositional information changes across spatial scales. We propose Functional Unification of Structural and Elemental Descriptors (FUSED), a multivariate functional representation that jointly models structural and elemental compositional information of ligand binding sites as functions of distance from the ligand. Structural information is captured through covariance-based descriptors derived from the CDPA framework, while chemical composition is represented through isometric log-ratio coordinates to account appropriately for compositional geometry. Treating distance from the ligand as a functional domain allows structural and compositional characteristics to be examined across distance thresholds rather than at a single prespecified value. The resulting representation can be used directly or combined with dimension-reduction, statistical-learning, and/or machine-learning procedures, with the distance interval tailored to the dataset, analytical task, and procedure. We evaluate FUSED on three benchmark datasets spanning complementary ligand binding-site analysis tasks: the Extended Kahraman dataset for multiclass ligand discrimination, TOUGH-C1 for binary binding-site classification, and TOUGH-M1 for pairwise matching of pockets associated with drug-like ligands. FUSED supports strong discrimination in the EK and TOUGH-C1 tasks using standard statistical-learning procedures, while a supervised Siamese neural network applied to the full FUSED representation achieves a mean ROC-AUC of 0.9375 on TOUGH-M1 under repeated sequence-cluster-disjoint evaluation, approaching the strongest reported benchmark performance. These results demonstrate that FUSED provides a flexible, alignment-free representation that supports both direct examination of threshold-dependent binding-site characteristics and competitive downstream classification and pocket matching while remaining computationally practical.
bioinformatics2026-09-09v2