Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Facilitating genome annotation using ANNEXA and long-read RNA sequencing
Hoffmann, N.; Besson, A.; Cadieu, E.; Lorthiois, M.; Le Bars, V.; Houel, A.; Hitte, C.; Andre, C.; Hedan, B.; Derrien, T.Abstract
With the advent of complete genome assemblies, genome annotation has become essential for the functional interpretation of genomic data. Long-read RNA sequencing (LR-RNAseq) technologies have significantly improved transcriptome annotation by enabling full-length transcript reconstruction for both coding and non-coding RNAs. However, challenges such as transcript fragmentation and incomplete isoform representation persist, highlighting the need for robust quality control (QC) strategies. This study presents ANNEXA, a pipeline designed to enhance genome annotation using LR-RNAseq data while also providing QC for reconstructed genes and transcripts. ANNEXA integrates two transcriptome reconstruction tools, StringTie2 and Bambu, applying stringent filtering criteria to improve annotation accuracy. It also incorporates deep learning models to evaluate transcription start sites (TSSs) and employs the tool FEELnc for the systematic annotation of long non-coding RNAs (lncRNAs). Additionally, the pipeline offers intuitive visualisations for comparative analyses of coding and non-coding repertoires. Benchmarking against multiple reference annotations revealed distinct patterns of sensitivity and precision for both known and novel genes and transcripts and mRNAs and lncRNAs. To demonstrate its utility, ANNEXA was applied in a comparative oncology study involving LR-RNAseq of two human and eight canine cancer cell lines. The pipeline successfully identified novel genes and transcripts across species, expanding the catalog of protein-coding and lncRNA annotations in both species. Implemented in Nextflow for scalability and reproducibility, ANNEXA is available as an open-source tool: https://github.com/IGDRion/ANNEXA.
bioinformatics2026-09-17v5Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme
Adamov, A.; Mueller, C. L.; Bokulich, N.Abstract
Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured, and predictive modeling from them typically relies on ad hoc feature representation choices that obscure their impact on performance and interpretation. We present ritme, an open-source Python package that jointly optimizes microbiome-specific feature representation and model selection - combined algorithm selection and hyperparameter optimization - tailored to these data. ritme systematically searches taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment together with model class and hyperparameters, using state-of-the-art optimizers that scale from a laptop to a compute cluster. Across three real-world use cases, ritme outperformed the original study pipelines by 7-29% on the primary task metric and surpassed three AutoML baselines in six of seven comparisons, while selecting substantially fewer features and exposing how feature and model choices drive performance. Open-source and modular, ritme supports reproducible, parsimonious predictive modeling, downstream biological investigation, and extension to other multi-omics modalities.
bioinformatics2026-09-17v4Locat: Joint enrichment and depletion testing identifies localized marker genes in single-cell transcriptomics
Lewis, W. R.; Aizenbud, Y.; Strino, F.; Kluger, Y.; Parisi, F.Abstract
Several methods identify marker genes that delineate cell populations in single-cell transcriptomic data, yet most emphasize enrichment within candidate populations without testing whether expression is significantly reduced elsewhere. We present Locat, a framework for identifying highly specific localized genes by testing whether expression is concentrated within compact regions of a cellular embedding and depleted outside them. For each gene, Locat fits weighted Gaussian mixture models to gene-specific and background densities, computes concentration and depletion statistics, and integrates them into a unified localization score. Across synthetic benchmarks with controlled ground truth, Locat detects uni-modal, multi-modal, and sparse localized patterns and loses significance when expression becomes indistinguishable from background structure. In developmental, perturbation, and differentiation datasets, Locat identifies compact marker sets that capture lineage organization, condition-specific programs, and temporal dynamics. These sets are often smaller than highly variable gene selections, while embeddings built from them preserve major cell populations and developmental programs in several cases. In murine dermis, interferon-treated PBMCs, and retinoic acid-induced embryonic stem cell differentiation, localized genes recover differentiation trajectories, stimulus-responsive programs, and reproducible stage-specific patterns. Together, these results show that jointly assessing concentration and depletion yields specific, interpretable marker genes.
bioinformatics2026-09-17v3FORGE audits residue-level information encoded in RNA tertiary structure geometry
Gow, L.; Li, J.; Tan, X.; Liang, K.; Gui, N.; Luo, B.Abstract
Coarse RNA coordinate representations are widely used, yet the biological information they encode remains unquantified. We introduce FORGE, which converts a seven-atom RNA geometry representation into 935 interpretable descriptors and reports which residue-level annotations this geometry supports. On 4,135 post-2025 RNA chains, FORGE recovered 64.6% of native nucleotides; a six-atom control lacking the glycosidic nitrogen retained 58.5%, locating most of this signal in phosphate-sugar geometry. Confidence was sharply graded: abstaining from the least-confident half of positions raised accuracy to 94.4%, yet many chains remained only partially identifiable. The same descriptors predicted base-pair state far better than a DMS-like proxy or protein-proximal context. Native-decoy, OpenKnot and solved-pseudoknot analyses showed that nucleotide identifiability, foldability and experimental design score are separable: AlphaFold3 reproduced the experimental fold for one of four AI-designed constructs and none of the sequences FORGE read from their geometry. FORGE provides a reproducible audit layer for RNA structural interpretation.
bioinformatics2026-09-17v3Lacuna: Cryptic Binding Pocket Discovery via Conformational Ensemble Analysis
Moore, C.Abstract
Lacuna, an open-source Python tool for discovering cryptic binding pockets: sites that are absent or too small to detect in a protein's unbound structure and open only during conformational fluctuation. Most binding-site predictors score a single static structure, which is precisely the structure in which a cryptic site is invisible. Lacuna instead generates a conformational ensemble from any input structure, detects pockets independently in every conformer, clusters the detections into persistent sites across the ensemble, and ranks those sites with a model fitted on within-structure pairs. Ensemble generation is pluggable: normal mode analysis by default, with implicit-solvent molecular dynamics, Boltz-2 diffusion sampling, or a user-supplied ensemble as alternatives. On the designated test fold of CryptoBench, Lacuna recovers 55.6% of cryptic sites in its top five predictions with the zero-dependency default and 66.1% with an optional PLM-assisted ranker; pooling the geometric detector with an optional learned surface detector recovers 73.9% while raising the fraction of sites found from 68.5% to 86.4%, measured on the held-out fold at five conformers. It recovers 73%, 45% and 87% on the PocketMiner set, a curated set of literature apo/holo pairs, and COACH420 respectively. The default backend completes in a median of 2.6 seconds per chain on one CPU core, so ensemble-based pocket finding does not require a simulation budget. Every site carries a continuous crypticity score, and outputs are emitted as docking-ready Boltz YAML constraints, AutoDock Vina boxes, pseudoatom PDB files, and the generated conformational ensemble as a multi-model PDB. Lacuna is MIT licensed and available at https://github.com/mooreneural/lacuna and on PyPI as lacuna-pockets.
bioinformatics2026-09-17v2Tractography from Serial Optical Coherence Tomography: How and Why?
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.Abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles, invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Due to its high resolution and its 3D nature, serial optical coherence tomography (S-OCT) offers promise for studying WM at the microscale. However, whether the reflectivity contrast from S-OCT supports tractography at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that multiscale Frangi filters outperforms structure tensor analysis for estimating ODF. We also show that anatomically-constrained particle filtering tractography enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing thalamocortical WM projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles that are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
bioinformatics2026-09-17v2iDriver: A patient-centric framework for genome-wide cancer driver discovery
Bahari, F.; Montazeri, H.Abstract
Tumor genomes harbor a mixture of neutral and positively selected mutations, yet distinguishing true cancer drivers remains a major challenge. Several factors can obscure the detection of selection signals, among which patient-specific variation in mutational burden plays a significant role. Current approaches often fail to account for the heterogeneity in mutation burden across different patients; in particular, no existing method explicitly accounts for it when integrating both mutation recurrence and functional impact. Here we present iDriver, a probabilistic graphical model that integrates both mutation recurrence and functional impact at the individual-patient level, enabling an enhanced estimation of positive selection across functional genomic elements. Applying iDriver to 29 cancer types, we identify both known and previously unrecognized drivers spanning coding and noncoding regions, and provide evidence for their clinical and biological relevance. In comprehensive benchmarks against 12 established driver discovery methods, iDriver consistently outperformed all competitors, achieving the highest rankings for known cancer drivers across both coding and noncoding elements.
bioinformatics2026-09-17v2LRP2: A proteogenomics pipeline for long-read informed protein isoform analysis and discovery
Schertzer, M. D.; Lewandowski, J. T.; Watts, E. F.; Rosenow, W.; Mehlferber, M. M.; Jeffery, E. D.; Adamson, S. I.; Bruand, J.; Tseng, E.; Neelamraju, Y.; Garrett-Bakelman, F. E.; Dolzhenko, E.; Knowles, D. A.; Sheynkman, G.Abstract
Most human genes produce multiple RNA isoforms, yet it remains unclear which isoforms are translated into stable, functional proteins. Long-read RNA sequencing resolves full-length transcript structures and, when paired with mass spectrometry, can provide empirical evidence of isoform translation. Despite this opportunity, comprehensive workflows integrating isoform discovery, open reading frame prediction, peptide identification, and protein inference remain limited, leaving users to handle these steps piecemeal. Here, we present LRP2, a modular, end-to-end long-read proteogenomics pipeline built in Nextflow. LRP2 scales transcript discovery to hundreds of samples via PacBio's latest Isocall tool, removes technical artifacts with SQANTI QC, generates and classifies predicted proteomes via CPAT and SQANTI Protein, performs multi-group differential expression and usage analysis via edgeR, DRIMSeq, and a long-read adaptation of LeafCutter, and integrates protein-level evidence from DDA and DIA MS data through FragPipe. For cross-dataset comparison of novel isoforms, LRP2 employs deterministic splice-junction, coordinate-based isoform identifiers. Used as an integrated pipeline, LRP2 enables the detection of novel peptides and improves the protein isoform inference to confirm protein isoform translation.
bioinformatics2026-09-17v2Unlocking Sensitive Data with SPHERE in the Age of AI
He, Z.; Park, J.; Pulgrossi, R. C.; Lee, J.; Butler, R. R.; Weber, A.; Tian, L.; Zhang, X.; Wang, J.; Sha, S.; Mormino, E. C.; Wyss-Coray, T.; Henderson, V. W.; Longo, F. M.; Zou, J.; Desai, M.; Altman, R.Abstract
Sensitive human data underpin discoveries across medicine, biology and the social sciences, yet privacy regulation often prevents sharing them with collaborators or artificial intelligence (AI) systems. We introduce SPHERE, a model-free method that makes sensitive datasets directly usable by AI and shareable for open science as a synthetic twin, while the original records never leave the local environment. Across 33 datasets spanning five scientific domains, SPHERE protects individual privacy against adversarial re-identification attacks while preserving the data's statistical structure: means, variances and correlations are reproduced exactly, effect size and P value in linear statistical analysis is numerically identical, nonlinear machine-learning utility is retained, and each twin is generated in seconds on a laptop. Frontier AI agents running on the twin reach the same scientific conclusions as on the original records. Analyses of the twin reproduce genome- and proteome-wide results at UK Biobank scale and recover the findings of landmark studies across three independent cohorts and consortia. The approach also extends to deep-learning embeddings across language, vision and time-series, with minimal utility loss. We make the Stanford Alzheimer's Disease Research Center cohort openly available for the first time, as a SPHERE twin spanning nine modalities that any registered researcher can analyze without an approval process. We release SPHERE with certification of each twin's privacy and fidelity, and an AI agent that autonomously executes research tasks on sensitive data without ever accessing it. Sensitive datasets that are currently closed to research could thus become routine inputs to open science and AI to enable key discoveries.
bioinformatics2026-09-17v2Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models
De Santis, F.; Malpetti, D.; Gualdi, F.; Mangili, F.Abstract
Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)
bioinformatics2026-09-17v1Targeting Outer Membrane β-barrel Proteins of Burkholderia mallei to design a Multi-epitope Vaccine Against Human Glanders
Kapoor, J.; Panda, A.; Kumar, S.; Bandyopadhyay, A.Abstract
Burkholderia mallei, a facultative intracellular Gram-negative pathogen, is the causative agent of glanders that primarily affects solipeds and is sporadically transmitted to humans. Current interventions mainly rely on antibiotics; however, increasing antimicrobial resistance and the lack of a licensed vaccine further complicate disease management. Surface exposed outer membrane {beta}-barrel (OMBB) proteins serve as excellent targets for vaccine development. In the present study, a consensus-based computational framework was employed on the B. mallei turkey2 proteome that identified 59 OMBB proteins - including porins, TonB receptors, autotransporters, and efflux components. These OMBB proteins were leveraged to predict B- and T-cell epitopes which were manually curated, and mapped onto the corresponding protein models to identify surface-exposed epitopes with direct accessibility to the host immune cells. These epitopes were linked together to construct a multi-epitope vaccine (MEV) that was predicted to be antigenic, and soluble upon overexpression. The tertiary structure of the MEV was generated which was used for molecular docking with TLR4 and TLR2. Molecular dynamics simulation and flexibility analysis confirmed the structural stability of the MEV-TLR4/TLR2 complexes. In-silico immune simulation showed the capability of MEV to induce a strong immune response. Codon optimization and in-silico cloning were performed to evaluate its efficient expression in the E. coli host. The findings suggest that surface exposed OMBB proteins can serve as promising antigenic candidates for designing an MEV construct.
bioinformatics2026-09-15v3Ryder: Epigenome normalization using a two-tier model and internal reference regions
Cao, Y.; Ge, G.; Zhao, K.Abstract
Motivation: Sequencing-based epigenomic profiling methods are powerful but suffer from technical variability that complicates cross-sample comparisons and can obscure true biological signals. While existing normalization methods using spike-in controls or computational approaches have been proposed, they often rely on assumptions that may not hold across diverse experimental conditions or require additional data types. Results: We present Ryder, a flexible and robust Python package for the normalization of epigenomic signal tracks. Ryder leverages stable internal reference regions, such as invariant CTCF binding sites, to correct for technical artifacts genome-wide. Our results show that it effectively adjusts both background noise and signal intensity, ensuring accurate signal alignment across samples while preserving genuine biological differences. We demonstrate that Ryder performs robustly across diverse assays including DNase-seq, CUT&RUN, ATAC-seq, MNase-seq, and ChIP-seq, with or without spike-in controls. By reducing technical noise, Ryder improves the detection of genuine biological changes, such as quantitative reduction of chromatin accessibility at key enhancer elements by depletion of BRG1, a key subunit of the chromatin remodeling BAF complexes. Availability and Implementation: The Ryder source code, documentation and test data are freely available at: https://github.com/YaqiangCao/ryder . The software version used in this study is archived at Zenodo: https://zenodo.org/records/21267457 .
bioinformatics2026-09-15v2Characterization of Outer Membrane β-barrel Proteins of Pasteurella multocida for Multi-Epitope Vaccine Design against Human Pasteurellosis
Panda, A.; Kapoor, J.; Kumar, S.; Bandyopadhyay, A.Abstract
Pasteurella multocida is a facultative anaerobic, Gram-negative coccobacillus that causes pasteurellosis in companion animals, livestock, and poultry and poses a significant zoonotic risk to humans through bite wounds, scratches, licking, and transfer of bodily fluids. Although vaccines are available for livestock and poultry, no vaccine is currently licensed for human use. In this study, we systematically identified and characterized 29 outer membrane {beta}-barrel (OMBB) proteins in P. multocida Past9 proteome and classified them into functional categories, including TonB-dependent receptors, porins, autotransporters, adhesins, and efflux pumps. B-cell, cytotoxic T-lymphocyte (CTL), and helper T-lymphocyte (HTL) epitopes were predicted from the identified proteins and screened based on antigenicity, non-allergenicity, and non-toxicity. Epitopes conserved across eight human-infecting P. multocida strains and located within the extracellular loop (ECL) region were incorporated into a multi-epitope vaccine (MEV) construct. The designed MEV was predicted to be antigenic, non-allergenic, and soluble. Its tertiary structural model was iteratively refined and validated. Molecular docking with human toll-like receptors 4/2 (TLR4/TLR2) predicted stable interactions, further supported by 100 ns molecular dynamics simulations. Immune simulation of the MEV construct predicted a strong simulated immune response. Furthermore, codon optimization and in silico cloning supported the feasibility of recombinant MEV expression in E. coli. The construct was benchmarked against OmpH, a well-known antigenic protein and exhibited broadly comparable predicted immune responses and receptor-binding energetics. This study proposes a designed MEV candidate against human pasteurellosis and highlights OMBB proteins as potential immunogenic targets for vaccine development.
bioinformatics2026-09-15v2Effects of UniProtKB restructuring and taxonomic database restrictions on downstream Unipept peptide-centric profiling
Vande Moortele, T.; Van de Vyver, S.; Binke, B.-B.; Van Den Bossche, T.; Dawyndt, P.; Martens, L.; Verschaffelt, P.; Mesuere, B.Abstract
Metaproteomics identifies the proteins present in a microbial community and, from them, which organisms are active and what functions they carry out. Tools such as Unipept assign peptides to taxa and functions by matching them to UniProtKB proteins and taking the lowest common ancestor of the matching taxa, deriving functional annotations from the same matches. This analysis therefore depends directly on the content of UniProtKB, which was substantially restructured in 2025-2026, reducing it by more than 100 million protein entries. We investigated how these changes affect taxonomic interpretation and how restricting the database to sample-specific taxa identified by rRNA profiling modifies that effect. Fixed, non-redundant peptide lists from human-gut and marine-hatchery studies were reanalysed with Unipept against UniProtKB releases 2025_03, 2025_04, and 2026_02. Peptide mapping coverage fell from 85.9% to 73.3% (gut) and from 82.3% to 68.7% (marine) between release 2025_03 and 2026_02, yet dominant taxonomic profiles remained robust: family- and genus-level distributions changed little and all 15 dominant gut species were retained. Root-level assignments dropped from 18.7% to 7.0% (gut) and 25.8% to 14.3% (marine). Restricting UniProtKB 2026_02 to taxa detected by large-subunit ribosomal RNA (LSU rRNA) profiling reduced coverage to 64.4% (gut) and 41.8% (marine), while genus- and species-level proportions changed by less than one percentage point. The major UniProtKB restructuring of 2025-2026 therefore narrowed mapping breadth without destabilizing the dominant taxonomic profiles in these two datasets.
bioinformatics2026-09-15v2TCRdenoise - an unsupervised similarity-based approach for denoising of TCR-pMHC specificity data
Lund, J. M.; Deleuran, S. N.; Nielsen, M.Abstract
Public repositories of T cell receptor (TCR)-peptide-MHC (pMHC) interactions constitute a critical resource for studying adaptive immunity and developing predictive models of TCR specificity. However, recent evidence suggests that a substantial fraction of reported TCR-pMHC interactions may be incorrectly annotated, limiting the quality of downstream analyses and machine learning applications. Here, we present an unsupervised sequence similarity-based framework for denoising peptide-specific TCR repertoires. The method combines pairwise TCR similarity metrics derived from TCRbase and TCRdist3 with hierarchical clustering and a novel adaptation of the silhouette score designed to address the prevalence of singleton clusters and highly imbalanced cluster structures. By incorporating a pseudo-cluster containing singleton and background TCRs, and by optimising both clustering distance thresholds and minimum cluster-size criteria, the proposed approach identifies TCRs likely to represent true antigen-specific binders while filtering putative noise. Using experimentally validated repertoires from TCRvdb, we demonstrate that the modified silhouette score closely tracks clustering solutions that maximise separation between binding and non-binding TCRs, achieving strong agreement with independent validation based on the Matthews correlation coefficient. Extension to a large collection of peptide-specific TCR data revealed a strong negative correlation between the percentage of TCRs classified as noise and the predictive performance of peptide-specific binding models. Further, the denoising classification labels on this data set were corroborated using structural modeling confidence scores of the peptide-TCR interface extracted from a refined AlphaFold 3 modeling pipeline. Additionally, retraining NetTCR on denoised data improved internal cross-validated performance compared with models trained on the full data set, whereas models trained exclusively on TCRs classified as noise performed close to random. Together, these results demonstrate that sequence similarity-based denoising can effectively enrich for biologically meaningful TCR-pMHC interactions and improve the quality of training data for predictive immunological models. The proposed framework provides a scalable strategy for improving the reliability of public TCR databases and facilitating the development of more accurate TCR specificity prediction methods.
bioinformatics2026-09-15v2SEAHORSE: A Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments
Quackenbush, A.; Kolluri, J.; Biju, R.; Nhong, S.; DeConti, D.; Wu, H.; Quackenbush, J.; Saha, E.; Eicher, T. D.Abstract
Large public molecular atlases such as the Genotype-Tissue Expression (GTEx) project and The Cancer Genome Atlas (TCGA) invite systematic discovery, yet most analyses remain hypothesis-driven and interrogate a tiny fraction of possible relationships among phenotypic, clinical, and molecular variables. We developed SEAHORSE (Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments), a discovery engine and accompanying R package that exhaustively precomputes all pairwise associations across heterogeneous data types and presents them as a searchable association landscape. Using GTEx (948 donors, 43 tissues, 154 phenotypes), SEAHORSE generated 341,008 phenotype-phenotype associations, 125,246,938 phenotype-gene associations, and 10,269,910,410 gene-gene correlations. In parallel analyses spanning 33 tumor types in TCGA, SEAHORSE generated 625,042 phenotype-phenotype associations, 183,080,369 phenotype-gene associations, and 12,096,948,950 gene-gene correlations. Across GTEx, height was repeatedly associated with enrichment of transcriptional programs, most strikingly the Kyoto Encyclopedia of Genes and Genomes (KEGG) term "Pathways in Cancer," significant in 17 tissues, offering a molecular entry point into long-reported links between stature and cancer risk. Height was also associated with immune, cardiovascular, and neurologic programs. In TCGA, age was consistently associated with WNT signaling, translation, cell differentiation, and cell cycle programs across tumors. These findings illustrate a new paradigm: large cohorts should be treated not merely as repositories for testing preconceived hypotheses but as association landscapes that can generate unexpected biological hypotheses.
bioinformatics2026-09-15v2Cell-level random splits leak group-owned answers in single-cell benchmarks
Sun, S.; Cang, H.Abstract
Machine learning models in single-cell biology increasingly forecast differentiation, reprogramming and therapeutic response from early transcriptomic profiles. Testing whether a model has learned real biology requires held-out cells. Single-cell data, however, are grouped: cells from the same clone, patient or batch share the same label. A random split therefore places relatives of each test cell, carrying its label, in the training set, and a model can score well by memorizing a relative instead of learning a transferable rule. Grouped validation removes this leakage but leaves far fewer independent units behind each error bar. Here we show how to estimate this leakage before training any model, from two properties of the data: exposure, the fraction of test cells with relatives in training, and retrievability, how often a nearest-neighbor search returns such a relative rather than an unrelated cell. Across lineage-barcoded and patient data, exposure determines whether a random split opens a leakage channel, and retrievability determines how much it can inflate the score. The inflation is negligible where cell state has decoupled from ancestry, much larger where clonal sisters remain close in expression space, and in a patient cohort large enough to overturn a clinical conclusion. We also provide eakcheck, which computes both properties in seconds, before the outcome model is fitted.
bioinformatics2026-09-15v1Accurate and scalable decontamination of imaging-based spatial transcriptomics via optimal transport
Chen, Y.; Liu, Y.; Chao, Z.; Han, S.; Zeng, Y.; Yu, B.; Zhang, F.; Wu, A.; Wang, J.; Chen, H.; Xiao, J.; Yang, C.Abstract
Imaging-based spatial transcriptomics enables molecule-resolved profiling of gene expression and tissue organization in situ. However, segmentation errors, transcript spillover and three-dimensional cell overlap can introduce misassigned transcripts into cell-level expression profiles, compromising biological interpretation and obscuring genuine signals. Existing methods either remove suspect expression at the cost of signal loss or lack a biologically grounded criterion for transcript assignment. Here we present CellDot, an optimal-transport framework that determines the fate of each transcript by retaining it in its host cell, reassigning it to a plausible neighboring cell or removing it as background. By integrating reference-guided expression compatibility with spatial information and data-adaptive constraints, CellDot enables accurate and traceable molecule-level correction while preserving biologically meaningful variation. In evaluations across multiple human tumor datasets, CellDot exhibited superior performance compared to existing decontamination methods, successfully restoring spatial expression patterns that matched independent cross-platform measurements. Moreover, it significantly enhanced the recovery of cellular states, intercellular communication, and spatial niche programs. Our experiments using real data demonstrated CellDot's scalability and established it as the only method applicable to a whole-transcriptome Atera dataset, underscoring its distinct advantages in the field of spatial transcriptomics.
bioinformatics2026-09-15v1Interpretable Machine Learning Reveals Complementary Age-Related Signatures in the Oral and Gut Microbiome
Ruthbah, C. A.; Sadi, T. H.; Jahan, N. E. S.; Adib, A. N. M. T.Abstract
Whether combining microbiome data from multiple body sites improves prediction, and whether different sites carry complementary or redundant information, are distinct questions that most studies conflate into a single accuracy metric. This work makes two contributions, one methodological and one biological, using paired stool and oral cavity microbiome samples from 44 subjects across two age groups, healthy adults and newborns (Ferretti et al., 2018). Methodologically, we show that a subject-matched fusion design combined with SHAP-based (SHapley Additive exPlanations) site attribution can detect complementary information between body sites even when no measurable accuracy gain results. This is a pattern that conventional model comparison would misread as a null result. Gut (stool) composition alone achieved near-perfect classification (area under the receiver operating characteristic curve, AUC = 1.00), and combined stool-oral models never exceeded this ceiling. A null baseline, bootstrap confidence intervals, and preprocessing sensitivity checks confirmed that this ceiling reflects genuine biological signal rather than an artifact. Despite the flat accuracy curve, SHAP analysis of the fused model showed that oral cavity features carried more total feature importance than stool features (58.1% versus 41.9%), indicating that the model draws on real, non-redundant information from both sites. Biologically, the taxa driving this pattern include Malassezia restricta, Staphylococcus epidermidis, and Prevotella melaninogenica. These taxa behave in a manner consistent with their established roles as early colonizers of the neonatal gut, skin, and oral cavity, once their model-specific behavior is verified directly against abundance data rather than inferred from the literature alone. An independent, substantially larger paired-cohort study using a different analytical method reports a compatible pattern. Together, these results support a model of oral-gut microbiome maturation as two distinct, complementary processes, and demonstrate that detecting this kind of relationship requires examining a model's internal reasoning rather than its accuracy alone.
bioinformatics2026-09-15v1Accounting for pseudo-replication of Linkage Disequilibrium for contemporary Ne estimation
Zhou, L.; Hui, T.-Y. J.; Burt, A.Abstract
The Linkage Disequilibrium (LD) of unlinked loci can be used to estimate contemporary effective population size (Ne) of one to a few generations ago. In genomic datasets loci on different chromosomes are considered unlinked, but there are many more pairs of unlinked loci than there are independent pairs of chromosomes, resulting to confidence intervals (C.I.) being too narrow if the non-independence is not taken into account. Simulations were run to investigate the correlation structure among LD of unlinked loci, which can be expressed by the LD of loci along the same chromosomes, based on a discovery of a novel Random Probe LD estimator. We classify the correlation into two categories: overlapping of loci and disjoint pairs. The former is induced from the same locus being considered twice and is the stronger form of correlation. These correlations feed into {rho}, a parameter to quantify the degree of pseudo-replication in a dataset, and further a correction formula from which C.I. can be properly inferred. We demonstrate the use of our method via an analysis of genomic data from the malaria-transmitting Anopheles gambiae s.s mosquitoes. Apart from the point and C.I. estimates, we find that Var((r^2 ) ) is inflated by about 550 times due to pseudo-replication, highlighting the danger of not handling genetic correlation properly.
bioinformatics2026-09-15v1Hierarchical temporal transformer for cancer grade prediction and cross cancer transfer learning from pathology reports
Brimo, N.; Anand, R.; Harb, H.; Serdaroglu, D. C.Abstract
Language models for cancer clinical reports carry two blind spots. They read each report in isolation, ignoring how a patients disease changes across visits, and they are evaluated only on cancer types present in their training data. We present the Hierarchical Temporal Transformer (HTT), a two-level architecture that addresses both. Level 1 encodes each report with BiomedBERT adapted by low-rank adaptation (LoRA). Level 2 is a temporal transformer that reads a patients full report sequence using a continuous-time positional encoding built from the measured number of days between visits, with learnable cancer-type embeddings supplying per-family conditioning. Two experiments test the two capabilities separately, since no fully open corpus contains longitudinal reports for many cancer types. On a controlled synthetic corpus of sequential radiology reports, in which progression phrases are inserted from templated trajectories, HTT reaches a validation AUROC of 0.942 against 0.881 for a single-report baseline and transfers to held-out pancreatic cancer at 0.995 against 0.949 while the two models are indistinguishable on a 60-patient test set. On 4,786 real pathology reports from the TCGA-Reports corpus spanning 14 cancer types HTT predicts tumor grade for three types withheld entirely from training, reaching AUROC 1.000 on thyroid carcinoma, 0.960 on sarcoma and 0.808 on lung squamous cell carcinoma. The mean held-out AUROC of 0.923 equals the in-distribution test AUROC of 0.923, so transfer to unseen cancer families incurred no measurable penalty. Ablation on the real corpus shows that the transfer is carried by the pre-trained encoder rather than by the temporal components, which, with one report per patient, contribute 0.39 AUROC points. Grade-related pathological language therefore appears to be learnable in a cancer-type agnostic way, which points toward unified cancer NLP systems that require no per-type retraining.
bioinformatics2026-09-15v1Codon language model scores provide information beyond protein language models for missense variant interpretation
Chen, R.; Palpant, N.; Foley, G.; Boden, M.Abstract
Predicting variant effects remains a central challenge in genomics. Protein language models (PLMs) capture amino-acid-level sequence constraints, whereas codon language models operate on coding sequences and may retain information that is lost upon translation. Here, we tested whether scores from the codon language model CaLM provide predictive information beyond protein-level representations for missense-variant interpretation. Across 71,436 ClinVar missense variants from 11,554 genes, adding CaLM to PLM baselines produced modest but reproducible improvements under gene-held-out cross-validation. PLM-only ensemble controls and explicit mutational-context analyses indicated that this improvement could not be explained solely by generic ensembling or simple codon-substitution features. Aggregating CaLM probabilities across synonymous codons attenuated codon-degeneracy-associated discordance while preserving most of the broader differences between CaLM and PLM scores. Gene-level analyses further showed that CaLM contribution varied continuously across genes and depended partly on the protein-model background. Across ClinMAVE functional assays, however, improvements were less consistent, indicating that codon-protein complementarity is context dependent rather than universal. Together, these results identify a modest but reproducible component of variant-effect information in CaLM-derived codon-level scores that is not fully captured by protein-level language-model representations.
bioinformatics2026-09-14v4Differential co-localisation analysis of multi-sample and multi-condition experiments with spatialFDA
Emons, M.; Scheipl, F.; Gunz, S.; Purdom, E.; Robinson, M. D.Abstract
Advances in spatial omics data generation have led to an explosion in new datasets that record the spatial location of transcripts and proteins. However, challenges remain in the analysis of spatial omics data. One important analysis is differential cellular co-localisation (CCoL): the quantification of the clustering, or spacing, of one or more cell types across multiple conditions. Our framework spatialFDA combines methodology from spatial statistics with functional data analysis to accurately quantify and test for differences between conditions in CCoL across spatial scales. Using two simulation studies, we show that spatialFDA performs well in controlled settings. Furthermore, spatialFDA recovers known biological processes in type-1 diabetes and adds insights about the CCoL strength in space. spatialFDA is readily available as an open-source Bioconductor R package.
bioinformatics2026-09-14v2Benchmarking niche identification via domain segmentation for spatial transcriptomics data
Wang, Y.; Chen, Y.; Yang, L.; Wang, C.; Cai, J.; Xin, H.Abstract
Tissue niches are spatially organized microenvironments in which coordinated multicellular interactions shape cellular states and biological functions. Currently, niche identification is routinely performed using domain segmentation frameworks. While interrelated, spatial domains and niches are not fundamentally equivalent. The former emphasizes intra-domain compositional consistency and transcriptomic homogeneity, whereas the latter is defined by the emergent properties of localized signaling gradients and the functional reciprocity between key cell lineages. Here, we present a high-resolution reference by thoroughly annotating single-cell resolution CosMx ST data of a human follicular lymphoid hyperplasia lymph node, a dynamic, non-compartmentalized tissue containing several critical immune niches defined by specific lineage architectures. We systematically benchmarked 16 contemporary domain segmentation algorithms, demonstrating that most methods in their default configurations fail to recapitulate biologically defined niche boundaries. Our analysis reveals that the definitive, disjoint spatial distributions of key functional lineages are frequently obscured by the stochastic infiltration of peripheral cell types. Such reduction in the spatial signal-to-noise ratio represents a primary bottleneck for existing algorithms, which prioritize local transcriptomic variance over global architectural logic. Following this observation, we demonstrate that strategic weighting of core functional lineages can restore the resolution of spatial niches in select domain segmentation frameworks. Cross-comparison against compartmentalized tissues further underscores the unique challenges of niche identification in non-mechanically separated environments and clarifies the fundamental divergence between structural domain segmentation and functional niche discovery. Our work delineates the limitations of current paradigms and advocates for the development of specialized computational approaches tailored specifically to the complexity of functional microenvironments.
bioinformatics2026-09-14v2SpaCoEx: Sparse Gene Selection for Spatially Varying Co-expression in Spatial Transcriptomics
Jian, M.; Pei, S.; Alterovitz, G.Abstract
Spatial transcriptomics enables gene expression to be measured while preserving tissue location, but most existing analyses focus on spatial variation in individual genes or expression-defined domains. Here, we introduce SpaCoEx, a sparse spatial representation framework that integrates gene-expression levels with spatially varying gene-gene co-expression. SpaCoEx first estimates local co-expression matrices from neighboring spatial spots, maps them into a log-Euclidean representation, and performs structured gene selection by retaining or removing the full row and column associated with each gene. The selected genes are then used to construct both expression-level features and local co-expression features, which are combined through an -weighted joint representation for downstream spatial analysis. We applied SpaCoEx to human cutaneous squamous cell carcinoma and annotated human breast cancer spatial transcriptomics datasets. In the cutaneous squamous cell carcinoma dataset, SpaCoEx selected 29 of 45 keratinocyte-related genes while preserving 96.74% of the spatial co-expression variation. In the breast cancer dataset, SpaCoEx identified spatially varying co-expression between B2M and HLA-C, a biologically meaningful major histocompatibility complex (MHC) class I antigen-presentation gene pair. Their local correlation was significantly higher in cancer-associated regions than in non-cancer regions (mean difference = 0.30, spatially adjusted SE = 0.045, P<1.0x10^(-10)), whereas B2M and HLA-C expression individually did not differ significantly between cancer and non-cancer. In benchmarking against manual tissue annotations, co-expression-only SpaCoEx achieved the strongest spatial coherence (percentage of abnormal spots [PAS] = 0.076), while the joint expression/co-expression representation achieved the highest annotation agreement, with an adjusted Rand Index (ARI) of 0.584 at = 0.60 and and normalized mutual information (NMI) of 0.663 at = 0.40. By integrating marginal gene-expression information with local gene-gene co-expression structure, SpaCoEx provides a sparse, low-dimensional, and interpretable representation of spatial transcriptomics data that captures complementary aspects of tissue organization beyond expression-based variation alone.
bioinformatics2026-09-14v2ForceFlowAb: physics-aware mixture-of-experts flow matching model for antibody CDRs design
Li, Z.; Lv, Z.; Zhang, G.Abstract
Abstract Motivation: Antibodies are a major class of therapeutic molecules, and their recognition of target antigens is largely mediated by complementarity-determining regions (CDRs), making antigen-conditioned CDR design a central problem in antibody engineering. Recent generative methods have enabled antigen-conditioned co-design of CDR sequences and structures, but their limited capacity to capture local interface heterogeneity and lack explicit energy-based guidance during sampling, which may result in unfavorable antibody-antigen interaction energies. Overcoming these limitations requires methods that better represent diverse interface environments via adaptive routing and incorporate physical guidance to steer sampling toward energetically favorable conformations. Results: We present ForceFlowAb, a physics-aware mixture-of-experts flow-matching framework for antigen-conditioned CDR sequence-structure co-design. The framework models heterogeneous interface environments through specialized expert routing and applies differentiable force-field guidance during sampling to guide CDR generation toward energetically favorable conformations. For CDR-H3 design, ForceFlowAb achieved more favorable antibody-antigen interaction energies than FlowDesign and Diffab, with improvement rates (IMP) of 46.5% versus 35.0% and 35.5%, respectively. For simultaneous six-CDR design, ForceFlowAb also outperformed Diffab, with IMP values of 16% versus 9%. These results suggest complementary roles for interface-adaptive modeling and energy-based guidance, with the former capturing binding-mode diversity and the latter leveraging physical constraints to ensure biophysical feasibility. Availability and implementation: The web server is freely available at http://zhanglab-bioinf.com/ForceFlowAb. The source code and implementation are available at https://github.com/iobio-zjut/ForceFlowAb. Contact: zgj@zjut.edu.cn Supplementary information: Supplementary data are available at Bioinformatics online.
bioinformatics2026-09-14v1LIGER2: Scalable Single-Cell Integration with On-Disk Datasets
Wang, Y.; Robbins, A.; Gadhvi, G.; Welch, J. D.Abstract
Correcting batch effects and integrating single-cell sequencing datasets has been a crucial step in large-scale biological studies. Many methods have been published for this task, and excel in various scenarios. Our previous work, LIGER, leveraging integrative non-negative matrix factorization (iNMF), stands out in providing an interpretable low-dimensional representation. To adapt to the modern need for integrating millions of cells, we developed a highly-optimized parallel factorization solution with on-demand loading from disk. The upgraded LIGER algorithm shows significant improvements in time and memory efficiency for single-cell data integration. We also developed a new downstream embedding alignment method significantly improved performance in conserving biological variation while still aligning corresponding cell types across datasets.
bioinformatics2026-09-14v1TransBind2: Improving Transcription Factor-DNA Binding Prediction with Multimodal Data and Bidirectional Cross Attention
Basnet, S.; Cheng, J.Abstract
Accurate genome-wide prediction of transcription factor (TF)-DNA binding remains challenging because many models focus mainly on DNA sequence and overlook chromatin context and TF structure. We previously developed TransBind, a protein-aware model that combines TF and DNA representations through cross-attention. Here, we introduce TransBind2, which improves on TransBind in several ways. It incorporates DNase-seq accessibility and genome mappability tracks as additional input, uses a biomodal protein language model (ProstT5) to capture both TF sequence and structure, and applies bidirectional cross-attention so DNA and protein features can refine each other. We also frame prediction as binary classification of individual <DNA bin, TF, cell type> triplets, allowing the model to generalize to new TFs and cell types. Across 690 human ChIP-seq experiments covering 161 TFs and 91 cell types, TransBind2 achieves a macro AUROC of 0.9648 and AUPR of 0.4215, outperforming TransBind and other baselines, with a [≥]12.67% relative AUPR gain. The model trained on human data also performs well in cross-species zero-shot prediction on mouse data. Saliency analysis shows that it can identify TF-binding peaks with a median error of 12-38 base pairs (bps) despite being trained on window-level labels. Ablation studies further show that TF structure, chromatin accessibility, and bidirectional attention each improve performance. Overall, these results show that combining TF structure with chromatin context leads to more accurate and generalizable TF-DNA binding predictions.
bioinformatics2026-09-14v1PHACTn enables training-free, context-independent inference of nucleotide variant tolerance across the genome
Yildirim, C.; Kuru, N.; Adebali, O.Abstract
Accurate prediction of single-nucleotide variant (SNV) tolerability across the entire human genome remains a fundamental challenge in computational genomics, particularly for non-coding regions where the regulatory landscape is vast and poorly understood. Machine learning classifiers suffer from data circularity and demographic bias, while genomic language models demand massive computational resources and offer little biological interpretability. Here, we present PHACTn (Phylogeny-Aware Computing of Tolerance for nucleotide variants), a training-free, parameter-minimal method that infers nucleotide variant tolerability by traversing the mammalian phylogenetic tree and explicitly modeling the evolutionary independence of observed substitutions and their distance from the query species. With only 4 interpretable parameters, no training and no GPU requirement, PHACTn outperforms all evaluated tools on non-coding variants curated from both the ClinVar, and on non-coding variants potentially responsible for selected Mendelian diseases curated from OMIM. Additionally, it achieves state-of-the-art performance on variants within the informative range of alignment-based inference. These results establish that principled probabilistic phylogenetic modeling captures evolutionary constraint signals that large-scale sequence models fail to recover, offering a powerful, accessible, and mechanistically transparent alternative for genome-wide variant effect prediction.
bioinformatics2026-09-14v1ImmuneLens: linking transcriptional states and TCR clonotypes through disentangled multimodal learning
Duan, Z.; Wang, Y.; Li, C.; Li, G.; Cao, Y.; Bai, X.; Yang, F.; Song, S.Abstract
Single-cell multi-omics technologies simultaneously capture the transcriptome and TCR sequence of T cells, providing an opportunity to study the relationship between transcriptional states and clonal architectures. However, jointly modeling the relationships between transcriptional states and TCR sequences while preserving modality-specific information remains challenging. Here, we present ImmuneLens, an interpretable multimodal representation learning framework designed for paired single-cell transcriptome and TCR sequence data. ImmuneLens supports the construction of a transferable multi-cohort immune reference atlas and enables unsupervised mapping of external query data. The complementarity between GEX and TCR information improves the stability of antigen-specificity prediction. In neoadjuvant immunotherapy cohorts, ImmuneLens resolves response-associated T cell heterogeneity and reveals links between clonal expansion and CD8 T cell functional states. Overall, ImmuneLens provides a
bioinformatics2026-09-14v1Heterogeneous graph neural networks with biological prior knowledge for interpretable drug repurposing in triple-negative breast cancer
Fernandez-Lozano, C.; Ferreiro, D.; V.-del-Rio, P.Abstract
Drug repurposing offers a cost-effective path to new therapies for triple-negative breast cancer (TNBC), a subtype with limited targeted treatment options. We present PRECISION, a framework integrating transcription factor (TF) regulatory networks, protein-protein interactions, and drug-target edges into a heterogeneous graph neural network (GNN) to identify TFs mediating drug sensitivity and prioritize repurposing candidates. The knowledge graph has 23,498 nodes and approximately 830,000 edges from CollecTRI, OmniPath and the PRISM screen. In a fair cell-line hold-out evaluation, the GNN matches ML baselines in global prediction (Pearson r = 0.76). Per-drug mechanistic attributions via Integrated Gradients on the trained GNN, replicated across three independent training seeds and robust to baseline choice (Spearman's rho = 0.94 between the mean and the random Gaussian baselines), highlight stress-response (CREB3L1), epithelial-mesenchymal transition (EMT; KLF8, ZEB1), stromal/TNBC-specific (AEBP1, MZF1), and epithelial (SPDEF) regulators as stable mediators of drug response. Multi-cohort validation in SCAN-B (n = 7,397), METABRIC (n = 1,979), and TCGA-BRCA (n = 1,072) shows that predicted drug sensitivity is associated with overall survival in 623 drugs in SCAN-B and 74 in METABRIC (FDR < 0.05). Fisher's meta-analysis identifies 551 drugs at FDR < 0.05, validated by positive controls paclitaxel (p_adj = 8.6x10-3), docetaxel (3.1x10-2), epirubicin (3.2x10-6), and talazoparib (1.4x10-4). Paired Wilcoxon tests in 39 AURORA-US patients with matched primary and metastatic samples confirm that 4 of the 10 IG-identified TFs (CREB3L1, KLF8, AEBP1, SPDEF) are significantly altered during metastatic progression after Bonferroni correction. An explicit rule applied to the PAM50-adjusted Cox multivariate results (penalizer = 0.1) selects seven candidates (osimertinib, saracatinib, erlotinib, brigatinib, pelitinib, entinostat, trametinib), revealing pharmacological convergence on the EGFR signaling axis.
bioinformatics2026-09-14v1Heterogeneous epigenetic regulatory patterns link mammalian aging, development, and mortality
Tikhonov, S.; Dmitriev, S. E.Abstract
Aging is often described as a monotonic accumulation of cellular damage, yet all-cause mortality follows a U-shaped trajectory with age, suggesting non-monotonic molecular changes. We investigated links between early childhood development, aging, and chronic diseases by analyzing DNA methylation in mammalian blood. A meta-analysis of 16 human chronic diseases revealed heterogeneous methylation signatures that formed 2 major disease clusters distinguished by their associations with development and sex-related methylation changes. Although epigenetic entropy increased monotonically across the lifespan, several diseases reduced blood DNA methylation entropy independently of blood cell composition. Across mammals, many CpG sites, particularly in intergenic regions, followed U-shaped age-related methylation changes that paralleled mortality curves. Based on these patterns, we developed epigenetic clocks that predict expected mortality across species and tissues and are effective in detecting a range of disease models. Overall, our findings reveal fundamental links between epigenetic regulation during development, aging, and chronic diseases.
bioinformatics2026-09-14v1Detection-Guided Beamforming for Efficient Bat Localisation
Alessandri, R.; Gilmour, L.; Tan, J. J.; Osses, A.; Stowell, D.Abstract
Passive acoustic monitoring is widely used to study wildlife, but current approaches provide limited insight into the spatial behaviour of animals. In bat ecology, reconstructing flight trajectories is essential for studying habitat use, movement patterns, and interactions, yet it remains difficult to achieve under field conditions. Acoustic cameras offer a potential solution by enabling sound source localisation, but their practical application is limited by the computational cost of beamforming and by the non-stationary, broadband, and transient nature of echolocation calls. In particular, exhaustive beamforming over wide ultrasonic bandwidths and dense spatial grids becomes infeasible for continuous monitoring. In this work, we propose a detection-guided beamforming framework for efficient localisation of free-flying bats. The method exploits the sparsity of echolocation signals by restricting beamforming to detector-identified time-frequency regions and combines this with physically consistent short-time analysis parameters and dense spatial sampling. The framework is evaluated on field recordings acquired with a Sorama CAM iV64s acoustic camera. Results show that detection-guided processing reduces the number of beamformer evaluations by approximately 77 times, corresponding to a reduction of about 98.7 % in computational effort, while preserving spatial resolution. At the same time, the proposed parameter configuration improves the stability and sharpness of reconstructed trajectories by avoiding artefacts associated with temporal averaging. These findings demonstrate that high-resolution acoustic localisation can be achieved under realistic computational constraints, supporting the integration of acoustic cameras into ecological monitoring workflows. The proposed framework enables the extraction of spatial trajectories from passive acoustic monitoring data, facilitating spatially resolved analyses of bat behaviour in field conditions.
bioinformatics2026-09-13v1Towards reconstruction of the human interactome from positive and negative experimental evidence
As, J.; Pelz, K.; Bernett, J.; Battini, F.; List, M.; Blumenthal, D. B.; Schaefer, M. H.Abstract
Protein-protein interactions (PPIs) have been detected and reported in the millions, but while they are used in many different contexts for better understanding cellular processes in health and disease, the knowledge of the human PPI network is far from complete, containing many false positive measurements and being highly biased. Both to chart the extent of those problems and to solve them requires not just a knowledge of high-confidence positive interactions, but also likely non-interacting protein pairs. However, this information is typically not reported in PPI studies. We developed a methodology to reconstruct this knowledge from existing PPI data. We reconstruct the experimental search space in which PPI screens have been performed and then create a model that informs how likely a PPI is real given its testing and observation frequency. We argue that negative protein pairs allow us to estimate the error rates of experimental and computational screens. We show how this knowledge could be incorporated for calibration. Finally, we evaluate a simple machine learning approach to PPI prediction and propose how such negative data can be used for training instead of random protein pairs. Together, our results show that reconstructing the experimental search space recovers a largely overlooked layer of information from existing PPI data that can help guide a more accurate and complete mapping of the human interactome.
bioinformatics2026-09-13v1OmniTCR: a foundation model unifying T cell receptor recognition prediction and conditional sequence generation
Zeng, F.; Feng, D.; Song, D.; Ding, L.; Tan, Z.; Lei, Q.; Lei, W.; Guo, A.-Y.Abstract
T cell receptor (TCR) recognition prediction and receptor generation are traditionally modelled separately, leaving vast TCR sequence collections disconnected from smaller TCR-peptide-MHC datasets. Here we present OmniTCR, a 113-million-parameter autoregressive foundation model pretrained on 328 million formatted human immune-sequence records. Sequence-type tokens and complementary component orders enable joint learning from individual TCR chains and partial or complete TCR-pMHC associations. On unseen epitopes, OmniTCR achieved AUPRCs of 0.7009 for peptide-TCR{beta}; recognition and 0.8235 for TCR-pMHC interaction prediction, exceeding the strongest evaluated comparators by 0.3396 and 0.3451, respectively. It distinguishes cancer from healthy repertoires across 11 independent pan-cancer cohorts (mean AUROC, 0.9436). The model achieved the highest sequence recovery on internal and external generation benchmarks. Structural modelling supported the plausibility of selected pMHC-conditioned CDR3{beta}; candidates. OmniTCR bridges heterogeneous immune sequence data, providing a foundation for computational immunology and receptor design.
bioinformatics2026-09-13v1pydreg: a fast Python package for identifying active cis-regulatory elements from nascent transcription
He, A. Y.; Danko, C. G.Abstract
Background: Active promoters and enhancers generate characteristic patterns of RNA transcription that can be measured through nascent RNA sequencing. dREG is a leading method that uses these patterns to identify active cis-regulatory elements across the genome, allowing regulatory activity and gene transcription to be profiled in the same experiment. However, its reference implementation was developed around an R-based workflow and a legacy GPU-accelerated support vector machine library that have become increasingly difficult to maintain and deploy. Findings: To improve future usability of dREG, we developed pydreg, a Python port of dREG. pydreg preserves the original pretrained models and peak calling procedure from dREG while using contemporary numerical libraries for CPU and GPU computation. pydreg achieves 4.5 and 5.4-fold reductions in runtime and peak host memory, respectively, compared to dREG while producing near identical peak calls. Conclusions: pydreg reduces practical barriers to running dREG locally, improves runtime and memory usage, integrates readily with Python-based genomics workflows, and provides a maintainable foundation on modern computing infrastructure. Availability and Implementation: pydreg is implemented in Python 3.11+ and is freely available under the GPL-3 license at https://github.com/adamyhe/pydreg and from PyPI via pip install pydreg[gpu] (for CUDA acceleration) or pip install pydreg[mlx] (for Apple Metal acceleration).
bioinformatics2026-09-13v1Benchmarking Twist Genotyping-by-Sequencing Against Whole-Genome Sequencing in Nuclear Families
Klugerman, J.; Iossifov, I.; Ye, K.Abstract
Genome-wide genotyping is widely used in human genetics research, and targeted sequencing-based approaches such as the Twist Bioscience genome-wide SNP capture platform (GxS) have emerged as alternatives to conventional SNP arrays. Here, we evaluated GxS genotype calls from 555 individuals in 184 nuclear families against matched whole-genome sequencing (WGS) calls and compared platform performance with that of the Illumina Infinium Global Screening Array-24 (GSA), which was evaluated in 987 individuals from 279 nuclear families. Genotype data were harmonized across platforms, and analyses were restricted to overlapping SNP loci. Across all callable positions, mean per-SNP call rates were 98.26% for GxS and 98.67% for GSA. Overall SNP concordance with WGS was 99.79% for GxS and 99.87% for GSA, and mean per-individual concordance was also 99.79% and 99.87%, respectively. Per-trio Mendelian violation rates of GxS are about 10 times those of WGS, while those of GSA are about 4 times those of WGS on average. These results indicate that GxS performs slightly worse than GSA by key concordance and inheritance metrics, while still showing strong overall agreement with WGS.
bioinformatics2026-09-12v3CasanovoGUI: a cross-platform desktop application for deep learning-based de novo peptide sequencing with Casanovo
Wen, B.; Li, K.; Riffle, M.; MacCoss, M. J.; Bittremieux, W.; Noble, W. S.Abstract
De novo peptide sequencing detects peptides directly from tandem mass spectra without a protein sequence database, and deep learning has substantially advanced its performance. Casanovo, one such widely used model, is distributed as a Python command-line program. Consequently, installation, GPU and dependency configuration, and manual parameterization can be challenging for many bench scientists and are a recurring source of errors. Interpreting and validating the resulting predictions poses a further challenge. We present CasanovoGUI, an open-source Java-based desktop application that makes all of Casanovo's main analysis functions available through a point-and-click interface on Windows, macOS, and Linux. On first use, CasanovoGUI automatically installs a private Python environment and Casanovo with a GPU-matched build, requiring no prior software setup. The GUI provides access to Casanovo's analysis functions and configuration parameters, streams live progress, and integrates results interpretation: annotated spectra with per-residue confidence scores in the PDV viewer, and mismatch-tolerant mapping of de novo peptides back to a reference proteome. CasanovoGUI is available at https://github.com/Noble-Lab/CasanovoGUI.
bioinformatics2026-09-12v2Leakage-Safe Cross-Cohort Transcriptomic Survival Analysis in Pancreatic Ductal Adenocarcinoma: A Corrected Reanalysis
Markarian, M. B.; Houssigian, G. V.Abstract
Background: Transcriptomic prognostic signatures in pancreatic ductal adenocarcinoma (PDAC) are vulnerable to endpoint misspecification, heterogeneous specimen inclusion, platform effects and information leakage. We re-audited and replaced an earlier batch-harmonized alive/dead classifier. Methods: GSE71729 was audited for assay and composition but excluded from prognosis because its GEO Series Matrix does not provide an auditable patient-level survival join. Development used 176 TCGA-PAAD primary tumors with 92 deaths. A penalized Cox model was evaluated by repeated nested 5-fold cross-validation with filtering, scaling, outcome-guided screening and tuning learned inside each training partition. Cross-platform transport was tested without joint ComBat by holding out 79 survival-labelled primary tumors from GSE85916 (57 deaths), using a per-sample percentile-rank representation. Discrimination was summarized by Harrell's concordance index (C-index) with patient-bootstrap 95% confidence intervals (CIs). The five genes highlighted in version 1 were re-examined by cohort-specific Cox models and random-effects meta-analysis. Results: Internal TCGA C-index was 0.630 (95% CI 0.567-0.694). TCGA-to-GSE85916 C-index was 0.527 (0.435-0.607); reverse-direction transport was 0.576 (0.508-0.647). A fixed five-gene Cox model trained in TCGA gave external C-index 0.569 (0.476-0.657). None of the five random-effects associations survived Benjamini-Hochberg correction across the panel; heterogeneity was high (I2 74%-93%). Conclusions: The corrected data support exploratory transcriptomic signal within TCGA but not a validated transportable prognostic model or validated five-gene biomarker panel. Cohort mixing after batch correction is not evidence of biological preservation or external validity. The revised open workflow provides a reproducible basis for larger, prespecified multi-cohort studies.
bioinformatics2026-09-12v2CRISPR-enhanced assessment of variants of unknown significance nominates oncology therapeutic targets and drug repositioning opportunities
Savino, A.; Oikonomou, A.; Perrone, F.; De Lucia, R. R.; De Pietri, L.; Belattar, Y.; Brown, L.; Grau, M. L.; McCarten, K.; Najgebauer, H.; Perron, U.; Azzolin, L.; Livanova, A.; Cremaschi, P.; Lopez-Bigas, N.; Sottoriva, A.; Coelho, M. A.; IORIO, F.Abstract
Interpreting infrequent somatic variants remains a challenge in cancer genomics. We developed CRISPR-VUS, a framework that uses public Cancer Dependency Map data to identify Dependency-Associated Mutations (DAMs) - variants linked to increased host-gene dependency - with resolution extending to singleton events. Analysis of 977 cell lines across 36 cancer types identified 2,376 DAMs in 1,383 genes, including 1,260 not established as cancer drivers. DAM-bearing genes converge on oncogenic networks, while recurrence in histology-matched tumours, functional-impact predictions, tractability and pharmacological associations enable systematic prioritisation. Prime editing showed that the prioritised NSCLC-specific RTN4IP1-A80T DAM conferred a significant competitive growth advantage in a lung epithelial model, nominating a candidate driver allele. Exploratory pharmacological testing showed a greater maximal response to istaroxime in ATP1B3-I189M-bearing than in ATP1B3-wild-type cells. CRISPR-VUS combines dependency-based rare-variant discovery with evidence-guided prioritisation to nominate candidate drivers, therapeutic targets and drug-repositioning hypotheses. Interactive results are available at https://vus-portal.fht.org/.
bioinformatics2026-09-12v2DNAharvester: A Nextflow Pipeline for Analysing Highly Degraded DNA from Ancient and Historical Specimens
Sharif, B.; Kutschera, V. E.; Oskolkov, N.; Guinet, B.; Lord, E.; Chacon-Duque, J. C.; Oppenheimer, J.; van der Valk, T.; Diez-del-Molino, D.; D. Heintzman, P.; Dalen, L.Abstract
Ancient DNA (aDNA) research has advanced rapidly with the development of high-throughput sequencing, enabling genome-wide analyses of large collections of prehistoric specimens. However, analysing palaeontological and archaeological material with highly degraded DNA constitutes a major bioinformatic challenge. DNA from such samples is characterised by short fragment lengths, low endogenous content, post-mortem damage, and cross-species contamination, which can increase spurious mapping and reference bias, affecting downstream population genetic inferences. We present DNAharvester, a modular and reproducible pipeline designed specifically for processing highly degraded DNA from ancient and historical specimens. DNAharvester integrates metagenomic filtering, competitive mapping, adaptive aligner selection (incorporating BWA-aln, BWA-mem, and Bowtie2), and systematic evaluation of reference bias and spurious mapping. By incorporating flexible mapping and filtering strategies, the pipeline can be adapted to varying sample preservation, focusing on maximising authentic data recovery. DNAharvester features subworkflows for iterative assembly of mitogenomes, identification of genomic repeats and CpG sites, taxonomic classification, microbial/pathogen screening, genetic sex determination, and variant calling. To accommodate varying sequencing depths, the pipeline supports diploid variant calling, genotype likelihood estimation, and pseudo-haploid random allele calling. Implemented in Nextflow, DNAharvester provides a highly scalable, containerised framework that enhances reproducibility, portability, and robustness in aDNA analyses. We validated the pipeline using simulated and empirical datasets, demonstrating its ability to systematically mitigate complex background contamination while preserving authentic genomic signals. By streamlining complex bioinformatic tasks through simple configuration files, DNAharvester establishes a standardised approach for analysing aDNA datasets and makes genomic analyses of ancient remains accessible to the broader research community.
bioinformatics2026-09-12v2VARION: A Network Propagation Framework for Individual Patient Somatic Mutation Interpretation in Cancer Molecular Subtyping
Kwon, T.; Park, Y.-G.; Choi, J.-G.Abstract
Accurate molecular subtyping of individual cancer patients from somatic mutation data remains a challenge in precision oncology research. Existing network-based stratification (NBS) methods treat all mutations equivalently, require full-cohort batch processing, and do not demonstrate generalization to independent datasets without retraining. To address this, we present variant interpretation via the adaptive network pRopagatION (VARION), which integrates population-level variant constraint scoring with protein-protein interaction (PPI) network topology. The Adaptive Topology-aware Random Walk with Restart (ATR-RWR) algorithm weights each mutated gene by {varphi}g = {surd}(GIS(g) x {rho}topo(g)), where GIS (Gene Intolerance Score) reflects population-level functional constraint, propagated across a shared PPI network; subtype assignment then uses cosine similarity to TCGA-derived reference centroids, enabling real-time single-patient classification. Across ten TCGA cancer cohorts (n = 2,417), VARION achieved 77.7% accuracy for ovarian cancer (OV), 69.5% for glioblastoma (GBM), 90.2% for cholangiocarcinoma (CHOL), and 75.4% for gastric cancer (STAD). A controlled benchmark applying two alternative clustering methods (PyNBS; a dense autoencoder) to identical ATR-RWR propagation matrices recovered no significant driver enrichment (OR = 1.79 and 1.52, n.s.), versus OR = 144.29 (p = 1.77x10^-12) for VARION, confirming that the GIS-weighted centroid architecture, not propagation alone, drives performance; generalization without retraining was further confirmed in two independent cohorts (ICGC CCA, n = 396; PCAWG, n = 110; OR = {infty}, p < 5x10^-9). Together, these results indicate that VARION's GIS-weighted centroid architecture enables individual-patient molecular subtyping that outperforms existing NBS and graph-learning clustering approaches, with high sensitivity for clinically actionable rare subtypes and robust cross-platform generalization.
bioinformatics2026-09-12v1FlashDeconv reveals resolution horizons in atlas-scale spatial transcriptomics
Yang, C.; Chen, J.; Zhang, X.Abstract
Coarsening Visium HD resolution from 8 to 64 m can flip cell-type co-localization from negative to positive (r = -0.12 [->] +0.80), yet many widely used compositional deconvolution workflows require coarsening or subsampling at million-bin scale. Here we introduce FlashDeconv, which combines leverage-score importance sampling with sparse spatial regularization to achieve competitive benchmark accuracy while processing 1.6 million bins in 153 seconds on commodity hardware. Systematic multi-resolution analysis of Visium HD mouse intestine reveals a tissue-specific resolution horizon (8-16 m), the scale at which this sign inversion occurs, validated by Xenium ground truth. Below this horizon, FlashDeconv provides, to our knowledge, the first sequencing-based quantification of Tuft cell chemosensory niches (15.3-fold stem cell enrichment). In a 1.6-million-bin human colorectal cancer cohort, FlashDeconv uncovers neutrophil inflammatory microdomains co-localized with immunoregulatory dendritic cells (mRegDC) at the tumor-stroma interface, spatial niches largely missed by discrete-label summaries, with RCTD doublet mode labeling only 2.3% of hotspot bins as neutrophil singlets.
bioinformatics2026-09-11v5Fast and accurate taxonomic domain assignment of short metagenomic reads using BBERT
Alekhin, D.; Alon, M.; Sidi, T.; Perez Mazeh, S.; Carmi, G.; Finkel, O. M.; Erez, A.Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses.
bioinformatics2026-09-11v4Histology-Aware Graph for Modeling Intercellular Communication in Spatial Transcriptomics
Wang, X.; Tao, C.; Jiang, Y.; Jiang, Y.; Liu, H.; Jiang, Z.; Zhu, P.; Que, N.; Xi, J.; Price, S.; Mou, Y.; Xu, J.; Li, C.Abstract
Cell-cell communication (CCC) is essential to how life forms and functions. Recent tools achieve single-cell-resolved CCC inference utilizing spatial transcriptomics (ST). However, most ignore the modeling of tissue contexts surrounding cells, causing high false-positive/negative rates. Here, we propose HARMONIC, a CCC inference method integrating multimodal ST and hematoxylin and eosin (H&E)-stained images. HARMONIC causally modeling the transcriptomic-to-contextual relationships for CCC inference. The state-of-the-art performance was verified across ST platforms, species and healthy/diseased status, on both synthetic and biological samples. HARMONIC was applied in various real-world scenarios, especially on tissues with clear morphological boundaries, including cortical layers in mouse brain, medullary-cortex structures in mouse kidney, as well as tumor-stromal/immune interface. Significant refinement of false-positive/negative predictions was observed compared to ST-only CCC tools.
bioinformatics2026-09-11v2Distinct geometries, comparable interfaces: binding modes and thermodynamic implications in conventional and single domain antibodies
Hauser, A.; Dangla-Pelissier, G.; Cazals, F.Abstract
Heavy-chain only antibodies, produced by the adaptive immune systems of camelids and cartilaginous fish, complement canonical antibodies comprising both heavy and light chain variable domains. Using an integrated interface model that unifies interfacial atoms--including solvent molecules, contacts, buried surface areas, and interface curvature measures, we shed light on two aspects of antibody binding that have so far remained elusive when comparing single domain (SdAb) and double domain (DdAb) antibodies. First, contrary to previous reports of smaller SdAb interfaces, we show that SdAb achieve an interface size comparable to that of DdAb despite using a single variable domain and fewer interface residues, a consequence of a geometric pattern driven by convexity and curvature effects. Second, we show that SdAb exhibit a broader diversity of binding modes than previously reported, with a prominent role played by FR regions. Finally, we discuss the thermodynamic implications of these findings for the design of high-affinity single domain antibodies, with particular relevance to protein engineering and design.
bioinformatics2026-09-11v2SenSASP: A Unified, Multi-Layer Database of Senescence and SASP Genes
Xuan, H.; Huang, Y.; Bian, J.Abstract
Research on cellular senescence and the senescence-associated secretory phenotype (SASP) draws on independently curated gene resources that differ in scope, identifiers, and update cycles, making cross-resource integration error-prone. We unified four widely used resources, CellAge, GenAge, the SenMayo signature, and the Reactome Cellular Senescence pathway, onto a single canonical identifier (the Ensembl gene ID) and enriched every gene with three annotation layers absent from all four inputs: cross-species conservation, tissue and cell-type expression, and high-confidence protein-protein interactions. Unification collapsed 1,460 summed source entries into 1,250 unique genes (210 redundant entries removed, 14.4%) while preserving full source provenance: 173 genes are corroborated by two or more resources and two (IL6, JUN) by all four. The three annotation layers reach 95.8%, 97.9%, and 93.0% of genes, with 89.4% annotated across all three. A 500-gene random sample of identifier mappings was validated against HGNC and Ensembl (98.0% exact match). The result, SenSASP, is a single, machine-readable, provenance-tracked database of harmonized identifiers and net-new functional context, illustrated here with a gene-prioritization score and a tissue-expression atlas. SenSASP is freely available at https://xuan13hao.github.io/sensasp/
bioinformatics2026-09-11v2scPyviewer: a Python-native interactive viewer from AnnData single-cell data
Xuan, H.; Huang, Y.; Bian, J.; Liu, X.Abstract
Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.
bioinformatics2026-09-11v2Kintsugi decides, gene by gene, where spatial transcriptomics borrows information
Yang, C.; Zhang, X.; Chen, J.Abstract
Subcellular spatial transcriptomics captures where RNA is in tissue, but a single location holds too few molecules of any one gene to estimate composition alone. Every current method fixes in advance where to borrow -- a smoothing scale, a cell outline or a factor model -- and the fixed choice shapes what is visible. Kintsugi removes the fixed choice and lets held-out molecules decide, gene by gene, how much to borrow from spatial neighbours and from other genes at the same location. On a lung section measured by both Xenium and Visium HD, the data-chosen allocation placed an epithelial programme where the Xenium molecules were, ahead of smoothing, cell segmentation and a factor model; the result replicated across tissues and against protein. Across a 45-core pulmonary fibrosis cohort, separating composition from captured amount shows that a fibroblastic focus is not a place with more RNA but a place with different RNA: 2.8-fold higher in activated-fibroblast composition while segmented nuclear density is at most 1.08-fold higher.
bioinformatics2026-09-11v2scOLAR: Ontology-Anchored Open-Set Annotation of Single-Cell RNA-seq Data
Liu, Y.; Yi, S.; Yin, H.; Ju, W.Abstract
Single-cell RNA sequencing profiles cellular heterogeneity at atlas scale, making automated annotation essential. However, target datasets often contain novel cell types missing from incomplete references. We present scOLAR, an ontology-guided open-set framework that learns prototypes over the Cell Ontology and uses both reference and target expression to annotate known classes while detecting unfamiliar populations. Guided by ontology hierarchies and decision-boundary regularization, scOLAR penalizes coarse-lineage misclassification and groups novel cells without requiring predefined cluster counts. Across benchmarks, scOLAR achieves a novelty-detection AUROC of 0.9726 and an average precision of 0.9871, enabling structured post-hoc lineage-level interpretation of populations absent from the reference.
bioinformatics2026-09-11v1