Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Fast and accurate taxonomic domain assignment of short metagenomic reads using BBERT
Alekhin, D.; Alon, M.; Sidi, T.; Perez Mazeh, S.; Carmi, G.; Finkel, O. M.; Erez, A.Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses.
bioinformatics2026-07-23v3GraTools, an user-friendly tool for exploring and manipulating pangenome variation graphs
Ravel, S.; Marthe, N.; Carrette, C.; Mohamed, M.; Sabot, F.; Tranchant-Dubreuil, C.Abstract
Background: Pangenome variation graphs (PVGs), which represent genomic diversity through multiple genomes alignment, are powerful tools for studying genomic variations in populations. However, current tools often lack integration, efficiency, or require format conversions, to use them, hindering their usability. Results: Here, we introduce GraTools, a fast and user-friendly command-line tool for manipulating PVGs using the original GFA file. After a one-time graph import, GraTools enables rapid subgraph extraction, fasta sequence retrieval, and comprehensive analyses, including core/dispensable genome ratio calculation or group-specific segment identification. The import step results in conversion in standard data formats (BAM/BED), enabling the reuse of well-optimized existing tools, allowing an efficient storage and the querying of the PVGs large complex data structures. Scalability is ensured by a modular architecture sup- porting parallel processing and asynchronous I/O operations. GraTools supports coordinates defined on both the primary reference as well as from alternative genomes within the graph without re-importing, and its outputs can be easily visualized or manipulated using external tools. Using an Asian rice pangenome graph (13 accessions), we demonstrate its ability to easily extract subgraphs, compute depth statistics, and identify subspecies-specific segments. An intuitive command-line interface, a real-time execution feedback and a detailed logging system make this tool suitable for a wide range of applications, from population genetics to breeding and genomic medicine, for both biologists and bioinformati- cians. Conclusions: Through its unified graph manipulation interface, GraTools offers an interesting alternative to the few existing tools for manipulating PVGs, facil- itating rapid, efficient and flexible downstream analyses. It is available as an open-source tool (GNU GPLv3), with its documentation available at https: //gratools.readthedocs.io.
bioinformatics2026-07-23v3GeneKnow: An auditable AI framework for source-grounded biological evidence synthesis
Zhang, H.; Sittipongpittaya, B.; Zang, C.Abstract
Biological function is context dependent, yet synthesizing evidence for a gene's function in a defined cell type, disease, or perturbation remains labor-intensive. Here we present GeneKnow, an auditable artificial intelligence (AI) framework for source-grounded biological evidence synthesis. GeneKnow separates evidence-critical operations, including literature retrieval, passage selection, provenance tracking, and bibliography construction, from generative steps that are constrained to semantic analysis and literature synthesis. GeneKnow supports multi-paper discovery and single-paper inspection while preserving links from synthesized claims to source passages, generating trustworthy syntheses without fabricated citations and minimizing hallucinations. Systematic benchmarking showed that GeneKnow achieved higher claim support and citation fidelity than leading general-purpose and scientific AI systems. These results demonstrate that an AI system with controlled division of labor between deterministic and generative components can substantially improve the fidelity and auditability of biomedical literature synthesis.
bioinformatics2026-07-23v2Hobrac: a reference-guided workflow for genome comparison and synteny visualization
Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.Abstract
Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO rthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.
bioinformatics2026-07-23v2Statistical tests for bivariate spatial association across multi-omics data with disjoint coordinates
Hawinkel, S.; Hu, W.; Velten, B.; Maere, S.Abstract
Spatial biology has entered a new era of multimodal profiling, with multiple, high-dimensional spatial omics types being measured on consecutive tissue slices, or co-assayed on the same slice. Interest then lies in statistical testing for spatial association between the features of the different modalities, to gain insight in biological processes. One major challenge is the multitude of bivariate combinations, leading to high computational demands. Another difficulty is the difference in spatial resolution between technologies, implying no one-to-one matching between the measurement spots of the different modalities, even after alignment. As a result, common statistical measures such as joint distributions and correlations are not defined, and tests need to rely on spatial vicinity only. Moreover, we argue that many existing bivariate association tests address an inappropriate null hypothesis, or make inappropriate assumptions, both implying absence of spatial autocorrelation in any of the features and leading to misleading conclusions. As a remedy, we modify tests for the detection of spatially variable genes (Moran's I, Gaussian processes and generalized additive models) to derive bivariate spatial association tests across modalities with non-overlapping coordinate sets, and provide variance estimators that do account for spatial autocorrelation. We develop inference methods for single sections as well as for replicated experiments with multiple sections, and compare their performance in nonparametric and parametric simulations. Finally, we apply the newly developed methods to two co-assayed spatial transcriptomics and metabolomics datasets from mouse and human. The full suite of tests is available from github.com/sthawinke/sbivar as the R-package sbivar.
bioinformatics2026-07-23v2GenoSim: A Forward-Time Genotype Simulator for Clinical and Population Genetics with Population Stratification
Bakar, A.; Gul, R.; Haq, W. u.; Afghani, T.Abstract
Motivation: Next-generation sequencing studies in clinical genetics are often limited by the scarcity of human genotype data, which stems from ethical, regulatory, and economic barriers. The shortfall is sharpest in consanguineous populations, which are common in South Asia and the Middle East, where family-based designs need large pedigrees that are rarely sequenced in full. Existing simulators do not combine pedigree-aware propagation, realistic population stratification, and clinical export formats in one tool. Results: We present GenoSim, an R package for forward-time simulation of diploid SNP genotypes. It runs in two modes: a population mode implementing inbreeding-adjusted Hardy-Weinberg sampling, Wright-Fisher drift, directional selection, recurrent mutation, and Haldane recombination across multiple generations; and a pedigree-constrained mode that ingests real family VCFs and a pedigree, reconstructs phase where the pedigree makes it identifiable, propagates genotypes through the observed family structure, and appends synthetic generations. Version 1.1.1 adds population stratification through the Balding-Nichols model parameterised by gnomAD v3.1 fixation indices (F_ST) for eight ancestry groups (AFR, AMR, EAS, EUR, FIN, MID, SAS, ASJ), empirical allele-frequency loading from external reference panels, and admixed-cohort simulation. Analysis functions cover Hardy-Weinberg testing, linkage disequilibrium, runs of homozygosity, principal component analysis, founder-referenced and between-generation F-statistics, and Nei gene diversity. Availability and implementation: GenoSim is available as an R package at https://github.com/malikbak/GenoSim under the MIT licence. It requires R [≥] 4.0.0 and depends only on base R packages (stats, utils, graphics, grDevices, tools).
bioinformatics2026-07-23v2EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-07-23v2A pan-cancer analysis of microRNA tissue specificity and its association with dysregulation
Poptsova, M.; Ismailov, A.; Belogurov, A.; Evpak, A.Abstract
MicroRNAs are frequently dysregulated in cancer, yet how their tissue-specificity is remodeled during malignant transformation remains poorly characterized. Here we systematically quantified the tissue-specificity of miRNAs across normal (GTEx) and tumor (TCGA) tissues using the Tau index, and compared its distribution between healthy and cancerous states. To robustly define dysregulation, we combined two independent analyses: a binomial test over per-project differential expression across 17 matched normal tissues within TCGA cohort, and a TCGA-GTEx pan-tissue comparison of mean expression. The change in specificity ({Delta}Tau) separated up- from down-regulated miRNAs, showing moderate agreement with the binomial signal and a strong correlation with the expression-based contrast. Finally, we identified 6 miRNAs that lose tissue-specificity upon transformation while remaining consistently upregulated (miR-519a-5p, miR-512-3p, miR-522-3p, miR-105-5p, miR-935, miR-1269a). Functional analysis of experimentally validated targets showed significant enrichment for converging on core oncogenic programs for miR-512-3p, miR-105-5p and miR-935, such as apoptosis and cellular-stress regulation, TP53, FoxO, PI3K-Akt/mTOR signaling, immune modulation. Collectively, integrating specificity dynamics with dysregulation evidence pinpoints candidate miRNAs with coordinated, cancer-relevant regulatory roles and highlights those with favorable tissue specificity profiles for therapeutic targeting.
bioinformatics2026-07-23v2Integrated Bioinformatics Analysis of PALB2 Reveals Expression Patterns, Molecular Interactions, and Prognostic Significance in Breast Cancer
Bithi, A. J.; Rahat, M. H.Abstract
Background: Partner and Localizer of BRCA2 (PALB2) is a key tumor suppressor gene involved in homologous recombination mediated DNA repair through its interactions with BRCA1 and BRCA2. Germline alterations in PALB2 have been associated with hereditary breast cancer risk; however, its broader molecular role in breast cancer progression and prognosis requires further investigation. Methods: A comprehensive in silico analysis of PALB2 was performed using publicly available databases and bioinformatics platforms. Differential expression of PALB2 in breast cancer were evaluated using GEPIA2. Prognostic significance was assessed through Kaplan Meier analyses for overall survival (OS) and disease free survival (DFS). Protein protein interaction (PPI) networks were constructed using STRING. Functional enrichment analyses of PALB2-associated genes were conducted using g. Mutational profiling of PALB2 in breast cancer was performed using cBioPortal with data from TCGA breast cancer cohorts. Results: PALB2 expression was elevated in breast tumor tissues compared with normal breast tissues. Survival analyses revealed no statistically significant association between PALB2 expression and either overall survival (HR = 0.88, p = 0.44) or disease free survival (HR = 0.74, p = 0.11). Protein interaction analysis revealed strong interactions between PALB2 and major DNA repair proteins including BRCA1, BRCA2, RAD51, RAD51C, FANCD2, and BRIP1. Functional enrichment analysis showed limited significant pathway enrichment, with only marginal transcription factor motif enrichment observed. Mutational analysis demonstrated diverse genomic alterations including missense mutations, truncating mutations, copy number gains, and shallow deletions. Conclusion: The findings support the biological relevance of PALB2 in breast cancer through its elevated expression and strong connectivity within DNA repair pathways. However, PALB2 expression alone does not appear to serve as an independent prognostic indicator. Further studies integrating genomic, transcriptomic, and clinical parameters are required to clarify its role in breast cancer progression and therapeutic response. Keywords: PALB2; Breast Cancer; Bioinformatics; Gene Expression Analysis; Protein Protein Interaction; Survival Analysis; Mutation Profiling; Homologous Recombination; TCGA; GEPIA2
bioinformatics2026-07-23v1PhytoFam: A Nextflow Pipeline for Genome-Wide Analysis of Plant Gene Families
Parajuli, S.; Adhikari, B.; Fennell, A.; Nepal, M. P.Abstract
Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complicates reproducibility and parameter tracking, and may require additional steps to remove redundant protein isoforms. To address these challenges, we developed PhytoFam, a Nextflow-based workflow that automates gene family identification from proteome input through phylogenetic reconstruction. The pipeline integrates HMMER for candidate sequence identification, isoform-aware deduplication, InterProScan for domain confirmation, BLAST reciprocal best hit (RBH) analysis for orthology assignment, MUSCLE for multiple sequence alignment with optional outgroup incorporation, TrimAl for alignment trimming, and IQ-TREE3 for phylogenetic reconstruction. PhytoFam is portable across local workstations and high-performance computing environments and supports deployment through Conda, Docker, and Singularity. We validated the workflow using the Morus alba MADS-box gene family, where the complete analysis finished in 1 h 10 min (9 CPU h). IQ-TREE3 accounted for most of the execution time, whereas InterProScan showed the highest memory requirement with a peak resident set size of 4.5 GB. PhytoFam provides a reproducible, automated, and scalable solution for plant gene family identification and phylogenetic analysis. The pipeline is freely available at https://github.com/sanamparajuli/PhytoFam.
bioinformatics2026-07-23v1A transparent multicriteria and fuzzy classification approach for genome-based probiotic candidate prioritisation
Ounissi, N. E.; Gomri, M. A.; El Hadef El Okki, M.Abstract
The identification of novel probiotic candidates with potential health-promoting properties remains a major challenge in food biotechnology and increasingly relies on in silico screening of genomic information. However, probiogenomic markers are heterogeneous, and safety, survival-colonisation, and functional-benefit traits do not contribute equally to probiotic potential. This study developed the Structured Probiotic Potential Index (SPPI), a fuzzy multicriteria system for genome-based probiotic candidate prioritisation. A hierarchical evaluation structure was established from probiogenomic evidence and organised into three main pillars and fourteen subcriteria. Expert judgements were collected using the Analytic Hierarchy Process, followed by consistency-based curation and weight aggregation. A curated dataset of 48 complete bacterial genomes, distributed into probiotic, potentially probiotic, neutral, and pathogenic groups, was taxonomically validated and analysed using an automated probiogenomic screening pipeline. Genome-wide screening generated 3,218 binary genomic features, from which curated probiogenomic markers were mapped to the scoring hierarchy. The resulting index was formulated as a normalised expert-weighted equation integrating safety, survival-colonisation, and functional-benefit components. SPPI prioritised genomes according to weighted probiogenomic profiles and separated pathogenic genomes from favourable probiotic and potentially probiotic profiles within the analysed dataset. Downstream Fuzzy Comprehensive Evaluation transformed the continuous score into probiotic/potentially probiotic, neutral, and pathogenic classes. Under internal leave-one-out evaluation, all 48 genomes were assigned to their expected reference classes, while confidence analysis distinguished high-confidence from borderline assignments. Post-classification comparison with ProbML showed concordant behaviour for most genomes and discordant predictions for selected cases. These results support SPPI as a transparent genome-based decision-support system for early probiotic candidate prioritisation before experimental validation.
bioinformatics2026-07-23v1An openly licensed benchmark and per-gene calibration map for missense pathogenicity predictors on activating cancer drivers
Lee, S.-G.Abstract
Missense pathogenicity predictors such as AlphaMissense are increasingly used in clinical variant interpretation, yet they are trained on germline labels dominated by loss-of-function (LOF) variants. Using an openly licensed, reproducible benchmark of 768 Cancer Gene Census genes scored with 49 predictors (labels from CIViC, COSMIC, cancerhotspots, ClinVar and gnomAD), we show that 42 of 49 tools (86%) score oncogene, gain-of-function (GOF) variants worse than tumour-suppressor variants. This under-scoring is mechanistically characterized: missed drivers occupy low-conservation, solvent-exposed, non-destabilizing positions (phyloP 2.51 versus 7.89; relative solvent accessibility 0.671 versus 0.185; gene-clustered p = 4.8x10-20 and 2.3x10-35), and, counter-intuitively, the unsupervised and protein-language models now entering clinical use are the most affected. Per-gene oncogenic thresholds span 0.07-0.99, so a single global cut-off is mis-calibrated for most genes; we provide a per-gene calibration map. A cancer-calibrated stack (OncoCal) modestly improves discrimination over the best single tool (AUROC {approx} 0.93 versus 0.87), rescues drivers such as JAK2 V617F (0.334[->]0.57), and generalizes to independent deep mutational scanning data. We provide an openly licensed framework to interpret and recalibrate these tools in the somatic setting rather than a replacement predictor.
bioinformatics2026-07-23v1pHaseMD4AI: Phase-Space Dynamics Dataset with Chemical and pH Perturbations for Physically and Kinetically Consistent Biomolecular AI
Song, T.; Guo, Y.; He, J.; Liu, Z.; Low, M.; Wang, K.; Zhang, Y.; Li, Z.; Huang, Y.; Wang, Y.Abstract
Protein function emerges from dynamic conformational ensembles and transitions that are challenging to characterize experimentally and computationally. Recent advances in generative AI have created new opportunities for learning molecular thermodynamics, kinetics, and conformational evolution directly from simulation data, but progress is limited by the availability of large-scale datasets that combine rigorous sampling, complete phase-space information, and diverse physicochemical perturbations. Here, we present pHaseMD4AI, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins. The peptide branch includes a complete set of canonical tripeptide and tetrapeptide systems together with post-translationally modified (PTM) and protonation-state datasets, providing synchronized atomic coordinates (R), velocities (V), forces (F), and Markov state model-based kinetic annotations. An accompanying web portal (https://isb.zju.edu.cn/md4ai/) enables users to browse, visualize, and download trajectories, annotations, and metadata. As an example application, we demonstrate a sequence-based model that can predict residue-level equilibrium dihedral distributions from sequence. pHaseMD4AI provides a resource for developing and benchmarking molecular machine learning methods while supporting broader studies of biomolecular dynamics under sequence, post-translational modification, and protonation-state perturbations.
bioinformatics2026-07-23v1scSAID: A Comprehensive Cross-Species Single-Cell Skin Atlas Reveals Species-Specific Responses to Psoriasis
Ren, Y.; Shen, Y.; Jin, L.; Huang, Y.; Deng, Y.; Xiao, Y.; Wang, C.Abstract
Single-cell RNA sequencing has rapidly expanded the scale of skin transcriptomic data, yet these datasets remain fragmented across studies spanning different species, diseases and experimental manipulations. An up-to-date, comprehensive, and queryable single-cell cross-species repository for skin is still lacking. Here, we present scSAID (skin-scsaid.com), a single-cell database with an interactive web portal offering a broad suite of in-depth analyses for human and mouse skin. It integrates more than 1.2 million high-quality cells collected from 252 samples, establishing a unified reference for cell-type annotation, cross-species comparison and pathological studies. Using psoriasis as a case study, we demonstrate how scSAID can be used to evaluate how faithfully mouse models reproduce human pathology. Systematic comparison with the imiquimod-induced mouse model revealed numerous species-specific molecular signatures of psoriasis, including human-specific NFKB1 activation and STAT1 involvement, indicating that the current mouse model captures only limited aspects of the disease. We further introduce psoSpotter, a disease-biomarker-selection algorithm that, coupled with in silico perturbation using scSAID data, uncovers PPIA as a novel psoriasis drug target, illustrating the potential of scSAID for identifying therapeutic approaches. Overall, scSAID delivers a large-scale, cross-species skin single-cell resource and analysis platform, opening new opportunities for the discovery of disease-relevant targets in skin diseases.
bioinformatics2026-07-23v1Transcriptional landscape of direct reprogramming toward the hematopoietic lineage
Cwycyshyn, J.; Stansbury, C.; Golts, S.; Lee, H.; Pickard, J.; Meixner, W.; Rajapakse, I.; Muir, L. A.Abstract
Direct reprogramming of human fibroblasts into hematopoietic stem cells (HSCs) offers a promising strategy for generating autologous cells to treat blood and immune disorders. Current protocols are limited by low efficiency and insufficient tools for evaluating reprogramming outcomes. Although functional assays are the standard for confirming cell identity, they require fully reprogrammed cells, limiting their utility during protocol development. To address this, we assembled a single-cell transcriptomic reference atlas of hematopoietic reprogramming and tested an algorithmically-predicted transcription factor recipe for HSC induction. Long-read single-cell RNA sequencing of CD34+ reprogrammed cells revealed progressive loss of fibroblast identity alongside induction of early hematopoietic and endothelial programs, with reference-atlas benchmarking placing reprogrammed cells in an intermediate transcriptomic state between fibroblasts, endothelial cells, and HSCs. Isoform-level analysis further revealed transcriptional remodeling not captured by gene-level analyses. This experimental-computational framework offers a generalizable strategy for characterizing partially reprogrammed states and guiding optimization of reprogramming protocols.
bioinformatics2026-07-22v5Benchmarking the Impact of Data Leakage on the Performance of Knowledge Graph Embedding Models for Biomedical Link Prediction
BRIERE, G.; STOSSKOPF, T.; LOIRE, B.; BAUDOT, A.Abstract
Motivation: Knowledge Graphs (KGs) are increasingly used to organize complex biomedical knowledge into structured representations of entities and relations. Knowledge Graph Embedding (KGE) models facilitate efficient exploration of KGs by learning compact representations, and are widely applied to biomedical link prediction, for instance to uncover new therapeutic uses for existing drugs. Despite extensive work on KGE models, current evaluations often overlook data leakage, which can artificially inflate performance and undermine benchmark validity. Data leakage can arise when (1) there is redundancy between training and test sets, (2) the model leverages illegitimate features, or (3) the test set does not reflect real-world inference scenarios. Results: We assess the impact of data leakage on KGE-based link prediction across three biomedical KGs, using both decoder-only and GNN-based models. We first demonstrate the impact of train-test redundancies and implement a systematic procedure to detect and remove them. Using permutation experiments, we investigate whether node degree acts as an illegitimate predictive feature, and find no evidence that predictions are driven by degree alone. Finally, we evaluate how well common test set sampling strategies reflect real-world inference in drug repurposing. Comparing random and cold-start splits with an independent Orphanet-derived test set, we observe a substantial performance drop on the latter, indicating that current practices may overestimate how well KGE models generalize. Overall, our findings highlight the importance of rigorous benchmark design and careful evaluation of the generalization ability of KGE models for biomedical link prediction. Availability and Implementation: All code and results are openly available on GitHub at https://github.com/galadrielbriere/data_leakage_kge_benchmark.git.
bioinformatics2026-07-22v3Graph Lens Lite: A browser-based tool for interactive visualization and exploration of biological networks
Ley, M.; Keska-Izworska, K.; Fillinger, L.; Walter, S. M.; Baumgärtel, F.; Bono, E.; Galou, L.; Andorfer, P.; Hauser, P.; Leierer, J.; Kratochwill, K.; Perco, P.Abstract
Biological network visualization together with graph-based analyses are key techniques in systems biology and network medicine to detect patterns and generate hypotheses regarding disease pathobiology, drug target identification, biomarker prioritization, and digital drug discovery. Network representations provide an intuitive way to communicate and share research findings. We have developed Graph Lens Lite, a browser-based tool that combines rich visualization with a streamlined interface for exploring and sharing biological networks. It offers an expressive query language, topological network analysis, interactive filtering, visual grouping, customizable layouts, a data editor, fine-grained property-based styling, animated edge-flow visualization, community detection, and a context-aware locally powered AI assistant particularly suited for exploring molecular models of disease pathobiology or drug mechanism of action. We demonstrate its utility on a curated network model of autosomal dominant polycystic kidney disease. Graph Lens Lite is open source, with a live web version available at https://delta4ai.github.io/GraphLensLite/.
bioinformatics2026-07-22v3CESAR: A R Package for High-Sensitivity Detection of Copy Number Variations in ctDNA Using Segmentation and Anchor Recalibration
Lu, J.; Ni, S.; Wang, L.; Wu, N.; Jiang, X.Abstract
Background: Detecting copy number variations (CNVs) in circulating tumor DNA (ctDNA) is crucial for the companion diagnosis and resistance monitoring of various solid tumors (e.g., NSCLC, Glioblastoma). However, when tumor-derived DNA fractions are extremely low (often <1%), traditional depth-based methods frequently fail due to non-linear sequencing depth fluctuations and probe-specific capture biases inherent to targeted Next-Generation Sequencing (NGS). Methods: We developed CESAR (CNV Estimation with Segmentation and Anchor Recalibration), a novel computational tool optimized for ultra-sensitive, tumor-only CNV detection in targeted NGS panels. CESAR utilizes Circular Binary Segmentation (CBS) to re-partition target regions based on relative capture efficiency. It then introduces a dynamic "anchor" selection algorithm that identifies a personalized set of genomic segments mirroring the non-linear coverage behavior of each target gene. By minimizing the Coefficient of Variation (CV) through iterative anchor selection, CESAR effectively recalibrates the baseline to suppress technical noise. Results: Validation using standard DNA reference materials demonstrated that CESAR successfully identified both amplifications (e.g., MET, ERBB2, EGFR) and relative copy number deletions at ultra-low tumor fractions. Notably, CESAR achieved stable detection of focal alterations as subtle as 2.18 copies (a mere 1.09x fold change relative to the diploid baseline), while maintaining zero false positives in control regions. Evaluation across distinct clinical biofluids, 36 clinical plasma samples and 41 glioma cerebrospinal fluid (CSF) samples, identified critical, previously undetected CNV events, including subtle ERBB2 gains and distinct MET deletions. Furthermore, comprehensive benchmarking revealed that CESAR consistently outperformed the widely used CNVkit, particularly in suppressing technical variance and resolving ultra-low-level copy number gains that CNVkit failed to distinguish from background noise. Conclusions: CESAR provides a highly stable and sensitive algorithmic framework for tumor-only CNV calling in liquid biopsies, facilitating precise therapeutic decision-making in precision oncology.
bioinformatics2026-07-22v2IntraTalker - Modelling Intracellular Signalling in1 Cellular Crosstalk
Kloker, V.; Nagai, J. S.; Feng, Z.; Mavrommatis, L.; Hermanns, L.; Moscoso, J. M. J.; Ruiz, M.; Kuppe, C.; Costa, I. G.Abstract
Single-cell sequencing has advanced the study of cell-cell communication, yet most methods focus on intercellular ligand-receptor interactions while neglecting downstream intracellular signalling cascades and the possibility that downstream target genes themselves encode ligands, thereby propagating communication across multiple cells. We present IntraTalker+CrossTalkeR that combines intracellular (IntraTalker) and intercellular (CrossTalkeR) signalling from multimodal single-cell data. IntraTalker infers cell-type-specific transcription factor activities and constructs receptomes that link receptors to downstream target genes, which are then integrated with ligand-receptor predictions in CrossTalkeR. To prioritize signalling receptors, the framework performs in silico receptor perturbation. In a murine bone marrow dataset, this recovered the known function of the Il1r1 receptor in driving myeloid progenitor cell-state changes. In a human myocardial infarction dataset, it predicted a novel role for IGF1R signalling in driving the differentiation of fibroblasts towards progenitor fibroblast states.
bioinformatics2026-07-22v1Integrative structure determination of a human mitochondrial contact site and cristae organizing system (MICOS) sub-assembly
Jindal, M.; Mahato, R.; Das, S.; Guha, A.; Majila, K.; Arvindekar, S.; Vaidya, A. T.; Viswanath, S.Abstract
The Mitochondrial contact site and Cristae Organizing System (MICOS) complex is an inner mitochondrial membrane (IMM) assembly present at the cristae junction. It is responsible for regulating cristae formation and remodeling. However, its structure is not known. We applied Bayesian integrative structure determination to characterize the structure of the Mic60, Mic19, Mic10, and Mic13-containing MICOS complex combining AlphaFold predictions with data from crosslinking mass spectrometry, biochemical assays, electron tomography, homology modeling, and sequence alignments. The integrative structure revealed novel mutual interfaces among Mic10N,C, Mic60LBS1,LBS2,mitofilin, and Mic13central,C, which were experimentally validated. Several likely-pathogenic missense mutations also localize to these novel interfaces, highlighting their importance. Our results indicate that Mic13 likely facilitates MICOS assembly by binding Mic10 in the IMM-proximal region and Mic60 in the intermembrane space. Taken together, our integrative approach sheds light on the structure and assembly of the MICOS complex.
bioinformatics2026-07-22v1AI4Life Open Calls and Public Challenges: why, how, and what we have learned.
Galinova, V.; Seifi, M.; Serrano Solano, B.; Lidayova, K.; Dalle Nogare, D.; Corbat, A. A.; Talks, J.; Giacomello, E.; Gomez-de-Mariscal, E.; Ferreira, M. G.; Fuster-Barcelo, C.; Battagliotti, J. M.; Garcia-Lopez-de-Haro, C.; Salmon, B.; Croft, M.; Yie, S. Y.; Rey-Paniagua, G.; Hu, X.; Cho, S.; Sheth, A.; Porwal, C.; Li, X.; AI4Life Consortium, ; Henriques, R.; Li, X.; Krull, A.; Klemm, A.; Munoz Barrutia, A.; Kreshuk, A.; Ouyang, W.; Jug, F.; Deschamps, J.Abstract
Within AI4Life, we ran three Open Calls and three Public Challenges (2023-2025), supporting 22 bioimage analysis projects from 151 applications and engaging 225 challenge participants, with the aim of applying FAIR deep learning in the life sciences. Our experience offers a view of the current state of bioimage analysis, the landscape of available tools, as well as the existing gaps between method developers, tool producers and potential users. It highlights that even after careful selection for AI-ready projects, most still require substantial effort to apply deep learning, and that the field still relies heavily on established, well-rounded methods to solve common problems. We come to the conclusion that for scientific AI in biology, the rate-limiting step is not methods and models but data, annotations, and shared infrastructure underneath them.
bioinformatics2026-07-22v1TrioNsight: Building a meta-predictor to evaluate the clinical impact of TrioN-like Dbl-homology domain variants
Taciroglu, A.; Aydin Son, Y.; Martin, A. C. R.; Orengo, C.Abstract
TRIO is a member of the Rho family of guanine nucleotide exchange factors (Rho GEFs), which promote the exchange of GDP for GTP to activate Rho GTPases and serve as key regulators of cellular signalling pathways. Mutations in TRIO are associated with neurodevelopmental disorders, including intellectual disability and autism spectrum disorders. TRIO contains two GEF units: one N-terminal and one C-terminal, each of which contains a Dbl-homology (DH) domain that drives its GEF activity to activate Rho GTPases. While the human proteome contains 70 highly conserved DH domains, the N-terminal DH domain of TRIO (TrioN) contains one-third of all reported DH domain pathogenic variants and has many variants of unknown significance. Numerous variant impact prediction tools exist, but most lack gene-specific considerations. Here, we describe TrioNsight, a meta-predictor designed to predict mutation impacts for TrioN and 12 highly similar human DH domains, including TrioC. TrioNsight exploits the naive-Bayes algorithm and leverages structural, evolutionary, and physiochemical features of approximately 1500 highly similar DH domains (TrioN-like DH domains) from 294 species. TrioNsight surpasses all available predictors, including AlphaMissense, achieving a Matthews' Correlation Coefficient of 0.890. Additionally, we provide a variant impact map that details the impacts of mutations at each position in the DH domain of these proteins, which can be valuable for clinical assessments. Furthermore, our approach establishes a standardised workflow adaptable for creating domain-specific variant predictors for other protein families, offering a template for improved variant interpretation.
bioinformatics2026-07-22v1IOBRpy enables agentic multi-omics decoding of anti-tumor immunity
Huang, H.; Li, X.; Liu, L.; Gu, W.; Wang, G.; Zeng, D.Abstract
Decoding the tumor immunity is pivotal for cancer immunotherapy, yet transcriptomic pipelines remain bottlenecked by fragmented tools and biased interpretations. Here we present IOBRpy, a Python toolkit driven by an innovative AI dual-agent layer for automated, highly standardized immuno-oncology workflows. Moving beyond conventional expression profiling, IOBRpy enables agentic multi-omics decoding. From raw FASTQ or TPM matrices, it seamlessly integrates upstream quality control, transcript quantification, and downstream TME parsing, encompassing signature scoring, ligand-receptor crosstalk, and cellular deconvolution. Crucially, IOBRpy expands data dimensions by incorporating complementary immunogenomic layers, empowering concurrent high-resolution SpecHLA typing and TRUST4-based TCR/BCR repertoire reconstruction from sequencing data. Deployed across two large-scale cohorts (IMvigor210 and OAKPOPLAR), IOBRpy successfully captured multi-dimensional prognostic insights. While broad HLA-I heterozygosity showed negligible impact, it precisely unmasked treatment-stratified, allele-specific survival associations (e.g., HLA-A*01 and HLA-DPA1*02) tightly coupled with distinct immunosuppressive ligand-receptor networks (such as HMGB1-THBD and EFNB2-EPHB6) and dynamic TCR clonal diversity shifts. Empowering this lifecycle is a paired agent framework: a workflow agent automatically audits project states to execute validated, path-aware commands, while a result agent evaluates tool provenance, handles method-aware adaptive visualizations, and organizes findings into evidence-constrained biological hypotheses. Collectively, IOBRpy provides a reproducible, scalable, and intelligence-augmented Python gateway to transform raw sequencing data into multi-omic, interpretation-ready discoveries for cohort-scale precision immunotherapy (https://iobr.github.io/IOBRpy/).
bioinformatics2026-07-22v1BioPhasor: Decoding Cellular State Tensors from Multi-Omics Phasor Dynamics for Quantum Ready Systems Biology
Sigdel, D.; Panday, N.Abstract
Integrating multi-omics data - transcriptomics, proteomics, metabolomics, single-cell - remains a fundamental challenge in systems biology. We present BioPhasor, a framework that encodes each measurement as a complex phasor z = e^(i{varphi}) on the compact N-torus T^N, modelling the cell as phase-coupled oscillatory programs whose dissipative dynamics generate limit cycles and an attractor landscape. From this geometry we derive the Cell State Tensor (CST), a rank-3 tensor whose axes we root in measured multi-omics quantities: a pathway/module atlas on the regulatory axis and a directional central-dogma modality axis. Across nine scenarios on open public data (GEO, CPTAC), loaded through one unmodified data layer, we report verdicts honestly: four reproduce, three are partial, two do not. A data-driven cell-cycle axis lifts agreement with a reference method from 0.34 to 0.69; an explicit circadian origin cuts peak-time error from 10.6 to 1.4 h; and central-dogma coupling mRNA phase organising protein amplitude clears a surrogate null and is tumour-specific. Grounding the quantum-ready claim, the CST maps to a density-matrix formalism whose coherence and entropy match quantum-information counterparts, and the phasor circuit transpiles gate-for-gate to a variational quantum circuit, though no empirical advantage emerges. One loader regenerates every reported number, and the code is released.
bioinformatics2026-07-22v1A Standardized Methodology for FAIRness Assessment and Multi-Dimensional Scoring in Agrosystem Research Data Infrastructures
Haleem, A. U.; Arend, D.; Etukala, J. R.; Mazon, E. R.; Schmidt, M.; Jung, J.; Martini, D.; Usadel, B.; Neidiger, C.; Ulrich, R.; Lange, M.Abstract
The NFDI-consortium FAIRagro has established a systematic framework for evaluating the FAIRness of Research Data Infrastructures (RDIs) within the German agrosystem research landscape. While FAIR principles are widely accepted, their practical implementation by research data infrastructures (RDI) remains challenging. By operationalizing the FAIR principles into a reproducible multi-dimentional scoring methodology, this initiative addresses the critical need for a transparent and citable benchmark of RDIs that moves beyond simple compliance. This paper details the underlying assessment criteria, comprising 20 aggregated core metrics, the iterative community-driven validation process, and the integration of these metrics into the FAIRagro Search Hub. This framework evaluates RDIs, like repositories or databases, instead of sampling hosted data sets., across the four distinct categories of FAIR independently, yielding granular, pillar-specific ratings. Unlike aggregate scoring models, which can inadvertently mask technical deficiencies by averaging performance across categories, this multi-dimentional approach ensures that a repositorys distinct strengths and bottlenecks remain fully visible. Our findings demonstrate that standardized scoring not only clarifies data accessibility for users but also highlights specific operational gaps, allowing repository providers to identify precisely where the service implementation can be enhanced. By establishing this data-driven service in the agronomy domain, we provide a scalable template for the broader NFDI and EOSC ecosystems to foster a culture of excellence in research data stewardship.
bioinformatics2026-07-22v1Hobrac: a reference-guided workflow for genome comparison and synteny visualization
Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.Abstract
Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO rthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.
bioinformatics2026-07-22v1scLEMBAS: Context-Aware Modeling of Signaling Pathway Activity at Single-Cell Resolution
Baghdassarian, H. M.; Meimetis, N.; Nordenstorm, O.; Joughin, B.; Nilsson, A.; Lauffenburger, D.Abstract
Cells sense and integrate extracellular cues through intracellular signaling networks that reshape transcription factor activity to dictate cellular responses. Signaling activity is difficult to decipher: it is non-linear, and it contains extensive feedback and crosstalk. Furthermore, the same perturbation can elicit markedly different responses depending on context (e.g., cell type, disease state, and tissue microenvironment) such that identical stimuli produce diverse responses in multicellular populations. Consequently, there is a vast combinatorial space of complex interactions and context-dependent responses that necessitate computational models. Computational models of single-cell perturbation responses are demonstrated to predict cellular responses, but are often limited in mechanistic insight. Prior knowledge networks offer a route to bridge predictive capability and interpretability. Here we present scLEMBAS, a context-aware, gray-box neural network that models signaling pathway activity at single-cell resolution while preserving mechanistic grounding. scLEMBAS encodes a prior-knowledge network of protein-protein interactions as a recurrent neural network whose learnable edge weights correspond to signaling interaction strengths. It also captures context and individual cell variance through compositional bias terms. An adversarial approach allows the model to answer a single-cell counterfactual - what a given cell's TF activity would be under a different perturbation or context - while involving mechanistic rather than simply relational information. Across two scRNA-seq datasets spanning single- and multi-perturbation settings, scLEMBAS accurately predicts out-of-distribution combinations of perturbation and context. Capturing population variance across individual cells enables the model to predict cell subtype specific perturbation responses, despite being agnostic to such labels. Beyond prediction, scLEMBAS learned parameters are biologically interpretable: learned edge weights carry information beyond network topology and "self-prune" spurious interactions, while the categorical bias nominates proteins associated with cell-type-specific perturbation states. Overall, scLEMBAS enables quantitative dissection of how signaling pathway activity is reshaped by perturbation within specific cellular contexts.
bioinformatics2026-07-22v1Learning Minimal Gene Programs for Disease-Aligned Representations
Madduri, A.; Patel, C. J.Abstract
Identifying small, interpretable gene sets that robustly capture disease-associated variation in single-cell transcriptomic data remains a central challenge for biological interpretation and experimental follow-up. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-held-out evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data.
bioinformatics2026-07-22v1Tangerine: A Python framework for dynamic gene regulation analysis from transcriptomic time series
Narendra, T.; Schweikert, G.Abstract
Motivation: Time-series single-cell transcriptomics enables the study of dynamic gene regulation. However, standard computational tools frequently aggregate temporal data into static, dense topologies, obscuring the precise regulatory rewiring that drives developmental transitions. Further, navigating the inherent noise of statistical inference without losing biological interpretability remains an important bottleneck. Results: We present Tangerine, a Python framework for the dynamic reconstruction and interactive exploration of time-varying gene regulatory networks. Tangerine integrates time-constrained metacell aggregation with regularized linear modelling and non-parametric correlation to infer dynamic topologies. To solve the interpretability gap, it features a browser-based visual analytics engine. Tangerine empowers researchers to track macroscopic gene module evolution, interactively filter effect sizes, and link topological rewiring directly to raw transcriptomic evidence. Availability and implementation: Tangerine is implemented in Python and Plotly Dash. The code is available on Github at https://github.com/ntanmayee/tangerine.
bioinformatics2026-07-22v1Target Preference Maps: A machine learning model generalizing transferable drug-receptor interactions and guiding drug discovery
Menezes, F.; Wahida, A.; Froehlich, T.; Grass, P.; Zaucha, J.; Napolitano, V.; Siebenmorgen, T.; Pustelny, K.; Barzowska-Gogola, A.; Rioton, S.; Didi, K.; Bronstein, M.; Czarna, A.; Hochhaus, A.; Plettenburg, O.; Sattler, M.; Nissen-Meyer, J.; Conrad, M.; Kurzrock, R.; Popowicz, G. M.Abstract
Modern AI models can decode the genomic landscape and protein structure world. Yet, they fail to generalize to one of the most important fields: small-molecule drug discovery. Since the late 1970s, the advent of macromolecular crystallography inspired the notion that structural knowledge alone could enable a lock-and-key approach to drug design. However, drug discovery continues to depend on costly, resource-intensive, and largely serendipitous screening campaigns that probe only an infinitesimal fraction of the drug-like chemical space. Despite some successful cases, our understanding of, and reasoning from, non-bonded interaction chemistry remains limited for general applicability. Furthermore, though structural databases contain hundreds of thousands of entries, a strong historical bias pervades protein-drug structures, hindering reliable advances through AI scaling. Here, we present a machine-learning framework that learns atom-type-specific spatial preference maps from local protein microenvironments in protein-ligand structures. By excluding ligand topology from the model input and learning from local atom-level environments, the framework is designed to reduce dependence on whole-ligand memorization and to capture transferable interaction preferences. The resulting maps recover chemically meaningful interaction patterns, including cases involving bridging waters and metal-dependent environments. The model was validated using retrospective and prospective real-world data in drug optimization when targeting a challenging protein-protein interface. This shows that the method can provide interpretable workflows to guide molecule optimization and provide input for downstream generative or docking workflows.
bioinformatics2026-07-21v11LeafRank: A phylodynamic framework for inferring relative fitness from single-cell phylogenies in chromosomally unstable tumors
Wu, C.; Leder, K.; Wang, Z.; Sun, R.Abstract
Tumors contain cancer cells with diverse growth potentials that shape evolutionary trajectories, yet this fitness diversity remains difficult to quantify in cases of whole-genome duplication (WGD) and chromosomal instability. We present LeafRank, a mathematical framework that leverages single-cell DNA-seq phylogenies to infer the relative fitness of individual cells. Using a multi-type branching process model, LeafRank integrates full tree topology, including branch lengths and bifurcation patterns, to estimate marginal fitness probabilities under punctuated evolutionary regimes driven by rare driver events. To account for elevated aberration rates following WGD, we introduce a tree-rescaling strategy that adjusts for lineage-specific genomic instability. Unlike methods focused on predefined subclones, LeafRank ranks all sampled cells, enabling flexible assessment of growth heterogeneity. Simulations demonstrate high accuracy across spatial and non-spatial virtual tumors. Applied to ovarian cancer, LeafRank reveals directional and parallel selection in WGD tumors and identifies recurrent copy number events enriched in high-fitness lineages. WGD lineages do not show immediate growth advantages but acquire fitness through subsequent alterations.
bioinformatics2026-07-21v4RNAStabFormer: Region-Aware Multi-Task Hybrid Learning for RNA Stability Prediction from Pulse-Chase Transcriptomics
Wang, S.; Zhang, C.Abstract
RNA stability is a major post-transcriptional regulator of gene expression, yet sequence-based prediction from pulse-chase transcriptomics remains difficult because labels depend on the time window, quantification region, and replicate quality. We present RNAStabFormer, a controlled RNA stability framework centered on a Region-Aware Multi-Task Hybrid Transformer (RAMHT). RAMHT encodes the nucleotide context of the 5-prime UTR, coding sequence (CDS), and 3-prime UTR; incorporates an additional CDS codon stream; upgrades engineered sequence features into a tabular interaction branch; and uses gated multi-task regression to predict four ENCODE BrU-seq/BruChase-seq RNA stability proxies, with Exon 6 h/0 h as the primary task. Across 26 outer data splits, including 23 chromosome holdout tests, a heterogeneous three-member RAMHT ensemble achieves a mean Pearson correlation of 0.773 on the primary task, statistically matching an engineered-feature XGBoost baseline, which also achieves 0.773. The mean paired difference is +0.000004, with a bootstrap 95 percent confidence interval from -0.003845 to +0.004077 and a Wilcoxon p-value of 0.8613. The ensemble improves upon the strongest individual RAMHT member, increasing the mean Pearson correlation from 0.768 to 0.773 and outperforming it on 23 of the 26 data splits. A strictly nested XGBoost-RAMHT blended model further increases the correlation to 0.775. Evaluations conducted on identical data splits also show that the ensemble outperforms frozen full-length mRNA language-model embeddings, which achieve a correlation of 0.760, and public LAMAR-DR transfer learning, which achieves a correlation of 0.180. Gate analysis, ablation experiments, and sequence recoding analyses indicate that engineered sequence grammar remains the dominant source of predictive information, whereas the nucleotide and codon branches provide complementary signals localized primarily within the CDS. RNAStabFormer narrows the performance gap between neural RNA sequence models and strong tabular baselines while retaining an extensible architecture for model interpretation and biological data integration.
bioinformatics2026-07-21v2Panomap: Unbiased Nanopore Signal Mapping with Pangenome Variation Graphs
Shih, P. J.; Sanghani, Z.; Guarracino, A.; Gamaarachchi, H.; Batten, C.Abstract
Motivation: Signal-space nanopore mappers enable real-time mapping and filtering decisions directly from raw nanopore signals. However, existing signal-space mappers are built around linear references, and using a single representative reference can introduce reference bias when the sample diverges from that reference. Pangenome reference collections can reduce this bias by representing diversity across related reference sequences, but linear-reference signal mappers must treat each sequence as a separate target, redundantly storing shared sequences. Pangenome variation graphs provide a more compact representation by storing shared sequences once and encoding variants as alternative paths through the graph. Although sequence-to-graph mapping is well established for basecalled reads, existing signal-space methods do not directly use pangenome variation graphs. Results: We present Panomap, the first signal-space mapper that operates on pangenome variation graphs. Panomap maps raw nanopore signals to graph references, allowing signal-space mapping to use pangenome diversity while representing shared sequences once. We evaluate Panomap in three settings. First, when a single reference already maps the sample well, Panomap preserves mapping accuracy as additional reference sequences are added to the reference collection, while state-of-the-art signal-space tools regress. Second, when the exact sample strain is absent from the reference collection, Panomap benefits from adding related assemblies from the same species to the pangenome reference. Third, using a highly polymorphic locus, we show that Panomap can map reads from alleles not represented in the reference collection by using related alleles in the pangenome, with the largest gains for more divergent alleles and for decisions made from short prefixes of the read signal. In addition, Panomap's graph index scales sublinearly with pangenome collection size. Together, these results show that Panomap brings population-aware reference representation into signal-space mapping. Availability and Implementation: Panomap is open source and available at https://github.com/cornell-brg/panomap.
bioinformatics2026-07-21v2Genomic, Transcriptomic, and Regulomic Analyses Do Not Support Profound Autism as a Distinct Biological Category
Eicher, T. D.; Ne'eman, A.; Quackenbush, J. D.Abstract
The Lancet Commission on the Future of Care and Clinical Research in Autism proposed the construct of "profound autism" as a recognizable subtype of autism. Supporters argue that this classification is necessary to ensure that autistic persons with severe impairment receive appropriate research attention and policy support, whereas critics contend that the construct lacks scientific validity and may reflect social or political considerations more than biological distinction. To inform this debate, we evaluate whether the proposed "profound autism" category represents a distinct genetic phenotype using multiple molecular data types collected in a large cohort. Across genomic, transcriptomic, and regulatory analyses, we find no evidence supporting "profound autism" as a biologically distinct phenotypic group. Instead, differences emerge primarily in inferred gene regulatory networks distinguishing nonspeaking from speaking autistic children, suggesting potential regulatory mechanisms contributing to speech ability. These findings suggest that future research into severe impairment may be more productive if focused on specific traits -- such as speech impairment -- rather than attempting to define a distinct biological subtype within the multidimensional phenomenon of autism.
bioinformatics2026-07-21v2SEEK-VEC: Augmenting topic modeling with spectral ensemble learning
Danning, R.; Ke, Z. T.; Ma, R.; Lin, X.Abstract
Count data are ubiquitous across many applications in which understanding latent patterns is of interest. Topic modeling is a powerful tool for detecting latent structure in count data. However, standard topic modeling methods are often constrained by their restrictive assumptions, susceptible to noise, and sensitive to misspecification of the number of topics. Here, we introduce SEEK-VEC (Spectral Ensembling of topic models with Eigenscore for K-agnostic Vocabulary Embedding and Classification), an ensemble topic modeling framework that integrates insights from multiple candidate topic models through a spectral ensembling procedure. SEEK-VEC produces a meta-structure matrix containing prioritization scores and grouping scores that enable variable classification, interactive pattern discovery, and model diagnostics. Through simulations, we demonstrate that SEEK-VEC augments the performance of standard topic models for identifying important vocabulary words and understanding the relationships among them, particularly when signal strength is weak. We apply SEEK-VEC to the MADStat dataset of statistical abstracts and demonstrate its utility for evaluating the proposed interpretation of a topic model.
bioinformatics2026-07-21v2Molecular Mimicry in Inflammatory Bowel Disease: Multi-layered Functional and Sequence-level Analysis of Gut Microbial Proteins Mimicking the Human Proteome
Anand, A. A.; Mishra, P.; Srivathsa, V. S.; Yadav, V.; Samanta, S. K.Abstract
Inflammatory bowel disease (IBD), encompassing Crohn's disease (CD) and ulcerative colitis (UC), is a chronic inflammatory disorder whose pathogenesis involves intricate host-microbiome interactions. Molecular mimicry (the structural or functional resemblance between microbial and host proteins) represents a plausible mechanism by which gut microbiota may trigger or perpetuate autoimmune responses. Here we present a comprehensive, multi-layered molecular mimicry in silico pipeline (MMIP) analysis of 39 baseline shotgun metagenomic samples from Human Microbiome Project 2 (HMP2/IBDMDB). Using DIAMOND-based homology search against the SwissProt followed by UniProt Retrieve/ID Mapping (URIM), we characterized microbial protein functional space through three complementary frameworks: (i) normalized GO term frequency comparison across diagnostic groups, (ii) protein family (PFAM) domain enrichment analysis, and (iii) sequence-level mimicry analysis identifying microbial proteins with direct homology to human proteins. Taxonomic profiling using MetaPhlAn3 pre-computed abundance profiles provided further biological context. CD microbiomes exhibited greater enrichment of immune-relevant biological processes and a higher per-sample sequence-level mimicry rate than UC. CD and UC also showed distinct pathobiont profiles, with CD enriched for oral-origin taxa including Haemophilus parainfluenzae and UC enriched for Fusobacterium nucleatum and Prevotella species. Notably, healthy gut microbiome maintains coordinated mimicry of host neuronal, signal-recognition-particle-associated, and antimicrobial peptide machinery, a repertoire dismantled in IBD and replaced by disease-subtype-specific signatures, alongside a candidate mimicry link between CD-enriched bacteria and NOD2. These results represent the first metagenome-wide characterization of molecular mimicry across IBD subtypes using shotgun metagenomic data, offering new mechanistic insight into how microbial dysbiosis may contribute to immune dysregulation in IBD. Keywords: Inflammatory bowel disease (IBD), Crohn's disease (CD), Ulcerative colitis (UC), Gastroenteritis, Molecular mimicry
bioinformatics2026-07-21v2COSMOS: A FAIR-aligned infrastructure for clinical trial data validation, warehousing, and interactive discovery
Roberts, C. A.; Nilsson-Takeuchi, A.; Stuart, C.; Chivers, M.; Soares, P.; Foure, V.; Griffiths, G.; Niazi, U.Abstract
The Clinical Omics System for Metadata and Outcome Storage (COSMOS) is an open-source, FAIR- and GCP aligned clinical trial unit (CTU) infrastructure designed to streamline the transition of academic clinical and multi-omic trial datasets into curated, analysis-ready repositories within Secure Data Environments (SDEs). By integrating an automated Data Quality and Data Validation (DQ&DV) "Trust Layer" with a relational Structured Query Language (SQL) schema, COSMOS enables programmatic and interactive data access via R Shiny applications. Dynamic integration of omics and clinical data is achieved through analytical data structures (e.g. ExpressionSet objects) linked via relational database identifiers. This lowers technical barriers for researchers and promotes governed data reuse and secondary discovery.
bioinformatics2026-07-21v1Bayesian Factor Analysis for Binary and Ordinal Phenotypes with Missingness
Shashaank, N.; Knowles, D. A.Abstract
Binary and ordinal phenotypes are common in clinical screening and self-reported questionnaires, but many factor analysis and matrix factorization methods are only applicable for quantitative phenotypes with real-valued and/or continuous data distributions. To address this, we propose FABOr (Factor Analysis of Binary and Ordinal data), a Bayesian framework for matrix factorization in which the low-rank matrices are latent variables with continuous priors while the phenotypes are observed variables modeled with appropriate binary/ordinal likelihoods. We also develop missing not at random (MNAR) extensions of FABOr for analyzing data with structured missingness. In experiments with simulated phenotypes, we found that FABOr performs similarly to the best-performing benchmark methods on binary data at imputation and exceeds the performance of all tested benchmark methods on ordinal data. We then applied FABOr to analyze real-world binary and ordinal phenotypes from the Simons Foundation SPARK dataset on autism spectrum disorder (ASD) and found that it improved imputation accuracy by up to 5% on binary data and up to 23% on ordinal data relative to the benchmark methods.
bioinformatics2026-07-21v1A new set of DNA methylation variants in the human genome show predominant tissue specificity and sensitivity to reprogramming with a potential for disease susceptibility.
Anne, A.; Kumar, L.; Singh, M.; Choudhury, S.; Das, S.; Zimmer-Bensch, G.; Bandyopadhyay, D.; K, N. M.Abstract
Analyses of 3,370 normal human tissues of ectodermal, endodermal and mesodermal origins identified 12,587 regions averaging ~585 bp with significant differences in DNA methylation levels within identical tissues. These methylation variants (MeVars) occurred in 8,037 genes enriched in neurological disorders and cancers of which, majority were tissue-specific rather than being systemic. This somatic variation was reduced by reprogramming in vitro into iPSCs and in vivo during spermatogenesis. Analysis of prefrontal cortices showed a higher incidence of MeVars in the candidate genes in controls than schizophrenia patients wherein a subset showed significantly altered transcript levels. Similar effects were observed for oral tissues and skin fibroblast cells. MeVars showed significant association with SINE1, simple and low complexity repeats, H3K27me3, H3k9me3 and H3K4me1 modifications and EZH2, SUZ12 and REST binding sites. Collectively, MeVars have postzygotic origins with an ability to reset during reprogramming, adding a new dimension in the form of epigenetic diversity and its relevance to disease susceptibility in humans.
bioinformatics2026-07-21v1ProtSyntax: a protein large language model for decoding post-translational modification syntax and function
Lin, Y.Abstract
Post-translational modifications (PTMs) regulate protein function through dependencies among residue chemistry, sequence context, three-dimensional microenvironments and modification states, yet most predictors model sites independently and do not connect modification propensity to functional consequences. Here we present ProtSyntax, a PTM-centered protein language model trained on 4.25 million examples spanning 40 PTM classes and supervised for kinase specificity, PTM crosstalk and enzyme kinetics. ProtSyntax integrates bidirectional long-range modeling with geometry-gated attention in a sparse mixture-of-experts architecture and uses adaptive multi-objective learning to couple residue-level PTM syntax to protein-level function. Across 40 PTM-site benchmarks, ProtSyntax improved mean MCC and AP by 12.7% and 10.7%, respectively, relative to the best-performing baselines. It also distinguished authentic sites from structurally incompatible motif decoys, transferred to rare PTMs, recovered crosstalk, linked PTM perturbations to enzyme-kinetic changes and identified disease-associated PTM disruptions. Together, ProtSyntax provides an interpretable framework for decoding PTM regulation across the proteome.
bioinformatics2026-07-21v1PepCL: A replay-based continual learning framework for updating peptide-MHC models
Chati, P. M.; Lashkari, V. D.; Salhotra, A.; Bruno, P. M.; Ntranos, V.Abstract
Understanding peptide-major histocompatibility complex (MHC) class I binding is critical for effective vaccine and immunotherapy design but is a combinatorially complex challenge for which prediction models have become essential. MHC ligands are typically identified at scale via untargeted mass spectrometry (MS), and this has built a strong base for peptide-MHC model training. However, MS incompletely captures the vast peptide-MHC space due to technical, sampling, and biological biases. Although recently developed experimental assays have queried such blind spots yielding complementary information, existing peptide-MHC predictors have not yet incorporated these orthogonal data and are not designed to be updated as new data are generated. Here, we introduce PepCL (Peptide-MHC Continual Learning), a continual learning framework for updating peptide-MHC predictors with new assay data while explicitly preserving prior MS knowledge. To enable PepCL, we also develop MHCPrime, a new state-of-the-art pan-allelic peptide-MHC prediction model, trained on publicly available MS data, that can be effectively updated under our framework. We demonstrate that PepCL allows MHCPrime to learn previously unseen, assay-specific information while preventing catastrophic forgetting that is typically observed with conventional fine-tuning. We evaluate PepCL and MHCPrime in a variety of biological contexts, including infectious disease and cancer, and show improved peptide-MHC prediction that transfers across alleles for broader applicability in clinical settings. Overall, our results establish PepCL as a flexible framework for extending the utility of peptide-MHC models by improving their predictive performance as immunopeptidomics assays continue to evolve and new data become available.
bioinformatics2026-07-21v1An automated platform for spatial functional modeling and fingerprint analysis of tissue molecular landscapes
Hajihosseini, M.; Patino-Martinez, E.; Ghosal, R.; Kaplan, M. J.; Pyne, S.Abstract
Spatial transcriptomics (ST) enables high resolution molecular profiling while preserving tissue architecture, creating new opportunities to investigate how disease-associated pathways are organized within tissues. However, existing analytical approaches largely focus on individual pathways or cell types and do not provide a unified framework for modeling spatially varying pathway interactions across tissue sections and anatomical planes. Here, we present an integrative framework, Spatial Fingerprints Analytics (SFinx), that combines reference-free deconvolution, pathway activity reconstruction, and Spatial Functional Data Analysis (SFDA) to map localized pathway activity and pathway phenotype interactions in complex tissues. Applying SFinx to10x Visium ST datasets from murine lupus nephritis, we reconstructed continuous spatial landscapes of pathway activity and disease-associated phenotypes across kidney sections. This approach identified anatomically restricted inflammatory domains characterized by coordinated activation of immune pathways and revealed substantial spatial heterogeneity in pathway crosstalk across renal compartments. Using generalized additive models with tensor-product splines, we quantified spatially varying associations between lupus nephritis and neutrophil activation pathways across tissue sections, uncovering regions with both positive and negative relationships that would be obscured by conventional bulk analyses. Multi-slice integration further demonstrated reproducible spatial interaction patterns while accounting for section-specific variability. Together, SFinx transforms mixed-spot transcriptomic measurements into interpretable spatial pathway landscapes and interaction maps, providing a general framework for identifying localized disease mechanisms. This approach reveals previously unrecognized spatial organization of inflammatory signaling in lupus nephritis and presents a broadly applicable strategy for studying spatially coordinated biological processes in autoimmune disorders, cancer, and neurodegenerative diseases.
bioinformatics2026-07-21v1Stratified Immune Profiling Uncovers Prognostic Heterogeneity Beyond MYCN Amplification and Age in Neuroblastoma
Magno, J. M.; Muzzi, J. C. D.; Resende, J. S. S.; Querne, L. B. P.; Alvarenga, L. M.; Cavalli, L. R.; Figueiredo, B. C.; Castro, M. A. A.Abstract
Neuroblastoma is the most common extracranial solid tumor in children, presenting remarkable clinical heterogeneity with survival outcomes ranging from spontaneous regression to aggressive progression. MYCN oncogene amplification and age at diagnosis are established prognostic factors that are typically treated as independent covariates in risk stratification, yet their joint influence on the tumor immune microenvironment remains poorly understood. Here we show that stratifying patients by both variables simultaneously reveals six reproducible immune subtypes with distinct transcriptional programs and prognostic significance. Consensus clustering of immunomodulatory gene expression profiles from 149 patients in the TARGET-NBL cohort identified subtypes whose survival trajectories differ significantly within clinical strata defined by MYCN status and age at diagnosis. A linear Support Vector Machine classifier trained on these subtypes, using immunomodulatory gene expression combined with MYCN amplification status and age at diagnosis as predictive features, achieved 97.2% accuracy and a Cohen's Kappa of 0.963 under 10-fold cross-validation, and generalized to an independent cohort of 493 patients (GSE62564). Kaplan-Meier analysis revealed significant survival differences across subtypes in both cohorts (TARGET-NBL: log-rank p = 0.0018; GSE62564: log-rank p < 0.0001). Single-sample gene set enrichment analysis identified differential activation of proliferative and immune response pathways across subtypes, consistent between both cohorts. These findings suggest that integrating MYCN amplification status and age at diagnosis as joint determinants of immune organization may reveal prognostic heterogeneity that is not fully captured when these factors are considered independently.
bioinformatics2026-07-21v1GeneAutomate: A Browser-Based, Integer-Indexed Platform for Dual-Gene-List Functional Annotation and Interactive Network Visualization
Singh, R. P.; Kumar, A.Abstract
Comparative interpretation of two gene lists, for example, two treatment arms, two tissues, or a discovery and a validation cohort, is a routine task in functional genomics. While several tools offer dual-list comparison (e.g., EnrichmentMap, RRHO packages), they typically require local software installation, R/Bioconductor, or manual reconciliation of separate single-list outputs. Most widely used web-based enrichment tools (DAVID, g:Profiler, Enrichr, ShinyGO, WebGestalt) are built around the analysis of a single gene list at a time, and those that support comparison often lack interactive, publication-ready visualization or depend on server-side query latency. Here we present GeneAutomate, a browser-based tool purpose-built for side-by-side comparison of two gene lists. GeneAutomate performs Over-Representation Analysis (ORA) against Gene Ontology (GO) and Reactome using an exact hypergeometric test with Benjamini-Hochberg false discovery rate correction, and Gene Set Enrichment Analysis (GSEA) when ranked (log2 fold-change) input is supplied, alongside Protein-Protein Interaction (PPI) subgraph extraction from BioGRID physical interactions. All reference data (Gene Ontology, Reactome, BioGRID, and NCBI/Ensembl identifier cross-references) are pre-compiled offline into a single integer-indexed database of approximately 32 MB for Homo sapiens, in which every gene identifier Ensembl ID, Entrez ID, official symbol, or alias is resolved to one canonical integer prior to any user query. This design removes live database round-trips from the runtime path, enabling fast, at-your-desk enrichment without installation or a server-side per-query bottleneck. The tool renders thirteen interactive, D3.js- and Cytoscape.js-based comparative visualizations, including a Rank-Rank Hypergeometric Overlap (RRHO) heatmap, a GO-slim "Radar/Spider" functional fingerprint, and chord/edge-bundled cross-talk diagrams that are, to our knowledge, not offered as an integrated set by any existing academic or commercial ORA/GSEA platform. GeneAutomate is an unfunded, individual student project developed with feedback from a professor, and is in its final stage of development. It requires no installation or login. We describe the tool's architecture, statistical methods, and comparative feature set relative to established academic tools (DAVID, ShinyGO, g:Profiler, Enrichr, WebGestalt, STRING, PANTHER, GeneMANIA, Cytoscape, clusterProfiler, GSEA, Metascape) and commercial platforms (IPA, MetaCore, Pathway Studio, iPathwayGuide, Partek Pathway), and we state candidly the current version's limitations, which are planned to be the added in next version: single-species (human-only) coverage, no upstream regulator analysis, and comparison currently limited to two (occasionally three) concurrent lists. GeneAutomate is available at https://geneautomate.tech/.
bioinformatics2026-07-21v1Extended t-cores for the de novo identification of transposable elements and other inexact repeats from short read RNAseq data
Darmon, S.; Mary, A.; Lacroix, V.Abstract
Transcribed repeats represent a major challenge in the de novo assembly of transcriptomes from short RNA-seq reads. Young transposable elements (TEs) and other inexact repeats create dense and ambiguous regions in the assembly graph, preventing the correct assembly of transcripts. In this paper, we introduce a fully de novo method based on the discovery of dense regions in the compacted De Bruijn graph (DBG) to identify such repeats directly from short reads RNA-seq data, without requiring a reference genome or repeat database. Our approach defines the extended t-cores, subgraphs of the DBG that capture the complex topology induced by highly expressed inexact repeats appearing in RNA-seq reads. Independently of its interest for transcriptome assembly, the proposed method appears to be effective for the de novo identification of repeats in transcriptomes. After classifying cores using sequence-based motifs to distinguish simple repeats from potential TEs, we demonstrate its potential for the de novo discovery of transposable elements. We validate the approach on a Mus musculus dataset using expressed TE consensus sequences, showing that extended t-cores correspond to known expressed TE families. We also illustrate its de novo discovery potential on a non-model species, Canis lupus familiaris, where the method was also able to recover known transposable elements.
bioinformatics2026-07-20v2FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations.
Clemente-Larramendi, I.; Hillion, S.; Cornec, D.; Jamin, C.; Foulquier, N.Abstract
Single-cell RNA sequencing (scRNA-seq) enables detailed characterization of cellular heterogeneity, yet understanding the full cellular and regulatory environment of complex tissues remains challenging. In the era of large single-cell atlases, this technology has become increasingly accessible, and datasets have grown in scale and statistical power. As a result, sample representation methods have emerged as a promising strategy to summarize patient-level biological variation. However, most existing approaches rely on unsupervised learning frameworks with ambiguous biological interpretability. Here we present a Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method. FloREN models single-cell data as a heterogeneous network integrating cells and genes together with gene regulatory and cell-cell communication relationships. Through condition-aware embeddings and interpretable attention networks, FloREN enables improved sample stratification and biomarker discovery. In addition, the framework supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
bioinformatics2026-07-20v1Metagenomic profiling of bacterial endosymbionts in wild mutant and permethrin-susceptible head lice shows an expanded microbiota in the resistant strain
Mohammadi, J.; Alipour, H.; Azizi, K.; Kalantari, M.; Moemenbellah-Fard, M. D.Abstract
Background Human head lice (Pediculus humanus capitis de Geer) are known to harbor diverse maternally inherited bacterial symbionts. These endosymbiotic bacteria may contribute to insecticide degradation, potentially helping lice withstand particular environmental pressures. Using next-generation sequencing (NGS), this study investigated the bacterial symbionts present in wild head-louse populations and characterized their phylogenetic relationships in mutant and permethrin-susceptible strains. Methods Head lice specimens were collected from 10 locations across Fars Province, Iran. Following DNA extraction, the samples were analyzed using polymerase chain reaction (PCR), and the resulting amplicons were sequenced to detect mutations in specimens from each location. Lice were then classified according to the presence or absence of mutations and subjected to NGS to characterize the symbiotic bacterial communities in mutant and putatively permethrin-susceptible strains. Results Mutant strains were detected at only three sampling stations. Bioinformatic analysis of the nucleic acid sequences revealed three exons and two introns, with an expected amplicon length of 582 bp. NGS analysis showed that Candidatus Riesia pediculicola (Arsenophonus), belonging to the phylum Proteobacteria, was the predominant bacterial genus. Actinobacteria and Firmicutes were the second- and third-most abundant phyla, respectively. Most of the remaining 25 bacterial taxa were associated with the mutant strain. Additionally, two previously unreported bacterial genera were deposited in GenBank. Conclusions The distinct distribution of Arsenophonus species between susceptible and mutant head-lice strains, along with the greater abundance of Escherichia, Shigella, Lawsonella, and Megamonas in mutant strains, highlights the need for advanced metagenomic analyses to determine how these endosymbionts may help their hosts withstand specific environmental disturbances.
bioinformatics2026-07-20v1Leveraging multiplicity in biologically informed neural networks to uncover disease heterogeneity
Gankin, D.; Beltrao, P.Abstract
Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map to named biological entities. Yet BINNs have been hard to train at biobank scale, and the reliability of their interpretations remains largely untested. Here we present a fast BINN implementation trained on UK Biobank genotype and plasma proteomics data from about 500,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph topology, because node degree and layer position influence the scores. Normalization reduces this bias but can weaken enrichment for known disease genes. Second, BINNs show substantial predictive multiplicity, that is, independently trained models with identical architecture and data reach similarly accurate solutions while prioritizing different genes and pathways. Although this multiplicity makes single-model explanations unstable, the range of interpretations can itself reveal disease biology. Across 100 replicate BINNs for type 2 diabetes, we find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways, mirroring known disease heterogeneity. Thus, analyzing the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms.
bioinformatics2026-07-20v1GDTR: Layer-wise Settling Depth Reveals Biological Grammar in Genomic Foundation Models
Cho, Y.; Kang, J.; Park, S.; Kim, S.Abstract
Genomic foundation models capture sequence regularities, yet existing interpretability tools rarely ask where in the layer stack a biological grammar becomes stable. We introduce GDTR, the Genomic Deep-Thinking Ratio, a training-free residual-stream lens that assigns each nucleotide token a settling depth : the first layer at which its representation stabilizes against the post-final-norm reference. On Evo 2 7B, splice donor and acceptor sites settle approximately two layers earlier than intronic contexts; enhancer-like cCREs show a smaller but measurable shift; and a chromosome 22 calibration transfers to held-out chromosome 17. Perturbing canonical splice donors shows that the signal is bidirectional: disrupting the central GT motif deepens settling, whereas shuffling the flanking grammar makes the preserved motif settle earlier. Differential GDTR further reveals consequence-associated peak-disruption depths across ClinVar variants, with synonymous substitutions peaking deepest but showing broad class overlap. GDTR therefore provides a layer-wise interpretability axis for genomic foundation models, complementary to existing prediction and variant-scoring tools.
bioinformatics2026-07-20v1A Label-Free Multi-Metric Pipeline for Benchmarking Single-Cell RNA-Sequencing Clustering and Testing the Reproducibility of Cell-Type Heterogeneity
Yousef, Z.; Simone, J.; Klein, D.; Cho, H.; Bhuyan, A.; Bhatt, P.; Wu, J.; Wang, H.; Cai, L.Abstract
A discovered sub-population from single-cell transcriptomic data is only meaningful if it is reproducible, yet clustering is usually done with one method on one embedding and rarely tested. We present a label-free, multi-metric pipeline that reframes clustering as an auditable, methods-blind decision and separates two notions of stability that are commonly conflated: reproducibility under cell resampling (bootstrap) and reproducibility under re-embedding (retraining the representation). The pipeline evaluates seven clustering configurations across cluster counts using five non-redundant quality metrics. As a whole-dataset control on a mouse retinal atlas, it recovers an eight-cell-type annotation at 96.3% accuracy (adjusted Rand index, ARI = 0.91) without labels. We then validate the discovery mode on two cell types with opposite ground truth. On bipolar cells, which have well-established subtypes, the pipeline accepts the sub-structure: across-embedding reproducibility rises with cluster number to a high plateau (mean pairwise ARI ~0.93 near the ~15 known bipolar subtypes), with quality metrics improving in parallel. On rod photoreceptors, treated as homogeneous, it rejects over-clustering: the metric-selected partition passes a bootstrap-stability check but is not reproducible when the embedding is retrained (mean pairwise ARI = 0.69), and the metrics do not improve with cluster number. On synthetic data, the test recovers real structure down to a 5% subpopulation while rejecting null data (high sensitivity and specificity). Bootstrap stability alone is therefore insufficient evidence for sub-population; the across-embedding test discriminates real sub-structure from over-clustering and applies to any cell type as a reproducible alternative to single-method, single-embedding clustering.
bioinformatics2026-07-20v1