Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Metagenomics analysis for microbial ecology investigation on historical samples: negligible effect of host DNA and optimal analysis strategies
Ng, S.-K.; Gutaker, R.Abstract
Microbiome composition and function are strongly influenced by environmental factors, with major shifts driven by intensified anthropogenic pressures over the past centuries. This timeframe extends beyond the scope of traditional experimental or longitudinal studies commonly used to investigate microbiome dynamics. The historical samples might provide important insights into the mechanistic consequences of anthropogenic pressures and the potential shift in microbial diversity and composition. Despite their vast potential, historical samples available in museums and herbaria worldwide remain underutilized for exploring host-microbiome interactions across broad temporal and spatial scales due to incompatibilities with standard analytical pipelines and limited understanding of optimal classification parameters. While host DNA removal has conventionally been considered essential for taxonomic assignment of metagenomic reads, and might be of particular importance when processing degraded DNA, this step is impractical for specimens with no reference genome available for host species. Here, we show that host DNA content has negligible impact on microbial data analysis with empirical and simulation datasets. Since DNA molecules from historical samples are highly fragmented and uneven in length, we further analysed the impact of k-mer value on the classification of metagenomic reads from historical samples. To improve recall rate, we proposed a simple two-step approach in which reads are classified with two annotation databases constructed with a long and a short k-mer values. Through simulation and published datasets, we demonstrated that this approach outperforms single-step workflows in effectively recovering microbial signals from reads in a wide range of length. Together, this study provides a solid foundation for incorporating natural history collections into host-associated microbiome research, offering valuable insights into the long-term effects of anthropogenic change on microbial communities.
bioinformatics2026-08-19v5Automating the Construction of Contextualized Biomedical Knowledge Graphs for Scientific Inference
Zheng, Y.; Liu, W.; Zeng, B.; Feng, Y.; Du, X.; Zhou, L.; Li, Y.Abstract
Biomedical interactions are inherently dynamic, often shifting or even reversing under specific physiological states. However, existing extraction methods simplify these complex mechanisms into context-agnostic binary associations, resulting in semantic loss and contradictory evidence. Here, we present AutoBioKG, an end-to-end framework that constructs context-aware knowledge graphs by leveraging composite triples to encode environmental conditions and entity attributes alongside core relationships. Powered by an open information extraction model trained on BioOpenIE and further refined through self-training with pseudo-labels from unlabeled literature, the framework exhibits broad generalization. Notably, AutoBioKG achieved the highest zero-shot F1 across DDI, ChemProt, and BioRED, outperforming the best-performing baseline on each benchmark by 3.6-17.8 percentage points. Furthermore, AutoBioKG-derived graphs outperformed existing approaches on yes/no, factoid, and list questions in the BioASQ biomedical question-answering evaluation, particularly for queries requiring fine-grained contextual information. Together, these results support AutoBioKG as a scalable framework for transforming unstructured literature into structured, context-aware biomedical knowledge.
bioinformatics2026-08-19v3PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking
Arora, R. K.; Chen, L. T.; Du, M.; Marks, D.; Church, G.Abstract
General-purpose frontier language models are being increasingly utilized for protein-design work, yet their ability to understand and evaluate variant effects remains unclear. Here, we introduce PG-LLM, a benchmark comprising 276 protein-variant prioritization tasks: 217 from ProteinGym and a temporally held-out set of 59 from recently published studies. Each task follows the same format: a language model is asked to rank a list of variant sequences given only the wild-type protein sequence and an assay description with no access to tools, multiple-sequence alignments, or protein structures. We evaluate thirteen language models and 95 published protein predictors on the same variants with the same evaluation metric. Claude Opus 5 (Max) and GPT 5.6 Sol (Max) are the best performing LLMs with Spearman correlations of {rho} = 0.406 and 0.402 respectively. Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at {rho} = 0.411, but remains below the leading predictor VenusREM at {rho} = 0.523. We observe that variant-ranking performance scales with test-time compute across GPT, Claude, and Gemini models, but gains taper before closing the gap to specialist protein predictors. To address contamination risk, we create a held-out evaluation set with 59 DMS assays from 19 studies whose scores first became public after January 2026. On this set, we observe performance and test time compute scaling trends similar to those on the 217 tasks derived from ProteinGym. PG-LLM shows that tool-free language models capture substantial protein-variant signal, outperforming many sequence-based predictors while remaining below the strongest specialized models.
bioinformatics2026-08-19v3Automating the Construction of Contextualized Biomedical Knowledge Graphs for Scientific Inference
Zheng, Y.; Liu, W.; Zeng, B.; Feng, Y.; Du, X.; Zhou, L.; Li, Y.Abstract
Biomedical interactions are inherently dynamic, often shifting or even reversing under specific physiological states. However, existing extraction methods simplify these complex mechanisms into context-agnostic binary associations, resulting in semantic loss and contradictory evidence. Here, we present AutoBioKG, an end-to-end framework that constructs context-aware knowledge graphs by leveraging composite triples to encode environmental conditions and entity attributes alongside core relationships. Powered by an open information extraction model trained on BioOpenIE and further refined through self-training with pseudo-labels from unlabeled literature, the framework exhibits broad generalization. Notably, AutoBioKG achieved the highest zero-shot F1 across DDI, ChemProt, and BioRED, outperforming the best-performing baseline on each benchmark by 3.6-17.8 percentage points. Furthermore, AutoBioKG-derived graphs outperformed existing approaches on yes/no, factoid, and list questions in the BioASQ biomedical question-answering evaluation, particularly for queries requiring fine-grained contextual information. Together, these results support AutoBioKG as a scalable framework for transforming unstructured literature into structured, context-aware biomedical knowledge.
bioinformatics2026-08-19v2GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations
Andrade, R. L.; Fiuza, T. d. S.; Ferraz, R. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.Abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearsons r improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTX-GUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
bioinformatics2026-08-19v2Method Choice, Not Biology, Determines In Silico Perturbation Results: A Systematic Evaluation of Eight Methods Across Four Datasets
Wenjie, G.; Wu, S.; Hu, G.; Yang, Z.; Wang, Z.; Cai, J.; Mao, J.Abstract
Most in silico perturbation methods for single-cell transcriptomics have been validated only on individual datasets, leaving their reliability and generalizability unknown. Through systematic cross-method, cross-dataset benchmarking of eight methods spanning six mathematical frameworks across four datasets, we find that six of eight methods--including widely used VAE-based and tensor decomposition approaches--fail to produce detectable transcription factor (TF)-to-pathway signals. Only CellOracle and DDIM consistently detected TF-to-glycolysis directional regulation. Cross-pathway analysis in PBMC monocytes revealed biologically coherent TF-pathway associations beyond glycolysis (SPI1[->]glycolysis 4.4x enrichment, FOS[->]AP-1 targets 4.4x), with SOX9 serving as a biological specificity control (no pathway enrichment). Method choice alone could reverse biological conclusions: DDIM and scTenifoldKnk rankings were significantly anti-correlated ({rho}=-0.811, p=0.027). CRISPRi Perturb-seq validation in K562 cells confirmed TF knockdown suppresses glycolysis gene expression (JUN {delta}=-1.72, CEBPB {delta}=-1.59, SPI1 {delta}=-1.57, FOS {delta}=-0.70), but CellOracle-predicted perturbation directions did not match experimental directions (40.9% agreement, not different from chance), revealing a fundamental gap between steady-state correlation and causal perturbation. Diagnostic analyses using VAE latent space profiling, correlation distribution comparison, and gene-gene graph analysis identified distinct failure modes in unsuccessful methods: VAE latent space competition (STAT3 signal-to-noise 0.44 vs. SPI1 4.25), correlation noise (TF-glycolysis |r|=0.038 indistinguishable from background |r|=0.047), and graph non-specificity (0.84x enrichment). A controlled ablation experiment showed that adding a GRN prior to DDIM did not improve target recall (delta=0 for all TFs), confirming that performance differences are multi-factorial. These findings establish preliminary guidance for method selection, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements ([≥]500 cells, [≥]1,000 HVGs).
bioinformatics2026-08-19v1reserBUGS: A reservoir computing framework for probabilistic forecasting of ecological abundance time series
Mohedano-Munoz, M. A.; Galeano, J.; Pastor, J. M.; de Aledo, J. G.; Bartomeus, I.; Allen-Perkins, A.Abstract
Forecasting species population dynamics is a central challenge in computational ecology, yet existing approaches rarely combine flexible nonlinear modelling, support for count-based ecological data, and systematic uncertainty quantification within a single, scalable framework. Here we introduce reserBUGS, an open-source Python framework for ecological forecasting based on reservoir computing, a recurrent neural network architecture in which only a simple readout layer is trained while a fixed high-dimensional dynamical system encodes temporal memory and nonlinear dependencies. reserBUGS integrates species abundance time series with environmental covariates retrieved automatically from global climate products, generates probabilistic ensemble forecasts, and provides tools for forecast evaluation and reliability assessment. We evaluated reserBUGS using insect abundance time series from available biodiversity monitoring datasets, comparing its performance against seven statistical and machine-learning baselines over one- to five-year forecast horizons. Reservoir-based models consistently outperformed alternatives in both predicting future abundance and capturing forecast uncertainty, with environmental predictors increasing the proportion of stable forecasts and contributing additional predictive value beyond historical abundance dynamics alone, particularly at 3-4-year forecast horizons. Probabilistic forecasts further enabled the identification of conditions associated with reduced predictive skill, providing a practical basis for communicating forecast confidence to end users. While default configurations already achieved competitive performance across a taxonomically and geographically diverse set of time series, hyperparameter optimisation revealed substantial room for performance gains through series-specific tuning. reserBUGS offers a computationally efficient and extensible framework for ecological forecasting that is well suited to the short, heterogeneous time series typical of biodiversity monitoring programmes. Its combination of flexible nonlinear modelling, probabilistic uncertainty quantification, and automated environmental data integration addresses key practical barriers to the adoption of modern forecasting methods in conservation and ecological research.
bioinformatics2026-08-19v1Tractography from Serial Optical Coherence Tomography: How and Why?
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.Abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Serial optical coherence tomography (S-OCT) is an imaging modality relying on the intrinsic contrast of a sample. When applied to brain tissues, the S-OCT contrast is primarily driven by the myelin reflectivity. Due to its high resolution, on the order of microns, and its 3D nature, S-OCT offers promise for studying WM connections at the microscale. However, while other microscopy imaging modalities have been shown to enable tractography, whether the reflectivity contrast from S-OCT supports the reconstruction of long-range WM fascicles at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We improve microscale orientation distribution functions (ODF) estimation by implementing a sliding-window formulation allowing the estimation of ODF at S-OCT resolution, and use apodized Dirac delta functions for reducing unwanted interference. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that using multiscale Frangi filters for estimating ODF outperforms structure tensor analysis. We also show that particle filtering tractography with anatomical constraints enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing the thalamocortical white-matter projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles visible at the microscale, and that these connections are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
bioinformatics2026-08-19v1Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling
Burq, M.; Stepec, D.; Kim, C.; Cimermancic, P.Abstract
Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.
bioinformatics2026-08-19v1Novel biologically relevant small RNA-sequencing alignment tool LevenMap for alignment to database of non-coding RNAs
Dlugas, H.; Dyson, G.; Dombkowski, A.; Kim, Y.; Gurdziel, K.; Boerner, J. L.; Bock, C.Abstract
A crucial aspect of the bioinformatics workflow in small RNA-sequencing is the alignment of reads to a database of reference ncRNAs. Alignment algorithms such as Bowtie, Burrows-Wheeler Aligner (BWA), and Spliced Transcripts Alignment to a Reference (STAR) - which are designed for aligning reads to a reference genome - are typically used. Aligning short RNA-sequenced reads to a database of non-coding RNAs (ncRNAs) is fundamentally a different task than aligning longer reads to a genome due to ncRNAs (i) having roughly the same number of nucleotides as the reads being aligned and (ii) being subsequences of other ncRNAs. To account for these differences, we developed the novel alignment algorithm LevenMap. Of all reads which exactly matched a reference ncRNA in a publicly available dataset, LevenMap aligned 100.0% of them to their respective ncRNA while all other aligners mapped less than 40% of these reads to their corresponding ncRNA. Furthermore, the mean ratio (length of read) / (length of corresponding reference ncRNA) of all aligned reads was 1.0 and 0.998 for LevenMap with at most zero and one mismatch(es) allowed, respectively; this ratio was no more than 0.51 for all other aligners. Overall, LevenMap is designed to account for the nuances of aligning small RNA-sequencing data to a database of reference ncRNAs and yields more biologically relevant counts compared to traditional aligners in this context. LevenMap is free and publicly available on GitHub: https://github.com/hdlugas/LevenMap.
bioinformatics2026-08-19v1Quantifying Crop Disease Trait Dynamics through Longitudinal Imaging and Temporal Analytics
Ewen, A.; Godoy Mendez, R.; Al-Shanoon, K.; Omoluabi, D.; Samarasinghe, A.; Oviedo-Ludena, M. A.; Chimbo Huatatoca, K.; Glor, K.; Nabetani, K.; Kutcher, R.; Wang, L.; Stavness, I.; Jin, L.Abstract
Reliable and objective phenotyping is essential for plant breeding programs to characterize genetic variation and accelerate crop improvement. Conventional disease assessment relies on expert visual scoring, which is labor-intensive, subjective, and prone to inter- and intra-rater variability. Although image-based phenotyping methods have been proposed, many require manual intervention, specialized imaging setups, or single time-point measurements, limiting their ability to capture disease progression over time. Here, we present a pipeline for longitudinal plant disease phenotyping that quantifies wheat stripe rust and leaf rust progression from time-series images. The pipeline performs semi-automated leaf and automated pustule segmentation from images acquired it in situ, enabling objective disease severity estimation with minimal user intervention and without requiring solid backgrounds or manual leaf manipulation or detachment. By extracting temporal traits, including disease severity trajectories and standardized area under the disease progress curve, the method provides a comprehensive characterization of disease development throughout infection. Association between automated and expert assessments was moderate for stripe rust (R2=0.58) and strong for leaf rust (R2=0.85), while expert inter-rater reliability was moderate for both diseases (ICC = 0.675 and 0.800, respectively). The proposed approach establishes a scalable and reproducible framework for longitudinal disease phenotyping in controlled environments, with broad applications in disease resistance screening and crop breeding.
bioinformatics2026-08-19v1An Exact-Residue Atlas of Opioid Receptor Wiring and Rewiring across Ligand and Transducer Contexts
Nael, M.; Alakonda, L.; Elokely, K.Abstract
Opioid-receptor structures span four human receptor subtypes, diverse ligands, signaling partners, and experimental constructs. We curated 86 human opioid-receptor structures representing 84 independent experimental maps, with one unique experimental data set counted once for structure-level inference, and analyzed them using our in-house StrucMind platform. StrucMind constructs exact Ballesteros-Weinstein (BW) contact graphs, meaning residue-contact networks restricted to unambiguous generic BW positions. Relative to active transducer-bound structures, structures classified as inactive showed 2.49% lower mean contact similarity and 34.84% more rewired contacts, where rewiring is the static set of contacts gained or lost between two structures. Among 77 maps with a resolved selected-ligand site, changed contacts were 15.28% direct to the site, 40.13% adjacent at one graph edge, and 44.59% connected-distal at a finite graph distance greater than one. The deposited-water analysis identified 137 receptor-proximal waters. Sixty-one contacted at least two protein residues, including 38 that bridged at least two exact-BW residues; a separate ligand-contact branch contained 10 waters contacting both selected ligand and receptor, only 3 of which belonged to the 38-water set. None of 32 component-association tests survived global correction. For peptide versus small molecule, the smallest nominal p value among four outcomes corresponded to 6.18% lower shared-contact distance root-mean-square deviation (p=0.00989; q=0.3165, where q is the adjusted p value). The atlas supports bounded, testable hypotheses, not causal component, hydration, or efficacy mechanisms.
bioinformatics2026-08-19v1Signature Recontextualization: Mapping perturbational signatures across biological contexts
Chen, A. D.; Girke, T.; Monti, S.Abstract
Perturbational transcriptomics is a powerful tool for understanding gene function and drug effects, yet predicting how perturbations manifest across different biological contexts remains a central challenge, limiting translation from model systems to clinically relevant tissues. Despite growing interest in this problem, benchmarking efforts have been hindered by inconsistent evaluation tasks, heterogeneous metrics, and limited assessment across perturbation types and biological systems. Here, we introduce a benchmarking framework for cross-context perturbation-signature prediction (a task we define as signature recontextualization), grounded in explicit definitions of the prediction task, target-data availability, and evaluation metrics centered on signature recovery. The framework evaluates prediction performance across three target-context data regimes: (1) control only, where only control profiles from the target context are measured; (2) low coverage, where a limited subset of perturbations in the target context are measured; and (3) high coverage, where most perturbations in the target context are measured. This design enables systematic assessment of how prediction performance depends on target-context sample size while providing a standardized basis for comparing methods. We evaluate newly developed projection-based (projectCor) and network-based (netProp) methods alongside deep learning-based foundation models (scGPT, STACK) and statistical baselines. The benchmark spans four diverse perturbational datasets: CRISPR knockdowns and drug perturbations in cell lines, plus in vivo chemical perturbations in rat tissues from DrugMatrix, extending evaluation beyond isolated cell-line models to tissue-level responses. Across tasks, projection and network propagation approaches show strong flexibility across perturbation types and biological contexts, and in several cases match or exceed the performance of deep learning and foundation models, suggesting that model complexity does not inherently improve cross-context generalization. We further show that perturbation predictability varies substantially with pathway conservation, transcriptional response strength, and baseline similarity between source and target contexts. All datasets, methods, and evaluation utilities are released as an open-source R package (sigRecon), providing a foundation for reproducible benchmarking and future method development.
bioinformatics2026-08-19v1Flexibility Drives Information Flow in Proteins: Fluctuation Potential Gradients Dictate Directional Entropy Transfer
Senguler Ciftci, F.; Erman, B.Abstract
Allosteric communication in biomacromolecules is fundamentally governed by thermal fluctuation gradients, yet standard Gaussian Network Models (GNMs) treat atomic contacts as uniform, binary couplings without differentiating core constraints from solvent-exposed surface flexibility. Here, we present an analytical matrix framework that incorporates continuous distance-dependent weighting into the Kirchhoff matrix L . This formulation captures the steep steric constraints of hydrophobic core packing versus peripheral surface loops while strictly recovering the classic unweighted GNM as a high-temperature limit (T[->]{infty} ). Using Schur complements of partitioned joint covariance matrices, we show that conditional fluctuation variances and higher-order entropy-transfer terms reduce analytically to exact ratios of submatrix determinants (covariance minors), eliminating the need for fitting parameters or molecular dynamics trajectories. Applied to KRAS (PDB: 6GOD), this framework constructs an integrated directional entropy-transfer asymmetry map. Order-1 minors (h(i)=Kii ) establish a single-node fluctuation potential gradient, while order-2 minors (Rij ) define pairwise channel bandwidths. Higher-order minors show multi-body spatial coupling: order-4 minors identify rigid core residues such as Phe156 as strategic interlobe relay hubs linking Lobe 1 and Lobe 2, and an order-3 triad cooperation index demonstrates that signal transmission from Switch II (Gln61) to Gly60 and Phe156 converges on a single, mechanically integrated allosteric sector. By deriving directional information flow directly from experimental atomic displacement parameters, this approach establishes a rigorous, computationally efficient framework for mapping allosteric networks across structural ensembles.
bioinformatics2026-08-19v1Modality-chain reasoning enables multimodal protein modelling and design
Liu, C.; Chao, L.; Ji, S.; Wang, H.; Zhou, G.; Zheng, J.; Hong, D.; Guo, Y.; Jiang, T.; Gao, Z.; Yang, M.; Pan, J.; Li, S. Z.; Zhang, X.Abstract
Reasoning has emerged as a central capability of large language models, yet how it should be formulated for scientific foundation models remains unclear because scientific knowledge is distributed across interdependent, domain-specific representations. Here we introduce modality-chain reasoning, which organizes representations into ordered computational chains, each conditioning prediction or generation of the next. Based on this principle, we develop ProteinReasoner, a multimodal generative protein foundation model that sequentially connects amino acid sequence, evolutionary constraints and three-dimensional structure within a shared autoregressive architecture. Across zero-shot structure prediction, inverse folding and fitness prediction, ProteinReasoner outperformed two multimodal protein foundation models, while controlled comparisons supported the functional contribution of the modality chain. We further extended this principle beyond pretraining: reorganizing the chain across successive structural states enabled multiple-conformation prediction, while introducing experimental feedback as an additional modality enabled an in-context learning paradigm for protein optimization without target-specific parameter updates. In particular, across thermostability and affinity-maturation evaluations, this paradigm improved over matched fine-tuned models and showed stronger mean performance than target-specific active-learning baselines. These results establish modality-chain reasoning as a unified and effective foundation-modelling strategy in protein science. More broadly, they suggest a general route towards reasoning across interdependent representations in other scientific domains.
bioinformatics2026-08-18v3antigen-prime: Simulating coupled genetic and antigenic evolution of influenza virus
Thornton, Z. T.; Tran, T.; Figgins, M. D.; Huddleston, J.; Bedford, T.; Matsen, F. A. D.; Haddox, H. K.Abstract
Seasonal influenza virus undergoes rapid antigenic drift to escape population immunity. Computational methods can be used to organize viral genetic diversity into antigenically similar variants and estimate variant-specific growth rates. However, benchmarking these methods is challenging because it can be di[ffi]cult to accurately quantify antigenicity and growth rates in nature. Simulating viral evolution using defined selective pressures can provide ground-truth data for benchmarking. But, existing simulators do not link genetic sequences to antigenic phenotypes under selection from host populations. Here, we present a forward-time epidemic simulator called antigen-prime that links these factors. We use it to simulate viral evolution over 30 years and validate the simulation recapitulates genetic and antigenic patterns observed in natural influenza evolution. We then use the simulated data to benchmark methods for assigning variants and estimating their growth rates. We evaluated a sequence-based and a phylogenetics-based method for variant assignment, finding the former was slightly more e[ff]ective at separating viruses into antigenically distinct groups. We also evaluated methods for estimating variant growth rates in one-year sliding windows. Estimates were accurate in most windows, but highly inaccurate in several others. Examining high- error windows revealed several examples of a previously unreported failure mode. In all, antigen-prime provides a simulation framework to benchmark models of influenza evolution, and could be used to help guide future development of these models. The source code is openly available at https://github.com/matsengrp/antigen-prime.
bioinformatics2026-08-18v2Protal: Ultra-fast metagenomic profiling and strain-resolved analysis
Fritscher, J.; Duncan, A.; Hildebrand, F.Abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal - profiling through alignment - an ultra-fast alignment-based method for species- and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom bench marks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler, respectively). At the strain level, protal reconstructs intraspecific phylogenetic relationships among detected bacteria with similar accuracy as StrainPhlAn 4; because protal produces precise alignments, the phylogenies can be de novo produced without reliance on reference strain collections. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
bioinformatics2026-08-18v2MOFTy: Multimodal Gaussian Process Factor Analysis with Numerical Information Field Theory
Neumann, M.; Arras, P.; Kaster, A.-K.; Ott, A.Abstract
Multimodal Gaussian process factor analysis provides a flexible framework for dimensionality reduction in temporally or spatially resolved omics data. Existing approaches, however, typically rely on pre-specified Gaussian process kernel families and do not explicitly separate each latent factor into a component capturing gradual, smooth variation and a complementary component capturing fine-scale, non-smooth variation. Here, we present MOFTy, a Bayesian multimodal factor analysis framework based on numerical information field theory (NIFTy) that replaces fixed kernel families with the flexible correlated field model in NIFTy and enables explicit additive component separation within each latent factor with quantified uncertainty. NIFTy has been successfully applied to high-resolution Bayesian imaging in astrophysics and facilitates scalable, curvature-aware variational inference for efficient posterior approximations. We validate MOFTy on simulated data; applications to published multi-omics data demonstrate that MOFTy disentangles latent spatial structures by separating smooth gradients from localized fine-scale heterogeneity in human glioblastoma and recovers cross-modal patterns in a mouse gastrulation dataset.
bioinformatics2026-08-18v2Population dynamics of ribozymes during in vitro selection
Brooks, E. M.; Kotlin, N. R.; Weiss, Z.; DasGupta, S.Abstract
Ribozymes are central to models of RNA-based primordial life. Understanding how ribozymes emerge from RNA populations under changing selection pressures can help delineate the biochemical constraints and capabilities of early RNA-based biology. In vitro selection has been instrumental in isolating ribozymes with diverse functions from combinatorial populations; however, their population dynamics during selection remain poorly characterized. By analyzing high-throughput sequencing data produced from all rounds of an in vitro selection experiment, consisting of over 5 million unique sequences, we tracked how sequences rose and fell in abundance during selection. We found that the most abundant sequences emerged early and maintained their dominance, suggesting that only a few cycles of selection-amplification may be sufficient to isolate well-adapted ribozymes. Characterization of the catalytic activities of the dominant sequences in each selection round indicated that abundance did not always predict activity. This observation was reinforced through nucleotide conservation analysis and activity assays of close sequence variants of the most dominant ribozyme. Our analyses show that sequence complementarity with the substrate emerged as a common feature as early as the first round and persisted throughout, highlighting early convergence of diverse sequence populations on a beneficial catalytic strategy. By combining bioinformatics and experimental approaches, our work presents a detailed description of ribozyme population dynamics during an in vitro selection experiment and provides a high-resolution view of how functional RNA sequences compete for survival and propagation under selection pressure.
bioinformatics2026-08-18v2Cell-Hub: a graphical interface for end-to-end single-cell RNA sequencing analysis
macaux, g.; Di Gallo, M.; Taglietti, V.; Amthor, H.; Maire, P.Abstract
Abstract Single-cell and single-nucleus RNA sequencing have become increasingly widespread, creating a significant demand for accessible analysis tools in research laboratories. Despite this need, the bioinformatics expertise required for such analyses remains rare. Cell-Hub addresses this gap by enabling single-cell data analysis for all researchers, regardless of computational background. Cell-Hub is a comprehensive, free, and open-source framework built on R/Shiny and distributed as a Docker image, integrating Seurat 5, CellChat 2, and Monocle 3 within a unified graphical interface. It supports all essential steps of single-cell RNA-seq analysis: data loading, quality control, normalization, clustering, multi-dataset integration, differential expression, and biomarker detection. Cell-Hub further incorporates ligand-receptor interaction inference powered by GaspouDB, a consolidated database of 11,563 mouse and 9,604 human interactions derived from CellChat, CellPhoneDB, CellTalkDB, and MultiNicheNet as well as trajectory inference via Monocle 3. All analyses produce publication-ready visualizations with flexible export options. By integrating these analytical frameworks into a single, intuitive interface requiring no programming expertise, Cell-Hub represents a significant step toward democratizing single-cell genomics for the broader research community.
bioinformatics2026-08-18v2HNSW-MS: Hierarchical Graph Indexing Enables Accurate Real-Time Mass Spectral Similarity Search at Repository Scale
Semenov, A.; Roberts, A. M. P.; Coler, E. A.; Gupta, S.; Kopylova, E.; Melnik, A. V.; Boginski, V.; Aksenov, A. A.Abstract
Spectral similarity comparison is the basis of mass spectrometry-based metabolomics, underpinning library matching, molecular network construction, and repository searches such as MASST. Public repositories now exceed billion spectra, and exhaustive pairwise comparison is increasingly limiting. We introduce HNSW-MS, which implements Hierarchical Navigable Small World graph indexing natively for mass spectra. It operates directly on GC-MS and LC-MS/MS spectra without embedding, so scores are identical to those of exhaustive comparison and results are fully reproducible. On the benchmark dataset of 8.4 million LC-MS/MS spectra, HNSW-MS is up to 900-fold faster than linear search, recall@1 and recall@10 achieve over 90%, with the recall-speedup tradeoff tunable at query. Such rapid retrieval allows for any query spectrum to be near-instantaneously integrated into a global molecular network. We demonstrate specific examples of use of global networking for molecular discovery, including uncovering a new pathway of stenothricin biosynthesis and a new dipeptide conjugate bile acid.
bioinformatics2026-08-18v2Assessing the Reliability of LLM-Generated Phenotype-Genotype Associations Through External Validation
Sun, C.; Xin, Y.; Zeng, S.; Sunthankar, S. D.; Su, W.-C.; Lynn, J.; Mundo, S.; Babanejad, M.; Feng, Q.; Wei, W.-Q.Abstract
Background: Phenotype-genotype associations underpin precision medicine by enabling disease prevention, early diagnosis, risk stratification, therapeutic target discovery, and personalized treatment. However, the rapid growth of scientific evidence has made manual curation of these associations increasingly labor-intensive, time-consuming, and incomplete. Large Language Models (LLMs) offer a potential path to scalable genomic generation and synthesis of this knowledge, but their ability to accurately identify phenotype-genotype associations and the extent to which these outputs are supported by established genomic knowledge bases remain unclear. Materials and Methods: Four LLMs, Claude Sonnet 4.6, DeepSeek V4 Flash, Gemini 3 Flash Preview, and GPT-5.5, were benchmarked on six zero-shot task categories covering forward and reverse phenotype-gene and phenotype-SNP generation. A total of 4,196 associations were identified from curated inputs and evaluated through a multistage external verification pipeline comprising phenotype normalization, ontology mapping, genomic identifier validation against Ensembl, and evidence verification using both the GWAS Catalog and OMIM. Associations were assigned a fused evidence level of strong, moderate, weak, or none. Results: Overall, 74.19% of generated associations were matched to at least one external genomic knowledge base; 9.15% received strong support and 54.46% moderate support. Phenotype-gene associations were more verifiable than phenotype-SNP associations (strong or moderate: 67.19% vs 54.06%). Among existing associations, Claude Sonnet 4.6 achieved the highest overall strong or moderate rate (69.2%), followed by GPT-5.5 (65.1%), DeepSeek V4 Flash (61.7%), and Gemini 3 Flash Preview (56.9%). Conclusion: LLMs can support scalable generation of candidate phenotype-genotype associations. Performance varied substantially by relation type and was lower for SNP-level and rare disease associations, highlighting both the limitations of current genomic resources and the need for rigorous validation pipelines. Key words: large language model; clinical genomics; phenotype-genotype association; external validation; genomic knowledge base
bioinformatics2026-08-18v1Efficient Game-Theoretic Explanations for Tree-Based Ensembles via Owen Values
Koh, H.Abstract
Shapley-value-based explanations, notably SHAP (SHapley Additive exPlanations), have gained prominence as a principled game-theoretic framework for local explanations and global feature importance. While exact Shapley value computation is exponential in feature count, TreeExplainer exploits the recursive structure of decision trees to achieve polynomial-time computation for tree-based ensembles. In many scientific applications, however, features are naturally organized into a priori groups reflecting domain knowledge, requiring explanations both across and within groups. The Owen value extends the Shapley value through a two-stage allocation rule that incorporates group structure while preserving fairness properties; yet, efficient algorithms for its computation remain limited. In this paper, we propose exact and Monte Carlo algorithms for computing Owen values in tree-based ensembles by combining hierarchy-guided group aggregation with tree-aware dynamic programming. The exact algorithm computes Owen values without sampling under the path-dependent characteristic function, which approximates the conditional expectation, whereas the Monte Carlo algorithm provides a scalable approximation that is unbiased for any prespecified sampling budget and converges almost surely as the sampling budget increases. We also provide global importance measures and visualization tools for structured, multi-resolution explanations. The proposed algorithms and tools are collectively referred to as TreeOwen. Through simulation experiments, we demonstrate the numerical accuracy and substantial computational gains of TreeOwen. We illustrate its practical utility using immunotherapy metagenomic data, showing how microbial genera (groups) and species (features) contribute to patient recovery.
bioinformatics2026-08-18v1Benchmarking Docking Protocols for GPCR Allosteric Modulators
Thompson, T. D.; Miao, Y.Abstract
G protein-coupled receptor (GPCR) allosteric modulators (AMs) offer significant therapeutic advantages over orthosteric drugs, yet structure-based virtual screening lacks validated protocols accounting for the conformational complexity of GPCR allosteric sites. We benchmark docking protocols using PDB experimental structures and structural ensembles derived from Gaussian accelerated Molecular Dynamics (GaMD) simulations across four Class A GPCRs (including the muscarinic M2 and M4 receptors, the {beta}2-adrenergic receptor, and the C-C chemokine receptor type 2) with four programs (Glide HTVS, AutoDock Vina, DOCK3.8, and Boltz-2) against experimentally validated modulator libraries and property-matched decoys. GaMD ensemble docking improved early AM enrichment across all four targets under at least one program. Glide ensemble docking was the only protocol to consistently improve early AM recovery across all four targets, ranking known actives almost exclusively within the top 0.5% of compounds at CCR2 and improving M2R active recovery nearly 9-fold relative to the PDB structure. GaMD free-energy landscape topology governed ensemble re-ranking strategy selection: population-skewed landscapes favored top binding energy ranking (BEmin) while flat, multi-populated landscapes favored average binding energy ranking (BEavg), and at targets with dominant low-energy states, a single GaMD cluster matched or exceeded full ensemble or PDB performance. Taking the union of top percentile hits identified by both ensemble re-ranking methods, BEmin / BEavg, maximizes chemical diversity at the earliest percentiles. Program-specific scaffold recovery biases further motivated a consensus BEmin / BEavg approach to maximize hit diversity. The Boltz-2 deep-learning program showed minimal sensitivity to GaMD templates and underperformed conventional docking, suggesting its affinity predictions complement rather than replace physics- and empirical-based docking approaches for GPCR AM screening.
bioinformatics2026-08-18v1A Generative Virtual Tissue Model Enables Computational Design of Therapeutic Perturbation Strategies
Lu, Y.; Zhang, W.; Chen, Y.-J.; Yin, J.; Chen, L.; Fleisher, K.; Gornet, J.; Liu, R.; Wang, Z. J.; Poon, Y.; You, Y.; Thomson, M.Abstract
Computational design has transformed many fields of engineering, where simulators can explore millions of candidate design configurations before experimental development and testing. Therapeutic design in biomedicine has resisted computational design approaches because disease progression and therapeutic response emerge from interactions among many cell types within human tissue, governed by biochemical parameters that are largely unknown and potentially unknowable. Here, we introduce the Cell Interaction Foundation Model (CIFM), a virtual tissue model that forward-simulates the transcriptional dynamics of cells in human tissue under arbitrary therapeutic conditions based upon a spatial transcriptomic seed. CIFM is a geometric graph neural network trained by self-supervised masked-transcriptome prediction on millions of cellular microenvironments spanning human tissue types and disease states; generative, auto-regressive, monte-carlo play-out, then, simulates transcriptional dynamics under combinatorial perturbations from a spatial transcriptomic seed. We validate CIFM by showing accuracy gains in gene expression prediction and imputation, disease classification, recapitulation of perturbation responses in prostate cancer models, and recovery of T cell-tumor signaling measured in cell-cell sequencing experiments. Beyond such conventional tasks, CIFM enables target identification and therapeutic design through generative tissue simulation play-outs. Analyzing over 10^6 single and combinatorial perturbations, CIFM designs immunotherapy strategies for cancer and autoimmune disease that exploit combinatorial manipulation of signaling pathways to induce or suppress immune activation. Broadly, CIFM shows how generative artificial intelligence methods can be applied to model emergent behavior in highly interacting biological systems, yielding new approaches to fundamental understanding of tissue behavior as well as large-scale therapeutic design.
bioinformatics2026-08-18v1From Lab Notes to Linked Data: MeSyTo for Ontology-Driven Metadata in Toxicological Omics
Pozhidaeva, M.; Schreiber, S.; Schubert, K.; Busch, W.; Hackermüller, J.; Canzler, S.Abstract
Toxicological omics studies require comprehensive metadata to support reproducibility, interoperability, and regulatory reuse. However, metadata requirements differ across public repositories, reporting frameworks, and laboratory workflows, resulting in inconsistent annotation and limited data integration. To address this challenge, we developed MeSyTo (Metadata for Systems Toxicology), an ontology-driven framework for harmonizing metadata across toxicological omics. Metadata concepts from public repositories, the OECD Omics Reporting Framework (OORF), community standards, and institutional workflows were semantically aligned and implemented as the MeSyTo Metadata Model (MMM). The MMM serves as the basis for the automatic generation of SHACL validation shapes and framework-specific metadata profiles, while curated value sets are represented as SKOS controlled vocabularies to support metadata collection and validation. The current implementation comprises 105 ontology classes and 527 data properties and supports transcriptomics, proteomics, and metabolomics. A prototype web application demonstrates ontology-driven metadata collection with integrated semantic validation and ontology-based term resolution. The ontology, validation shapes, controlled vocabularies, generation scripts, and software are publicly available as open-source resources. MeSyTo provides a reusable semantic foundation for harmonized, machine-actionable metadata and facilitates repository submission, regulatory reporting, and interoperable data exchange across toxicological omics studies.
bioinformatics2026-08-18v1Single-Cell Mapping of Malignant Signaling Networks Guides Drug Combinations
Yavuz, B. R.; Jang, H.; Nussinov, R.Abstract
Single-cell breast cancer atlases reveal malignant, immune, and stromal diversity; however, how the recurrent signaling pathways drive untreated malignant-cell states and could inform combination therapy remains unclear. Here, we analyzed 15,753 malignant cells from untreated primary breast tumors using a cell-resolved network framework. Individual transcriptomes were projected onto a protein-protein interaction network, partitioned into Leiden communities, and annotated by pathway enrichment. Pathway recurrence was evaluated against matched null models preserving community size, protein-network degree, and gene detection rate. Before null correction, recurrent pathways included PI3K/AKT, MAPK, JAK/STAT, and HIF-1 (hypoxia-inducible factor 1) signaling. After correction, HIF-1 emerged as the dominant recurrent signal across patients, indicating convergence of diverse upstream pathways on a shared hypoxia- and stress-adaptive malignant-cell program. The recurrent JAK/STAT, cAMP, glucagon, oxytocin, and thyroid hormone signaling suggest inflammatory, metabolic, and endocrine crosstalk. These findings support rational drug combinations targeting HIF-1 together with upstream PI3K/AKT/mTOR, MAPK, or JAK/STAT signaling.
bioinformatics2026-08-18v1Causally-inspired meta-representation learning framework for predicting patient-specific clinical responses to drug combinations
Zhang, Q.-Q.; Zhang, S.-W.; Shi, M.-H.; Li, J.-N.; Qiang, Y.-R.; Zhang, T.-H.Abstract
Large-scale prediction and assessment of clinical patient responses (i.e., RECIST class) to drug combinations remains challenging due to scarce patient-derived data. The existing prediction methods mainly rely on cancer cell line models. However, substantial biological heterogeneity between cancer cell lines and cancer patients within same tissues, as well as the heterogeneity between one tissue and another, often limit the generalizability of these methods in clinical patients. To overcome these limitations, here we present CaMeRe, a Causally-inspired Meta-representation learning framework designed to predict patient-specific clinical Response to drug combinations. In situations where stable causal factors and domain-specific response-modulating factors are unobservable, explicit discrete domain labels are unavailable, and data is scarce, CaMeRe designed a domain-invariant causal representation learning (DICRL) model guided by the invariant information bottleneck theory and causal intervention invariance principle, and also built a meta-learning framework with bi-level domain generalization to optimize DICRL model for achieving multi-domain generalization within and across tissues. By integrating the causal representation learning and meta learning framework, CaMeRe not only exhibited robust multi-domain generalization performance across multiple clinical drug combination response datasets and PDXs drug combination response datasets and generalization scenarios, but also had better interpretability. We applied CaMeRe to predict drug-combination response scores for 3,423 patients across 542,080 drug combinations. The predicted scores were significantly associated with biomarkers of known drug combinations and enabled the prioritization of candidate drug combinations across 11 cancer types, with stronger support from literature and clinical trial evidences than random baselines. We believe that CaMeRe can be a useful tool for predicting large-scale clinical individual drug combination responses and it has broad clinical applications.
bioinformatics2026-08-18v1LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Liu, A.; Ho, A.; Droste, A. M.; Martin, D.; Wong, E.; Zhou, E.; Zhou, I.; Park, J.; Jiao, J.; Skelly, K.-R.; Kim, K.; Li, J.; Rao, K.; Uehara, M.; Marion, M.; Fitzgerald, N.; Dias, R.; Shringarpure, S.; Yuan, Y.; Wang, Y.Abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.
bioinformatics2026-08-18v1From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores
Barbosa Araujo, P. V.; Fiuza, T. d. S.; Ferraz, R. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.Abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
bioinformatics2026-08-18v1IUPAC Consensus References Improve Short-Read Variant Detection in Clinically Challenging Regions: A Stratified Benchmarking Study with BurdenBench
Saidin, A.; Ricos, M.; Dibbens, L.Abstract
Abstract Motivation: Reference bias depresses variant detection in low-mappability regions, segmental duplications and the major histocompatibility complex (MHC); precisely the regions of greatest clinical relevance. Existing benchmarks rely on aggregate precision, recall and F1 metrics that obscure the absolute true-positive and false-positive counts that determine laboratory workload. No study has systematically evaluated IUPAC consensus references for short-read whole-genome sequencing (WGS) variant calling across Genome in a Bottle (GIAB) stratifications, multiple allele-frequency thresholds and multiple variant callers. Results: We aligned 30x WGS from three GIAB samples to IUPAC consensus references (allele frequency >=10% and >=30%) using the ambiguity-aware aligner novoAlign, benchmarking against BWA-MEM/GRCh38 and novoAlign/GRCh38 baselines across BCFtools, FreeBayes and GATK HaplotypeCaller. SNV recall increased by 3.1-3.9 percentage points (pp) in low-mappability regions and 1.8-3.1 pp in segmental duplications; INDEL recall rose by 4.5-5.8 pp and 2.4-3.8 pp, respectively, with similar gains in the MHC and challenging medically relevant genes (CMRG). Decomposition analysis showed that the aligner change drove most INDEL gains, while IUPAC encoding contributed additional SNV-specific improvement. We introduce BurdenBench, an open-source framework that computes net benefit and region-size-normalised metrics directly from standard hap.py outputs, revealing divergent caller-specific trade-off profiles that are invisible to aggregate F1: FreeBayes showed the most favourable precision-recall balance in low-mappability regions, while GATK achieved positive net benefit in the MHC. A controlled comparison using an identical variant set showed severe recall and precision losses for SALT (a published SNP-aware dual-index aligner) across all three callers, supporting the value of preserving linear reference structure. Pan-human and population-specific consensuses performed within 0.2 pp of one another. All findings are descriptive and hypothesis-generating from three samples. Availability and implementation: To mitigate potential bias associated with software developed by an author's employer, primary hap.py outputs and derived burden metrics were independently verified by co-authors with no affiliation to that employer. BurdenBench (v1.0.0) is implemented in Python (pandas, numpy; Python >=3.7) and freely available under the MIT licence at https://github.com/akzam/BurdenBench, including raw hap.py outputs and an audit trail enabling independent recomputation without a novoAlign licence. novoAlign and novoUtil (version 4, Novocraft Technologies) are commercial software with no-cost academic trial licences.
bioinformatics2026-08-18v1Celldega: Integrated Toolkit for Visualization and Analysis of Spatial Data
Fernandez, N.; Ishar, J.; Wang, H.; Saad, A. B.; Lipinski, M.; Farhi, S. L.Abstract
Spatial-transcriptomics integrates high-dimensional single-cell data with microscopy to reveal cellular states, communication, and tissue organization. Analyzing this data requires a combination of multi-modal data processing, high-dimensional data analysis, spatial analysis, and integrated visualization. However, computational analysis is increasingly becoming a bottleneck as approaches mature and dataset sizes increase. Additionally, visualization can be challenging as open-source visualization tools struggle to scale to large datasets (exceeding 1 billion transcripts), and commercial visualization tools are costly, closed source, and inflexible. We present Celldega, an open-source Python and JavaScript library for scalable, interactive visualization and analysis of spatial-omics data. Celldega integrates custom analyses, performs neighborhood analysis, implements an efficient visualization-specific file format, and enables interactive exploration in notebooks and web galleries. We demonstrate Celldega across multiple technologies, tissues, and datasets, including 3D reconstructions of the developing whole mouse head comprising over four million cells. Finally, we demonstrate how Celldega can be utilized throughout the entire lifecycle of spatial data analysis, from quality control to building a public shareable gallery.
bioinformatics2026-08-18v1Simulating Iron Deficiency in Plant Plastidial Metabolism With a Flexible Neural-Mechanistic Hybrid Approach
El Alaoui, S.; Henry, C. S.; Blaby-Haas, C.; Paape, T.; Xie, M.; Seaver, S. M.Abstract
Flux balance analysis has proven to be a successful approach in metabolic engineering and systems biology to predict intracellular fluxes of large genome-scale networks and the essentiality of genes encoding enzymes and regulatory factors. Flux balance analysis (FBA) relies on a key assumption of a metabolic state being persistent ("steady") over a given time frame. This assumption works well for microbial growth because of the ease with which microbial media can be fixed, biomass can be decomposed, and growth rates can be measured. However, the assumption is far less tenable for the cells and tissues of complex multicellular organisms, particularly if any integrated data is sampled from a heterogeneous collection of developing cells continually interacting between and across tissues. These will likely exhibit transient metabolic states equilibrating over varying timescales, and many FBA studies in complex organisms typically either ignore time as a parameter, or integrate data taken over long timescales (days/weeks). In this work, we adopt and modify a previously published machine learning approach originally developed to hybridize several aspects of a constraint-based approach with machine-learning in order to predict growth. This approach accommodates transient state dynamics, at the cost of tolerating a controlled amount of slack in the steady-state assumption, in order to enable transcript-constrained flux estimation in plant tissues without optimizing a growth objective. For our case study, we reconstruct the metabolism of the plastid of Poplar and Sorghum, integrating data sampled from leaf tissue under varying levels of iron bioavailability. The two species diverge sharply: Sorghum suppresses its photosynthetic electron transport and Calvin cycle abruptly at day 7 and its carbon delivery to biomass collapses to a third of control by day 21, whereas Poplar declines gradually and retains roughly 70%. Normalizing each reaction score to the plastid proteome pool makes the two species comparable and exposes reallocation as the pool contracts, including a split within Sorghum's sulfate assimilation that tracks which enzymes carry iron cofactors.
bioinformatics2026-08-17v2Long-Read epigenetic clocks identify improved brain aging predictions
Grant, S. M.; Eger, S. J.; Makarious, M. B.; Meredith, M.; Moller, A.; Grant-Peters, M.; Hicks, A.; Mandal, A.; Auluck, P.; Villegas-Lanau, A.; Mejia-Cupajita, B.; Acosta-Uribe, J.; Aguillon, D.; Leonard, H.; Kuznetsov, N.; Weller, C.; Reed, X.; Catching, A.; Jain, M.; Ferrucci, L.; Kosik, K. S.; Cookson, M. R.; Ryten, M.; Nalls, M. A.; Billingsley, K. J.Abstract
Epigenetic clocks are widely used to estimate biological aging, yet most are built from array-based data from peripheral tissues of predominantly European-ancestry individuals, limiting their generalizability. Here, we present aging clocks on DNA methylation from Oxford Nanopore long-read sequencing (LRS), leveraging over 28 million CpG sites from prefrontal cortex samples across individuals of African and European ancestry. These models were developed using GenoML, an automated machine learning platform for multi-omics data that leverages a diverse catalog of existing model architectures. Our long-read-informed clocks were developed using promoter-based and whole-genome window-based features, yielding models for each individual cohort as well as a combined-cohort clock. Each of these models demonstrated favorable performance compared to existing methylation clocks and was externally validated in a cohort of Colombian individuals. We further performed enrichment analyses and nominated both shared and cohort-specific pathways, cell types, and transcription factor binding motifs which may be implicated in aging and were not fully explained by cell type proportions or postmortem interval. Altogether, our findings highlight the power of long-read methylation data for constructing accurate, ancestry-aware aging clocks and emphasize the importance of inclusive training datasets.
bioinformatics2026-08-17v2The extended Split-ORF pipeline: Prediction and evaluation of Split-ORFs using Ribo-seq data
Kalk, C.; Murtagh, J.; Despic, V.; Mueller-McNicoll, M.; Schulz, M.Abstract
Split Open Reading frames (Split-ORFs) occur in transcripts containing at least two open reading frames, each encoding a part of the same full-length protein. These multiple open reading frames arise from alternatively spliced transcript isoforms. Understanding which genes make Split-ORFs, and in which cell types and under which conditions, would generate new insights into gene regulation. We previously published the Split-ORF pipeline, a computational tool that predicts candidate Split-ORFs from transcript sequences. Here, we present a new and improved version of the Split-ORF pipeline adding modules to analyze Ribo-seq data, calculate regions unique to the Split-ORF candidates, quantitatively assess Ribo-seq coverage in these regions, and perform candidate prioritization. Using this pipeline, we predicted more than 14,000 candidate Split-ORF transcripts from alternatively spliced human transcripts containing premature termination codons or retained introns. Hundreds of candidate Split-ORFs show significant Ribo-seq coverage across diverse cell types and diseases in at least one of the Split-ORFs, and 120 transcripts in both Split-ORFs. The candidate Split-ORF genes with significant Ribo-seq coverage are enriched for RNA-binding and RNA-processing functions and the majority of them encode RNA-binding proteins.
bioinformatics2026-08-17v2In situ Discovery of Immune Repertoire Reveals Antitumor Immunity and Therapeutic Antibodies
Zhang, H.; Wang, P.; Zhao, Y.; Yang, L.; Xue, T.; Liu, L.; Zhao, Y.; Zhang, Z.; Ma, J.; Zeng, B.; Zhang, P.; Wang, C.; Pan, D.; Gao, Z.; Liu, Z.; Zeng, Z.Abstract
Spatial transcriptomics offers a glimpse into the immunology of tissues. However, limitations in spatial transcriptomics preclude the detection of highly diverse, low-abundance, and previously unknown sequences, including immune repertoires and microbiota. Here, we introduce Archimap, a spatial transcriptomic platform that simultaneously profiles spatial transcriptomes, immune repertoires, and microbiota from formalin-fixed paraffin-embedded (FFPE) tissues. Using Archimap, we profile the spatial localization of TCRs, BCRs, and the microbiota landscape in archived clinical tissues at single-cell resolution. Through comprehensive benchmarking, we validate Archimap's performance and fidelity. Archimap in situ assembles the immune complex and reconstructs the clonal evolution of antibodies. Together, Archimap shows the power of in situ discovery of functional immune repertoires for their antitumor immunity.
bioinformatics2026-08-17v1Conformal Uncertainty Quantification for BayesAge Epigenetic Age Predictions
Mboning, L.; Pellegrini, M.; Bouchard, L.-S.Abstract
Epigenetic clocks predict chronological age from DNA methylation (DNAm) profiles, yet most provide point estimates without calibrated uncertainty. This limits their use when quantified error bounds are required. We apply conformal prediction to BayesAge, a maximum-likelihood clock that models nonlinear DNAm--age relationships using a small set of CpG loci and a count-based likelihood. Split conformal prediction yields distribution-free prediction intervals with finite-sample marginal coverage guarantees under exchangeability and requires only a single model fit per split. We also evaluate a locally scaled variant that produces age-dependent interval widths using a locally estimated error scale. In our targeted bisulfite sequencing cohort, split conformalized BayesAge attains near-nominal empirical coverage while preserving BayesAge point-prediction accuracy. The locally scaled variant yields wider intervals at older ages, but its coverage is less stable in small-calibration regimes, consistent with additional uncertainty from estimating the local scale. Relative to Monte Carlo intervals that propagate read-sampling variability and to higher-dimensional linear baselines (conformalized linear quantile regression and conformalized Lasso regression), conformalized BayesAge provides calibrated uncertainty using substantially fewer CpG sites and with weaker age-dependent structure in residuals. These results support conformal prediction as a practical approach for uncertainty quantification in DNAm-based age estimation.
bioinformatics2026-08-17v1Metabolites form a globally connected chemical network across protein families
Skolnick, J.; Srinivasan, B.Abstract
Metabolites are generally viewed as substrates, products, cofactors, or regulators of individual proteins, whereas metabolites recurring across many protein families are often regarded as promiscuous binders. Here, we analyzed 989,058 BioLiP2 protein - ligand binding sites and assigned 929,546 sites to ECOD v295 homologous groups to quantify ligand specificity, cross-fold scatter, structural breadth, and metabolite-mediated connectivity across protein-family space. Many ancient metabolites preferentially occupied cognate structural groups, demonstrating that broad evolutionary reuse can coexist with local structural discrimination. After excluding elemental metals, BioLiP potential-artifact/dual-use ligands, and metabolites containing fewer than six heavy atoms, 32 ancient metabolites occupied a mean of 185.38 ECOD F-groups per metabolite, compared with 6.32 F-groups for 2,540 mapped filtered non-ancient metabolites - a 29.35-fold enrichment (bootstrap 95% CI, 18.66 - 43.46). The complete 40-ancient-metabolite network connected all 6,798 associated F-groups into a single giant connected component (GCC). Even after stringent filtering, all 3,135 ancient-metabolite-associated F-groups remained in one GCC. Degree-preserving configuration-model randomizations and maximum-degree capping showed that this connectivity follows from the broad, recurrent distribution of metabolite binding rather than dependence on a few extreme hubs or a specialized higher-order topology. Differences between ancient and filtered non-ancient networks were not explained by metabolite size, whereas generic crystallization additives preferentially occupied smaller pockets. These results indicate that a limited ancient chemical repertoire established a globally connected protein-family architecture that subsequent metabolite diversification expanded while preserving its basic organization.
bioinformatics2026-08-17v1Using GIS Dashboards to highlight AMR data disparities in Africa for Policy, Research, and Public Health
Dogbegah, W. A.; Opiyo, S. O.; Proscovia, P. A.; Tiambo, C. K.Abstract
Antimicrobial resistance (AMR) continues to pose a major public health threat across Africa, yet available surveillance data remain highly fragmented across private, public, and academic sources. This study analysed continent-wide AMR surveillance patterns by integrating datasets from multiple independent repositories and visualising them through interactive Geographic Information System (GIS) dashboards. The objective was to generate an integrated evidence base that highlights resistance patterns, surveillance disparities, reporting gaps, and opportunities for improved AMR monitoring across Africa. Data were compiled from major private AMR surveillance programmes including Pfizers ATLAS, GSKs SOAR, Johnson & Johnsons DREAM, Venatorxs GEARS, and Shionogis SIDERO-WT covering the period 2004-2022. Public datasets from the WHO Global Antimicrobial Resistance and Use Surveillance System (GLASS) and the Fleming Funds Mapping Antimicrobial Resistance and Antimicrobial Use Partnership (MAAP) were incorporated for 2016-2020, together with published AMR studies conducted between 2010 and 2024. Datasets were harmonised to align key variables including bacterial species, isolate identifiers, antibiotics tested, surveillance source, geographical location, and categorical AMR outcomes while preserving the original structure of the contributing datasets. Interactive dashboards were developed using R Shiny to support spatial visualisation and dynamic analytical exploration of resistance patterns, temporal trends, species distribution, and country-level surveillance coverage. Descriptive analyses including means, standard deviations, medians, interquartile ranges (IQR), frequency distributions, Gini coefficients, Shannon entropy, Herfindahl-Hirschman Index (HHI), and Lorenz curves were used to assess inequality and concentration in country-level AMR reporting across surveillance systems. The integrated analyses revealed substantial heterogeneity and concentration in AMR surveillance reporting across Africa, reflecting major differences in surveillance intensity, laboratory infrastructure, reporting systems, and diagnostic capacity across countries. Private datasets demonstrated broader antibiotic panels and longer temporal coverage, whereas public datasets exhibited substantial gaps in country participation and pathogen-antibiotic representation. Published AMR studies additionally highlighted important surveillance information absent from formal surveillance databases. By integrating multiple streams of AMR evidence, this study demonstrates the value of interactive GIS dashboards as exploratory and updateable surveillance-support tools for improving visibility of fragmented AMR datasets, identifying surveillance disparities, supporting geographically informed interpretation of resistance trends, and strengthening future AMR surveillance harmonisation efforts across Africa.
bioinformatics2026-08-17v1Integration of proteomic data from cell lines and tumors
Ta, C. Q.; Auth, J. M.; Schilling, M.; Klingmueller, U.; Raue, A.Abstract
Cancer cell lines are widely used in preclinical research, yet the clinical translation of findings from cell lines remains limited. Identifying cell lines that best resemble patient tumors requires integration of molecular profiles across biologically distinct sample types. Recent advances in transcriptomic integration have demonstrated the potential of deep learning for aligning data across different sample types. However, comparable approaches for proteomic data integration remain lacking, potentially because of the prevalence of missing values in proteomic datasets. Here, we introduce ProtInt, a deep learning-based framework that integrates proteomic data from cell lines and patient tumors by combining principles from proteomic imputation and transcriptomic integration methods. We applied ProtInt to integrate label-free proteomic profiles from 771 cancer cell lines and 550 treatment-naive tumors. ProtInt outperformed batch correction and transcriptomic integration methods in aligning cell line and tumor proteomes. Comparison of the cell line proteomes before and after integration revealed recurrent increase of proteins associated with immune reaction, cell-cell communication, and interaction with the extracellular matrix, and reduction of proteins involved in transcription, post-transcriptional processing, and mitochondrial gene expression as proteomes of cell lines were adapted to resemble tumors. These results establish ProtInt as a framework for joint analysis of proteomic datasets across distinct sample types and may facilitate the identification of cell lines best suited for clinically relevant studies.
bioinformatics2026-08-17v1Degree-ranked gene lists omit the cross-module connectors, and a partition-free centrality recovers them
Zhao, Q.; Zheng, H.; Zhang, Y.; Bi, J.; Sun, T.Abstract
Network centrality is the workhorse of gene prioritisation, yet what a ranking omits is rarely audited. Scoring each selection against an annotation-count-matched maximum-entropy reference--asking whether a selected gene set covers the genome's functional space or collapses it--reveals that the criterion in standard use has a measurable blind spot in exactly the class it is meant to surface. Degree, the most widely used criterion, returns the cross-module bridges that are also locally dominant--connector hubs--and omits the non-hub connectors: where 26% of the genome occupies these coordinating roles, a degree-ranked list holds 18% and an EDVS-ranked list 55%, and degree's top-1% collapses functional coverage below the reference on all five networks tested. We repurpose EDVS (Entropy of Degree-Vector Sums), an information-theoretic diversity measure, as an annotation-free, partition-free centrality that recovers this omitted class. The coverage it preserves is carried by cross-module participation P, which cannot be computed without a community partition; EDVS matches P-level coverage on all five networks using none, and retains 0.84 of its selection under edge perturbation that leaves partition-based selections at 0.21-0.46. The deficit is general: the collapse holds in the same direction on the two networks built without functional annotation (0.5-1.1 bit; co-expression, physical interaction) as on the three supervised by it (1.6-3.3 bit; RiceNet, AraNet, STRING), so supervision amplifies it rather than creates it. The remedy is bounded: EDVS ceases to preserve coverage on the sparse physical-interaction network. And the class EDVS isolates is organizational, not an importance signal: pre-registered probes--essentiality, transcription-factor identity, tissue-specificity, date/party-hub character, phenotype co-localisation--return null or reversed throughout. The conclusive ones are equivalent to their degree-matched nulls within {+/-}5 percentage points (demonstrated, not merely undetected), and the classical coupling of centrality to importance itself holds only network-dependently.
bioinformatics2026-08-17v1Phenotype-associated spatial biomarker discovery in spatial transcriptomics with spHOT
Kim, H.; Kim, D.; Jung, S.; Lee, S.; Kim, K.Abstract
Spatial transcriptomics now profiles patient cohorts at single-cell resolution, enabling analysis of disease-associated cell organization in situ. However, discovering such spatial biomarkers remains challenging because relevant structures occur at unknown scales and cell- or niche-level annotations are rarely available. We present spHOT, a deep learning framework that localizes phenotype-associated spatial biomarkers from sample-level labels. spHOT combines spatial foundation model embeddings, a hierarchical domain tree for multi-resolution tissue representation, and a teacher-student multiple instance learning architecture that converts sample labels into cell-level biomarker scores. In controlled simulations and real-tissue benchmarks, spHOT outperformed existing spatial and single-cell methods in localizing ground-truth biomarkers. Across fibrotic, metabolic, and autoimmune disease datasets, spHOT recovered disease-relevant niches and tissue states reported by supervised analyses in the original studies. Cross-disease application of spHOT transferred biomarkers across chronic lung diseases without retraining. spHOT enables scalable, annotation-efficient spatial biomarker discovery in cohort-scale spatial transcriptomics.
bioinformatics2026-08-17v1Cryptic binding sites are detected but not ranked: coverage, conversion, and detector consensus
Moore, C. W.Abstract
Methods for predicting cryptic binding sites are compared almost exclusively on top-n recovery, a number that conflates two independent abilities: proposing a candidate at the right location, and ranking it highly enough to be seen. We separate them by retaining the per-candidate overlap of every proposal, rather than only the top five, for four structurally different detectors spanning 2009 to 2026, across the CryptoBench benchmark. The separation is large and it reorders the field. On the designated test fold of 178 structures, fpocket, a purely geometric method from 2009, proposes a qualifying candidate for 74.2% of targets, the highest coverage of any tool tested, yet surfaces one in its top five for only 43.8%. P2Rank proposes qualifying candidates for 66.3% and surfaces 63.5%, and IF-SitePred, a 2024 method built on protein language model embeddings, proposes 70.8% and surfaces 61.8%. Coverage across tools varies by 8 points while conversion, the share of a tool's own coverage that reaches the top five, varies from 59% to 96%. Unioning the four detectors reaches 92.1% coverage, and only 7.9% of cryptic sites are invisible to all of them. The field's headroom is therefore predominantly in ranking and in combination, not in detection: perfect ranking of a single tool's existing proposals would reach 74.2%, and of the union 92.1%, against the 66.3% currently achieved. We show the practical consequence is governed by candidate budget. Added coverage converts to recovery at about 85% while a structure carries fewer than roughly fifteen candidates and at about 51% above it, which explains a series of interventions that raised coverage and returned nothing. Working within that budget, proposing pockets from a protein language model at locations where geometry finds no concavity improves single-structure recovery by 8.5% (95% CI +4.0 to +13.6) on test-fold data, and lets a five-conformer ensemble match a twenty-conformer one at a third of the wall clock. We release per-candidate overlaps for all tools so that coverage and conversion can be reported separately without re-running any method.
bioinformatics2026-08-17v1scDRP: Disentangled representation learning for predicting single-cell responses to perturbations and estimating individual treatment effects
Sun, J.; Stojanov, P.; Zhang, K.Abstract
Dissecting cell-state-specific changes in gene regulation induced by perturbations is crucial for understanding biological mechanisms. However, single-cell sequencing provides only unmatched snapshots of cells under different conditions. This destructive measurement process hinders the estimation of individualized treatment effects (ITEs), which are essential for pinpointing these heterogeneous mechanistic responses. We develop scDRP, a generative framework that leverages disentangled representation learning with asymptotic correctness guarantees to separate perturbation-induced and perturbation-invariant latent variables via a sparsity regularized beta-VAE. Assuming quantile-preserving effects of perturbations conditional on confounders, scDRP performs conditional optimal transport in the disentangled latent space to infer counterfactual states and estimate ITEs. Applied to simulated and real single-cell perturbation data, scDRP accurately estimates treatment effects and individual counterfactual responses, with subsequent biclustering analysis further elucidating cell-type-specific functional gene module dynamics. Specifically, it captures distinct cellular patterns under rhinovirus and cigarette-smoke extract exposures, reveals heterogeneous responses to interferon stimulation across diverse immune cell types, and identifies distinct functional module activation in chronic myeloid leukemia cells following CRISPR activation targeting different genes. scDRP also generalizes to unseen perturbation doses and combinations. Our framework provides a principled computational approach to extracting heterogeneous causal relationships from single-cell perturbation data, enabling a deeper understanding of cellular and molecular mechanisms.
bioinformatics2026-08-13v3FlyPredictome: A structural atlas of predicted protein-protein interactions in Drosophila
Kim, A.-R.; Comjean, A.; Veal, A.; Rodiger, J.; Han, M.; Hu, Y.; Perrimon, N.Abstract
Protein-protein interactions (PPIs) are fundamental to cellular function. Yet most Drosophila PPIs remain structurally uncharacterized despite the wealth of genetic and biochemical data available for this organism. Here we present FlyPredictome, a structural interactome based on 1.5 million pairwise AlphaFold-Multimer predictions. Using a local confidence metric that performs robustly for interactions involving flexible and disordered proteins, we systematically assess experimentally reported Drosophila PPIs and predict direct binding interfaces at residue-level resolution. Testing their functional relevance, we find that phenotype-associated missense mutations are enriched at predicted interaction interfaces. Building on these validated predictions, we construct an evidence-supported PPI network, revealing modular organization from signaling pathways to individual protein complexes. FlyPredictome is available as an open database, providing a structural foundation for interaction discovery in Drosophila.
bioinformatics2026-08-13v3Reproducible Community States Resolve the Nasopharyngeal Diversity Paradox
Song, K.; Brochu, H. N.; Bustos, M. L.; Zhang, Q.; Thomas, M. J.; Harris, A. B.; Icenhour, C. R.; Letovsky, S. N.; Iyer, L.Abstract
The nasopharyngeal microbiome is critical for respiratory health, yet its compositional architecture remains largely uncharacterized, underlying conflicting reports of increased, decreased, or unchanged diversity during infection. Here, we demonstrate that these contradictions arise from a fundamentally overlooked factor: the nasopharynx is organized into six reproducible nasopharyngeal community state types (NPCSTs) that are associated with intrinsic diversity baselines independent of disease status. By uniformly reprocessing 7,790 16S rRNA gene sequencing samples from 28 studies, we show that NPCSTs explain 52% of community variance, four-fold more than study effects, and that disease-diversity associations are attenuated by over 92% after NPCST adjustment. To identify genuine disease markers, we applied NPCST-aware differential abundance testing to SARS-CoV-2 as a case study. Leave-one-study-out (LOSO) cross-validation across 10 cohorts confirmed 13 reproducible taxa with opposing shifts: obligate anaerobes typically studied in other body compartments were enriched, while resident hub genera were depleted with high directional consistency across LOSO iterations. Co-occurrence network analysis independently validated this pattern, identifying the depleted taxa as structural backbone hubs across all NPCSTs. To bridge microbiome ecology and clinical application, we developed the Nasopharyngeal Microbiome Health Index (NMHI), a continuous community-level wellness score achieving AUC of 0.90 and 0.92 in internal and external validations, respectively, with robust performance across all six NPCSTs (AUC: 0.848-0.953). An independent clinical cohort of 147 specimens confirmed that these NPCST and NMHI characteristics are reproducible. Unlike binary classifiers, NMHI quantifies nasopharyngeal health along a spectrum, establishing NPCST-aware analysis as a foundation for precision respiratory microbiome research.
bioinformatics2026-08-13v3RobustCell: A Model Attack-Defense Framework for Robust Transcriptomic Data Analysis
Liu, T.; Kang, Q.; Luo, X.; Xiao, Y.; Zhao, H.Abstract
Computational methods should be accurate and robust for tasks in biology and medicine, especially when facing different types of attacks, defined as perturbations of benign data that can cause a significant drop in method performance. Therefore, there is a need for robust models that can defend attacks. In this manuscript, we propose a novel framework named RobustCell to analyze attack-defense methods in single-cell and spatial transcriptomic data analysis. In this biological context, we consider three types of attacks as well as two types of defenses in our framework and systemically evaluate the performances of the existing methods on their performance of both clustering and annotating single cells and spatial transcriptomic data. Our evaluations show that successful attacks can impair the performances of various methods, including single-cell foundation models. A good defense policy can protect the models from performance drops. Finally, we analyze the contributions of specific genes toward the cell-type annotation task by running the single-gene and group-genes attack methods. Overall, RobustCell is a user-friendly and extension-flexible framework for analyzing the risks and safety of analyzing transcriptomic data under different attacks.
bioinformatics2026-08-13v2Discrete Inverse Rendering: Biological Image Analysis with Integer Programming
Kirkegaard, J. B.; Zdyb, F. O.Abstract
Biological image and video analysis is full of discrete decisions: whether an object is present, which multi-hypothesis detections are real, whether two detections are tracking the same object, or whether a cell divides or not. Standard pipelines resolve these locally and in sequence, e.g through non-max suppression, per-frame segmentation, distance-based linking, or dedicated lineage rules. Image evidence and temporal evidence are thus rarely weighed against each other in a single objective. We recast these discrete steps as \emph{discrete inverse rendering}: candidate renderings are generated, and then a integer programming solver selects the subset that best reconstructs the observed video subject to problem-specific constraints. We demonstrate that the same solver and the same reconstruction principle can handle three otherwise separate motifs: suppression of overlapping detections in dense \emph{C.\ elegans} tracking, extraction of a single connected curve in sperm tracking, and event-structured tracking in single cell tracking.
bioinformatics2026-08-13v2Human-supervised Agentic AI for Hypothesis Generation and Experimental Assistance in Drug Repurposing
Huynh, D.-L.; Asp, E.; Ballante, F.; Puigvert, J. C.; DeGrave, A.; Karki, R.; Nader, K.; Östling, P.; Pokharel, B.; Rietdijk, J.; Schlotawa, L.; Schmidt, L.; Seal, S.; Seashore-Ludlow, B.; Aittokallio, T.; Spjuth, O.Abstract
Computational drug repurposing has largely been focused on rapid hypothesis generation, yet real-world applications span a far broader lifecycle, from drug candidate suggestion to designing experiments, analyzing assay data, and iteratively refining candidates. Here, we demonstrate that agentic AI can operate throughout this lifecycle. To this end, we developed RepurAgent, a hierarchical multi-agent AI system comprising a supervisor agent and a planning agent that coordinate four specialized sub-agents (research, prediction, data, and report), through a human-in-the-loop design, with episodic memory and retrieval-augmented generation. The system is grounded in data, tools, and standard operating procedures specific for drug repurposing, developed within the REMEDi4ALL consortium. We validated the agentic system across three scenarios spanning the various stages within the repurposing lifecycle: in Acute Myeloid Leukemia, a blinded expert evaluation indicated that RepurAgent produced substantially more novel and mechanistically credible candidates compared to a vanilla LLM baseline; in a retrospective COVID-19 antiviral screen, RepurAgent acted as an adaptive experimental collaborator, prioritizing compounds with AUC-ROC up to 0.99 without predefined thresholds and flagging confounders missed in manual review; and for Multiple Sulfatase Deficiency, it prioritized 81 high-confidence candidates from 5000 compounds, which were further corroborated by domain experts. These results demonstrate that agentic AI can support across the drug repurposing lifecycle, from hypothesis generation to experimental analysis. RepurAgent is open source and deployed at https://repuragent.serve.scilifelab.se/.
bioinformatics2026-08-13v2scDIVA: semi-supervised integration and fine-grained annotation of tumor-immune single-cell atlases
Leslie, C. S.; Rapolu, V.; Lee, B.; Wong, W.; Karbalayghareh, A.Abstract
Single-cell atlases of the tumor-immune microenvironment have defined numerous fine-grained immune cell states, but each study uses its own nomenclature and procedure for annotating cell types. Transferring annotations from a reference atlas to a query dataset is complicated by both batch effects and by the presence of query populations that the reference does not contain. Here we present scDIVA, a semi-supervised deep generative model that adapts the Domain Invariant Variational Autoencoder to scRNA-seq for fine-grained tumor-immune label transfer. Three encoders disentangle each cell's expression profile into separate latent subspaces for cell type, batch, and residual variation; a single decoder reconstructs the cell's expression profile from all three latent embeddings, and auxiliary classifiers on the cell type and batch embeddings encourage each encoder to capture the respective source of variation; scDIVA's cell type embeddings are batch-invariant by construction rather than through explicit or adversarial correction. Benchmarked against four established reference-mapping approaches -- Harmony/Symphony, scANVI with scArches, scPoli with scArches, and Seurat with label transfer -- across six tumor-immune atlases spanning five cancer types, scDIVA achieved the highest mean macro-F1 in five of six atlases and the highest biological conservation scores. To detect query-enriched populations, we adapted the Milo differential abundance (DA) framework, added a directional test, corrected spatialFDR weighting, and parallelized neighborhood distances for atlas-scale data, and applied it to scDIVA's cell type embeddings with reference-versus-query membership as the condition. This design correctly identified cell types held out from the reference as OOR and flagged exhausted CD8 T cells from a tumor-immune atlas as OOR relative to a healthy pan-tissue immune reference; conversely, this procedure confirmed a conserved immune landscape between two independent colorectal cancer cohorts. scDIVA thus couples fine-grained annotation and integration with an FDR-controlled test for states the reference lacks, solving both tasks required for accurate label transfer in tumor-immune scRNA-seq atlases.
bioinformatics2026-08-13v1