Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
OTRec: Deep learning recommender for prospective druggable disease-target associations
Ofer, D.; Linial, M.Abstract
Identifying druggable disease--target associations remains a central challenge in translational medicine, limiting therapeutic discovery and repurposing. Here, we present OTRec, a deep learning--based recommender system that ranks such associations at scale and evaluates them in a temporal hold-out setting. Unlike approaches that rely on manually curated or aggregated evidence scores, OTRec employs a two-tower architecture to learn latent representations from 663,351 disease--target pairs. The model integrates heterogeneous inputs, including textual descriptions, ontology-derived features, and biological annotations such as tractability, Gene Ontology (GO) terms, and pathway information. We perform temporal validation by training on the 2022 Open Targets (OT) release and evaluating on clinical trial data from 2025. OTRec improves on the retrospective OT association score (ROC-AUC: 0.872 {+/-} 0.005 vs 0.559; PR-AUC: 0.288 {+/-} 0.009 vs. 0.08). In 5x5 target-disjoint cross-validation, OTRec reaches ROC-AUC 0.950 and PR-AUC 0.844) improving on the OT evidence score (ROC-AUC 0.91; PR-AUC 0.45). We rank the druggable genome across ~19,000 OT platform (OTP) diseases and release ~282,500 candidate associations above a 0.65 score threshold (in-distribution CV precision 0.92), covering 4,346 diseases including 2,322 orphan diseases, through an interactive prediction platform.
bioinformatics2026-08-20v3Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT
Hossain, M. S.; Sojib, M. R.; Tahmid, M. T.; Rahman, M. S.Abstract
Motivation: RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair tokenization breaks the one-token-one-nucleotide correspondence that nucleotide-level attribution assumes. Results: We present SPIRAL, a layer-wise SAE analysis of BiRNA-BERT. Independent SAEs at layers 0, 5, and 11 expand each 768-dimensional hidden state into 6,144 features while preserving model behaviour (explained variance above 0.99997; masked-language-model sequence recovery near 99.7%). Tokenizer-aware offset propagation aligns features to nucleotides: at layer 5, 44.3% of tested features are significantly associated with bpRNA secondary-structure classes (mean enrichment 1.61x), and all 1,237 eligible features with RNAcentral RNA types. Sparse profiles raise k-nearest-neighbour balanced accuracy from 0.328 to 0.359 over dense embeddings at layer 5. Availability and Implementation: Source code is available at https://github.com/SadatHossain01/SPIRAL; the code, evaluation data, and trained SAE checkpoints are archived at https://doi.org/10.5281/zenodo.21891845. Contact: mrahman@cse.buet.ac.bd
bioinformatics2026-08-20v2Cryptic binding sites are detected but not ranked: coverage, conversion, and the limits of detector consensus
Moore, C. W.Abstract
Methods for predicting cryptic binding sites are compared almost exclusively on top-n recovery, a number that conflates two independent abilities: proposing a candidate at the right location, and ranking it highly enough to be seen. We separate them by retaining the per-candidate overlap of every proposal, rather than only the top five, for four structurally different detectors spanning 2009 to 2026, across the CryptoBench benchmark. The separation is large and it reorders the field. On the designated test fold of 178 structures, fpocket, a purely geometric method from 2009, proposes a qualifying candidate for 74.2% of targets, the highest coverage of any tool tested, yet surfaces one in its top five for only 43.8%. P2Rank proposes qualifying candidates for 66.3% and surfaces 63.5%, and IF-SitePred, a 2024 method built on protein language model embeddings, proposes 70.8% and surfaces 61.8%. Coverage across tools varies by 8 points while conversion, the share of a tools own coverage that reaches the top five, varies from 59% to 96%. Unioning the four detectors reaches 92.1% coverage, and only 7.9% of cryptic sites are invisible to all of them. The fields headroom is therefore predominantly in ranking and in combination, not in detection: perfect ranking of a single tools existing proposals would reach 74.2%, and of the union 92.1%, against the 66.3% currently achieved. We show the practical consequence is governed by candidate budget. Added coverage converts to recovery at about 85% while a structure carries fewer than roughly fifteen candidates and at about 51% above it, which explains a series of interventions that raised coverage and returned nothing. Working within that budget, proposing pockets from a protein language model at locations where geometry finds no concavity improves single-structure recovery by 8.5% (95% CI +4.0 to +13.6) on test-fold data, and lets a five-conformer ensemble match a twenty-conformer one at a third of the wall clock. We release per-candidate overlaps for all tools so that coverage and conversion can be reported separately without re-running any method.
bioinformatics2026-08-20v2scUnify: a unified framework for training and inference across multiple single-cell foundation models
KIM, D.; Hong, A.; Jeong, K.; KIM, K.Abstract
Single-cell foundation models (scFMs) differ in software requirements and performance across downstream tasks and adaptation strategies, complicating comparison and reuse. We present scUnify, a framework that preserves each backbone's required processing while separating model-specific trainers, downstream tasks, and adaptation strategies as reusable components. Across five scFMs, scUnify reproduced original inference and training workflows, extended model-native tasks with multiple parameter-efficient fine-tuning methods, and demonstrated extensibility by connecting a newly implemented custom trainable task to multiple backbones and adaptation strategies. Together, these capabilities enable researchers to systematically compare these combinations and extend custom tasks across heterogeneous scFMs within a common workflow.
bioinformatics2026-08-20v2PyMOL plugin for Protein Circuit Topology
Dimins, M.; Bazba, A.; Mogyorosi, A.; Kennon, E.; Fiol, T. D.; Hagen, L. A.; Sheikhhassani, V.; Akulov, V.; Mashaghi, A.Abstract
Circuit Topology (CT) provides a fundamental framework for analysing folded polymer chains, with applications in functional annotation, protein engineering and drug development. We present a protein CT analysis plugin for PyMOL v3.1.6.1 with a graphical user interface (GUI), automatic installation, and novel features developed through integration with PyMOL's application programming interface (API). The plugin integrates various previously developed CT methodologies for studying structured proteins and their complexes as well as the dynamics of disordered proteins. Analysis of a representative protein and a molecular dynamics trajectory demonstrates the plugin's three analysis modes and their outputs. The plugin reproduces the reference ProteinCT implementation exactly on the structures tested, and is distributed with a versioned release, a pinned environment and a one-command reproduction of every result reported here.
bioinformatics2026-08-20v2A comprehensive quality control pipeline in human microbiome research for large population studies
Li, R.; A.M. Verlouw, J.; G. Boer, C.; Arp, P.; van Meurs, J.; G. Uitterlinden, A.; Kraaij, R.; Medina-Gomez, C.Abstract
The widespread application of high-throughput Next Generation Sequencing (NGS) technologies has made microbiome research an emerging field in public health and biomedical sciences. However, there are still many challenges that need to be addressed in this field. Pipelines available to generate microbiome data across cohorts are diverse, and sources of variation to be recorded and evaluated during microbiome profiling have not been standardized. Moreover, meticulous quality control of the microbiome data processing, from collection to computational quantification is still challenging, especially in large population studies. Innovative approaches are required to handle samples and to minimize the potential bias introduced by logistic hurdles in biobanking. In this paper, we describe the methodological steps surrounding the optimization of the 16S rRNA gut microbiome profiling in two large prospective cohorts the Generation R Study (mean age 9.83 {+/-} 0.32 years) and the Rotterdam Study (mean age 62.67 {+/-} 5.66 years). This paper also highlights potential solutions to sample mislabeling in large-scale microbiome analysis. To summarize, our study addresses common problems in human microbiome research. It aims to improve the research quality and reliability by integrating more stringent quality control standards into microbiome research.
bioinformatics2026-08-20v2Toward a Minimal Amino Acid Alphabet for Protein Design
Pubal, K.; Kushnir, K.; Spiwok, V.; Louzecka, K.; Setnicka, V.; Lipovova, P.Abstract
Proteins are built from 20 canonical amino acids. It is interesting to explore whether proteins can be formed from significantly reduced amino acid alphabets. Our bioinformatics survey of UniProt (more than 250 M sequences) revealed that proteins composed of reduced amino acid alphabets (< 10) are extremely rare among existing proteins. Next, we used computational protein design to design proteins composed of all 1,013 possible alphabets of 2-10 early amino acids (Ala, Asp, Glu, Gly, Ile, Leu, Pro, Ser, Thr, and Val). The length of all proteins was 100 amino acid residues. Small amino acid alphabets preferred simple helices or helix bundles. Larger amino acid alphabets allowed for the design of more complex structures. A protein composed of 8 amino acid types (Ala, Asp, Gly, Leu, Val, Ser, Thr, and Pro) was successfully experimentally verified. It adopts the {beta}-sheet-rich fibronectin type III domain architecture. Attempts to experimentally verify designs composed of 6 and 4 amino acid types were unsuccessful. We show by a computational experiment with an experimental validation that inverse folding models, namely ProteinMPNNsol, can stabilize a designed protein within the same eight-amino-acid alphabet. Our results show that globular proteins may have formed early in evolution. Furthermore, we show that it is possible to design proteins with interesting properties for biotechnology and synthetic biology.
bioinformatics2026-08-20v2BART-spatial unravels biologically significant transcriptional regulators from spatial omics data
Wang, J.; Zhang, H.; Wang, Z.; Zang, C.Abstract
Transcriptional regulators (TRs) are crucial regulators of cell fate decisions by activating or repressing lineage-specific genes and integrating environmental signals with intrinsic networks. Identifying functional TRs is essential for understanding development, tissue organization, and disease. Emerging spatial transcriptomics and epigenomics technologies now provide near-single-cell resolution mapping of genomic features while preserving information of each cell's physical location and microenvironment which influence TR activity. Despite these advances, identifying active TRs in spatial data remains challenging due to low TR expression and the fact that TR activity often does not correlate directly with mRNA levels. Moreover, existing tools mainly designed for non-spatial single-cell data overlook spatial heterogeneity. To bridge this gap, we developed BART-spatial (Binding Analysis for Regulation of Transcription for spatial omics data), an innovative computational method to infer functional TRs from spatial omics data. BART-spatial integrates spatial variability and pseudotemporal information with publicly available TR binding profiles. Applied to multiple spatial datasets from diverse platforms, including 10x Visium, Visium HD, Atera, and spatial RNA-ATAC-seq, BART-spatial consistently outperforms existing methods, identifying state-specific TRs and revealing regulators undetectable by expression alone. Its compatibility with spatial epigenomics data further strengthens its utility and enables cross-validation. Overall, BART-spatial provides a powerful and robust tool for decoding spatially resolved gene regulatory programs.
bioinformatics2026-08-20v2PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring
Panda, P. K.Abstract
We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docking, and an SE(3)-equivariant graph neural network scoring function trained at scale. Ligand flexibility is represented as a torsion tree and pose parameters are optimized by Monte Carlo with Metropolis acceptance refined by L-BFGS, with rotational gradients obtained in closed form through the derivative of the SO(3) exponential map rather than by finite differences. Affinity grids are built by a blocked neighbor-selection scheme that is exact and 5.6-9.7x faster than dense evaluation, and may be cached across ligands sharing a receptor and site, reducing a six-ligand series from 29.3 s to 10.4 s. On 814 protein-ligand complexes spanning 14 target families, PandaDock recovers a pose within 2 Angstroms of the crystal geometry in 33.7% of cases at rank 1 and in 57.0% of cases within the returned ensemble. The GNN scoring function is trained on 741,706 co-folded complexes from SAIR under target-disjoint splits, reaching a Pearson r of 0.407 on 90,219 held-out complexes and transferring to 202 independent crystal structures with measured Ki, Kd, IC50 or EC50 at r = 0.467. We report the model against three controls, a target-mean predictor, a ligand-descriptor-only baseline, and within-target correlations, and document both where it performs and where it does not, including its unsuitability for pose rescoring. On an independent 30-compound series against a single GABAA receptor target, PandaDock's empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR. At full scale on the PDBbind v2020 refined set (n = 4,640, native crystal poses), the fully independent SAIR model reaches r = 0.531, and a dedicated model trained on PDBbind alone under a target-disjoint split reaches r = 0.690 on its own held-out test complexes, the strongest evidence in this work that PandaDock's affinity predictions generalize. PandaDock is distributed under an open-source license at https://github.com/pritampanda15/PandaDock with a complete command-line interface and a reproducible benchmarking harness.
bioinformatics2026-08-20v1A generalizable normalization framework to decouple protocol and instrument effects: Application to high-sensitivity proteomics multicentric study (PME13)
Arauz-Garofalo, G.; Ciordia, S.; Gonzalez de Peredo, A.; Chaoui, K.; Rijal, J. B.; Gaxotte, V.; Folch-i-Casanovas, I.; Azkargorta, M.; Almey, R.; Aloria, K.; Kirim, B. A.; Barderas, R.; Braga-Lagache, S.; Calvo, E.; Chicano-Galvez, E.; Clemente, F.; Chiritoiu, G.; Chiva, C.; Decourcelle, M.; Dhaenens, M.; Diaz, R.; Douche, T.; Duran-Cortines, A.; Duran-Ruiz, M. C.; El Koulali, K.; Escobar-Nino, A.; Fernandez Acero, F. J.; Fernandez-Irigoyen, J.; Garcia-Garcia, C.; Gil, C.; Goetze, S.; Gonzalez Vidal, E.; Gutierrez, M.; Hernaez, M. L.; Lopez, C. M.; Marin-Vicente, C.; Mateos-Martin, M. L.; MatoAbstract
Multicenter studies are essential for benchmarking analytical workflows, yet their interpretation is often confounded by the combined effects of experimental protocols and instrumentation. To address this challenge, we introduce a simple normalization-based analytical framework, the recovery metric ({rho}), designed to decouple protocol driven effects from instrument dependent variability. We applied this framework to the 13th Proteomics Multicentric Experiment (PME13), a large multicentric proteomics dataset generated across 27 laboratories using high sensitivity workflows and varying sample preparation protocols. By leveraging a common digested reference sample, {rho} enables direct cross-comparison of all datasets on a unified scale, effectively minimizing instrument-related biases. Using this approach, we demonstrate that apparent instrument dependent trends are largely removed when evaluated through {rho}, revealing consistent protocol driven effects across laboratories. Statistical modeling identified key variables influencing {rho}, including sample input amount, reduction and alkylation, and the use of n-dodecyl-{beta}-D-maltoside (DDM). While DDM was associated with improved {rho}, reduction and alkylation and additional handling steps led to reduced performance, particularly at low input levels. We further highlight practical considerations for the application of ratio based normalization, including the occurrence of values exceeding theoretical bounds, which reflect deviations from underlying assumptions and require appropriate filtering. Overall, this work establishes a generalizable analytical strategy for disentangling confounding factors in multicentric datasets and provides practical guidelines for optimizing high sensitivity proteomics (HSP) workflows. The proposed framework is broadly applicable to other analytical fields where cross laboratory comparability is required.
bioinformatics2026-08-20v1bulk2scDiff: A Pseudobulk-Conditioned Diffusion Model for Bulk-to-Single-Cell RNASeq Generation
Xiao, J.; Raue, A.Abstract
Bulk RNA sequencing remains the predominant profiling strategy for large clinical cohorts, but it aggregates transcriptional signals across cell populations, thereby masking the underlying cellular heterogeneity. Inferring this heterogeneity from existing bulk transcriptomic data could extend large cohort-based studies that have already been profiled, but constitutes an underdetermined inverse problem, as one bulk profile can be compatible with multiple underlying cellular populations. Existing computational deconvolution methods address this problem primarily by estimating cell-type proportions or cell-type-averaged expression profiles rather than resolving expression at the level of individual cells. Here, we present bulk2scDiff, a proof-of-concept conditional diffusion framework that reformulates bulk-to-single-cell inference as conditional generation of single-cell expression profiles from pseudobulk transcriptomic input. We evaluated bulk2scDiff on two cancer single-cell RNA sequencing datasets, breast cancer and acute myeloid leukemia, where pseudobulk profiles were derived from the single-cell data and used as conditioning inputs, with the matched single-cell populations providing ground truth for controlled evaluation. Across both cases, bulk2scDiff closely reconstructed populations from training samples and generated biologically coherent single-cell populations for held-out samples, generalizing most consistently to recurrent immune features. A pseudobulk-swap control further confirmed sample-specific conditioning, with each sample corresponding pseudobulk yielding the closest agreement with its observed population in nearly all cases. Overall, our work establishes the feasibility of conditional diffusion for generating single-cell populations from pseudobulk transcriptomic profiles, providing a foundation for future evaluation with clinical bulk RNA sequencing data.
bioinformatics2026-08-20v1CryoForge: A Self-Correcting Agent for Cryo-EM Model Building That Learns When to Act and When to Stop
Feng, W.; Jiang, y.; Sun, F.; Yang, J.; Gao, X.; Zhang, F.; Han, R.Abstract
Automated atomic model building has accelerated cryo-EM structure determination, but different builders leave distinct residual error profiles requiring expert inspection. The post-building challenge is to decide which local interpretations are sufficiently supported by experimental evidence to be retained, corrected or rejected. Here we introduce CryoForge, an evidence-gated post-builder agent that separates repair proposal from repair acceptance. Rule and learning-based components identify candidate regions and prioritize legal actions, whereas an independent evidence gate evaluates each edit using map and half-map support, stereochemistry, connectivity and local structural context. Supported edits are retained; unsupported or conflicting modifications are rejected, rolled back, stopped or escalated for expert review. Across a resolution-stratified benchmark, 84.9% of 26,153 released trajectories yielded standard validated improvements and 3.7% yielded low-confidence partial improvements, with no quality-degrading edit retained in the final promoted models. Relative to rule-only control, learned prioritization reduced non-improving candidates and harmful actions while preserving global structural stability. External evaluations using an alternative initializer, same-team automated/manual-assisted challenge submissions and three recently released complex assemblies showed that CryoForge adapts to distinct residual error phenotypes and performs bounded, evidence-supported correction without uncontrolled remodeling. CryoForge provides a builder-independent, scalable and auditable correction layer between automated model generation and expert structural interpretation.
bioinformatics2026-08-20v1A multi-agent molecular optimization framework leads to a rapid-recovery intravenous anesthetic candidate with an improved safety margin
Xue, Z.; Liu, X.Abstract
Lead optimization, the systematic refinement of therapeutic compounds through iterative structural modification, faces a dual challenge in modern drug discovery: navigating astronomically vast molecular design spaces while balancing conflicting demands on potency, pharmacokinetics, and safety. We present MASCOT (Multi-Agent SearCh for molecular OpTimization), a role-specialized multi-agent framework for molecular optimization. Integrated with a chemically constrained graph-editing search, MASCOT coordinates three specialized agents: a trade-off agent that reprioritizes competing objectives, a strategy agent that adapts how molecular edits are proposed, and a reflection agent that distills lessons from previous decisions. Computational experiments showed that MASCOT achieved the best performance over competing methods on six benchmark settings. On the SARS-CoV-2 main protease task, its mean docking-score improvement was 3.6 times that of the strongest baseline. Applied to the clinically used anesthetic remimazolam (RM), MASCOT prioritized RM-1, which showed a shorter liver microsomal half-life, higher brain exposure, and a larger therapeutic index than RM. Subsequent derivative design yielded RM-7. Extensive animal studies established RM-7 as a rapid-recovery intravenous anesthetic candidate with greater potency, faster functional recovery, a wider safety margin, and preserved flumazenil reversibility. These results demonstrate that multi-agent coordination can link adaptive molecular search to medicinal chemistry and experimental pharmacology.
bioinformatics2026-08-20v1ETAP-CLF: an ESM3-based transformer attention framework for binary protein classification
Ren, J.; Jiang, H.; Li, P.; Yang, X.; Mei, L.; Tong, H.; Lin, L.Abstract
Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but downstream classifiers commonly compress them using fixed pooling operations that may discard task-relevant sequence context. We present ETAP-CLF, a compact framework that combines pretrained per-residue ESM3 embeddings with lightweight transformer contextualization and learned attention pooling to classify variable-length proteins and generate residue-level attention scores. The ESM3 parameters remained frozen, and the same ETAP-CLF architecture and hyperparameter configuration were used across ferroptosis-, senescence-, and pyroptosis-associated protein prediction. ETAP-CLF achieved AUROCs of 0.98, 0.95 and 0.91 for these tasks, respectively. In the ferroptosis benchmark, ETAP-CLF outperformed the evaluated published models. These results demonstrate that a common downstream design can adapt to multiple process-associated classification tasks without fine-tuning the task-specific model architecture. ETAP-CLF provides a generalizable approach for sequence-based protein prioritization and a basis for broader evaluation across binary protein-classification problems.
bioinformatics2026-08-20v1PlantOmicsGWAS: An end-to-end, reproducible framework for plant genome-wide association and genomic prediction using linear and pan-genome references
Khan, F. S.; Yassin, A.; Rehman, S. u.; Sun, T.; Wang, X.; Sun, H.; Abe-Kanoh, N.; Su, Y. H.; Guo, L.; Ye, W.Abstract
Genome-wide association studies (GWAS) play a crucial role in unraveling the genetic foundations of complex traits in plants but are also hampered by the application of heterogeneous tools, incompatible file formats and disparate computational environments. Existing GWAS frameworks are often restricted to a single linear reference genome, limiting the capacity for the analysis of structural variations and presence/absence variations (PAV) within plant populations. These issues pose obstacles to reproducibility, scalability, and comprehensive investigations. Here, we present PlantOmicsGWAS, an open-source Python framework for reproducible plant genome-wide association analysis and genomic prediction. It integrates reference indexing, FASTQ quality control, alignment, variant calling, VCF normalization, PLINK conversion, linkage disequilibrium analysis, population-structure estimation, association testing, marker scoring, genomic prediction, and visualization within a unified Linux and HPC workflow. The framework supports conventional linear-reference analyses and includes an optional pangenome-oriented module for working with multiple assemblies and graph-derived variation. Using a Vitis benchmark dataset containing 120 accessions and 118,247 graph-derived variants, PlantOmicsGWAS reduced manual workflow fragmentation and generated standardized association outputs. This tool provides a modular and extensible platform for plant GWAS and pan-GWAS workflows while retaining compatibility with established command-line tools and common genotype formats. The GWAS workflow described herein is adaptable to a range of sequencing methods and plant genomes, bridging research on crop related issues across various biological levels, from the individual organism to entire populations. PlantOmicsGWAS implements Bayesian sparse linear mixed modeling (BSLMM) through GEMMA for multi-trait association discovery, while also supporting FaST-LMM, regression-based approaches, and machine-learning algorithms (Random Forest, XGBoost) as benchmarking alternatives. The PlantOmicsGWAS, a versatile toolkit is available at GitHub https://github.com/plantomicsgwas1-boop/PlantOmicsGwas_V1 and on Linux and HPC platform (https://pypi.org/project/PlantOmicsGwas/1.0.2/).
bioinformatics2026-08-20v1nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality
Wyatt, C. D. R.; Duarte Frutos, F.; Turner, S. D.; Cerqueira De Araujo, A.; Rashid, U.; Begley, V.; Sumner, S.Abstract
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.
bioinformatics2026-08-20v1A single-antigen, multi-epitope subunit vaccine candidate against lumpy skin disease virus designed by conservation-guided reverse vaccinology
Koirala, B.; Ghimire, U.Abstract
Lumpy skin disease virus causes devastating economic losses in cattle, and current live attenuated vaccines carry risks of reversion and cannot differentiate infected from vaccinated animals. We applied conservation-guided reverse vaccinology (vaccine design from genomic sequence data) to identify conserved epitopes (immune-recognized protein fragments) and design a single-antigen, multi-epitope subunit vaccine candidate. A nine-taxon phylogenetic supermatrix of three candidate antigens was built, and BLA-restricted T cell and linear B cell epitopes were predicted. Conservation was quantified via Shannon entropy and Fisher's exact tests; structural disorder via AlphaFold2 and IUPred2A, followed by codon optimization and in silico cloning. All six selected epitopes mapped to the ankyrin locus. The 134 epitope columns showed significantly higher constraint than 2,376 background columns (mean entropy 0.008 vs. 0.243 bits; odds ratio 29.74). The 87-aa, 9.31 kDa construct was predominantly disordered and cloned in silico into pET28a. This work is entirely computational; wet-lab validation of immunogenicity is required before translational claims.
bioinformatics2026-08-20v1A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences
Motta, J. A.; Motta, M. d. M.; Fernandez, C.Abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
bioinformatics2026-08-20v1A Five-Gene Stromal-EMT Signature Predicts Prognosis, Immunotherapy Resistance, and Therapeutic Vulnerability in Bladder Cancer
Zhang, W.; Ji, S.Abstract
Background: Bladder cancer has entered an era in which immune checkpoint blockade (ICB) and antibody-drug conjugate (ADC)-based combinations are reshaping clinical management. However, transcriptomic scores that connect prognosis, tumor microenvironment state, and treatment response are incompletely defined. Methods: Open-access TCGA-BLCA RNA-seq, clinical, mutation, copy-number, and RPPA data were downloaded from the Genomic Data Commons (GDC). Tumor-normal differential expressions, survival screening, LASSO-Cox modeling, train-test validation, GEO validation, pathway enrichment, immune signature scoring, mutation/CNV/RPPA support, drug sensitivity prediction, single-cell/spatial localization, and ICB validation were performed using reproducible Python and R scripts. A reduced model was derived using only genes shared by TCGA, GSE13507, and GSE31684. The fixed formula was then applied without refitting to IMvigor210 and GSE176307. Results: A five-gene model composed of EMP1, AHNAK, TNFRSF14, CLEC2D, and GSDMB retained TCGA internal prognostic value (train C-index 0.693, test C-index 0.605, all-sample C-index 0.667; TCGA test log-rank p = 0.015), although GEO survival validation in GSE13507 and GSE31684 was modest. High-risk tumors were enriched for epithelial-mesenchymal transition (EMT), TNF-alpha/NF-kB signaling, inflammatory response, hypoxia, complement, CAF, macrophage, checkpoint, and cytotoxic programs. Single-cell and spatial analyses localized the score to basal tumor, endothelial, fibroblast, and perivascular compartments. In IMvigor210, risk scores were higher in ICB non-responders than responders (Wilcoxon p = 0.044; AUC for non-response = 0.580), high-risk tumors had a lower responder rate (17.6% vs. 28.0%), and high risk predicted poorer OS (log-rank p = 0.016; multivariate continuous risk HR = 3.15, p = 0.044). GSE176307 showed directionally consistent but non-significant response results (AUC = 0.576). Conclusions: The five-gene score is best interpreted not as a standalone universal prognostic classifier, but as a compact stromal-EMT and immune-suppression phenotype associated with inferior ICB response. These findings support a framework linking prognosis, microenvironment biology, immunotherapy resistance, and therapeutic hypotheses in bladder cancer.
bioinformatics2026-08-20v1Microbial bioprospecting for benzoxazolinate-like molecules: unleashing the potential of genome mining
Paliyal, S.; Kaur, B.; Rao, L.; Chakrabortty, A.; Singh, L.; Sehgal, I.; Sharma, M.; Singh, D.; Chaudhry, V.; Mantri, S. S.Abstract
The benzoxazolinate moiety is a key functional group found in a few natural products (NPs), exhibiting diverse bioactivities, including antitumor, antibacterial, and cytotoxic activities. Despite their clinical importance, only a few bacterial strains and NPs have been reported harboring this rare bis-heterocyclic moiety, underscoring a largely unexplored chemical space. Here, we performed large-scale genome mining and identified 277 putative biosynthetic gene clusters (BGCs) across diverse bacterial hosts, including previously unreported bacterial genera and strains. The BGCs were grouped into three compound classes: benzoxazolinate, benzobactin, and ashimides based on sequence similarity network clustering. Bioactivity predictions of the identified BGCs revealed the predominance of antibacterial and cytotoxic potential, highlighting promising candidates for future experimental validation and functional studies. This study also presents a neural network-based bioprospecting model that efficiently detects rare BGCs encoding benzoxazolinate-containing molecules from genomic sequences. Overall, our findings expand the known repertoire of bacterial hosts with the potential to produce benzoxazolinate-containing NPs and provide a comprehensive framework for the discovery and identification of candidate BGCs.
bioinformatics2026-08-20v1MSMICA: computational metabolite identification in untargeted metabolomics by integrating MS, retention time, and biological evidence
Zhan, J.; Weinberg, J.; Crandall, W. J.; Qin, Z.; Jarrell, Z. R.; Preston, J. D.; Nellis, M.; Teeny, S.; Liang, D.; Martin, G. S.; Price, N. L.; de Cabo, R.; Master, V.; Cohn, B. A.; Go, Y.-M.; Jones, D. P.Abstract
Mass Spectrometry Metabolomics Identification Connection Algorithm (MSMICA) is an algorithm for automated metabolite identification in untargeted liquid chromatography-high-resolution mass spectrometry (LC-HRMS) analyses. Limitations in metabolite identification can occur due to the availability and cost of standards and prevent recognition of metabolic factors impacting human health and disease. MSMICA performs mass-to-charge-ratio matching with chemical structures and clusters of LC-HRMS features for adduct and isotope forms. A local optimization is then used to integrate retention time prediction, metabolite precursor-product and transporter correlations, and biospecimen-specific abundance information for metabolite identification. Applying MSMICA to various internal and external mammalian datasets, validation results showed a 96.2 +- 5.1% correct rate of metabolite identification. When multiple LC-HRMS datasets were used, MSMICA enabled greater metabolite identifications, expanded metabolic pathway coverage, and data harmonization. Thus, MSMICA applies multiple pieces of evidence to substantially improve metabolite identification coverage and accuracy for known metabolites.
bioinformatics2026-08-20v1PASTRI: Resolving Stage-Specific Cell-State Dynamics from Annotated Cell Lineage Trees
Yang, W.; Li, Z.; Yu, X.; Wu, P.; Zhang, X.; Ren, C.; Liu, K.; Chen, J.; Chen, F.; He, X.; Zhang, J.; Chen, X.; Yang, J.Abstract
Cellular transitions between phenotypic states are fundamental to development and disease, yet quantitative analysis of their dynamics remains challenging. Here we present PASTRI (Phylogenetic Adjacency-based State Transition Rate Inference), a computational framework that infers transition rates from cell lineages/phylogenies annotated with terminal phenotypic states, such as single-cell transcriptomes. We validate PASTRI using simulated lineages and the Caenorhabditis elegans embryonic lineage. Importantly, by leveraging cell pairs at varying phylogenetic distances, PASTRI accurately resolves stage-specific transition rates, circumventing the issue of developmental changes in dynamics. Applied to three cell phylogeny datasets from our lineage-tracing experiments spanning diverse developmental/disease models and tracing systems, PASTRI uncovers rate-limiting steps in the activation of hepatic stellate cells and the differentiation of primordial lung progenitors, as well as attractor states that support cancer cell proliferation. PASTRI thus opens up a venue for dissecting cell state transition dynamics from annotated cell lineage/phylogeny.
bioinformatics2026-08-20v1How do noise and precision contraints jointly determine the minimum sketch size for fixed-size MinHash Jaccard estimation?
Ebou, A. E. T.; KOUA, D. K.Abstract
MinHash-based genome comparison is governed by two statistical constraints: a noise floor on k-mer size, controlling chance collisions between unrelated sequences, and a precision floor on sketch size, controlling uncertainty in the estimated Jaccard similarity. These constraints are best characterized for fixed-size bottom-sketch MinHash, the classical construction underlying tools such as Mash and still the standard basis for comparing genomes of similar size. Nevertheless, while, the noise floor has an established closed-form solution, sketch size is commonly selected using fixed defaults or heuristic choices independent of genome length and target precision. Moreover, recent work on scaled (FracMinHash) sketching has highlighted the difficulty of obtaining an analogous closed-form confidence interval for the Jaccard similarity. Here, we show that under the fixed-size bottom-sketch MinHash model, the shared-hash count follows a binomial model, which lets classical proportion-estimation theory be applied directly. Combining this with Fofanov's k-selection criterion and the exact identity-Jaccard relationship yields a single closed-form design rule, s* (L, {rho}* , ), that unifies the noise and precision floors into one operating envelope. The resulting design rule was validated against exact Clopper-Pearson interval inversion, idealized binomial sampling, realistic overlapping k-mer simulation and 54 bacterial genome pairs spanning species-to-order taxonomic levels. In realistic simulations, empirical coverage remained within 1-3 percentage points of the idealized reference in 14 of 15 tested cells. In real genomes, the classical identity-Jaccard relationship showed increasing positive bias with taxonomic divergence, from a median of -1.4% within species to +27% at genus and +77% at family level. We further show that s* (L) is not smooth in genome length but a discontinuous, previously unreported staircase caused by the ceiling function used for k-mer selection, with practical consequences concentrated at specific genome-size boundaries.
bioinformatics2026-08-20v1Predicting the operon structure of the Mycococcus xanthus genome using the novel software DiscOperon
Brunet, T.; Habermann, B. H.Abstract
ABSTRACT Motivation: Myxococcus xanthus is a predatory soil bacterium with a large genome of 9.14 MB due to a genome duplication event. While complete genome sequences of M. xanthus are available, gene annotation remains challenging due to its size and the resulting large number of duplicated genes. Operons, so syntenic block of genes that are co-regulated in bacterial genomes, are an important resource to help predict gene function accurately. Results: In order to help improve the annotation of complex genomes such as the one from M. xanthus, we developed a novel operon prediction tool, DiscOperon, which combines gene expression data with homology searches to identify syntenic blocks: Co-expression data of neighbouring genes across the genome are first used to define gene clusters, which are then used to search for conserved syntenic blocks in fully sequenced bacterial genomes using sequence homology searches. This strategy enables DiscOperon to account for gene insertions, rearrangements and deletions, which is its most distinguishing feature. We have tested DiscOperon against ground truths gene pair information on 3 different species from ODB and RegulonDB and compared it to state-of-the-art and still available operon prediction software and we demonstrate its general usability for operon prediction of any bacterial complete genome. We have applied DiscOperon to predict the operons of M. xanthus, which we are making available for the research community. Availability and implementation: DiscOperon is lightweight, user-friendly python tool with minimal dependencies. It is freely available at https://gitlab.com/habermann_lab/discoperon for general usage. Contact: Theo Brunet (theo.brunet@univ-amu.fr); Bianca Habermann (bianca.habermann@univ-amu.fr). Supplementary information: The operon-structured and annotated M. xanthus genome is available from this manuscript, as well as from Zenodo (https://doi.org/10.5281/zenodo.21976167). We furthermore plan to submit the M. xanthus operon information to the operon database OBD.
bioinformatics2026-08-20v1Relational Graph Convolutional Networks for Glioblastoma Biomarker Discovery via ceRNA and Copy Number Variation Analysis
Khandelwal, S.; Jarvis, N.; Zhan, J.Abstract
Glioblastoma (GBM) is a highly aggressive brain tumor with an extremely poor 5-year survival rate of 6.9%, largely attributable to the lack of reliable biomarkers. While competing endogenous RNA (ceRNA) and copy number variation (CNV) analyses offer unique biomarker identification potential, current approaches neglect the integration of multiple regulatory mechanisms for biomarker detection. To address this limitation, we applied relational graph convolutional networks (RGCNs) to ceRNA and CNV knowledge graphs through a novel late fusion ensemble architecture. The proposed architecture outperformed baseline models and identified five novel biomarkers, including hsa-miR-196a and hsa-miR-224. Kaplan-Meier survival analysis and Cox regression indicated that the identified genes hold significant prognostic and diagnostic power. The early stratification of the Kaplan-Meier curves indicates the potential these genes hold for patient survival prediction. The results illustrate that a late fusion RGCN ensemble effectively captures complex gene interactions, overcoming limitations of existing models and providing a framework for biomarker discovery. The novel biomarkers serve as prospective targets for future GBM therapeutic development and candidates for non-invasive diagnostic assays.
bioinformatics2026-08-20v1BoYueGRN: Zero-shot causal discovery of directed gene regulatory networks from single-cell transcriptomes via amortized inference over synthetic structural causal models
Wu, J.; Shen, Y.-Q.Abstract
Gene regulatory network (GRN) inference from single-cell RNA-seq conventionally relies on per-dataset optimization. Existing tools must be refit for every new dataset, and the majority fail to infer causal regulatory directions. Here we present BoYueGRN, an amortized causal discovery framework trained exclusively on 10,000 synthetic structural causal models. For any unseen dataset, a single forward pass returns edge probabilities and regulatory directions, while TF-centric sliding windows with asymmetric fusion extend this fixed-size model to full-transcriptome coverage. BoYueGRN demonstrates strong zero-shot performance across BEELINE benchmarks. On two independent genome-wide CRISPRi Perturb-seq screens, directional accuracy on retained edges reaches 0.86 and 0.95. Reconstructed cell-type- and stage-specific GRN dynamics across five diseases spanning more than 270,000 cells yield experimentally testable biological hypotheses. BoYueGRN reframes directed GRN inference as a train-once, reuse-across-datasets paradigm. By decoupling network reconstruction from per-dataset optimization, this paradigm opens the door to systematic, atlas-scale mapping of regulatory dynamics across human diseases.
bioinformatics2026-08-20v1GlycoMeSH: linking glycan structures to biomedical context for systematic enrichment analysis
Kitani, A.; Zhang, B.; Himori, K.; Matsui, Y.Abstract
Glycan identification has advanced, but glycan structures remain difficult to translate into reproducible biomedical context because reusable glycan-level annotations are sparse. We present GlycoMeSH, a resource that links glycans to Medical Subject Headings (MeSH) through an inference model, a traceable association database and a glycan-set enrichment workflow. GlycoMeSH-BERT recovered ~60% of literature-derived associations at recall@30 and expanded open-vocabulary MeSH coverage beyond closed-label baselines, without higher per-prediction accuracy. At matched candidate counts, its predictions showed motif-level semantic agreement comparable to those baselines, independently of the training labels. GlycoMeSH-DB contains 789,627 associations between 26,954 glycans and 20,302 MeSH terms. GlycoMeSH-EA returned enriched MeSH terms for glycan sets from glycomics and glycoproteomics datasets. Each association represents a biomedical context rather than a validated mechanism, and retains its source PMID or prediction score for audit. GlycoMeSH supplies the missing, evidence-traceable annotation layer that makes glycan sets directly analyzable by enrichment across glycoscience datasets.
bioinformatics2026-08-20v1Metagenomics analysis for microbial ecology investigation on historical samples: negligible effect of host DNA and optimal analysis strategies
Ng, S.-K.; Gutaker, R.Abstract
Microbiome composition and function are strongly influenced by environmental factors, with major shifts driven by intensified anthropogenic pressures over the past centuries. This timeframe extends beyond the scope of traditional experimental or longitudinal studies commonly used to investigate microbiome dynamics. The historical samples might provide important insights into the mechanistic consequences of anthropogenic pressures and the potential shift in microbial diversity and composition. Despite their vast potential, historical samples available in museums and herbaria worldwide remain underutilized for exploring host-microbiome interactions across broad temporal and spatial scales due to incompatibilities with standard analytical pipelines and limited understanding of optimal classification parameters. While host DNA removal has conventionally been considered essential for taxonomic assignment of metagenomic reads, and might be of particular importance when processing degraded DNA, this step is impractical for specimens with no reference genome available for host species. Here, we show that host DNA content has negligible impact on microbial data analysis with empirical and simulation datasets. Since DNA molecules from historical samples are highly fragmented and uneven in length, we further analysed the impact of k-mer value on the classification of metagenomic reads from historical samples. To improve recall rate, we proposed a simple two-step approach in which reads are classified with two annotation databases constructed with a long and a short k-mer values. Through simulation and published datasets, we demonstrated that this approach outperforms single-step workflows in effectively recovering microbial signals from reads in a wide range of length. Together, this study provides a solid foundation for incorporating natural history collections into host-associated microbiome research, offering valuable insights into the long-term effects of anthropogenic change on microbial communities.
bioinformatics2026-08-19v5Automating the Construction of Contextualized Biomedical Knowledge Graphs for Scientific Inference
Zheng, Y.; Liu, W.; Zeng, B.; Feng, Y.; Du, X.; Zhou, L.; Li, Y.Abstract
Biomedical interactions are inherently dynamic, often shifting or even reversing under specific physiological states. However, existing extraction methods simplify these complex mechanisms into context-agnostic binary associations, resulting in semantic loss and contradictory evidence. Here, we present AutoBioKG, an end-to-end framework that constructs context-aware knowledge graphs by leveraging composite triples to encode environmental conditions and entity attributes alongside core relationships. Powered by an open information extraction model trained on BioOpenIE and further refined through self-training with pseudo-labels from unlabeled literature, the framework exhibits broad generalization. Notably, AutoBioKG achieved the highest zero-shot F1 across DDI, ChemProt, and BioRED, outperforming the best-performing baseline on each benchmark by 3.6-17.8 percentage points. Furthermore, AutoBioKG-derived graphs outperformed existing approaches on yes/no, factoid, and list questions in the BioASQ biomedical question-answering evaluation, particularly for queries requiring fine-grained contextual information. Together, these results support AutoBioKG as a scalable framework for transforming unstructured literature into structured, context-aware biomedical knowledge.
bioinformatics2026-08-19v3PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking
Arora, R. K.; Chen, L. T.; Du, M.; Marks, D.; Church, G.Abstract
General-purpose frontier language models are being increasingly utilized for protein-design work, yet their ability to understand and evaluate variant effects remains unclear. Here, we introduce PG-LLM, a benchmark comprising 276 protein-variant prioritization tasks: 217 from ProteinGym and a temporally held-out set of 59 from recently published studies. Each task follows the same format: a language model is asked to rank a list of variant sequences given only the wild-type protein sequence and an assay description with no access to tools, multiple-sequence alignments, or protein structures. We evaluate thirteen language models and 95 published protein predictors on the same variants with the same evaluation metric. Claude Opus 5 (Max) and GPT 5.6 Sol (Max) are the best performing LLMs with Spearman correlations of {rho} = 0.406 and 0.402 respectively. Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at {rho} = 0.411, but remains below the leading predictor VenusREM at {rho} = 0.523. We observe that variant-ranking performance scales with test-time compute across GPT, Claude, and Gemini models, but gains taper before closing the gap to specialist protein predictors. To address contamination risk, we create a held-out evaluation set with 59 DMS assays from 19 studies whose scores first became public after January 2026. On this set, we observe performance and test time compute scaling trends similar to those on the 217 tasks derived from ProteinGym. PG-LLM shows that tool-free language models capture substantial protein-variant signal, outperforming many sequence-based predictors while remaining below the strongest specialized models.
bioinformatics2026-08-19v3Automating the Construction of Contextualized Biomedical Knowledge Graphs for Scientific Inference
Zheng, Y.; Liu, W.; Zeng, B.; Feng, Y.; Du, X.; Zhou, L.; Li, Y.Abstract
Biomedical interactions are inherently dynamic, often shifting or even reversing under specific physiological states. However, existing extraction methods simplify these complex mechanisms into context-agnostic binary associations, resulting in semantic loss and contradictory evidence. Here, we present AutoBioKG, an end-to-end framework that constructs context-aware knowledge graphs by leveraging composite triples to encode environmental conditions and entity attributes alongside core relationships. Powered by an open information extraction model trained on BioOpenIE and further refined through self-training with pseudo-labels from unlabeled literature, the framework exhibits broad generalization. Notably, AutoBioKG achieved the highest zero-shot F1 across DDI, ChemProt, and BioRED, outperforming the best-performing baseline on each benchmark by 3.6-17.8 percentage points. Furthermore, AutoBioKG-derived graphs outperformed existing approaches on yes/no, factoid, and list questions in the BioASQ biomedical question-answering evaluation, particularly for queries requiring fine-grained contextual information. Together, these results support AutoBioKG as a scalable framework for transforming unstructured literature into structured, context-aware biomedical knowledge.
bioinformatics2026-08-19v2GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations
Andrade, R. L.; Fiuza, T. d. S.; Ferraz, R. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.Abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearsons r improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTX-GUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
bioinformatics2026-08-19v2Method Choice, Not Biology, Determines In Silico Perturbation Results: A Systematic Evaluation of Eight Methods Across Four Datasets
Wenjie, G.; Wu, S.; Hu, G.; Yang, Z.; Wang, Z.; Cai, J.; Mao, J.Abstract
Most in silico perturbation methods for single-cell transcriptomics have been validated only on individual datasets, leaving their reliability and generalizability unknown. Through systematic cross-method, cross-dataset benchmarking of eight methods spanning six mathematical frameworks across four datasets, we find that six of eight methods--including widely used VAE-based and tensor decomposition approaches--fail to produce detectable transcription factor (TF)-to-pathway signals. Only CellOracle and DDIM consistently detected TF-to-glycolysis directional regulation. Cross-pathway analysis in PBMC monocytes revealed biologically coherent TF-pathway associations beyond glycolysis (SPI1[->]glycolysis 4.4x enrichment, FOS[->]AP-1 targets 4.4x), with SOX9 serving as a biological specificity control (no pathway enrichment). Method choice alone could reverse biological conclusions: DDIM and scTenifoldKnk rankings were significantly anti-correlated ({rho}=-0.811, p=0.027). CRISPRi Perturb-seq validation in K562 cells confirmed TF knockdown suppresses glycolysis gene expression (JUN {delta}=-1.72, CEBPB {delta}=-1.59, SPI1 {delta}=-1.57, FOS {delta}=-0.70), but CellOracle-predicted perturbation directions did not match experimental directions (40.9% agreement, not different from chance), revealing a fundamental gap between steady-state correlation and causal perturbation. Diagnostic analyses using VAE latent space profiling, correlation distribution comparison, and gene-gene graph analysis identified distinct failure modes in unsuccessful methods: VAE latent space competition (STAT3 signal-to-noise 0.44 vs. SPI1 4.25), correlation noise (TF-glycolysis |r|=0.038 indistinguishable from background |r|=0.047), and graph non-specificity (0.84x enrichment). A controlled ablation experiment showed that adding a GRN prior to DDIM did not improve target recall (delta=0 for all TFs), confirming that performance differences are multi-factorial. These findings establish preliminary guidance for method selection, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements ([≥]500 cells, [≥]1,000 HVGs).
bioinformatics2026-08-19v1reserBUGS: A reservoir computing framework for probabilistic forecasting of ecological abundance time series
Mohedano-Munoz, M. A.; Galeano, J.; Pastor, J. M.; de Aledo, J. G.; Bartomeus, I.; Allen-Perkins, A.Abstract
Forecasting species population dynamics is a central challenge in computational ecology, yet existing approaches rarely combine flexible nonlinear modelling, support for count-based ecological data, and systematic uncertainty quantification within a single, scalable framework. Here we introduce reserBUGS, an open-source Python framework for ecological forecasting based on reservoir computing, a recurrent neural network architecture in which only a simple readout layer is trained while a fixed high-dimensional dynamical system encodes temporal memory and nonlinear dependencies. reserBUGS integrates species abundance time series with environmental covariates retrieved automatically from global climate products, generates probabilistic ensemble forecasts, and provides tools for forecast evaluation and reliability assessment. We evaluated reserBUGS using insect abundance time series from available biodiversity monitoring datasets, comparing its performance against seven statistical and machine-learning baselines over one- to five-year forecast horizons. Reservoir-based models consistently outperformed alternatives in both predicting future abundance and capturing forecast uncertainty, with environmental predictors increasing the proportion of stable forecasts and contributing additional predictive value beyond historical abundance dynamics alone, particularly at 3-4-year forecast horizons. Probabilistic forecasts further enabled the identification of conditions associated with reduced predictive skill, providing a practical basis for communicating forecast confidence to end users. While default configurations already achieved competitive performance across a taxonomically and geographically diverse set of time series, hyperparameter optimisation revealed substantial room for performance gains through series-specific tuning. reserBUGS offers a computationally efficient and extensible framework for ecological forecasting that is well suited to the short, heterogeneous time series typical of biodiversity monitoring programmes. Its combination of flexible nonlinear modelling, probabilistic uncertainty quantification, and automated environmental data integration addresses key practical barriers to the adoption of modern forecasting methods in conservation and ecological research.
bioinformatics2026-08-19v1Tractography from Serial Optical Coherence Tomography: How and Why?
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.Abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Serial optical coherence tomography (S-OCT) is an imaging modality relying on the intrinsic contrast of a sample. When applied to brain tissues, the S-OCT contrast is primarily driven by the myelin reflectivity. Due to its high resolution, on the order of microns, and its 3D nature, S-OCT offers promise for studying WM connections at the microscale. However, while other microscopy imaging modalities have been shown to enable tractography, whether the reflectivity contrast from S-OCT supports the reconstruction of long-range WM fascicles at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We improve microscale orientation distribution functions (ODF) estimation by implementing a sliding-window formulation allowing the estimation of ODF at S-OCT resolution, and use apodized Dirac delta functions for reducing unwanted interference. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that using multiscale Frangi filters for estimating ODF outperforms structure tensor analysis. We also show that particle filtering tractography with anatomical constraints enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing the thalamocortical white-matter projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles visible at the microscale, and that these connections are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
bioinformatics2026-08-19v1Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling
Burq, M.; Stepec, D.; Kim, C.; Cimermancic, P.Abstract
Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.
bioinformatics2026-08-19v1Novel biologically relevant small RNA-sequencing alignment tool LevenMap for alignment to database of non-coding RNAs
Dlugas, H.; Dyson, G.; Dombkowski, A.; Kim, Y.; Gurdziel, K.; Boerner, J. L.; Bock, C.Abstract
A crucial aspect of the bioinformatics workflow in small RNA-sequencing is the alignment of reads to a database of reference ncRNAs. Alignment algorithms such as Bowtie, Burrows-Wheeler Aligner (BWA), and Spliced Transcripts Alignment to a Reference (STAR) - which are designed for aligning reads to a reference genome - are typically used. Aligning short RNA-sequenced reads to a database of non-coding RNAs (ncRNAs) is fundamentally a different task than aligning longer reads to a genome due to ncRNAs (i) having roughly the same number of nucleotides as the reads being aligned and (ii) being subsequences of other ncRNAs. To account for these differences, we developed the novel alignment algorithm LevenMap. Of all reads which exactly matched a reference ncRNA in a publicly available dataset, LevenMap aligned 100.0% of them to their respective ncRNA while all other aligners mapped less than 40% of these reads to their corresponding ncRNA. Furthermore, the mean ratio (length of read) / (length of corresponding reference ncRNA) of all aligned reads was 1.0 and 0.998 for LevenMap with at most zero and one mismatch(es) allowed, respectively; this ratio was no more than 0.51 for all other aligners. Overall, LevenMap is designed to account for the nuances of aligning small RNA-sequencing data to a database of reference ncRNAs and yields more biologically relevant counts compared to traditional aligners in this context. LevenMap is free and publicly available on GitHub: https://github.com/hdlugas/LevenMap.
bioinformatics2026-08-19v1Quantifying Crop Disease Trait Dynamics through Longitudinal Imaging and Temporal Analytics
Ewen, A.; Godoy Mendez, R.; Al-Shanoon, K.; Omoluabi, D.; Samarasinghe, A.; Oviedo-Ludena, M. A.; Chimbo Huatatoca, K.; Glor, K.; Nabetani, K.; Kutcher, R.; Wang, L.; Stavness, I.; Jin, L.Abstract
Reliable and objective phenotyping is essential for plant breeding programs to characterize genetic variation and accelerate crop improvement. Conventional disease assessment relies on expert visual scoring, which is labor-intensive, subjective, and prone to inter- and intra-rater variability. Although image-based phenotyping methods have been proposed, many require manual intervention, specialized imaging setups, or single time-point measurements, limiting their ability to capture disease progression over time. Here, we present a pipeline for longitudinal plant disease phenotyping that quantifies wheat stripe rust and leaf rust progression from time-series images. The pipeline performs semi-automated leaf and automated pustule segmentation from images acquired it in situ, enabling objective disease severity estimation with minimal user intervention and without requiring solid backgrounds or manual leaf manipulation or detachment. By extracting temporal traits, including disease severity trajectories and standardized area under the disease progress curve, the method provides a comprehensive characterization of disease development throughout infection. Association between automated and expert assessments was moderate for stripe rust (R2=0.58) and strong for leaf rust (R2=0.85), while expert inter-rater reliability was moderate for both diseases (ICC = 0.675 and 0.800, respectively). The proposed approach establishes a scalable and reproducible framework for longitudinal disease phenotyping in controlled environments, with broad applications in disease resistance screening and crop breeding.
bioinformatics2026-08-19v1An Exact-Residue Atlas of Opioid Receptor Wiring and Rewiring across Ligand and Transducer Contexts
Nael, M.; Alakonda, L.; Elokely, K.Abstract
Opioid-receptor structures span four human receptor subtypes, diverse ligands, signaling partners, and experimental constructs. We curated 86 human opioid-receptor structures representing 84 independent experimental maps, with one unique experimental data set counted once for structure-level inference, and analyzed them using our in-house StrucMind platform. StrucMind constructs exact Ballesteros-Weinstein (BW) contact graphs, meaning residue-contact networks restricted to unambiguous generic BW positions. Relative to active transducer-bound structures, structures classified as inactive showed 2.49% lower mean contact similarity and 34.84% more rewired contacts, where rewiring is the static set of contacts gained or lost between two structures. Among 77 maps with a resolved selected-ligand site, changed contacts were 15.28% direct to the site, 40.13% adjacent at one graph edge, and 44.59% connected-distal at a finite graph distance greater than one. The deposited-water analysis identified 137 receptor-proximal waters. Sixty-one contacted at least two protein residues, including 38 that bridged at least two exact-BW residues; a separate ligand-contact branch contained 10 waters contacting both selected ligand and receptor, only 3 of which belonged to the 38-water set. None of 32 component-association tests survived global correction. For peptide versus small molecule, the smallest nominal p value among four outcomes corresponded to 6.18% lower shared-contact distance root-mean-square deviation (p=0.00989; q=0.3165, where q is the adjusted p value). The atlas supports bounded, testable hypotheses, not causal component, hydration, or efficacy mechanisms.
bioinformatics2026-08-19v1Signature Recontextualization: Mapping perturbational signatures across biological contexts
Chen, A. D.; Girke, T.; Monti, S.Abstract
Perturbational transcriptomics is a powerful tool for understanding gene function and drug effects, yet predicting how perturbations manifest across different biological contexts remains a central challenge, limiting translation from model systems to clinically relevant tissues. Despite growing interest in this problem, benchmarking efforts have been hindered by inconsistent evaluation tasks, heterogeneous metrics, and limited assessment across perturbation types and biological systems. Here, we introduce a benchmarking framework for cross-context perturbation-signature prediction (a task we define as signature recontextualization), grounded in explicit definitions of the prediction task, target-data availability, and evaluation metrics centered on signature recovery. The framework evaluates prediction performance across three target-context data regimes: (1) control only, where only control profiles from the target context are measured; (2) low coverage, where a limited subset of perturbations in the target context are measured; and (3) high coverage, where most perturbations in the target context are measured. This design enables systematic assessment of how prediction performance depends on target-context sample size while providing a standardized basis for comparing methods. We evaluate newly developed projection-based (projectCor) and network-based (netProp) methods alongside deep learning-based foundation models (scGPT, STACK) and statistical baselines. The benchmark spans four diverse perturbational datasets: CRISPR knockdowns and drug perturbations in cell lines, plus in vivo chemical perturbations in rat tissues from DrugMatrix, extending evaluation beyond isolated cell-line models to tissue-level responses. Across tasks, projection and network propagation approaches show strong flexibility across perturbation types and biological contexts, and in several cases match or exceed the performance of deep learning and foundation models, suggesting that model complexity does not inherently improve cross-context generalization. We further show that perturbation predictability varies substantially with pathway conservation, transcriptional response strength, and baseline similarity between source and target contexts. All datasets, methods, and evaluation utilities are released as an open-source R package (sigRecon), providing a foundation for reproducible benchmarking and future method development.
bioinformatics2026-08-19v1Flexibility Drives Information Flow in Proteins: Fluctuation Potential Gradients Dictate Directional Entropy Transfer
Senguler Ciftci, F.; Erman, B.Abstract
Allosteric communication in biomacromolecules is fundamentally governed by thermal fluctuation gradients, yet standard Gaussian Network Models (GNMs) treat atomic contacts as uniform, binary couplings without differentiating core constraints from solvent-exposed surface flexibility. Here, we present an analytical matrix framework that incorporates continuous distance-dependent weighting into the Kirchhoff matrix L . This formulation captures the steep steric constraints of hydrophobic core packing versus peripheral surface loops while strictly recovering the classic unweighted GNM as a high-temperature limit (T[->]{infty} ). Using Schur complements of partitioned joint covariance matrices, we show that conditional fluctuation variances and higher-order entropy-transfer terms reduce analytically to exact ratios of submatrix determinants (covariance minors), eliminating the need for fitting parameters or molecular dynamics trajectories. Applied to KRAS (PDB: 6GOD), this framework constructs an integrated directional entropy-transfer asymmetry map. Order-1 minors (h(i)=Kii ) establish a single-node fluctuation potential gradient, while order-2 minors (Rij ) define pairwise channel bandwidths. Higher-order minors show multi-body spatial coupling: order-4 minors identify rigid core residues such as Phe156 as strategic interlobe relay hubs linking Lobe 1 and Lobe 2, and an order-3 triad cooperation index demonstrates that signal transmission from Switch II (Gln61) to Gly60 and Phe156 converges on a single, mechanically integrated allosteric sector. By deriving directional information flow directly from experimental atomic displacement parameters, this approach establishes a rigorous, computationally efficient framework for mapping allosteric networks across structural ensembles.
bioinformatics2026-08-19v1Modality-chain reasoning enables multimodal protein modelling and design
Liu, C.; Chao, L.; Ji, S.; Wang, H.; Zhou, G.; Zheng, J.; Hong, D.; Guo, Y.; Jiang, T.; Gao, Z.; Yang, M.; Pan, J.; Li, S. Z.; Zhang, X.Abstract
Reasoning has emerged as a central capability of large language models, yet how it should be formulated for scientific foundation models remains unclear because scientific knowledge is distributed across interdependent, domain-specific representations. Here we introduce modality-chain reasoning, which organizes representations into ordered computational chains, each conditioning prediction or generation of the next. Based on this principle, we develop ProteinReasoner, a multimodal generative protein foundation model that sequentially connects amino acid sequence, evolutionary constraints and three-dimensional structure within a shared autoregressive architecture. Across zero-shot structure prediction, inverse folding and fitness prediction, ProteinReasoner outperformed two multimodal protein foundation models, while controlled comparisons supported the functional contribution of the modality chain. We further extended this principle beyond pretraining: reorganizing the chain across successive structural states enabled multiple-conformation prediction, while introducing experimental feedback as an additional modality enabled an in-context learning paradigm for protein optimization without target-specific parameter updates. In particular, across thermostability and affinity-maturation evaluations, this paradigm improved over matched fine-tuned models and showed stronger mean performance than target-specific active-learning baselines. These results establish modality-chain reasoning as a unified and effective foundation-modelling strategy in protein science. More broadly, they suggest a general route towards reasoning across interdependent representations in other scientific domains.
bioinformatics2026-08-18v3antigen-prime: Simulating coupled genetic and antigenic evolution of influenza virus
Thornton, Z. T.; Tran, T.; Figgins, M. D.; Huddleston, J.; Bedford, T.; Matsen, F. A. D.; Haddox, H. K.Abstract
Seasonal influenza virus undergoes rapid antigenic drift to escape population immunity. Computational methods can be used to organize viral genetic diversity into antigenically similar variants and estimate variant-specific growth rates. However, benchmarking these methods is challenging because it can be di[ffi]cult to accurately quantify antigenicity and growth rates in nature. Simulating viral evolution using defined selective pressures can provide ground-truth data for benchmarking. But, existing simulators do not link genetic sequences to antigenic phenotypes under selection from host populations. Here, we present a forward-time epidemic simulator called antigen-prime that links these factors. We use it to simulate viral evolution over 30 years and validate the simulation recapitulates genetic and antigenic patterns observed in natural influenza evolution. We then use the simulated data to benchmark methods for assigning variants and estimating their growth rates. We evaluated a sequence-based and a phylogenetics-based method for variant assignment, finding the former was slightly more e[ff]ective at separating viruses into antigenically distinct groups. We also evaluated methods for estimating variant growth rates in one-year sliding windows. Estimates were accurate in most windows, but highly inaccurate in several others. Examining high- error windows revealed several examples of a previously unreported failure mode. In all, antigen-prime provides a simulation framework to benchmark models of influenza evolution, and could be used to help guide future development of these models. The source code is openly available at https://github.com/matsengrp/antigen-prime.
bioinformatics2026-08-18v2Protal: Ultra-fast metagenomic profiling and strain-resolved analysis
Fritscher, J.; Duncan, A.; Hildebrand, F.Abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal - profiling through alignment - an ultra-fast alignment-based method for species- and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom bench marks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler, respectively). At the strain level, protal reconstructs intraspecific phylogenetic relationships among detected bacteria with similar accuracy as StrainPhlAn 4; because protal produces precise alignments, the phylogenies can be de novo produced without reliance on reference strain collections. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
bioinformatics2026-08-18v2HNSW-MS: Hierarchical Graph Indexing Enables Accurate Real-Time Mass Spectral Similarity Search at Repository Scale
Semenov, A.; Roberts, A. M. P.; Coler, E. A.; Gupta, S.; Kopylova, E.; Melnik, A. V.; Boginski, V.; Aksenov, A. A.Abstract
Spectral similarity comparison is the basis of mass spectrometry-based metabolomics, underpinning library matching, molecular network construction, and repository searches such as MASST. Public repositories now exceed billion spectra, and exhaustive pairwise comparison is increasingly limiting. We introduce HNSW-MS, which implements Hierarchical Navigable Small World graph indexing natively for mass spectra. It operates directly on GC-MS and LC-MS/MS spectra without embedding, so scores are identical to those of exhaustive comparison and results are fully reproducible. On the benchmark dataset of 8.4 million LC-MS/MS spectra, HNSW-MS is up to 900-fold faster than linear search, recall@1 and recall@10 achieve over 90%, with the recall-speedup tradeoff tunable at query. Such rapid retrieval allows for any query spectrum to be near-instantaneously integrated into a global molecular network. We demonstrate specific examples of use of global networking for molecular discovery, including uncovering a new pathway of stenothricin biosynthesis and a new dipeptide conjugate bile acid.
bioinformatics2026-08-18v2Population dynamics of ribozymes during in vitro selection
Brooks, E. M.; Kotlin, N. R.; Weiss, Z.; DasGupta, S.Abstract
Ribozymes are central to models of RNA-based primordial life. Understanding how ribozymes emerge from RNA populations under changing selection pressures can help delineate the biochemical constraints and capabilities of early RNA-based biology. In vitro selection has been instrumental in isolating ribozymes with diverse functions from combinatorial populations; however, their population dynamics during selection remain poorly characterized. By analyzing high-throughput sequencing data produced from all rounds of an in vitro selection experiment, consisting of over 5 million unique sequences, we tracked how sequences rose and fell in abundance during selection. We found that the most abundant sequences emerged early and maintained their dominance, suggesting that only a few cycles of selection-amplification may be sufficient to isolate well-adapted ribozymes. Characterization of the catalytic activities of the dominant sequences in each selection round indicated that abundance did not always predict activity. This observation was reinforced through nucleotide conservation analysis and activity assays of close sequence variants of the most dominant ribozyme. Our analyses show that sequence complementarity with the substrate emerged as a common feature as early as the first round and persisted throughout, highlighting early convergence of diverse sequence populations on a beneficial catalytic strategy. By combining bioinformatics and experimental approaches, our work presents a detailed description of ribozyme population dynamics during an in vitro selection experiment and provides a high-resolution view of how functional RNA sequences compete for survival and propagation under selection pressure.
bioinformatics2026-08-18v2Cell-Hub: a graphical interface for end-to-end single-cell RNA sequencing analysis
macaux, g.; Di Gallo, M.; Taglietti, V.; Amthor, H.; Maire, P.Abstract
Abstract Single-cell and single-nucleus RNA sequencing have become increasingly widespread, creating a significant demand for accessible analysis tools in research laboratories. Despite this need, the bioinformatics expertise required for such analyses remains rare. Cell-Hub addresses this gap by enabling single-cell data analysis for all researchers, regardless of computational background. Cell-Hub is a comprehensive, free, and open-source framework built on R/Shiny and distributed as a Docker image, integrating Seurat 5, CellChat 2, and Monocle 3 within a unified graphical interface. It supports all essential steps of single-cell RNA-seq analysis: data loading, quality control, normalization, clustering, multi-dataset integration, differential expression, and biomarker detection. Cell-Hub further incorporates ligand-receptor interaction inference powered by GaspouDB, a consolidated database of 11,563 mouse and 9,604 human interactions derived from CellChat, CellPhoneDB, CellTalkDB, and MultiNicheNet as well as trajectory inference via Monocle 3. All analyses produce publication-ready visualizations with flexible export options. By integrating these analytical frameworks into a single, intuitive interface requiring no programming expertise, Cell-Hub represents a significant step toward democratizing single-cell genomics for the broader research community.
bioinformatics2026-08-18v2MOFTy: Multimodal Gaussian Process Factor Analysis with Numerical Information Field Theory
Neumann, M.; Arras, P.; Kaster, A.-K.; Ott, A.Abstract
Multimodal Gaussian process factor analysis provides a flexible framework for dimensionality reduction in temporally or spatially resolved omics data. Existing approaches, however, typically rely on pre-specified Gaussian process kernel families and do not explicitly separate each latent factor into a component capturing gradual, smooth variation and a complementary component capturing fine-scale, non-smooth variation. Here, we present MOFTy, a Bayesian multimodal factor analysis framework based on numerical information field theory (NIFTy) that replaces fixed kernel families with the flexible correlated field model in NIFTy and enables explicit additive component separation within each latent factor with quantified uncertainty. NIFTy has been successfully applied to high-resolution Bayesian imaging in astrophysics and facilitates scalable, curvature-aware variational inference for efficient posterior approximations. We validate MOFTy on simulated data; applications to published multi-omics data demonstrate that MOFTy disentangles latent spatial structures by separating smooth gradients from localized fine-scale heterogeneity in human glioblastoma and recovers cross-modal patterns in a mouse gastrulation dataset.
bioinformatics2026-08-18v2Assessing the Reliability of LLM-Generated Phenotype-Genotype Associations Through External Validation
Sun, C.; Xin, Y.; Zeng, S.; Sunthankar, S. D.; Su, W.-C.; Lynn, J.; Mundo, S.; Babanejad, M.; Feng, Q.; Wei, W.-Q.Abstract
Background: Phenotype-genotype associations underpin precision medicine by enabling disease prevention, early diagnosis, risk stratification, therapeutic target discovery, and personalized treatment. However, the rapid growth of scientific evidence has made manual curation of these associations increasingly labor-intensive, time-consuming, and incomplete. Large Language Models (LLMs) offer a potential path to scalable genomic generation and synthesis of this knowledge, but their ability to accurately identify phenotype-genotype associations and the extent to which these outputs are supported by established genomic knowledge bases remain unclear. Materials and Methods: Four LLMs, Claude Sonnet 4.6, DeepSeek V4 Flash, Gemini 3 Flash Preview, and GPT-5.5, were benchmarked on six zero-shot task categories covering forward and reverse phenotype-gene and phenotype-SNP generation. A total of 4,196 associations were identified from curated inputs and evaluated through a multistage external verification pipeline comprising phenotype normalization, ontology mapping, genomic identifier validation against Ensembl, and evidence verification using both the GWAS Catalog and OMIM. Associations were assigned a fused evidence level of strong, moderate, weak, or none. Results: Overall, 74.19% of generated associations were matched to at least one external genomic knowledge base; 9.15% received strong support and 54.46% moderate support. Phenotype-gene associations were more verifiable than phenotype-SNP associations (strong or moderate: 67.19% vs 54.06%). Among existing associations, Claude Sonnet 4.6 achieved the highest overall strong or moderate rate (69.2%), followed by GPT-5.5 (65.1%), DeepSeek V4 Flash (61.7%), and Gemini 3 Flash Preview (56.9%). Conclusion: LLMs can support scalable generation of candidate phenotype-genotype associations. Performance varied substantially by relation type and was lower for SNP-level and rare disease associations, highlighting both the limitations of current genomic resources and the need for rigorous validation pipelines. Key words: large language model; clinical genomics; phenotype-genotype association; external validation; genomic knowledge base
bioinformatics2026-08-18v1Efficient Game-Theoretic Explanations for Tree-Based Ensembles via Owen Values
Koh, H.Abstract
Shapley-value-based explanations, notably SHAP (SHapley Additive exPlanations), have gained prominence as a principled game-theoretic framework for local explanations and global feature importance. While exact Shapley value computation is exponential in feature count, TreeExplainer exploits the recursive structure of decision trees to achieve polynomial-time computation for tree-based ensembles. In many scientific applications, however, features are naturally organized into a priori groups reflecting domain knowledge, requiring explanations both across and within groups. The Owen value extends the Shapley value through a two-stage allocation rule that incorporates group structure while preserving fairness properties; yet, efficient algorithms for its computation remain limited. In this paper, we propose exact and Monte Carlo algorithms for computing Owen values in tree-based ensembles by combining hierarchy-guided group aggregation with tree-aware dynamic programming. The exact algorithm computes Owen values without sampling under the path-dependent characteristic function, which approximates the conditional expectation, whereas the Monte Carlo algorithm provides a scalable approximation that is unbiased for any prespecified sampling budget and converges almost surely as the sampling budget increases. We also provide global importance measures and visualization tools for structured, multi-resolution explanations. The proposed algorithms and tools are collectively referred to as TreeOwen. Through simulation experiments, we demonstrate the numerical accuracy and substantial computational gains of TreeOwen. We illustrate its practical utility using immunotherapy metagenomic data, showing how microbial genera (groups) and species (features) contribute to patient recovery.
bioinformatics2026-08-18v1