Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Stability-driven multi-omics integration for reproducible latent structure
Guan, H.; Gerwen, M. v.; Kim-Schulze, S.; Colicino, E.; Dolios, G.; Petrick, L.Abstract
High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We propose a Stability-driven framework for multi-omics integration that combines sparse generalized canonical correlation analysis with repeated cross-validation, out-of-sample projection, and systematic evaluation of both component-level and feature-level stability. We apply this framework to untargeted metabolomic and Olink targeted inflammation proteomic profiles in a thyroid cancer case-control cohort (n = 162). Our Stability-driven integration identified reproducible metabolomic and proteomic latent components that showed consistent out-of-sample disease associations and tracked temporally structured changes relative to time to diagnosis. The proposed framework provides a generalizable strategy for identifying reproducible latent structures that improve robustness of biological inference in multi-omics studies.
bioinformatics2026-08-25v3CCIDeconv: Hierarchical model for deconvolution of subcellular cell-cell interactions in single-cell data
Jayakumar, R.; Panwar, P.; Yang, J. Y. H.; Ghazanfar, S.Abstract
Cell-cell interaction (CCI) underlies several fundamental biological processes, including development, homeostasis and disease progression. Subcellular spatial transcriptomics (sST) provides an opportunity to examine whether CCI-associated signals show compartment-specific patterns within cells. Assessing CCI at subcellular level can help us gain insights into the distinct pathway activation and signalling patterns. We developed a novel approach that deconvolutes CCI into subcellular CCI (sCCI) information from non-spatial single-cell transcriptomics (scRNA- seq) based CCI using a modified CellChat-derived communication score. By estimating communication scores separately for cytoplasmic and nuclear compartments, we identified compartment-associated sCCI. We then deconvolved whole-cell communication scores into subcellular compartments using a hierarchical classification and regression framework, which we call CCIDeconv. To ensure biological fidelity, we integrated protein localization data from the Human Protein Atlas in our deconvolution model. Across nine publicly available human sST datasets, leave-one- dataset-out validation achieved a median composite score of 0.75, with mean R2 values of 0.87 and 0.80 for cytoplasmic- and nuclear-associated scores, respectively. Performance without spatial features approached that of spatial models as the number of training datasets increased, supporting application to non-spatial scRNA-seq data. This highlighted the potential for prediction of sCCI from scRNA-seq, given a sufficiently large number of training datasets. Overall, our method can attribute whole-cell CCI to its subcellular compartments, allowing researchers to dissect sCCI patterns and gain insights into the underlying biology of healthy and disease tissues. Keywords Cell-Cell Communication, Single Cell RNA-seq, Predictive Modeling, Bioinformatics, Transcriptomics, Machine Learning
bioinformatics2026-08-25v3GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations
Andrade, R. L.; Fiuza, T. d. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.Abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearson's $r$ improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTX-GUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
bioinformatics2026-08-25v3OMIO: A policy-driven Python library for reproducible microscopy image I/O
Musacchio, F.; Antony, H.; Baijal, A.; Crux, S.; Fuhrmann, F.; Gockel, N.; Hoffmann, D. M.; Mercan, D.; Nebeling, F. C.; Wolff, K.; Fuhrmann, M.Abstract
Modern fluorescence and multiphoton microscopy workflows operate within a heterogeneous ecosystem of file formats, partially overlapping metadata standards, and reader-specific conventions. In practice, this frequently leads to silent axis misinterpretations, loss or corruption of physical voxel size information, and laboratory-specific glue code that is fragile, poorly documented, and difficult to reproduce. OMIO, short for Open Microscopy Image I/O, addresses these issues by providing a lightweight, policy-driven image I/O layer for Python that enforces a canonical, OME-compatible data representation at the API boundary. The central contribution of OMIO is the explicit separation of low-level format access from semantic normalization. Existing reader libraries are used as interchangeable backends for extracting pixel data and available metadata, while OMIO enforces axis conventions, metadata interpretation, and fallback decisions in a centralized and auditable policy layer. This design allows heterogeneous microscopy inputs to be converted into a stable representation without propagating backend-specific assumptions into downstream analysis code. The core design principles of OMIO include canonical axis semantics (TZCYX), robust metadata normalization with explicit and auditable fallbacks, memory-aware operation via optional Zarr-based backends, and workflow-level semantics that extend beyond individual files to folder stacks and BIDS-like project structures. This architecture allows OMIO to orchestrate existing reader libraries into a coherent and reproducible I/O pipeline without replacing or duplicating their functionality. OMIO is implemented as an open-source and community-oriented system in which support for additional file formats and metadata conventions can be added incrementally through modular reader backends. By encouraging the contribution of example datasets, backend extensions, and feature requests, OMIO is designed to evolve alongside emerging acquisition systems while preserving strict semantic guarantees at the interface level. The resulting standardized OME-TIFF outputs are immediately suitable for downstream quantitative analysis and interactive inspection in scientific Python workflows, including workflows based on ImageJ and Napari. By standardizing image data at the I/O boundary, OMIO supports FAIR-aligned data sharing and reproducible microscopy analysis while facilitating the development of interoperable downstream tools.
bioinformatics2026-08-25v2From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores
Barbosa Araujo, P. V.; Fiuza, T. d. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.Abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
bioinformatics2026-08-25v2Differential Effects of Incomplete Lineage Sorting and Gene Tree Estimation Error on Gene Tree Distributions and Species Tree Inference
Tahmid, N.; Rhythm, S. I.; Bayzid, M. S.Abstract
Accurate species tree inference from genome-scale data is complicated by gene tree discordance, which can arise both from biological processes such as incomplete lineage sorting (ILS) and from technical factors such as gene tree estimation error (GTEE). While both factors reduce the accuracy of summary methods, their relative impact and characteristic patterns remain poorly understood. Here, we systematically compare the effects of ILS and GTEE by simulating gene tree datasets with comparable overall discordance levels, but with discordance arising exclusively from either ILS or GTEE. Using widely employed summary methods such as ASTRAL and wQFM, we show that GTEE typically has a stronger detrimental effect on species tree accuracy than ILS, even at matched discordance levels. We further characterize the structure of gene tree distributions under these two sources of discordance and show that ILS induces a structured, constrained skew in quartet distributions, whereas GTEE generates more uniform, high-entropy noise that does not diminish with additional genes. Our case study on a widely used avian phylogenomic dataset reveals similar distributional patterns across exons, introns, and ultraconserved elements (UCEs), which differ substantially in their levels of phylogenetic signal. A quartet-based analysis of these gene trees further shows that prioritizing loci with stronger and more consistent quartet support can improve the recovery of established avian clades. Overall, these results provide an empirical framework for a nuanced understanding of how ILS and GTEE shape gene tree distributions and influence species tree inference from limited or noisy gene tree datasets.
bioinformatics2026-08-25v2eSkip2 prioritizes exon-skipping antisense oligonucleotide target regions across exon--intron contexts
Chiba, S.; Kunitake, K.; Shirakaki, S.; Haque, U. S.; Wilton-Clark, H.; Shah, M. N. A.; Leckie, J. N.; Matsui, K.; Uno-Ono, F.; Yokota, T.; Aoki, Y.; Okuno, Y.Abstract
Exon-skipping antisense oligonucleotides (ASOs) can restore productive transcripts, but identifying effective binding regions remains difficult because splicing regulation extends across exons, introns and splice junctions. Here we develop eSkip2, a genome-informed framework that ranks target regions within a unified exon-intron sequence context. eSkip2 combines a genome-pretrained sequence model with ASO-induced exon-skipping data and single-nucleotide-variant splicing perturbations, followed by target-locus adaptation that requires no experimental ASO labels from the locus being designed. Across benchmarks comprising canonical exons and pseudoexons, multiple cell types and chemistries, and exonic, intronic and exon-intron-spanning targets, eSkip2 prioritized active regions and showed a higher median AUROC than applicable exon-restricted models. Prospective application to the combinatorial design of dual-targeting ASOs for DMD exon 46 enriched active candidates near the top of the ranking: the two most active new ASOs ranked within the top three and produced dose-dependent dystrophin restoration in patient-derived cells. These results support eSkip2 as a practical first-pass strategy for reducing experimental search space in exon-skipping ASO discovery.
bioinformatics2026-08-25v2Detecting CYP2C19 deletions from genotyping array signals using neural networks
Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.Abstract
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
bioinformatics2026-08-25v1AFP-R: An Open Resource Dedicated to Antifreeze Proteins
Liu, W.; Zhang, Y.; Xiu, D.; Liu, Y.; Wang, T.; Chai, X.; Qu, H.; Min, Y.; Zhang, Z.Abstract
Antifreeze proteins (AFPs), lower the freezing point via thermal hysteresis activity and/or ice recrystallization inhibition, playing a crucial role in protecting organisms from freezing damage under sub-zero milieu. This property endows them with promising applications in biomedicine and agriculture, ranging from tissue-organ cryopreservation to the development of frost-resistant crops. However, the lack of comprehensive resources dedicated for AFPs hinders further progress in elucidating their functional mechanisms and advancing their applications. Here, we report AFP-R, an online resource comprising AFP-DB and AFP-Predictor. AFP-DB is a comprehensive database with manually curated proteins bearing experimentally validated antifreeze activity derived from published literature, whereas AFP-Predictor is a sequence-based machine-learning model to identify AFPs. AFP-DB stores diverse AFP-related information, including sequences, structures, post-translational modifications, taxonomy and annotations of antifreeze-activity experimental assays. It now holds 186 entries, 607 sub-entries, and 1444 experimental records. AFP-Predictor, an AFP-identification algorithm built on protein language model ESM2 (Evolutionary Scale Modeling2), is trained on data in AFP-DB and outperforms several existing models. This work offers a valuable resource for systematically dissecting the mechanisms underlying AFP antifreeze activity and will facilitate their broader applications.
bioinformatics2026-08-25v1PhageLysData: an evidence-aware and AI-ready dataset of phage lytic enzymes and depolymerases
Medina-Ortiz, D.; Olivera-Nappa, A.; Lienqueo, M. E.; Opazo, R.; Romero, J.Abstract
Bacteriophage lytic enzymes and depolymerases are relevant to phage biology, antimicrobial development, and protein engineering, but their sequence and annotation data remain dispersed across general databases, specialized resources, genome-centred collections, and prediction-oriented datasets. We present PhageLysData, an evidence-aware and AI-ready resource constructed through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The release integrates 807,366 source observations from seven primary resources into 759,105 unique exact-sequence entities, comprising an evidence-supported Core of 11,867 entities, a Prediction Extension of 745,092 prediction-only candidates, and 2,146 Context entities retained for provenance and reference. This architecture preserves broad sequence-space coverage while maintaining a clear distinction between non-predictive and prediction-derived support. Core entities are enriched with harmonized biological annotations, physicochemical properties, independent InterProScan-derived functional annotations, mapped PDB and AlphaFold DB structural assets, and reusable numerical representations. For 11,259 eligible Core sequences, PhageLysData provides embeddings from 11 protein language models together with one-hot encoding under a common representation contract. Release-facing examples demonstrate latent-space exploration, unsupervised clustering, supervised classification, and evidence-aware candidate retrieval without defining a universal predictive benchmark. PhageLysData provides a traceable, versioned, and computationally accessible foundation for protein retrieval, comparative analysis, task-specific dataset construction, and machine-learning applications involving phage lytic enzymes and depolymerases.
bioinformatics2026-08-25v1Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles
Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.Abstract
Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.
bioinformatics2026-08-25v1UELer: a Jupyter-based framework for interactive exploration of multiplexed imaging datasets
Wu, Y.-L.; Liu, C.-S.; Lenoir, B.; Merz, K.; Dill, M. T.; Hartmann, F. J.Abstract
Summary Multiplexed imaging and spatial proteomics generate complex datasets that require both computational analysis and visual inspection. However, these tasks mostly occur in separate environments because interactive viewers generally require a local display or an additional data server beyond the remote Jupyter sessions itself where large datasets are computationally analyzed. We here present UELer, an interactive viewer that links multi-channel image views with quantitative analysis results directly within Jupyter notebooks, requiring no dedicated infrastructure beyond the notebook session. Cells selected through computational analysis and summary plots can be inspected directly in their tissue context, and selections made in the image can be made available to any downstream analysis. Together, these capabilities support interactive data exploration, iterative cell annotation, and reproducible retrieval of selected regions. Availability and Implementation UELer is a Python package built on ipywidgets and runs in Jupyter environments supporting ipywidgets 8.1 or later, tested in JupyterLab and Visual Studio Code on Linux, macOS, and Windows. It is freely available under GPL-3.0 license and can be installed via pip. Source code and documentation are available at https://github.com/HartmannLab/UELer and https://hartmannlab.github.io/UELer/. An online, no-install version runs remotely via BinderHub (https://mybinder.org/v2/gh/HartmannLab/UELer/main), accessible through the script/run_ueler_binder.ipynb notebook.
bioinformatics2026-08-25v1An inflammation-associated five-gene expression signature stratifies survival and immune states in lung adenocarcinoma: an integrative public-cohort analysis
Zhou, X.; Le, Z.; Song, P.; Xu, Q.; Chen, M.; Liu, X.; Cao, M.; Zhan, S.; Liu, Y.; Zhang, L.Abstract
Background: Inflammation and the tumor immune microenvironment contribute to lung adenocarcinoma (LUAD) progression, but the relationship among inflammation-linked transcriptional heterogeneity, patient survival, and immune-state variation remains incompletely defined. Objective: We aimed to identify inflammation-associated LUAD subtypes, derive a parsimonious survival-stratification signature, and characterize its immune and pathway context across public transcriptomic cohorts. Methods: Expression profiles and clinical data were obtained from TCGA-LUAD, GTEx normal lung, and GEO datasets GSE11969, GSE30219, GSE31210, and GSE40791. A curated set of 596 inflammation-related genes was used for consensus clustering. Differential-expression analysis, functional enrichment, univariate Cox regression, and LASSO-Cox modeling were integrated to construct a gene-expression risk score. The prognostic dataset comprised 730 cases and was randomly divided into training (n=502) and internal-validation (n=228) sets; 85 GSE30219 cases formed an external-validation cohort. Immune-cell enrichment, gene set enrichment analysis (GSEA), gene set variation analysis (GSVA), and pan-cancer analyses were used for biological contextualization. Results: The LUAD-versus-control comparison identified 1,305 differentially expressed genes, including 498 upregulated and 807 downregulated genes. Consensus clustering resolved two inflammation-associated subtypes and 67 subtype-associated genes, of which 64 were higher and 3 were lower in Cluster 1 relative to Cluster 2. Thirty-three genes overlapped between the tumor-control and subtype contrasts. LASSO-Cox regression selected CHRDL1, FDCSP, CXCL13, CYP4B1, and S100P. The 1-, 3-, and 5-year areas under the time-dependent receiver operating characteristic curve were 0.6625, 0.6581, and 0.6658 in the training set; 0.7422, 0.6537, and 0.6761 in internal validation; and 0.6560, 0.6387, and 0.6753 in external validation. Risk groups differed across multiple T-cell, B-cell, natural-killer-cell, myeloid, dendritic-cell, macrophage, and granulocyte signatures. Positive GSEA signals included cell cycle (normalized enrichment score [NES]=2.67; adjusted P=1.42 x 10-), DNA replication (NES=2.52; adjusted P=2.52 x 10-), and mismatch repair (NES=2.20; adjusted P=1.77 x 10-). Conclusions: The five-gene expression score separated LUAD survival groups and captured coordinated proliferative and immune transcriptional states. Its moderate discrimination supports further biological and clinical validation rather than immediate clinical application.
bioinformatics2026-08-25v1MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data
Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.Abstract
Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.
bioinformatics2026-08-25v1MultiFlow: coupled flow matching for predicting single-cell multiomic perturbation responses in unseen cellular contexts
Wang, H.; Zhang, C.; Zhang, M.; Nie, X.; Liu, Q.Abstract
Predicting cellular responses to perturbation requires resolving coordinated changes across molecular layers, yet most single-cell perturbation models focus on transcriptional responses alone. Here we present MultiFlow, a coupled flow-matching framework that unifies generation and perturbation prediction of paired gene expression and chromatin accessibility. By learning coupled RNA-ATAC flows conditioned on perturbation and control-derived cellular-state representation, MultiFlow enables prediction of coordinated multiomic responses in unseen cellular contexts. Across multiomic generation benchmarks, MultiFlow accurately reproduced paired RNA-ATAC states and their population distributions. In multiomic perturbation benchmarks, MultiFlow achieved the strongest overall performance in predicting both gene-expression and chromatin-accessibility responses, outperforming competing modality-specific perturbation-prediction methods. Joint multiomic modeling further preserved perturbation-induced RNA-ATAC coordination, including concordant peak-gene effects and cross-modal cellular neighborhood structure. These results establish coupled flow matching as a unified generative framework for modeling paired multiomic states and predicting coordinated perturbation responses across cellular contexts. Code and tutorial for MultiFlow are available at https://github.com/liuq-lab/MultiFlow.
bioinformatics2026-08-25v1ClustoCell reveals cell states and their markers from single-cell transcriptomes
Salavaty, A.; Foroutan, M.; Pretel, N. P.; Egelberg, J.; Parish, I. A.; Huntington, N. D.; Beltran, H.; Sandhu, S.; Molania, R.Abstract
Accurate identification of cell types and states is essential for reliable single-cell RNA-sequencing analyses, yet current methods remain sensitive to continuous biological states, data preprocessing choices, and reference selection. Here we present ClustoCell, a reference-free method that resolves cell identity using within-cell transcriptional architecture. By stratifying gene expression of each cell into high and medium tiers, ClustoCell constructs cell-cell similarity graphs that prioritize intrinsic expression structure over global variance. Across 450 datasets spanning over 24 million cells, ClustoCell recovered expert annotations with high concordance (92%). Benchmarked against state-of-the-art methods, ClustoCell identifies more stable and coherent cell types and states, avoids excessive partitioning of closely related cells, and improves the identification of cell type-specific markers. From transcriptional structure alone, ClustoCell resolves rare and transitional cell states, distinguishes malignant from non-malignant cells, and refines expert cell annotations. Applied to immunotherapy datasets, ClustoCell uncovered coordinated pre-treatment immune circuits linking T cell states to PD-1 responsiveness in a tumour-type-specific manner. ClustoCell provides an interpretable and scalable foundation for single-cell analysis and translational profiling.
bioinformatics2026-08-25v1ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation
Moran, J.; Freda, P. J.; Ghosh, A.; Hernandez, M. E.; Moore, J. H.Abstract
Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.
bioinformatics2026-08-25v1HuMMANet: A Harmonized Cross-Study Resource for Integrative Analysis of Human Gut Microbiome Metabolome Associations
Verma, S.; Arora, N.; Ajay, C. P.; Singh, P.; Mallick, H.; Ghosh, T. S.Abstract
Deciphering gut microbiome to host metabolome interaction is critical for understanding how microbial communities generate bioactive signals that shape host physiology and disease. Progress, however, has been hindered by inconsistent metabolite annotations, poor interoperability across studies, and the absence of integrated resources placing microbiome-derived metabolites within their functional, microbial, physiological, and clinical context. Here we present HuMMANet (Human Microbiome Metabolome Annotation Network), a harmonized resource integrating 46 paired gut microbiome metabolome studies (59 study-units; 14,405 samples; 13 disease categories plus a healthy/control reference category) with a scalable metabolite-harmonization framework. HuMMANet resolves heterogeneous annotations through a multi-stage workflow spanning RefMet, HMDB, PubChem, Metabolomics Workbench, SMPDB, MiMeDB 2.0, GNPS/microbeMASST, DrugBank, and DrugCentral, yielding a reference atlas of 54,914 unique metabolites, annotated with standardized chemical identifiers, biochemical pathways, microbial producer associations, physiological distributions, disease links, and structural relationships to approved therapeutics, a unified reference framework for microbiome metabolome research. Applying HuMMANet to a multi-cohort integration of adult serum and fecal metabolomes, we identified 519 serum and 322 fecal metabolites reproducibly associated with gut microbial community composition (PERMANOVA, P < 0.05 in at least 50% of studies in which detected), enriched for specific biomolecular classes and pathways. Cross-referencing these against Health Associated Core Keystone (HACK) taxa revealed 58 serum and 25 fecal metabolites (HACK positive) whose taxon-level associations tracked positively with the taxon specific HACK indices. These reproducible metabolomic signatures of microbiome health included indole3propionic acid, a gut barrier-protective microbial tryptophan metabolite, and 3phenylpropionate. Drug similarity annotation within HuMMANet linked 16 of this serum and 13 fecal HACK positive metabolites to therapeutics used in neurological, inflammatory, and vascular disease. Conversely, 38 serum and 65 fecal metabolites, including imidazole propionate and long-chain acylcarnitines such as ACar 18:0, showed HACK negative signatures previously associated with dysbiosis-linked disease. GNPS/microbeMASST and MiMeDB 2.0 annotations further traced subsets of these metabolites to putative bacterial producers. HuMMANet thus provides a standardized framework for reproducible microbiome metabolome integration, enabling cross study discovery and translational prioritization of conserved microbiome derived metabolic signatures across human populations and disease states.
bioinformatics2026-08-25v1Addressing technical variations in ATAC-seq data and improving motif accessibility analyses
Wang, J.; Sonder, E.; Domcke, S.; Robinson, M. D.; Germain, P.-L.Abstract
Tagmentation-based methods such as ATAC-seq and Cut&Tag have provided easy ways to profile the epigenome in low-input samples and even single cells. In this contribution, we discuss forms of bias (i.e. technical variations) in tagmentation-based data, in particular ATAC-seq, and introduce three R/bioconductor packages to facilitate bulk and single-cell epigenomic data analysis, with a special focus on motif accessibility analysis. The weightedMotifAccess package uses weight models to enable motif accessibility analysis, including transcription factor footprint information. The betterChromVAR package provides a novel, analytical re-implementation of the popular chromVAR method that offers substantial speed improvements, eliminates stochasticity, and offers additional features. Based on this, we also propose a method, CVnorm, that outperforms alternatives in normalizing technical bias in peak count data. The computational efficiency of these tools further enables a new framework for systematically investigating synergistic and antagonistic interactions between transcription factor motifs. Finally, the epiwraps package streamlines the visualization, normalization, and summarization of epigenomic data.
bioinformatics2026-08-25v1NetSyn: prokaryotic genomic context exploration of protein families
Stam, M.; Langlois, j.; Chevalier, C.; Mainguy, J.; Reboul, G.; Bastard, K.; Medigue, C.; Vallenet, D.Abstract
Background: The growing availability of large prokaryotic genomic datasets presents an opportunity to discover new metabolic pathways and enzymatic reactions useful for industrial or synthetic biological applications. Efforts to identify new enzyme functions in this vast number of sequences cannot be achieved without bioinformatics tools and the development of new strategies. Standard methods for assigning a biological function to a gene are based on sequence similarity. However, complementary approaches rely on mine databases to identify conserved gene clusters (i.e. syntenies). In prokaryotic genomes, genes involved in the same pathway are frequently encoded in a single locus with an operonic organisation. This genomic context conservation is considered as a reliable indicator of functional relationships, and is therefore a promising approach for improving the gene function prediction. Methods. Here we present NetSyn (Network Synteny), a tool to group protein sequences based on the conservation of their genomic context rather than solely on sequence similarity. From a list of protein sequence identifiers, NetSyn searches corresponding genome entries to retrieve neighboring genes. Corresponding protein sequences are grouped into families to define homology relationships and compute a synteny conservation score between the different extracted genomic contexts. A network is then created in which the nodes represent the input proteins and the edges indicate that two proteins share a conserved synteny. Finally, the network is partitioned into clusters grouping proteins with similar genomic contexts, using a community detection algorithm. Results. As a proof of concept, we used NetSyn on two different datasets. The first one is the BKACE protein family (formerly named DUF849) which has previously been divided into isofunctional sub-families. NetSyn was able to go a step further by providing additional sub-families beyond those already described. The second dataset corresponds to a set of non-homologous proteins belonging to three different glycoside hydrolase (GH) families. These GHs are known to work cooperatively in a Polysaccharide-Utilization Loci (PUL) and are therefore grouped together in the same genomic contexts. NetSyn was able to identify a locus grouping 3 GHs, involved in the degradation of xyloglucan, in 162 prokaryotic genomes. Discussion. By highlighting conserved synteny in distantly related prokaryotic species, NetSyn enables functional links between proteins to be established beyond sequence similarity alone. We showed that NetSyn is efficient for exploring large prokaryotic protein families, enabling the definition of isofunctional groups and the identification of functional interactions between non-homologous enzymes. These features enable the prediction of new genomic structures that have not yet been experimentally characterized. Finally, NetSyn is also useful for pinpointing annotation errors that have been propagated across databases, and for suggesting annotations on proteins lacking functional prediction. NetSyn is freely available at https://github.com/labgem/netsyn.
bioinformatics2026-08-24v5Scaling genome annotation across the eukaryotic tree of life with OrionGeno
Liu, L.; Cai, X.; Wang, S.; Deng, Y.; Wu, Y.; Pan, Y.; Wang, J.; Zhang, C.; Xia, H.; Tan, N.; Su, K.; Liu, Y.; Zhou, X.; Liu, L.; Wei, T.; Zhang, Y.; Li, Q.; Li, Y.; Yin, P.; Xu, X.Abstract
The rapid expansion of eukaryotic genome sequencing has created an urgent demand for accurate and scalable genome annotation. Existing ab initio methods often struggle to reconstruct complex gene architectures and generalize across distant lineages, limiting their use for large-scale annotation. Here we present OrionGeno, a phylogeny-aware deep learning model for end-to-end eukaryotic genome annotation. OrionGeno integrates phylogenetic context, long-range sequence modeling and joint prediction of gene structures and repetitive elements to annotate exons, introns, untranslated regions and repeats directly from genomic sequences. Applied to chromosome-level eukaryotic genomes from NCBI that lack annotations, OrionGeno generates annotations for more than 5,300 genomes, substantially expanding public annotation resources. Across diverse eukaryotic lineages, OrionGeno outperforms state-of-the-art methods at the exon, gene, protein-sequence, and protein-structural levels. It also identifies candidate protein-coding loci absent from reference protein-coding annotations in well-curated genomes. Together with a web platform and integrated annotation database, OrionGeno provides a scalable and accessible framework for translating genome assemblies into functional biological resources and supporting large-scale biodiversity initiatives such as the Earth BioGenome Project.
bioinformatics2026-08-24v2Interpolating and Extrapolating Node Counts in Colored Compacted de Bruijn Graphs for Pangenome Diversity
Parmigiani, L.; Peterlongo, P.Abstract
A pangenome is a collection of taxonomically related genomes, often from the same species, serving as a representation of their genomic diversity. The study of pangenomes, or pangenomics, aims to quantify and compare this diversity, which has significant relevance in fields such as medicine and biology. Originally conceptualized as sets of genes, pangenomes are now commonly represented as pangenome graphs. These graphs consist of nodes representing genomic sequences and edges connecting consecutive sequences within a genome. Among possible pangenome graphs, a common option is the compacted de Bruijn graph. In our work, we focus on the colored compacted de Bruijn graph, where each node is associated with a set of colors that indicate the genomes traversing it. In response to the evolution of pangenome representation, we introduce a novel method for comparing pangenomes by their node counts, addressing two main challenges: the variability in node counts arising from graphs constructed with different numbers of genomes, and the large influence of rare genomic sequences. We propose an approach for interpolating and extrapolating node counts in colored compacted de Bruijn graphs, adjusting for the number of genomes. To tackle the influence of rare genomic sequences, we apply Hill numbers, a well-established diversity index previously utilized in ecology and metagenomics for similar purposes, to proportionally weight both rare and common nodes according to the frequency of genomes traversing them.
bioinformatics2026-08-24v2Integration of proteomic data from cell lines and tumors
Ta, C. Q.; Auth, J. M.; Schilling, M.; Klingmueller, U.; Raue, A.Abstract
Cancer cell lines are widely used in preclinical research, yet the clinical translation of findings from cell lines remains limited. Identifying cell lines that best resemble patient tumors requires integration of molecular profiles across biologically distinct sample types. Advances in transcriptomic integration have demonstrated the potential of deep learning for aligning data across different sample types. However, comparable approaches for proteomic data integration remain lacking, potentially because of the prevalence of missing values in proteomic datasets. Here, we introduce ProtInt, a deep learning-based framework that integrates proteomic data by combining principles from proteomic imputation and transcriptomic integration methods. We applied ProtInt to integrate label-free proteomic profiles from 771 cancer cell lines and 550 treatment-naive tumors, and showed that ProtInt outperformed batch correction and transcriptomic integration methods. Comparison of the cell line proteomes before and after integration revealed recurrent increase of proteins associated with immune reaction and reduction of proteins involved in mitochondrial gene expression as proteomes of cell lines were adapted to resemble tumors. These results establish ProtInt as a framework for joint analysis of proteomic datasets across distinct sample types and may facilitate the identification of cell lines best suited for clinically relevant studies.
bioinformatics2026-08-24v2Virtual-cell models compress unseen intervention geometry through a target-specific generalization bottleneck
Huang, Y.; Wang, H.; Wilson, P. C.Abstract
Predictive models of cellular perturbation are often judged by how closely they reconstruct molecular states after unseen interventions. We show that high state-level similarity can coexist with loss of the relationships that distinguish perturbations, a failure we term Intervention Geometry Compression (IGC). Across established models and perturbation settings, unseen interventions show weakened global and local geometry, reduced between-intervention variance and spectral collapse. The failure is not primarily explained by response-space capacity. Instead, diagnostic projections localize much of the missing geometry to a small number of residual response directions learned from seen interventions; these directions outperform complexity-matched random subspaces and replicate in an independent Jiang perturbation resource. Polarity captures part, but not all, of this continuous orientation signal. Time-resolved analyses further show that correct trajectory entry markedly improves downstream propagation, while a held target's own early empirical response rapidly reveals endpoint orientation. Finally, same-target empirical anchoring transfers intervention identity across contexts far more effectively than increasing exposure to other interventions. These results identify intervention-coordinate assignment as an information bottleneck in virtual-cell generalization and support a design principle: empirically anchor intervention identity, then use models to generalize anchored effects across cellular contexts.
bioinformatics2026-08-24v1Thal-Kak: unifying biomolecular structure predictors reveals a sampling-selection gap
Bae, J.; Jo, S.; Kim, Y.; Kim, D.; Kim, K.; Park, S.; Park, S.; Myung, S.; Shin, H.; Kim, M. H.; Kang, M.; Baek, M.Abstract
Complementary all-atom structure predictors sample different solutions, but how to allocate a fixed sampling budget across them and select the best output remains unclear. Thal-Kak unifies five released predictors under shared upstream inputs and a common schema. Across FoldBench and CASP16, model mixing improves oracle sampling over single-model runs, but selection remains a bottleneck because confidence scores do not transfer across models and existing quality-assessment methods cannot resolve this gap.
bioinformatics2026-08-24v1Point-in-time evidence and cross-area clinical precedent anticipate clinical entry across 100 focal areas: retrospective validation of the Intangia triage layer
Elliott, T. O.; Molnar, S.; Peeters, G.; Collart, O.Abstract
Early-opportunity teams face a combinatorial problem: once a focal target, mechanism or indication is fixed, the space of plausible partners runs to thousands of candidates per area. Intangia's triage layer ranks that space from point-in-time evidence (how much literature, patent and clinical activity a candidate pairing has accumulated, and whether the partner already has clinical precedent in other contexts) so that review starts where clinical activity is most likely to begin next. This preprint validates that capability retrospectively across 100 focal areas spanning drug targets, mechanisms and disease indications, replaying 24.1 million historically scored combination-years with every area scored by a model trained on the other 99 and never on itself. The headline is operational. At a twenty-partner review shortlist per focal area, the median area's four-year first-alert precision is 0.234, against a matched random-ranker median of 0.008: roughly one in four shortlisted partners subsequently entered the focal clinical context within four years, about 38 times each area's own background rate (95% CI 31 to 45). A panel-level permutation puts the result at p = 0.0005. Discrimination generalises: the full 13-feature specification reaches a median leave-one-focal-out ROC-AUC of 0.922 (95% CI 0.911 to 0.929), with no area below chance and all 100 areas beating their strongest count-based baseline. Shortlisted entrants are anticipated with a median observed lead of two years within the evaluation window, and three years (interquartile range one to five) once the window cap is removed and every realised entrant is counted. The core ranking is carried by two interpretable signal families: cumulative co-occurrence counts and leave-one-area-out clinical precedent. Burst detection serves a complementary role: it supplies the time-stamped, source-specific momentum evidence attached to every recommendation (what is accelerating, and why now) rather than additional ranking power. A conditional view of the same landscape ranks candidates with no cross-area precedent against one another, enriched relative to matched random ranking, supporting a lower-yield emerging-opportunities capability. Two worked examples, PD-1 combination immunotherapy and CTLA-4, are point-in-time historical replays of the same architecture in familiar territory, showing what an alert looked like with the dated evidence behind it. The endpoint throughout is first clinical entry, not clinical success; prospective validation is the next stage.
bioinformatics2026-08-24v1Erosion of regenerative regulation: age-associated shifts in the skeletal muscle fiber epigenome and transcriptome
Moo, K. G.; Orchard, P.; Varshney, A.; D'Oliveira Albanus, R.; Manickam, N.; Kinnunen, L.; Lakka, T.; Saramies, J.; Laakso, M.; Tuomilehto, J.; Mohlke, K.; Boehnke, M.; Scott, L.; Koistinen, H.; Collins, F.; Parker, S.Abstract
Skeletal muscle aging is characterized by the deterioration of muscle function, which can lead to negative quality-of-life outcomes including frailty and sarcopenia. While understanding the mechanisms of this process is increasingly important as the global population ages, previous molecular studies of skeletal muscle aging have been limited by statistical power and cell type resolution. In this study, we analyzed single-nucleus gene expression and chromatin accessibility data from 287 human skeletal muscle samples from individuals aged 20-79 years to explore sex- and cell type- specific aging effects. Across 467,126 nuclei from 13 cell types, we identify 384 age-associated genes and 4,061 age-associated chromatin regions. These age-associated molecular features are enriched for functional pathways, including metabolic processes, cell-to-cell communication, and senescence Kyoto Encyclopedia of Genes and Genomes KEGG terms. Age-associated closing chromatin was more common across fiber types and sexes than opening chromatin, and was enriched in active enhancer regions while depleted for active transcription start sites. We observe enrichment for specific transcription factor motifs in closing chromatin, including those of glucocorticoid and androgen receptors, both of which play a key role in the maintenance of healthy skeletal muscle. Together, these findings identify an age-associated regulatory shift, largely invisible in matched transcriptomic data, characterized by closing chromatin which reduces accessibility to hormone receptor binding sites and enhancer regions in the muscle fiber epigenome.
bioinformatics2026-08-24v1Click-Prep: An Interactive Data Preparation Tool for Click-qPCR
Kubota, A.; Tajima, A.Abstract
Click-qPCR is a browser-based application for relative qPCR analysis that requires a tidy-format CSV file containing four columns: sample, group, gene, and Cq. Preparing this input from qPCR instrument output typically requires manual reformatting and calculation of mean Cq values for technical replicates. To simplify this process, we developed Click-Prep (https://kubo-azu.shinyapps.io/Click-Prep/), an interactive web-based application designed specifically to create Click-qPCR input files. Click-Prep imports CSV, TXT, TSV, and XLS/XLSX files and supports skipping of instrument-generated metadata rows, interactive column mapping, and manual assignment of experimental groups. Users can review technical-replicate measurements, exclude selected rows according to predefined quality-control criteria, and calculate mean Cq values for each sample-group-target combination. Missing or nonnumeric Cq values are flagged for review and must be resolved before the mean is calculated. Click-Prep can also combine compatible formatted CSV files, such as datasets obtained from separate qPCR plates. The resulting dataset is exported as a standardized CSV file containing the four fields required by Click-qPCR. By integrating these operations into a guided browser-based workflow, Click-Prep enables users to prepare Click-qPCR input files rapidly and consistently without programming.
bioinformatics2026-08-24v1RevPert: ranking candidate drivers of transcriptomic state transitions via gallery-native reverse perturbation
Liang, S.; Yang, C.; Wang, J.; Li, y.Abstract
Cellular state transitions underlie adaptation, ageing and disease, yet prioritizing catalogued genetic perturbations whose expression signatures match an observed transcriptomic shift remains difficult. Most models predict phenotype from a nominated intervention, whereas genetic inverse benchmarks are largely restricted to within-screen identity recovery. Here we introduce RevPert, a gallery-native reverse perturbation model that ranks a fixed genetic catalog for a query contrast {Delta}Y* = YB - YA by combining signed Pearson connectivity with a learned residual. Across Replogle Essential Perturb-seq (four lines) and LINCS-KO screens (ten lines), RevPert recovered held-out interventions at leading performance relative to matched baselines. Applied to public drug-resistance contrasts in HCC and CML, dual-arm ranking placed pre-specified disease anchors far higher on the expected arms than ranking the same signatures by differential-expression magnitude alone (Essential residual model for HCC; a transductive GWPS residual for CML). RevPert therefore couples within-screen reverse ranking to a screen-external signed-geometry check; the latter calibrates literature anchors and is not claimed as held-out recovery.
bioinformatics2026-08-24v1Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery
Seal, S.; Zalte, A. S.; Araripe, D. A.; Gomes, R. A.; Korani, D.; Shekhar, M.; Siramshetty, V. B.; Patra, A.; Mou, Z.; Yu, X.; Kuhn, D.; Weskamp, N.; Ash, J.; Cheng, A. C.; Fang, C.; Price, D.; Aldeghi, M.; Rodriguez-Perez, R.; Clevert, D.-A.; Engkvist, O.; Deibler, K.; Rouquie, D.; Reutlinger, M.; Richmond, N. J.; Ainsley, J.; Ledeboer, M.; Green, W. H.; Bender, A.; Wognum, C.Abstract
Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.
bioinformatics2026-08-24v1IMMF: An Interpretable Multi-Modal Framework for Hypothesis-Driven Biomarker Discovery in Triple-Negative Breast Cancer Using Public Data
Imran, A.; Rahat Hossain, K. M.; Islam, S. M. R.; Rahman, M. S.Abstract
Triple-Negative Breast Cancer (TNBC) is characterized by high heterogeneity, poor prognosis, and limited targeted treatment options. Bridging the gap between molecular alterations and histopathological morphology remains a major challenge in precision oncology. We propose an interpretable, multi-modal framework that integrates histopathological image analysis with multi-omics profiling (somatic mutations, DNA methylation, copy number alterations), leveraging U-Net-based nuclei segmentation, vision-language models (BLIP), biomedical language models (BioGPT), and explainable AI (SHAP, LIME). Our framework achieves strong predictive performance (AUC = 0.989) and provides transparent, biologically grounded interpretations by integrating morphological features with genomically prioritized biomarkers. Cross-modal analysis confirms established TNBC drivers and generates novel, testable hypotheses associating specific epigenetic alterations with distinct morphological phenotypes. While causal validation requires future wet-lab experiments, our framework accelerates hypothesis-driven biomarker discovery by integrating complementary data modalities with language-based reasoning, providing a transparent foundation for hypothesis generation and clinical translation.
bioinformatics2026-08-24v1A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models
Li, L.; Duan, S.; Zha, X.; Ye, F.; Zhang, Y.; Zhang, X.; Cao, Y.; Liu, C.Abstract
Single-cell drug perturbation models are increasingly used to predict how compounds remodel cellular states, but they are still largely assessed by expression reconstruction. Whether high expression similarity reflects preservation of drug-response signatures remains unclear. Here we present scDrugPerturb-Bench, a mechanism-annotated benchmark that links matched control and drug-treated single-cell RNA-sequencing profiles to literature-curated directional key-gene evidence. The resource covers 181 datasets, 423 annotated response cases, 717 unique key genes and 2.5 million cells. We introduce the Mechanism Fidelity Score (MFS) to evaluate key-gene direction, effect-size recovery, gene-set coherence, mechanism specificity and pathway-level response polarity. Across 12 perturbation-prediction models, 3 baselines and 10 data splits, expression-similarity metrics were weakly aligned with MFS and selected different model configurations. Mechanism-aware selection improved early drug retrieval in a transcriptome-based drug design evaluation, indicating that MFS provides practical information beyond benchmark reporting. Systematic benchmarking revealed limited fidelity to drug-response signatures across cell-line and source-integrated settings. Frozen single-cell foundation model embeddings produced local, metric-dependent gains rather than universal improvements, and source context substantially reshaped model assessment. Hard-negative tests further showed that plausible perturbation responses can arise from non-specific transcriptional shortcuts. These results show that expression reconstruction is an insufficient proxy for preserving drug-response signatures and establish scDrugPerturb-Bench as a benchmark for mechanism-aware evaluation of single-cell drug perturbation models.
bioinformatics2026-08-24v1Gene identity, not variant effect, dominates ClinVar benchmarks of missense pathogenicity predictors
Arnoult, D.; Harrizi, S.; Nait Irahal, I.; Mostafa, K.Abstract
Missense pathogenicity predictors are routinely benchmarked against ClinVar, whose labels are strongly structured by gene: genes under diagnostic scrutiny accumulate pathogenic submissions while incidentally sequenced genes accumulate benign ones. We asked how much of a benchmark score this structure alone can produce. On 197,904 ClinVar missense variants validated against UniProt canonical sequences, a null model using no variant-level information, scoring each variant only by the pathogenic fraction of its own gene, reaches an area under the receiver operating characteristic curve (AUROC) of 0.921 under a random 10-fold split. On a common intersection of 169,989 variants, four current predictors exceed it by only 0.036 to 0.044. The inflation is not uniform, so it does not cancel when predictors are compared: under within-gene evaluation the ranking inverts, AlphaMissense rising from third to first and gMVP falling to third (p < 0.0001). The inversion survives removal of ceiling genes and replicates on an independently curated benchmark. Because both rankings derive from the same ClinVar labels, we arbitrated between them using data with no gene-level structure: agreement with 47 human deep mutational scanning assays matches the within-gene ranking and inverts the conventional one (p = 0.027, 0.0023). Across twenty-two dbNSFP predictors scored on one common intersection of 112,248 variants, with each tool's exposure to clinical labels registered before any score was extracted, predictors never trained on such labels sit 0.051 AUROC behind supervised ones globally but only 0.026 behind within genes (difference +0.025 [+0.023, +0.027], p < 0.0001). Leave-one-out correction, the standard remedy, is worth 0.002 AUROC. Much of ClinVar benchmark performance reflects gene identity rather than variant effect, and the distortion changes which predictor a benchmark ranks first, in a direction experimental data contradicts. We release genenull, a single-file implementation, so reporting this baseline costs one function call.
bioinformatics2026-08-24v1A Beta-Binomial Model for Estimating Zero- or One-inflated Pain Trajectories
Liu, Y.; Harris, R. E.; Clauw, D.; Bayman, E.; Leroux, A.; Lindquist, M. A.Abstract
Chronic pain is a widespread public health issue that imposes substantial health, emotional, and economic burdens on individuals and communities. Because pain is subjective and lacks objective biomarkers, it is typically measured using patient-reported scores, often on a numerical scale from zero to ten. Increasingly, pain studies use ecological momentary assessment, with multiple daily assessments over days and across study phases (e.g., a series of baseline and post-intervention assessments). These data frequently show many ratings at the extremes (i.e., at minimum or maximum pain scores), commonly referred to as zero- and one-inflation in the statistical literature, along with considerable within-person variability both within and across days. These phenomena present challenges for statistical analyses, as they violate assumptions of most commonly used statistical techniques (e.g., the normality assumption of linear mixed models). We propose a Bayesian beta-binomial mixed-effects model for modeling potential zero- or one-inflated pain scores while accounting for variability using random effects on the mean and variance parameters across subjects. A simulation study demonstrates that the method accurately estimates model parameters across realistic sample sizes, time points, and zero- and one-inflation levels. An application to data from two longitudinal pain studies demonstrates that the model fits the data better and, when correctly specified, yields accurate uncertainty intervals for longitudinal changes in pain compared to existing models, especially for zero- and one-inflated outcomes. Additionally, the model directly estimates the probability of clinically meaningful pain events. The proposed method provides a powerful statistical framework for studying the patient-reported pain trajectories.
bioinformatics2026-08-23v3EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-08-23v3Sparse autoencoder features from InterPLM predict neuropeptide precursors among secreted proteins
Kulikova, A. V.; Bookout, A. L.; Koch, T. L.; Safavi-Hemami, H.Abstract
Neuropeptides are a diverse class of short, secreted signaling molecules that regulate key physiological processes in animals. Despite their important biological roles and increasingly recognized therapeutic value, the discovery of new neuropeptides remains challenging, largely because their short length and high sequence heterogeneity limit the effectiveness of motif- and homology-based approaches. Here, we present a pipeline for neuropeptide precursor prediction that leverages sparse autoencoders (SAEs) from InterPLM to decode dense ESM-2 protein language model embeddings into sparse, disentangled features. We identify a small subset of features strongly associated with neuropeptide precursors that achieve high discriminative performance. A logistic regression classifier trained on this reduced feature set, accurately separates human neuropeptide and non-neuropeptide sequences. We then applied this classifier to important model organisms: mouse (Mus musculus), zebrafish (Danio rerio), nematode (Caenorhabditis elegans), and fruit fly (Drosophila melanogaster ) and show that the approach generalizes across diverse species. Overall, InterPLM SAE features provide an interpretable and effective strategy for neuropeptide prediction and enable a trained classifier to predict neuropeptides from large datasets. A web tool for this classifier is freely available at https://biolib.com/ATGCACTGTTCAGGCCTC/SAE Neuropeptide-Predictor
bioinformatics2026-08-23v2RapidMACS: MACS3-identical peak calling, 50x faster
Hung, L.-H.; Yeung, K. Y.Abstract
Motivation: MACS3 is a comprehensive peak-calling toolkit whose subcommands span single- and paired-end data, narrow and broad peaks, and a range of signal-track utilities. However, many ATAC-seq and multiomic pipelines, including our own, use just one of those capabilities: narrow peak calling. A single analysis may call peaks under different conditions, so a per-call saving is multiplied and the time recovered can be substantial. Furthermore, we wanted an efficient and embeddable narrow peak caller that could be integrated directly with our Chromap Suite aligner, so we wrote RapidMACS. Results: RapidMACS is a narrow peak caller optimized for this purpose, building its signal tracks in a single lazy sweep that avoids a global sort, processing chromosomes in parallel, and keeping intermediates in memory rather than creating temporary files. In benchmarks involving single-cell ATAC-seq, bulk ATAC-seq, ChIP-seq, and CUT&RUN, it is 3.5 to 51 times faster than MACS3 v3.0.3 depending on the applications and number of threads used. More importantly, the output is byte-identical. This means that any downstream analysis using RapidMACS will produce identical results to those using MACS3. The speed gains are largely due to these algorithmic changes rather than the choice of language, because MACS3's peak-calling code is itself compiled (Cython). RapidMACS depends only on htslib and zlib and links as a small static archive without a Python or Cython runtime. Availability and implementation: RapidMACS is open source under the MIT license at https://github.com/morphic-bio/rapidmacs, with both a standalone CLI executable (rapidmacs) and a linkable C++ library. Prebuilt containers for x86-64 and arm64 are published as biodepot/rapidmacs on Docker Hub.
bioinformatics2026-08-23v2Delta Marches: Generative AI based image synthesis to decode disease-driving morphologic transformations.
Nguyen, T. H.; Panwar, V.; Jarmale, V.; Perny, A.; Dusek, C.; Cai, Q.; Kapur, P. H.; Danuser, G.; Rajaram, S.Abstract
Deep learning reveals that tissue morphology contains rich pathophysiological information beyond human understanding. However, approaches to convert these spatially distributed signals into subcellular insights informing disease mechanisms are lacking. We introduce Delta-Marches, an interpretability-first approach that nominates distinguishing morphological features rather than explaining existing models' decisions. Delta-Marches simulates idealized morphological changes between classes by coupling latent-space traversals with generative AI. Comparing each image to its class-shifted counterpart allows feature extractors to infer the most affected aspects, reducing sample-to-sample variability and resolving transformations at subcellular resolution. Prototyped in renal carcinoma grading, Delta-Marches generates realistic grade transitions and pinpoints tumor-cell nuclear phenotypes as key determinants. It also reveals reduced vasculature with increasing grade, a known pattern absent from standard rubrics. Applied to dysplastic progression in colorectal tissue, it recovers glandular remodeling and goblet-cell loss, generalizing across tissue types and scales. These results show Delta-Marches parses complex image phenotypes and catalyzes hypothesis generation.
bioinformatics2026-08-22v5Population dynamics of ribozymes during in vitro selection
Brooks, E. M.; Kotlin, N. R.; Weiss, Z.; DasGupta, S.Abstract
Ribozymes are central to models of RNA-based primordial life. Understanding how ribozymes emerge from RNA populations under changing selection pressures can help delineate the biochemical constraints and capabilities of early RNA-based biology. In vitro selection has been instrumental in isolating ribozymes with diverse functions from combinatorial populations; however, their population dynamics during selection remain poorly characterized. By analyzing high-throughput sequencing data produced from all rounds of an in vitro selection experiment, consisting of over 5 million unique sequences, we tracked how sequences rose and fell in abundance during selection. We found that the most abundant sequences emerged early and maintained their dominance, suggesting that only a few cycles of selection-amplification may be sufficient to isolate well-adapted ribozymes. Characterization of the catalytic activities of the dominant sequences in each selection round indicated that abundance did not always predict activity. This observation was reinforced through nucleotide conservation analysis and activity assays of close sequence variants of the most dominant ribozyme. Our analyses show that sequence complementarity with the substrate emerged as a common feature as early as the first round and persisted throughout, highlighting early convergence of diverse sequence populations on a beneficial catalytic strategy. By combining bioinformatics and experimental approaches, our work presents a detailed description of ribozyme population dynamics during an in vitro selection experiment and provides a high-resolution view of how functional RNA sequences compete for survival and propagation under selection pressure.
bioinformatics2026-08-22v2Protal: Ultra-fast metagenomic profiling and strain-resolved analysis
Fritscher, J.; Duncan, A.; Hildebrand, F.Abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal -- profiling through alignment -- an ultra-fast alignment-based method for species-and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom benchmarks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler, respectively). At the strain level, protal reconstructs intraspecific phylogenetic relationships among detected bacteria with similar accuracy as StrainPhlAn 4; because protal produces precise alignments, the phylogenies can be de novo produced without reliance on reference strain collections. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
bioinformatics2026-08-22v2Real-time repository-scale spectral search and global molecular networking with HNSW-MS
Semenov, A.; Roberts, A. M. P.; Coler, E. A.; Gupta, S.; Kopylova, E.; Melnik, A. V.; Boginski, V.; Aksenov, A. A.Abstract
Spectral similarity comparison is the basis of mass spectrometry-based metabolomics, underpinning library matching, molecular network construction, and repository searches such as MASST. Public repositories now exceed billion spectra, and exhaustive pairwise comparison is increasingly limiting. We introduce HNSW-MS, which implements Hierarchical Navigable Small World graph indexing natively for mass spectra. It operates directly on GC-MS and LC-MS/MS spectra without embedding, so scores are identical to those of exhaustive comparison and results are fully reproducible. On the benchmark dataset of 8.4 million LC-MS/MS spectra, HNSW-MS is up to 900-fold faster than linear search, recall@1 and recall@10 achieve over 90%, with the recall-speedup tradeoff tunable at query. Such rapid retrieval allows for any query spectrum to be near-instantaneously integrated into a global molecular network. We demonstrate specific examples of use of global networking for molecular discovery, including uncovering a new pathway of stenothricin biosynthesis and a new dipeptide conjugate bile acid.
bioinformatics2026-08-22v2Mapping pathogenic patterns in membrane transporters from the GLUT transporter family
Kadasova, N.; Martinat, D.; Spackova, A.; Hutarova Varekova, I.; Berka, K.Abstract
Significance Missense mutations can lead to pathological effects in human cells. Predictive methods that account for structural context, such as AlphaMissense, can provide pathogenicity scores. The accumulation of pathogenicity hotspots can reveal important structural features within individual proteins of protein families, such as GLUT transporters. Mapping pathogenicity scores onto the structure can thus provide a mechanistic explanation of the protein function necessary for its role in the cell. Abstract Non-synonymous amino acid substitutions (missense mutations) are common in the general population; some are causative of serious disease. Depending on their structural context, they can disrupt protein function, folding, or dynamics. Computational predictive methods developed in recent years, such as AlphaMissense, provide new insights into how missense mutations affect protein structure by predicting and mapping their pathogenicity across each amino acid in the human proteome. In this study, we identify recurring patterns of pathogenicity prediction across the GLUT family membrane transporters encoded by genes SLC2A1-14. Within the GLUT transporter family, we observe higher pathogenicity profiles in the transmembrane domains, particularly in pore-lining and binding-site residues. Predicted missense pathogenicity is elevated throughout residues assigned to the central cavity, suggesting sensitivity of the transport pathway. Another finding shows higher pathogenicity in specific transmembrane helices of the protein, with the same pattern across all proteins. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family. We validated these predictions against clinically observed missense variants from ClinVar and benchmarked three prediction methods. AlphaMissense showed the strongest concordance with clinical classifications (AUC = 0.88), and the structural distribution of clinically reported pathogenic variants broadly recapitulated the predicted pathogenicity landscape. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family, and in select cases, clinical and predicted data diverged, suggesting more localized functional constraint than genome-wide prediction alone would indicate. These findings show that the pathogenicity of glucose transport within the GLUT family may be shaped by functional redundancy and physiological essentiality across GLUT groups, supported by convergent evidence from structural, computational, and clinical variant data.
bioinformatics2026-08-22v2AlphaConformers: Structure-guided sampling enables prediction of multiple protein conformations
Daniel, J.; Vitoriano De Queiroz Lira, L.; Zea, D. J.Abstract
Proteins are dynamic molecules capable of adopting multiple conformations. However, AlphaFold2 predominantly generates models around a single conformation, usually representing a ligand-bound state. To address this limitation, we developed AlphaConformers, a structure-guided pipeline that steers AlphaFold2 toward alternative conformations. It is based on the idea that protein structure databases can capture the structural space accessible to members of a protein family. Given a target protein, AlphaConformers retrieves structures from structurally similar proteins. These structures are organized into structure-based alignments and template sets, which are supplied to AlphaFold2 as conformational hypotheses. The resulting models are clustered and filtered, facilitating their analysis. Evaluated on a curated benchmark of 88 proteins with known ligand-bound and unbound conformations, AlphaConformers expanded AlphaFold2 conformational sampling and recovered alternative states missed by AlphaFold2 and other state-of-the-art methods. AlphaConformers ranked first for modelling subtle conformational changes commonly observed between ligand-bound and unbound states. These results show that structural information from protein databases can be leveraged to steer AlphaFold2 toward alternative conformations.
bioinformatics2026-08-22v1Detecting and typing Chlamydia trachomatis strains in metagenomes using the MetaChlam pipeline
Sharma, P.; Dean, D.; Read, T. D.Abstract
The Gram negative bacteria Chlamydia trachomatis (Ct), an obligate intracellular human pathogen, is a predominant cause of sexually transmitted infections and ocular trachoma globally, exerting a significant impact on public health. Ct "strains" (major lineages within the species) are known to have different tissue tropisms and be associated with different disease outcomes. Metagenome samples from typical sites where Ct infects (e.g., endocervix, conjunctiva, rectum) rarely contain enough reads for traditional genotyping methods such as Multi-Locus Sequence Typing (MLST) or ompA genotyping. To overcome these limitations, we implemented an ensemble tool called MetaChlam that can accurately classify Ct strains with as few as 250 Ct reads. Using 109 publicly available Ct genomes from naturally circulating strains, we established that an ANI-based threshold of 99.75% was capable of distinguishing Ct strains from each other. We implemented metagenome-based typing using the previously developed LINtax, Strainscan, StrainGE, and Sourmash softwares. MetaChlam integrated the four tools along with custom databases into an automated nextflow pipeline. Using simulated metagenomic reads, we found that our pipeline accurately identified the correct strains in both single strain and multi-strain mixtures of samples. Finally, we showed that MetaChlam had higher specificity for the true presence of Ct reads in NCBI SRA metagenomic datasets than NCBI PebbleScout software. A surprising finding of these analyses was that reads from Ct, an obligate human intracellular pathogen, can be found as contaminants in samples from sites where the organism is almost certainly not present. Overall, our study enhances the characterization and classification of Ct strains and provides protocols for identification and typing of Ct in shotgun metagenome data. The MetaChlam pipeline is available on Github: https://github.com/parul-sharma/MetaChlam.
bioinformatics2026-08-22v1CLEAR-ST: Physics-informed probabilistic decontamination of spatial transcriptomics by modeling mRNA lateral diffusion
Ma, K.; Huang, Y.; Ho, J. W. K.Abstract
Spatial transcriptomics is a rapidly evolving technology that allows for the measurement of gene expression in a spatially resolved manner. However, one technical problem that occurs for many sequencing-based spatial transcriptomics platforms is the presence of mRNA lateral diffusion, where mRNA from one spot can bind to probes in another spot, leading to contamination and inaccurate gene expression measurements. In Visium-like assays, this artifact is often visible as structured out-of-tissue signal and boundary-associated expression halos, yet its magnitude, spatial decay, and directional bias vary substantially across samples. Here, we present CLEAR-ST, a physics-informed probabilistic framework for correcting diffusion-like contamination in spatial transcriptomics data. CLEAR-ST infers a latent clean expression field using a denoising autoencoder and links it to the observed counts through a graph-Laplacian forward contamination model with learnable diffusion parameters, finally evaluated with a selectable count likelihood. We first conducted a comprehensive comparison between 10X official and independently generated Visium samples, demonstrating that out-of-tissue count profiles are highly related to nearby in-tissue expression, more concentrated near tissue boundaries, and diffusion directions across genes are likely coherent. Across real samples with varying contamination burden, CLEAR-ST improved spatial domain recovery, increased gene-level spatial autocorrelation, and enhanced the biological specificity of downstream analyses such as marker gene discovery, pathway identification and cell type deconvolution. Compared to benchmark methods, CLEAR-ST showed consistent gains in clustering quality and concordance with manual annotations. Together, CLEAR-ST provides an interpretable and practical approach for diffusion-aware correction of capture-based spatial transcriptomics data.
bioinformatics2026-08-22v1Cross-Kingdom Multi-Omics Harmonization Uncovers Coordinated Host Defense and Vector Small RNA Regulatory Networks in Begomovirus Transmission
Badeli, G.; Kaboosi, K.; Mohebbi, A.; Nasrollanejad, S.Abstract
Begomoviruses present severe threats to global crop production through complex vector-mediated transmission by the whitefly Bemisia tabaci to host plants such as tomato (Solanum lycopersicum). Unraveling the molecular dialogue between host immune activation and vector non-coding RNA networks is essential for identifying key drivers of virus persistence and transmission. Public host transcriptomic (GSE309527) and vector small RNA (sRNA) sequencing datasets (GSE111343) were processed through a multi-omics harmonization and signal calibration pipeline. Differential expression analysis was performed using empirical Bayes moderated linear models, followed by non-parametric Spearman rank correlation modeling ({rho}) to infer cross-kingdom co-expression dynamics and pathway enrichment profiling across host and vector bio-systems. Harmonized principal component analysis showed clear separation by infection status across host plant and vector cohorts. Differential expression analysis identified 138 significantly altered host genes (69 upregulated, 69 downregulated) and 130 differentially expressed vector sRNAs (65 upregulated, 65 downregulated). Host responses were dominated by significant upregulation of gene-silencing machinery, including Suppressor of Gene Silencing 3 (SGS3; log2 FC = 3.67, q = 7.47 x 10-5), and pathway enrichment in Jasmonate defense (q = 0.0004) and RNA Interference & Silencing (q = 0.0001). Vector sRNAs exhibited targeted dynamic alterations, with pathway enrichment in Salivary Gland Secretion ($q = 0.0030) and Gut Endosymbiont Response (q = 0.0210). Cross-kingdom correlation modeling revealed two distinct, highly anticorrelated regulatory modules (mean |{rho}| = 0.76). Host SGS3 expression strongly correlated with vector sRNA VEC_0080 ({rho} = 0.9762) and virus-derived siRNA Bt-vsiRNA-01 ({rho} = 0.7619). These findings demonstrate a tightly synchronized tripartite molecular crosstalk between host antiviral immunity, viral siRNA accumulation, and vector small RNA remodeling. These cross-kingdom regulatory modules highlight promising targets for dual-action RNA interference strategies aimed at controlling Begomovirus transmission.
bioinformatics2026-08-22v1LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Liu, A.; Ho, A.; Droste, A. M.; Martin, D.; Wong, E.; Zhou, E.; Zhou, I.; Park, J.; Jiao, J.; Skelly, K.-R.; Kim, K.; Li, J.; Rao, K.; Uehara, M.; Marion, M.; Fitzgerald, N.; Dias, R.; Shringarpure, S.; Yuan, Y.; Wang, Y.Abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.
bioinformatics2026-08-22v1A Persistent Fleet of AI Scientists Exhibits Cooperative and Autopoietic Behavior
Patel, M. S.; Wierson, W. A.; Ekker, S. C.Abstract
Scientific work depends on memory, provenance, and continuity across projects, yet most agentic scientist systems are evaluated in bounded workflows or short benchmark runs. We describe a persistent fleet of cooperative AI scientist agents that operated continuously for nearly six months using shared memory, tools, and cross-agent communication. Critically, failures identified during longitudinal scientific research in this fleet prompted an advanced and recursively improving persistent memory architecture (MoE) that, with agentic science workflows, induced the generation of a novel trust architecture for the enablement of full provenance across all agentic scientific operations. An identity-level fabrication constraint reduced delusion-reinforcement probe failures from 91.7% to 0%, and a verification pipeline reduced wrong-topic citation hallucination more than 14-fold in companion benchmarks. This high provenance enabled the use of project memory systems to improve a local open-weight model on internal benchmarks from 44% to [~]90% through the deployment of fleet-specific institutional knowledge. This high-fidelity data environment also supported to date 104 recurring multi-phase reasoning cycles and produced 43 manually curated hypotheses, including cross-domain convergence events and a self-correcting rare-disease pharmacological chaperone-design case. While by design the fleet did not achieve unconstrained autonomous self-improvement or full autopoiesis, we term this bounded pattern AI Autopoietic Behavior due to the recurring operational improvement mediated by internal feedback and retained through institutional records with high confidence. Together, persistent memory, trusted provenance, and recursive learning shifted these agents from episodic assistants toward accountable, long-term scientific collaborators.
bioinformatics2026-08-22v1Design-informed Size Factor Estimation
Pocuca, T.; Pare, G.; Bolker, B. M.Abstract
Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.
bioinformatics2026-08-22v1Bridging Biomedical Atlas Ecosystem: Cross-Atlas Alignment And Scalable Tissue Specimen Registration
Jain, Y.; Desai, B.; Qaurooni, D.; Bhavsar, A.; Kienle, P.; Pouch, A. M.; ONeill, K.; Apte, S.; Herr, B. W.; Fisher, S. A.; Börner, K.Abstract
Over the last five years, over 13,000 tissue datasets with 200+ million cells from 20 consortia have been spatially registered into the Human Reference Atlas (HRA) common coordinate framework (CCF). The shared 3D spatial and semantic reference system enables exploration of datasets in the context of all other data across organs, assay types, and spatial scales. However, manual registration of individual samples remains resource intensive, posing feasibility challenges exacerbated by the proliferation of samples, assays, and atlasing efforts. This paper presents two approaches to scale up HRA construction: (1) projecting data across biomedical reference atlas systems and (2) using millitomes to bulk register tissue blocks into a reference organ. Both methods use the AutoMated Alignment and Projection (AMAP) pipeline to align 3D mesh models using point cloud registration. We demonstrate the evolving HRA-aligned atlas ecosystem for 6 models from the SPARC Program (heart), Gut Cell Atlas (large intestine), 500-subject consensus kidneys, and the Julich Brain Atlas. Additionally, we used AMAP to project 7 millitome models across 5 organs onto the HRA ecosystem, integrating 300+ tissue extraction sites. AMAP enables scalable tissue registration of data across atlas ecosystems enabling the construction of detailed reference maps of the human body.
bioinformatics2026-08-22v1