Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
GlyComboCLI enables command line-based FAIR workflows for glycan composition assignment in mass spectrometry data
Kelly, M. I.; Thang, W. C. M.; Pang, C. N. I.; Gustafsson, O. J. R.; Ashwood, C.Abstract
Glycans are integral biomolecules whose presence cannot be predicted from genomic data alone, necessitating experimental characterisation through approaches including mass spectrometry. Assignment of glycan compositions to observed mass to charge ratios is computationally challenging due to the potential monosaccharide diversity and existing tools lack the required flexibility for integration into automated bioinformatic workflows. Here, we present GlyComboCLI, an open-source command-line application for the assignment of glycan compositions to mass spectrometry data which expands upon the previous GUI application, GlyCombo. GlyComboCLI accepts mass lists and vendor-neutral mzML files, supports diverse monosaccharides, derivatisation states, reducing-end modifications and adducts, and is validated through automated tests spanning supported search parameters, input formats, and instrument vendors. Outputs are compatible with downstream tools including Skyline and GlycoWorkBench while GlyTouCan accessions support persistent identification of registered base compositions. Deployment as a standalone executable, a Docker container, and a Galaxy tool, supports FAIR and reproducible workflows. Applied to published mouse and human glycomics datasets, GlyComboCLI reproduced major qualitative glycomic motifs and relative abundance trends, while Galaxy workflow reproduced expected effects of sialidase treatment. These results demonstrate a flexible, scalable, and reproducible approach for glycan composition assignment within automated glycomics workflows.
bioinformatics2026-08-31v3A self-supervised DNA foundation model with collapse-resistant multimodal fusion
Chen, Y.Abstract
Genomic foundation models pretrained on DNA sequence have achieved strong performance across many tasks, but sequence-only representations cannot fully capture regulatory information from additional DNA-centric modalities. Existing multimodal genomic models are optimized for specific prediction tasks rather than reusable embeddings. Directly fusing heterogeneous modalities is challenging because sparse, peak-shaped regulatory signals and dense sequence embeddings have markedly different statistical structures, making naive alignment prone to near-zero solutions. We present a self-supervised DNA-centric multimodal foundation model integrating DNA sequence embeddings with local and global chromatin accessibility in a shared encoder to produce reusable window-level embeddings. We show that global normalization alleviates this collapse, enabling effective joint learning. The resulting embeddings improve regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection, achieving a 4.6-fold AUPRC improvement over the DNA-only baseline, with further gains on external ClinVar, GTEx eQTL and PBMC caQTL datasets.
bioinformatics2026-08-31v2Decoding heterogeneous aging clocks and disease risk stratification using MetAgeFormer
Xu, Y.; Zou, B.; Xie, G.; Chen, T.; Jia, W.; Zhang, L.Abstract
Metabolomic aging clocks estimate biological age by modeling metabolite concentrations, thereby capturing aging signals from healthspan and adverse outcomes. However, existing clocks generally assume homogeneous aging trajectories and yield only a single age acceleration metric, limiting their capacity to capture inter-individual metabolic heterogeneity and characterize nuanced individual-level representations. To address these limitations, we proposed MetAgeFormer, a transformer-based metabolomic model pre-trained on nuclear magnetic resonance (NMR) metabolomic profiles from over 430,000 participants in UK Biobank via self-supervised learning. This large-scale pre-training enables MetAgeFormer to learn a metabolomic representation space that captures the complex, nonlinear structure of systemic metabolism as reflected in NMR data. Building on MetAgeFormer, we developed a mortality-informed metabolomic aging clock by fine-tuning an attached survival module, deriving age acceleration that demonstrates significant associations with multiple age-related diseases and factors. We further validated zero-shot transfer in the independent Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort. More importantly, we utilized embeddings generated by MetAgeFormer to identify 13 distinct metabolic subtypes and consolidated them into four meta-subtypes with markedly divergent susceptibility profiles for major age-related diseases, particularly type 2 diabetes and neurodegenerative disorders. This finding empirically demonstrated substantial metabolic heterogeneity across populations, persisting even at comparable levels of age acceleration. To enhance clinical applicability, we further employed contrastive learning to distill a lightweight model that approximates the learned metabolomic representation space using only 14 routine clinical blood test measurements as inputs. Both hold-out testing within UK Biobank and external validation in the China Health and Retirement Longitudinal Study replicated similar disease onset patterns across the identified subtypes, underscoring the robust generalizability of MetAgeFormer and supporting its translational potential as a scalable framework for metabolomic aging assessment and early disease risk stratification.
bioinformatics2026-08-31v2JMod: Joint modeling of mass spectra for empowering multiplexed DIA proteomics
McDonnell, K.; Geiszler, D. J.; Wamsley, N.; Derks, J.; Sipe, S.; Cohen, Z. A.; Warinner, L. K.; Yeh, M.; Koo, E.; Leduc, A.; Zwang, T. J.; Specht, H.; Slavov, N.Abstract
Parallelization of data acquisition substantially increases the throughput of mass spectrometry-based proteomics. However, parallelization also increases the density of mass spectra and consequently the overlap between ions, frustrating their analysis. To improve sequence identification and quantification from such spectra, we developed an open-source software for Joint Modeling of mass spectra (JMod). JMod models overlapping peaks as linear superpositions of their components in both MS1 and MS2 space, which permits multiplexed DIA with smaller mass offsets to increase the multiplexing capacity and thus proteomics throughput for a given plexDIA tag. This enables 9-plexDIA using 2 Da offset PSMtags, increasing throughput 9-fold while preserving quantitative accuracy and coverage depth. Furthermore, we use JMod to deconvolve simultaneous labeling by mass tags and heavy amino acids, thus increasing the throughput of metabolic pulse experiments measuring protein synthesis and degradation rates in single cells from mouse liver. By supporting enhanced decoding of highly multiplexed DIA spectra, JMod provides an open and flexible software that increases the throughput of sensitive proteomics.
bioinformatics2026-08-31v2OmniSplice: detection of non-canonical splicing events from RNA-seq
Lannes, R.; Li, R. Y.; Fingerhut, J. M.; Cummings, R. A.; Salagean, A. D.; Yamashita, Y. M. M.Abstract
Splicing generates mature mRNA by removing introns from nascent transcripts and is widely studied using RNA sequencing. However, most RNA-seq analysis pipelines classify RNA-seq reads according to predefined splice-junction structures and discard those that do not conform to such predefined models, potentially obscuring biologically meaningful splicing events. In this study, we developed OmniSplice, a computational framework that captures and analyzes RNA-seq reads that overlap annotated exon ends without assuming predefined splicing architectures. This approach enables systematic detection of non-canonical splicing events that are often overlooked by conventional analyses. Applying OmniSplice to Drosophila splicing factor mutants and mouse TDP-43 mutant datasets, we found widespread splicing defects with non-canonical junctions that were not previously recognized, including back-splicing and trans-splicing. Together, these results demonstrate that RNA-seq datasets may contain a substantial reservoir of overlooked splicing information, warranting more comprehensive approaches for analyzing RNA-seq data for splicing events.
bioinformatics2026-08-31v2Universal physical principles of protein structural genesis emerge in language-model representation space
Chuanyang, L.; Liu, J.; Qiu, X.; Wu, X.; Li, W.; Min, L.; Zhang, G.; Zhang, S.; Zhu, L.Abstract
Protein structure is usually treated as the endpoint of sequence by most AI models, yet its true biological emergence is a process of ordered change, whose logic remains hidden in opaque black box. ProtGenesis creates a bidirectional mirror world: it maps amino acid assembly, elongation and mutation in biological space into quantitative trajectories and ensembles in protein language model representation space, then translates spatial geometry into testable biophysical hypotheses. Across peptides, reporter proteins and protein families, three universal general principles emerged: hierarchical directional assembly (Principle I), quantitative and reproducible structural-emergence trajectories (Principle II) and discrete topological transitions (Principle III) from short- to long-range order. Three novel metrics of spatial features [D,{rho} ,{delta} ], complemented by squared Gaussian Wasserstein-2 measures, located structural anchors, sensitive regions and state transitions, enabling split-protein engineering and programmable protein design. ProtGenesis makes latent representations mechanistically interpretable, offering AI a route to discovering, rather than merely predicting, scientific principles. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=173 HEIGHT=200 SRC="FIGDIR/small/706798v2_ufig1.gif" ALT="Figure 1"> View larger version (60K): org.highwire.dtl.DTLVardef@18f690org.highwire.dtl.DTLVardef@e36903org.highwire.dtl.DTLVardef@35f15org.highwire.dtl.DTLVardef@1577b81_HPS_FORMAT_FIGEXP M_FIG C_FIG
bioinformatics2026-08-31v2The Metabarcoding Analysis Pipeline (MAP): Simple, accurate, and flexible metabarcoding
Prosser, S. W.; Bard, N. W.; Thompson, K. A.; Floyd, R. A.; Padhye, S.; Ozsahin, E.; Jafarpour, S.; Hebert, P. D. N.Abstract
Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP.
bioinformatics2026-08-31v2S2F-Agent: Harnessing sequence-to-function models for verifiable genome interpretation
Li, J.; Qin, T.; Li, J. G.; Bao, Z.Abstract
Sequence-to-function (S2F) models offer a revolutionary paradigm for genotype-phenotype mapping, yet their broader application is bottlenecked by the need for reliable orchestration and interpretation across a fragmented model ecosystem. While general-purpose language models can automate scientific workflows, they are not inherently grounded in the model-specific execution constraints required for robust S2F analysis. Here, we present S2F-Agent, a human-in-the-loop framework designed for the verifiable orchestration of the heterogeneous S2F ecosystems. The framework employs a contract-based harness to bridge model-specific capabilities (Skills) and model-agnostic biological objectives (Playbooks), seamlessly translating free-form biological requests into reliable execution and rigorous downstream interpretation. Evaluated on a benchmark of 54 query cases derived from published S2F workflows, S2F-Agent systematically outperformed general-purpose LLMs, demonstrating superior reliability accuracy in routing, groundedness, and end-to-end task execution success. We further demonstrate the robustness and scalability of S2F-Agent across model adaptation, variant interpretation, genome-scale functional profiling and personal-genome analysis. First, the agent autonomously adapts a genomic foundation model to quantitative chromatin profiles, resolving sequence features associated with primed and active regulatory states. Second, integrating multi-perspective variant effect predictions prioritized 42 high-priority candidate variants among CAD-associated variants (>16,000), and identified tissue-resolved regulatory mechanisms including the hepatic SORT1 axis. Third, genome-scale profiling of multiple traits GWAS atlas variants (>250,000) revealed pervasive context dependence in molecular consequences and regulatory architecture, highlighting the analytical focus toward fine-grained, tissue-specific regulatory variants. Finally, evidence-gated analysis of personal genomes expanded functional hypothesis generation beyond clinically annotated variants to thousands of prioritized candidates per individual while imposing explicit evidence-dependent boundaries on clinical claims. Collectively, these results establish S2F-Agent as a general framework for converting heterogeneous sequence-to-function capabilities into verifiable, scalable, and evidence-aware genomic analyses. By bridging the chasm between LLMs, specialized S2F ecosystems and rigorous genomic science, this framework democratizes the S2F paradigm for unlocking the full potential of these advanced models in real-world discoveries.
bioinformatics2026-08-31v2HESTIA: Scalable Multimodal Integration of Histology and High-Resolution Spatial Transcriptomics for Robust Spatial Domain Identification
Zhong, Z.; Zhu, X.; Guo, J.; Liao, S.; Chen, A.Abstract
Spatial omics has revolutionized molecular biology by providing invaluable insights into how native tissue microenvironments regulate cellular functions and disease mechanisms. Accurately capturing this structural complexity and decoding the underlying biological processes requires effectively integrating data from multiple modalities. However, transitioning to subcellular resolutions introduces massive data scales and severe transcriptomic sparsity, which challenge current analytical frameworks. To address this, we present HESTIA (Histology-Enhanced Scalable cross-Resolution inTegration for spatial trAnscriptomics), a highly efficient multimodal algorithm designed for identifying spatial domains in large-scale, high-resolution spatial omics data. By circumventing memory-intensive computations, HESTIA efficiently processes massive datasets on which existing algorithms fail due to memory constraints. HESTIA outperforms current multimodal methods in clustering accuracy and spatial continuity, accurately delineating fine structural boundaries. Furthermore, applying HESTIA to large-scale pathological samples successfully dissects clinically relevant intratumoral heterogeneity and maps distinct immune microenvironments in lung and colorectal cancers.
bioinformatics2026-08-31v2Prioritizing peptides for targeted mass spectrometry experiments using deep learning
Sonthalia, S.; Wen, B.; Dasgupta, P.; Hsu, C.; MacCoss, M. J.; Noble, W. S.Abstract
One critical step in any targeted mass spectrometry experiment is selecting, from each protein of interest, a small number of peptides that respond well in the mass spectrometer and can serve as reliable proxies for protein quantification. Existing methods select target peptides either by relying on prior empirical measurements, limiting their applicability to previously observed peptides, or using machine learning to predict peptide behavior from sequence alone. However, current machine learning tools suffer from various limitations, including using detectability as an indirect proxy for intensity, relying on small training sets, or ignoring the precursor charge state. In this study, we introduce Bromo, a transformer-based deep learning model that ranks peptide precursors from a given protein by their relative response, taking charge state into account. Trained on millions of annotated peptide pairs derived from large-scale, publicly available data-independent acquisition mass spectrometry data, Bromo consistently outperforms existing sequence-based methods across diverse, independent datasets. Furthermore, we show that fine-tuning Bromo on experiment-specific data can account for differences in sample preparation, sample matrix, and instrument platform, all of which influence which peptides serve as optimal targets. This adaptability makes Bromo a practical tool for selecting target peptides for selected reaction monitoring and parallel reaction monitoring assay development across a wide range of experimental conditions.
bioinformatics2026-08-31v2seqproc: An efficient, flexible, and concise tool for sequence geometry description and transformation
Cape, N.; Fisher, E.; Liu, D.; Patro, R.Abstract
Complex sequencing protocols encode technical information in read structure and require accurate, e[ff]icient preprocessing. We introduce seqproc, which compiles concise sequence-geometry descriptions into execution graphs. Across four single-cell RNA-sequencing protocols, seqproc has the lowest mean runtime at every tested thread count and uses substantially less memory than the next-fastest tool. It has the highest F1 agreement with conservative structural references on all three discriminative chemistries and ties both alternatives on the 10x length-filter control. By separating protocol description from execution, seqproc makes complex read transformations compact, reusable, and efficient.
bioinformatics2026-08-31v2scPyviewer: a Python-native interactive viewer from AnnData single-cell data
Xuan, H.; Huang, Y.; Bian, J.; Liu, X.Abstract
Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.
bioinformatics2026-08-31v1LRSPAT: A low-rank framework for spatial omics statistics
Frost, H. R.Abstract
We describe LRSPAT (low-rank spatial toolkit), a fast and memory-efficient framework for approximating measures of spatial association for high-dimensional data. While LRSPAT can be applied to any multivariate spatial dataset, development was motivated by the computational challenge of identifying spatially variable genes in high-resolution spatial transcriptomics (ST) data generated by technologies such as 10x Visium HD, Xenium and Atera. LRSPAT leverages a truncated SVD of the expression data and a thresholded spatial weights matrix to perform reduced-rank reconstruction of spatial statistics in the quadratic form family, including global and local versions of Moran's I, Geary's C, and Getis-Ord G. A regularization approach is leveraged to account for the inflated null distribution of spatial statistics computed on latent variables. By performing key operations on the low-dimensional embeddings, LRSPAT is orders of magnitude faster than standard implementations with significantly lower memory requirements. Because the low-rank approach denoises and desparsifies ST data, LRSPAT is also more accurate than standard techniques at identifying genes with true spatial expression patterns. The dramatic improvements in execution time and memory consumption enable the genome-wide analysis of spatially variable genes (SVGs) and exploration of the full range of hyperparameters including spatial scale, distance metric, and embedding rank. This preprint outlines the background and mathematical details of the approach with limited preliminary results and a short conclusion.
bioinformatics2026-08-31v1Lineage-specific X chromosome inactivation escape and skew underlie sex-biased immune gene dosage and deleterious variant exposure
Kavanagh, D.; Steel, A.; King, H. E.; Vieira, H. G. S.; Kumar, K. R.; Masle-Farquhar, E.; King, C.; Skvortsova, K.; Weatheritt, R. J.Abstract
The X chromosome carries an unusually high density of immune genes and is a major contributor to sex differences in immune function and autoimmune diseases. In females, X-chromosome inactivation (XCI) has two major functional consequences: it shapes X-linked gene dosage through XCI escape and determines the cellular exposure of heterozygous X-linked variants through XCI skew. Yet because XCI creates a mosaic of cells expressing different parental X chromosomes, these properties have remained largely inaccessible in individual women, becoming measurable only where XCI is non-random or after aggregation across large cohorts. Consequently, how X-linked variation contributes to sex-biased immunity and differs between individual women has remained unresolved. Here we present scDaisyChain, a graph-based framework that reconstructs chromosome-scale X haplotypes directly from heterozygous SNPs and single-cell long-read transcriptomes. scDaisyChain achieves near-ground-truth accuracy in highly polymorphic mouse hybrids and shows strong concordance with orthogonal long-read whole-genome phasing in human samples. Applied to peripheral blood immune cells from healthy women, it reveals a lineage-specific escape program in which lymphoid cells escape XCI more broadly than monocytes, with corresponding gains in the inactive X chromatin accessibility and female-biased expression. Lineage-specific skew further alters the proportion of cells expressing each heterozygous X-linked variant, a property we term variant exposure. Predicted deleterious variants are preferentially found in low-exposure states, exemplified by a splice-altering TLR8 variant expressed in few cytotoxic T cells. In rheumatoid arthritis (RA), the monocyte compartment - which has the lowest escape in health - shows reproducible inactive X dysregulation converging on a trained-immunity programme linked to disease flare and synovial macrophage activation, with elevated escape of IL13RA1 and HDAC8. These findings establish lineage-specific escape, skew and variant exposure as quantifiable, patient-resolved determinants of sex-biased immune gene dosage and X-linked variant penetrance in health and autoimmune disease, resolving a dimension of female biology that has been previously inaccessible in individual donors.
bioinformatics2026-08-31v1Survey of the human proteostasis network: the ubiquitin-proteasome system
Elsasser, S.; Powers, E.; Stoeger, T.; Sui, X.; Kurtzbard, R. D.; Martinez-Botia, P.; Wangaline, M. A.; Gama, A. R.; Huttlin, E. L.; Elia, L. P.; Kelly, J. W.; Gestwicki, J. E.; Frydman, J. E.; Finkbeiner, S.; Clerico, E. M.; Morimoto, R.; Prado, M. A.; Vertegaal, A. C. O.; Hofmann, K.; Finley, D.Abstract
Modification by ubiquitination governs the half-lives of thousands of proteins that are fated for elimination by either the proteasome or autophagy pathways, depending on the intricate architectures of ubiquitin modification. This system mediates quality control for individual proteins, protein complexes, and organelles, as well as myriad purely regulatory functions. Here we provide a comprehensive survey of the ubiquitin-proteasome system (UPS), the scope of which is at present poorly defined. The UPS, with the inclusion of pathways involving ubiquitin-like modifiers, comprises in our estimate over 1430 distinct proteins in humans, a vast set of activities whose collective impact on the biology of the cell is pervasive. The UPS is an integral component of the proteostasis network (PN), the remainder of which we have also surveyed in recent studies. With the addition of molecular chaperones, proteins from autophagy-lysosome pathway, and related activities, the PN includes in total over 3150 components by our estimates. Comprehensive and systematic definition of these pathways should support a range of ongoing investigations in the areas of genomics, proteomics, biochemistry, cell biology, and disease research.
bioinformatics2026-08-30v3Trends in Machine Learning and Feature Selection Stability for Human Gut Microbiome (Shotgun Metagenomics) and Metabolomics Matched Datasets
Palmer, S. N.; Mishra, A. A.; Zarek, C. M.; Gan, S.; Wang, R.; Kim, J.; Liu, D.; Koh, A.; Zhan, X.Abstract
Microbiome research is often limited by methodological inconsistencies that reduce reproducibility and functional insight. Traditional taxonomy-based profiling is limited by sparse data, variable resolution, and reliance on gDNA sequencing which provides only indirect links to microbial function. Multi-omics integration offers a framework for linking community composition to functional outputs, but progress has been hindered by the lack of standardized frameworks and inconsistent use of machine learning. In particular, feature selection stability, which is central to biomarker discovery and experimental validation, remains underexplored. Here, we systematically benchmarked three widely used algorithms (Elastic Net, Random Forest, XGBoost) across seven multi-omics integration strategies and single-omics models. Additionally, we evaluated the impact of transforming metabolomics and taxonomic abundance data. Using human gut microbiome datasets that integrate metagenomic taxonomic profiles with metabolomics, we evaluated models for 9 binary and 8 continuous outcomes across 20 train/test splits per dataset. We further assessed the effect of feature reduction on both predictive accuracy and feature selection stability. Nonlinear learners were most consistently competitive: continuous outcomes favored metabolomics-dominant models, whereas binary outcomes favored stacked multi-omics models. Random Forest and XGBoost also yielded greater feature selection stability, particularly for full dimensional metabolomics data. Together, these findings demonstrate how integration strategy, algorithm choice, and data preprocessing jointly shape predictive performance and feature selection reproducibility in multi-omics microbiome modeling.
bioinformatics2026-08-29v4CIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations
Reisle, C.; Grisdale, C. J.; Krysiak, K.; Danos, A. M.; Khanfar, M.; Pleasance, E.; Saliba, J.; Hanos, M.; Patel, N. V.; Jain, A.; Seifi, M.; McMichael, J. F.; Venigalla, A. C.; Griffith, M.; Griffith, O. L.; Jones, S. J. M.Abstract
Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. Large language models (LLMs) offer a potential mechanism for accelerating biomedical claim verification, but their rapid turnover, variable availability, and known risks of unsupported reasoning require standardized and reproducible evaluation before integration into curation workflows. To address this, we developed CIViC-Fact, an expert-curated, full-text benchmark and evaluation framework. Domain experts linked structured cancer-variant claims to sentence-level evidence from source publications, including evidence from full-text articles, tables, and non-abstract sections that are commonly omitted from existing biomedical question-answering and scientific fact-checking datasets. Claim-verification reference labels were derived from CIViC records, revision histories and controlled data augmentation. A major finding of CIViC-Fact is that abstracts are insufficient for realistic biomedical claim verification. In the evaluated development subset of text-verifiable entries with full-text access, fewer than 30% could be fully validated from the abstract alone, highlighting the importance of full-text evaluation for biomedical curation. Upon the application of our fact-checking pipeline to newly submitted CIViC entries, after excluding entries requiring supplementary material or images for validation, automated retrieval successfully identified appropriate evidence for most cases (93%), supporting low-incremental-effort evaluation of future systems. Fine-tuning improved agreement with CIViC-Fact reference labels on the static benchmark, but larger general-purpose models performed better on a heterogeneous post-cutoff cohort. These findings support CIViC-Fact primarily as a reproducible framework for comparing evolving retrieval and verification systems rather than as validation of a single deployment-ready model. These findings suggest that, in a rapidly changing model landscape, the durable contribution is not a single optimized model but a reproducible benchmark framework that enables continual testing, model substitution, and lightweight updating through small high-quality few-shot exemplar sets.
bioinformatics2026-08-29v3FuncSeek: Multi-PLM contrastive learning for protein functional similarity search
Cloete, L. J.; Patterton, H. G.Abstract
Below 30% pairwise sequence identity, alignment-based methods struggle to reliably distinguish true homologs from chance, and enzyme function prediction degrades accordingly: on proteins in this regime, even advanced methods (CLEAN) achieves only 55.1% accuracy at full EC specificity on the CARE benchmark. To this end Protein Language Models (PLMs) have gained favor as alternatives. However, PLMs often encode only a subset of the biology, whereas the understanding of enzyme function requires among other things a combination of sequence, structure and functional-context simultaneously. In this work, we describe FuncSeek, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities). Using SwissProt data, these 2816-D embeddings are labelled with Enzyme Commission numbers (EC) and are trained through a supervised contrastive head into a 256-D space. FuncSeek attains 64.6% nearest-neighbour EC4 accuracy on the CARE out-of-distribution benchmark set (ood30; proteins below 30% identity to training set), outperforming CLEAN (55.1%) and Diamond BLASTp (51.4%), and obtains 93.7% nearest-neighbour EC4 accuracy on the promiscuous, multi-functional enzymes benchmark (CLEAN, 69.4%). We also show that the learned representations transfer without retraining to the TrEMBL database, achieving 97.3% nearest-neighbour EC4 accuracy on an 8,031 BRENDA-validated enzyme set, never seen during training. Because only projected embeddings are stored in the target index, and function is inferred from an annotated reference set, we propose this paradigm for rapidly searching extremely large metagenomic databases, bypassing costly sequence alignment and annotation pipelines.
bioinformatics2026-08-29v2HANSEN: An Integrated Structural and Functional Proteome Resource for Structure-Guided Drug Discovery in Mycobacterium leprae
Vedithi, S. C.; Rees, R.; Malhotra, S.; Munir, A.; Matusevicius, M.; Alsulami, A. F.; Beaudoin, C. A.; Sunkara, K. S.; Das, M.; Blundell, T. L.; Floto, R. A.Abstract
Leprosy remains a leading infectious cause of preventable disability, yet its causative agent, Mycobacterium leprae (M. leprae), is structurally under-characterised. Only ten Protein Data Bank (PDB) entries represent seven of its 1,603 protein-coding genes. We present HANSEN, a proteome-wide structural and functional resource for M. leprae. Monomeric and oligomeric models were generated with AlphaFold 3, Boltz-1, Boltz-2 and Chai-1, and annotated with per-residue confidence, predicted aligned error and, for assemblies, interface confidence. Ligand-binding pockets were predicted with AF2BIND, P2Rank and fpocket, template-derived ligands were modelled within oligomeric complexes, residue-level B-cell epitope propensity was estimated with DiscoTope-3.0, and gene essentiality was transferred from Mycobacterium tuberculosis transposon-sequencing labels. These features are integrated in a relational database with interactive visualisation and combined into a calibrated Target Priority Score that ranks all 1,603 proteins into four tiers and recovers established antimycobacterial targets. HANSEN (https://hansen-leprosy.medschl.cam.ac.uk/home) provides a practical basis for target prioritisation and structure-guided drug discovery in leprosy.
bioinformatics2026-08-29v2Genome-Wide in silico analysis reveals activation of a silent resistome driving imipenem resistance in Pseudomonas aeruginosa
Anwar, S.; Aromal, A. R.; Anurag Anand, A.; Samanta, S. K.Abstract
Resistance to imipenem in Pseudomonas aeruginosa relies on multiple factors that remain poorly understood. In our work, we performed a systemic analysis of genome-wide changes involved in resistance in a set of 95 clinically unrelated strains, including 41 resistant (MIC [≥] 64 mg/L) and 54 susceptible (MIC [≤] 2 mg/L) isolates. Our approach is based on the pan-genomics analysis, combining the use of core-genome phylogenetic analysis, MLST (Multilocus Sequence Typing), GWAS (Genome-Wide Association Studies) and variant level profiling of the blaOXA genes. Higher-order structure within the set was studied using methods of the co-occurrence networks and WGCNA (weighted gene co-expression network analysis) specifically adjusted to handle presence/absence data. Despite having a broader and more diverse resistome, no clonal grouping of the resistant isolates was observed indicating independent evolutionary origins. The LASSO model using a lineage-aware approach showed robust predictive capability (AUC = 0.836) that validates the polygenic characteristic of resistance. Twelve accessory genes were found to be significant determinants of resistance; however, only four genes (group_10880, group_10887, group_4947, and phzB) were identified using both GWAS and gene network analysis, showing involvement in protein folding, metal stress response, genome plasticity, and metabolic adaptation. Interestingly, some carbapenemase-active variants of blaOXA were also found in imipenem-susceptible strains, showing that gene presence alone does not ensure resistance. We therefore propose the Silent Resistome Activation Model, where resistance genes become functional only with support from identified accessory genes and coordinated interactions at both the genomic and network levels.
bioinformatics2026-08-29v2Quantifying the Rearrangement Complexity of Pangenomes
Bohnenkaemper, L.; Stoye, J.Abstract
The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.
bioinformatics2026-08-29v1Bacterial metagenome in plaque, saliva, and tumor samples from individuals with and without OSCC by next-generation sequencing
ERIRA, A.; ROBAYO, D. A. G.; GAMBOA, F.; CHALA, A.; MORENO, A.; ARREGUI, A. C.; MUNOZ, E.; NOGUERA, J.; TOBAR-TOSSE, F.Abstract
Background: Oral dysbiosis has been associated with oral squamous cell carcinoma (OSCC); however, most microbiome studies rely on 16S ribosomal RNA (rRNA) gene sequencing, limiting species-level taxonomic resolution. Methods: Dental plaque, saliva, and tumor tissue samples from 10 patients with OSCC and dental plaque and saliva samples from 10 healthy controls were analyzed in this exploratory cross-sectional study. DNA was extracted and subjected to shotgun metagenomic sequencing using the Illumina MiSeq platform. Sequence reads were quality filtered with fastp, taxonomically classified using Kraken2 v2.1.3, and species-level abundances were re-estimated with Bracken v2.9 following the removal of human reads and low abundance taxa. Relative abundances were compared using the Mann Whitney U test with the Benjamini Hochberg false discovery rate correction, while the Bray Curtis principal coordinate analysis was used as an exploratory approach to visualize microbial community patterns. Results: Shotgun metagenomic sequencing revealed distinct bacterial community profiles across the oral microenvironment. Dental plaque exhibited the highest taxonomic diversity and relative abundance. The control plaque was enriched in Streptococcus koreensis, Capnocytophaga sp. oral taxon 878, Treponema sp. Marseille Q4132, and Leptotrichia sp. oral taxon 498, whereas the plaque from patients with OSCC showed a higher relative abundance of Pyramidobacter piscolens, Parvimonas parva, and Gemella sanguinis. Salivary samples displayed lower diversity and a more homogeneous composition, predominantly comprising Capnocytophaga endodontalis, Prevotella jejuni, Aggregatibacter aphrophilus, and Gemella sanguinis. The tumor tissue showed relatively higher abundance of Sellimonas catena, Escherichia coli, Solobacterium moorei, and Lacrimispora sp. HJ 01. Conclusions: This exploratory study provides species-level characterization of the oral microbiome across multiple oral microenvironments in OSCC and generates hypotheses for future integrative metagenomic and functional studies investigating the potential contribution of oral bacterial communities to OSCC pathogenesis.
bioinformatics2026-08-29v1When AI encounters natural history: Morphological OTUs reshape our understanding of Earth's life
Zhan, Z.; Ye, M.; Orr, M. C.; Chen, W.; Liu, X.; Yue, L.; Sun, X.; Zhang, F.Abstract
Biodiversity can be quantified only after organisms are assigned to reproducible units, yet most individuals encountered in nature lack reliable species-level identifications. Molecular operational taxonomic units can organize unnamed diversity, whereas image-based approaches generally depend on predefined species levels. Here, we show that operational biodiversity units can be derived directly from phenotypes. We developed morphOTU, a framework combining self-supervised representation learning, operational metric supervision, and adaptive hierarchical clustering to organize specimen images in continuous phenotypic space. Across five benchmark datasets comprising flowers, wood anatomy, and beetle habitus, morphOTUs recovered coherent species-level structure and produced -diversity estimates close to those obtained from expert identifications. This structure remained informative when species were excluded from representation learning, when labeled data were sparse, and per-species sampling was limited. In a heterogeneous field-survey dataset of 4,717 insects representing 269 species across 12 orders, fine-tuning on only 28 common species produced diversity estimates close to expert labels (Shannon index, 3.75 versus 3.53). Visual explanations localized variation to biologically meaningful structures, including body outliers, surface sculpture, floral symmetry, and wood vessels. Morphological units therefore provide an operational layer for organizing and quantifying biodiversity before, alongside, or in the absence of formal species names.
bioinformatics2026-08-28v3Can SMILES be fragmented into a concatenable ordered sequence of retrosynthetically interesting string block ?
Reboul, E.; Prabakaran, H.; Waldispuhl, J.; Taly, A.Abstract
Molecules generated by deep learning models are often difficult to synthesize. Their synthetic accessibility can be improved with automated retrosynthetic analysis, which allows for identifying synthons. However, synthons in a SMILES can be scattered throughout the string depending on the path taken through the molecular graph used to generate the SMILES. We tested whether the ensemble of possible SMILES for a molecule can be used to generate a concatenable ordered sequence of string fragments (blocks) from SMILES that match potential synthons obtained through automated retrosynthetic analysis. We found that exhaustively sampling the SMILES space of a molecule improves the coverage of retrosynthetic breaks. We achieved full coverage of retrosynthetic bonds in string form for 85\% of the 1.9 million molecules in the MOSES dataset. Doing so allowed us to test our block SMILES in an unconditional de novo drug design test case with MolGPT and Monte Carlo Tree Search (MCTS). We found that using blocks as an LLM's token did degrade MolGPT performance due to the curse of dimensionality. However, using the SMILES selected by our blocking algorithm with the default SMILES tokenizer improved the reproduction of physico-chemical properties of samples and also improved uniqueness, novelty, and validity. The MCTS outperforms our MolGPT models in terms of validity and novelty. However, samples generated by the MCTS had physico-chemical properties that were further away from the MOSES baseline than the samples produced by molGPT, with an improved distribution of quantitative estimation of drug-likeness (QED).
bioinformatics2026-08-28v2Evaluating Aggregated Gene Level eQTL Scores
Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.Abstract
Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.
bioinformatics2026-08-28v2Gene expression inference from cell-free DNA using uncertainty-aware deep learning
Patton, R. D.; McDeed, A. P.; Netzley, A.; Pawar, A.; Persse, T. W.; Nair, A.; Galipeau, P. C.; Coleman, I. M.; Itagi, P.; Chandra, P.; Sayar, E.; Adil, M.; Vashisth, M.; Hiatt, J. B.; Dumpit, R.; Kollath, L.; Demirci, R. A.; Ghodsi, A.; Lam, H.-M.; Morrissey, C.; Chen, D. L.; Schweizer, M. T.; Iravani, A.; Hsieh, A. C.; MacPherson, D.; Haffner, M. C.; Nelson, P. S.; Ha, G.Abstract
Tumor gene expression profiling provides crucial diagnostic information for guiding therapy, but standard tissue biopsies are invasive, spatially biased, and may inadequately sample metastatic disease. Cell-free DNA (cfDNA) provides a minimally invasive alternative for tumor genotyping, yet reconstructing robust, transcriptome-wide expression from standard-depth cfDNA whole-genome sequencing (WGS) remains a major challenge. We developed a deep learning framework comprising Triton, for comprehensive cfDNA feature extraction, and Proteus, a probabilistic model that infers single-gene expression from standard-depth cfDNA WGS. Proteus outperformed prior cfDNA approaches in reconstructing molecular phenotypes from matched tumor transcriptomes across multiple cancer types, including prostate, lung, and bladder cancer cohorts, with uncertainty-guided withholding improving model reliability. Proteus further enabled assessment of therapeutic target activity, prognostic transcriptional programs, and candidate treatment-emergent resistance states, establishing a generalizable framework for minimally invasive functional genomics in precision oncology.
bioinformatics2026-08-28v2Deciphering the Gut-Brain Dialogue: A Survey-Based and In-Silico Comparative Analysis of Gut Microbial Dysbiosis in Common Neurological Disorders
Goyal, S.; Kalra, A.Abstract
The gut microbiome maintains a complex, bidirectional communication network with the central nervous system, commonly referred to as the gut-brain axis and its disruption has been implicated in several neurological disorders. This study combines a survey-based assessment of public awareness with an in-silico comparative analysis of gut microbial dysbiosis across four prevalent neurological disorders as observed in the current study: depression, anxiety, schizophrenia and autism spectrum disorder (ASD). A structured, anonymous online survey (n = 230) captured perceptions of the gut-brain connection along with dietary, lifestyle and gastrointestinal correlates of stress in a predominantly young, health-sciences-affiliated Indian cohort. In parallel, disorder-specific lists of elevated and reduced faecal microbial taxa were retrieved from the Disbiome database, compared using a multiple list comparator and taxonomically classified using the NCBI Taxonomy tool to construct phylogenetic trees in iTOL. Approximately three-quarters of respondents were aware of a potential gut-mental health link, yet about half reported no specific dietary practice and roughly 60% experienced stress-related digestive symptoms while rarely seeking medical consultation for them. Comparative analysis showed that depression, anxiety and schizophrenia shared a substantially overlapping dysbiosis signature, with common elevation of Actinomyces, Bacteroidaceae, Blautia, Eggerthella, Oscillibacter, Parasutterella and Veillonella and common reduction of Coprococcus, Lachnospiraceae, Ruminococcaceae, Clostridium, Faecalibacterium and Sutterella. In contrast, ASD displayed a distinct microbial signature with limited overlap with the other three disorders. Phylogenetic clustering confirmed that the shared taxa belonged predominantly to the phyla Bacillota (formerly Firmicutes), Bacteroidota (formerly Bacteroidetes), Actinomycetota (formerly Actinobacteria) and Pseudomonadota (formerly Proteobacteria). Notably, this phylum-level pattern parallels recent comparative analyses of microbial dysbiosis in neurodegenerative diseases, suggesting that broad phylogenetic shifts may be a relatively general correlate of chronic neurological disease, while disorder specificity emerges at the level of individual taxa. These findings support a shared microbial pathway linking depression, anxiety and schizophrenia that is distinct from the dysbiosis pattern observed in ASD and they underscore the value of microbiome-informed, disorder-specific therapeutic strategies.
bioinformatics2026-08-28v1CoTRA: an integrated R/Shiny framework for transparent bulk and single-cell RNA-seq analysis
Seemab, U.; Vainionpaa, K.; Tanoli, Z.; Leinonen, H. O.Abstract
Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.
bioinformatics2026-08-28v1OmicsFM brings proteomics into the foundation model era
Heyndrickx, S.; Gabriels, R.; Ramadasan, H.; Martens, L.; Claeys, T.Abstract
While foundation models have been shown to learn biological representations from large transcriptomic atlases, it remained unknown whether proteomics data allow the same. We here therefore introduce OmicsFM, a modality-agnostic transformer pretrained through masked abundance reconstruction on an unprecedented proteomics data corpus of 48,837 quality-filtered proteomics profiles from 1,397 reprocessed PRIDE projects. Interestingly, despite training on 14- to 93-fold fewer profiles than matched bulk- and single-cell transcriptomic models, respectively, our proteomics model rivals both. On held-out projects, OmicsFM attention networks recovered more molecular relationships than co-expression methods and existing single-cell foundation models across nine reference databases that reveal pathway-level organization. Sample-level embeddings preserved biological structure across independent studies, and its representations transferred successfully to cell-type classification, gene-essentiality prediction, and perturbation-response prediction, while consistently outperforming task-specific models. Moreover, our results show that proteomics and transcriptomics representations capture complementary biology. OmicsFM thus firmly establishes the possibility of training highly performant proteomics-based foundation models, and their importance in modelling and uncovering fundamental biology.
bioinformatics2026-08-28v1A pretrained unified model enables cellular functional profile prediction and multi-objective virtual drug screening
Chen, R.; Huang, L.; Qiao, Y.; Mandal, S.; Mo, L.; Li, L.; Leshchiner, D.; Zhang, X.; Pu, J.; Xie, Y.; Girgis, R.; Ellsworth, E.; Huang, L.; Chen, X.; Li, X.; Zhou, J.; Chen, B.Abstract
Cells are characterized by molecular states, coordinated molecular interactions, regulatory programs, and responses to perturbations. Systematic mapping of these cellular functional profiles across biological contexts remains experimentally costly and fragmented. Here we present InsilicoCell, a pretrained multi-modal, multi-task model that unifies prediction of cellular functional profiles spanning molecular states, molecular interactions, and perturbation-induced responses. Built on a supervised transformer architecture and pretrained on more than 88 million measurements across seven tasks, including drug sensitivity, drug-induced gene expression, and drug-protein binding, InsilicoCell learns a shared representation that links molecular profiles to cellular phenotypes, improves performance over task-specific models, and generalizes to unseen entities, contexts, and conditions. InsilicoCell extends beyond cell line systems to patient, spatial and single-cell settings, and enables multi-objective virtual drug screening. It identifies novel candidate compounds with experimental validation, including c-Myc activity inhibitors, antifibrotic agents and stemness-inducing compounds. Together, InsilicoCell provides a scalable framework for predictive cellular biology and therapeutic discovery.
bioinformatics2026-08-28v1Label Noise Limits TCR-pMHC Specificity Prediction: Improved Performance Through AlphaFold3-Based Structural Modeling and Data Denoising
Ballesteros-Cuartero, P.; Lund, J.; Nielsen, M.Abstract
T cell receptor (TCR) binding to peptides presented by major histocompatibility complex (MHC) molecules is a key step in T cell activation, and forms the basis of adaptive immunity. Predicting this specificity is therefore essential to developing effective TCR-based immunotherapies and vaccines. Despite its clinical relevance, predicting TCR-pMHC specificity for previously unseen peptides remains an open problem, with structural modeling so far the only strategy showing any predictive power in this setting. In this study, we find that this limited performance is substantially driven by label noise in the data used to train and evaluate these methods, an effect that has so far been largely underexplored. Using an AlphaFold3-based pipeline adapted for TCR-pMHC structural modeling, we achieve state-of-the-art specificity prediction, outperforming AlphaFold2.3-based and sequence based methods, and performing at par with the leading Immrep2025 competition submission. Combining this pipeline with a cluster-based denoising algorithm, we show that removing mislabeled points from a large specificity dataset increased binder ranking accuracy by more than 70% relative to the full dataset. Together, these results highlight label noise as a major factor limiting the performance that any method in this field can achieve, and show that combining structural modeling with label denoising substantially improves TCR-pMHC specificity prediction, making such approaches an attractive complement to current sequence-based approaches for refining TCR target selection.
bioinformatics2026-08-28v1FOCUS-3D: Robust, generalizable volumetric cell segmentation for three-dimensional fluorescence microscopy
Zhang, Q.; Mu, Z.; Liu, B.; Chi, Y.; Li, D.; Wang, W.; Ni, J.-Q.; Wan, Y.; Yu, L.; Navajas Acedo, J.; Yu, G.Abstract
Understanding how cells establish spatial organization within tissues is a fundamental question in life sciences. While modern three-dimensional fluorescence microscopy captures large-volume tissue architecture, extracting quantitative cellular insights from complex volumetric datasets remains a major barrier. Here, we introduce FOCUS-3D, a robust, broadly generalizable volumetric cell segmentation framework built on a large, diverse manually annotated cell resource and advanced AI designs. Integrating volumetric representation learning, multi-scale feature extraction, and query-based mask prediction, FOCUS-3D achieves state-of-the-art performance across diverse species, tissues, fluorescent reporters and imaging modalities. During zebrafish (Danio rerio) development, FOCUS-3D uncovers three successive phases of notochord morphogenesis. We disentangle early motility-driven rearrangements from later cell shape remodeling and tissue repacking, and further link these morphological states to spatial and developmental transcriptional programs across independent datasets.
bioinformatics2026-08-28v1Erosion of regenerative regulation: age-associated shifts in the skeletal muscle fiber epigenome and transcriptome
Moo, K. G.; Orchard, P.; Varshney, A.; D'Oliveira Albanus, R.; Manickam, N.; Kinnunen, L.; Lakka, T. A.; Saramies, J.; Laakso, M.; Tuomilehto, J.; Mohlke, K. L.; Boehnke, M.; Scott, L. J.; Koistinen, H. A.; Collins, F. S.; Parker, S. C.Abstract
Skeletal muscle aging is characterized by the deterioration of muscle function, which can lead to negative quality-of-life outcomes including frailty and sarcopenia. While understanding the mechanisms of this process is increasingly important as the global population ages, previous molecular studies of skeletal muscle aging have been limited by statistical power and cell type resolution. In this study, we analyzed single-nucleus gene expression and chromatin accessibility data from 287 human skeletal muscle samples from individuals aged 20-79 years to explore sex- and cell type- specific aging effects. Across 467,126 nuclei from 13 cell types, we identify 384 age-associated genes and 4,061 age-associated chromatin regions. These age-associated molecular features are enriched for functional pathways, including metabolic processes, cell-to-cell communication, and senescence Kyoto Encyclopedia of Genes and Genomes KEGG terms. Age-associated closing chromatin was more common across fiber types and sexes than opening chromatin, and was enriched in active enhancer regions while depleted for active transcription start sites. We observe enrichment for specific transcription factor motifs in closing chromatin, including those of glucocorticoid and androgen receptors, both of which play a key role in the maintenance of healthy skeletal muscle. Together, these findings identify an age-associated regulatory shift, largely invisible in matched transcriptomic data, characterized by closing chromatin which reduces accessibility to hormone receptor binding sites and enhancer regions in the muscle fiber epigenome.
bioinformatics2026-08-27v2HIDE-Deconv: A hierarchical deconvolution framework for multiscale characterization of cellular remodeling
Voelkl, D.; Bolz, S.; Rayford, A.; Sterr, T.; Mensching-Buhr, M.; Seifert, N.; Arp, J.; Tausche, J.; Engel, L.; Schuster, C.; Stevenson, T.; Zacharias, H. U.; Altenbuchinger, M.; Goertler, F.Abstract
Most deconvolution methods estimate cellular composition at a single level of cellular resolution despite biological processes often manifesting within fine-grained cellular subpopulations. We present HIDE-Deconv, a hierarchical deconvolution framework that jointly optimizes cellular compositions across multiple levels of a cell-type hierarchy while maintaining consistency between resolutions. In benchmark experiments, HIDE-Deconv achieved the highest overall predictive performance among evaluated methods. Analyses of lung adenocarcinoma, sepsis, COVID-19 and systemic lupus erythematosus revealed biologically relevant cellular remodeling that remained concealed at broader levels of cellular resolution. HIDE-Deconv is available as an open-source framework at https://github.com/dvoelkl/HIDE-deconv.
bioinformatics2026-08-27v2ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation
Moran, J.; Freda, P. J.; Ghosh, A.; Walker, C. T.; Hernandez, M. E.; Moore, J. H.Abstract
Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.
bioinformatics2026-08-27v2NeuroMesh: A Bottleneck Topology Controller for Missing-Modality Brain Tumor Segmentation - A Mechanistic Pilot Study on BraTS
Kamalakannan, N. K.; Kamalakannan, J.Abstract
Deep segmentation networks can degrade sharply when an expected MRI sequence is unavailable at inference. We present NeuroMesh, a bottleneck controller that combines a gated recurrent unit (GRU) with a graphconvolutional edge-activation mask, designed to adapt a U-Net-style segmentation backbone to missing input. We evaluate NeuroMesh in a pilot study using a 30-patient subset of the BraTS 2020 benchmark (22 training, 4 validation, and 4 held-out test patients) under a prespecified frozentest protocol. On the frozen test set, NeuroMesh has higher tumor-core and enhancing-tumor Dice than a plain U-Net in most evaluated missing-modality conditions, but wholetumor Dice falls from 0.596 to 0.108 when FLAIR is missing, compared with 0.604 to 0.545 for the plain U-Net. Direct analysis of the predicted edge-activation mask shows negligible change across modality-availability conditions. A parameter-light static-gating control reproduces the FLAIR failure mode without recurrence, a failure-signal input, or graph-structured machinery. These results do not support the intended interpretation that the trained controller performs input-conditional topology rewiring at the scale of this pilot. Instead, they expose a discrepancy between architectural intent and realized behavior and identify a specific missing-modality failure mode that warrants further investigation. Given the small validation and test sets, the findings are descriptive and do not establish clinical or population-level generalization.
bioinformatics2026-08-27v1AntiSite: Modality Dropout Enables Antibody Paratope Prediction With or Without Structure From a Single Model
Papadopoulos, A. M.; Alvarez, F.; Daras, P.Abstract
Summary: Reliable paratope identification is central to understanding antibody antigen recognition and advancing therapeutic antibody discovery. AntiSite is a unified antibody paratope prediction framework that combines protein language-model sequence embeddings with structure-derived molecular-surface features and, through modality dropout, trains a single checkpoint to predict both with and without a structure. This lets one model support sequence-only inference when no structure is available and structure-aware inference when an antibody structure is provided. Availability and implementation: Source code, trained models and evaluation scripts are freely available at https://github.com/aggelos-michael-papadopoulos/AntiSite. Processed benchmark structures and corrected split metadata are archived on Zenodo at https://doi.org/10.5281/zenodo.21705412.
bioinformatics2026-08-27v1DeMoP: A Language-Model-Guided Mixture-of-Experts Framework for Cancer Prognosis
Tang, C.; Yu, L.; Li, Q.; Xu, L.Abstract
Integrating heterogeneous clinical and molecular data for cancer prognosis remains challenging because their dimensionality, semantics and distributions differ across patients and cohorts. Here we present DeMoP, a language-model-guided mixture-of-experts framework that serializes structured patient profiles as natural-language sequences and learns adaptive prognostic representations from clinical variables, copy-number alterations, and gene descriptions. DeMoP combines a fine-tuned DeBERTa-v3-large encoder, attention-based token pooling, and a residual mixture-of-experts prediction head. In held-out tests from two independent pan-cancer cohorts, GENIE (63,090 patients) and TCGA (4,123 patients), DeMoP outperformed the conventional machine-learning and deep-learning baselines evaluated, achieving AUROCs of 0.939 and 0.805 and class-1 F1 scores of 0.72 in both cohorts. A GENIE-trained model transferred directly to TCGA with an overall class-1 F1 score of 0.62. Gene-level ablations recovered established cancer-associated genes and highlighted less-studied candidates. DeMoP provides a unified approach to heterogeneous biomedical data integration, cross-cohort outcome prediction, and model interpretation.
bioinformatics2026-08-27v1Identification of novel HDAC11 inhibitors: In silico & in vitro studies
Paul, M.; Kumar, D. S.; Mishra, S.; Kalle, A. M.Abstract
Histone deacetylases (HDACs) are pivotal epigenetic regulators that modulate diverse cellular pathways by removing acetyl groups from lysine residues on both histone and non-histone proteins. Histone deacetylase 11 (HDAC11), the sole member of class IV HDACs, exhibits both deacetylation and fatty acid deacylation activities. Accumulating evidence implicates HDAC11 as a key epigenetic regulator of fundamental cellular processes, including metabolism, immune responses, and tissue development. Dysregulation of HDAC11 activity has been associated with inflammatory diseases, metabolic disorders, neurodegenerative conditions, and cancer, highlighting its potential as a therapeutic target. Although several HDAC11-specific inhibitors have been identified, none have progressed to clinical development. In this study, we aimed to discover HDAC11-selective inhibitors by integrating in silico and in vitro validation approaches. Homology modelling of the HDAC11 structure was conducted, followed by model validation, structure-based virtual screening, molecular dynamics (MD) simulations, and binding free energy calculations. We identified and validated three lead compounds and their intermediates using biochemical and cell-based assays. Fluorescence-based and HPLC-based enzymatic assays demonstrated potent inhibition of both the deacetylase and deacylase activities of HDAC11, with Inhibitor 6 and Inhibitor 3 exhibiting the strongest effects among the six compounds tested. Further, a decrease in lipid accumulation, reduced stability of the HDAC11 substrate SHMT2, as determined by immunoblot analysis and decreased cell viability, as assessed by MTT assay, confirmed HDAC11 inhibition in cellular models. The study shows that new HDAC11 inhibitors significantly reduce the viability of breast cancer cells and induce apoptosis; inhibitor 6, in particular, showed high potency, similar to the reference compound SIS-17. Flow cytometry showed that treated MDA-MB-231 cells exhibited cell-cycle arrest and increased apoptosis, a finding further confirmed by Annexin V/PI staining. Molecular analysis showed that BAX increased while BCL2 decreased, indicating that apoptotic pathways were activated in novel compound-treated MDA-MB-231 cells. The results suggest that inhibiting HDAC11 is an effective way to induce cancer cell death and provide a basis for further assessment of these compounds as potential treatments for breast cancer. Collectively, this study identifies novel zinc-chelating HDAC11 inhibitors containing a nitro-sp2 group, providing promising candidates for further therapeutic development.
bioinformatics2026-08-27v1PMPNN-DDG: an accurate machine learning-based {triangleup}{triangleup}G prediction pipeline trained on a novel interpretable feature set extracted from ProteinMPNN
Jani, R.; Ahmed, S.Abstract
An accurate and tractable approximation of the single-point mutation-induced change in protein thermodynamic stability, denoted by DDG, is critical for understanding the genotype-phenotype relationship. Several computational methods have been proposed for this problem; however, limited and error-prone training data and the difficult-to-predict magnitude of structural perturbations make this a challenging task. Consequently, the computational predictors proposed throughout the past decade incrementally improved prediction performance by proposing novel features, combining existing features, task-adapted neural network architectures, loss functions, data augmentation techniques, and pre-training procedures. In this work, we propose PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN. On the S669 independent test set, PMPNN-DDG achieves rF +R = 0.64 and RMSE = 1.45, outperforming all compared baseline methods across the reported evaluation measures. On the Ssym independent test set, it achieves rF +R = 0.81, rF -R = -0.99, and RMSE = 1.10, showing competitive performance relative to the compared baselines. PMPNN-DDG is publicly available at https://github.com/dRanger666/PMPNN-DDG.
bioinformatics2026-08-27v1Proteome modulation by opposite inotropic drugs in human engineered cardiac tissue revealed by topology-driven cross-modal integration
Staykova, D. K.; Snippert, D.; Wessels, H. J. C. T.; Passier, R.; Conte, F.Abstract
Engineered heart tissues (EHTs) represent an innovative platform enabling physiologically relevant in vitro evaluation of drug-induced cardiac responses. While functional characterization remains central to EHTs, molecular profiling is increasingly used to elucidate mechanisms underlying drug-induced phenotypes. Proteomics provides broad molecular characterization of drug responses at the protein level, yet the complexity, heterogeneity, and high dimensionality of proteomics datasets challenge conventional statistical approaches, which are not designed for cross-modal integration and streamlined multi-omics analysis. In this study, we developed an innovative framework based on topological data analysis (TDA) for the integration of large proteomics profiles and functional readouts to investigate system-level responses to drugs with opposing inotropic effects, epinephrine and doxorubicin. Samples were organized into a topological connectivity network according to multimodal similarity enabling simultaneous exploration of treatments, cardiac function and proteome alterations. Highly correlated features were then used for pathway enrichment analysis, which revealed strong similarities between the enrichment profiles associated with contractile force and epinephrine. These findings are consistent with the positive inotropic effect of epinephrine, whereas doxorubicin exhibited an opposing enrichment profile. Energy homeostasis, mitochondrial translation and proteostasis emerged as the major cellular processes displaying opposite associations with the two inotropic drugs, highlighting a link between cardiac contractility and perturbations in these processes. In conclusion, our TDA-based framework successfully integrated functional and proteomic data to uncover treatment-specific remodeling in EHTs, offering a modular and scalable approach that could be adapted to other in vitro organ models for systems-level mechanistic studies and next-generation drug development.
bioinformatics2026-08-27v1Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients
Ramesh, P.; Fyta, M.Abstract
Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.
bioinformatics2026-08-27v1DeepTMHMM2 enables accurate prediction of transmembrane protein topology and subcellular location
Teufel, F.; Hallgren, J.; Nielsen, H.; Krogh, A.; Tsirigos, K. D.; Winther, O.Abstract
Transmembrane -helical and {beta}-barrel proteins are a ubiquitous component of proteomes. Topology prediction infers how proteins are embedded in lipid bilayers, identifying membrane-spanning segments and their orientation. While recent methods achieve high performance for membrane-spanning segments, they cannot predict re-entrant regions and interfacial helices - membrane-associated segments that partially insert but do not cross the bilayer - nor identify which biological membrane a protein resides in. Here, we present DeepTMHMM2, the first predictor to include re-entrant regions and interfacial helices in its topologies and jointly predict localization across 17 biological membranes. Benchmark results show that DeepTMHMM2 successfully learns to predict the additional elements, while achieving strong performance on canonical -helical and {beta}-barrel topology prediction. Applying DeepTMHMM2 to Swiss-Prot reveals that non-crossing segments are a ubiquitous feature of the transmembrane proteome, with interfacial helices present in nearly a quarter of all -helical transmembrane proteins.
bioinformatics2026-08-27v1reactifpTM: an accessible reimplementation of actifpTM
Simpkin, A. J.; Johnson, E.; Rigden, D.Abstract
Motivation: The actual interface pTM score (actifpTM) is a modified version of the ipTM score that limits the calculation to only those residues at the interface. Whilst actifpTM provides an effective interface quality score, a limiting factor is that it makes use of the predicted aligned error (PAE) with probabilities, information that is generated during a ColabFold run, but not output by the package or other model prediction software. The consequent inability to generate actifpTM scores for the results of software such as AlphaFold 2 or AlphaFold 3 has limited its adoption. With reactifpTM we address this problem by providing a standalone tool that can be run on the standard outputs of most model prediction packages. Results: Using the same underlying principles as actifpTM, reactifpTM has been developed to use standard output files from model prediction software (a model and corresponding PAE) to perform an actifpTM-like calculation. ColabFold models were generated for a dataset of 1079 known interfaces in the PDB. A strong correlation was shown between actifpTM and reactifpTM for this dataset. Availability and implementation: reactifpTM is coded in Python. All scripts and associated documentation are available from https://github.com/hlasimpk/reactifptm or https://pypi.org/project/reactifptm.
bioinformatics2026-08-27v1SPC-Clean: A napari Plugin for Reducing Speckle and Isolated Pixel Noise in Fluorescence Microscopy Images
Alirezazadeh, P.; Kirsch, E. M.; Tian, Y.; Bewersdorf, J.; Rittscher, J.; Mergenthaler, P.Abstract
Speckle artifacts and isolated foreground pixels are common in fluorescence microscopy and can interfere with segmentation and subsequent quantitative image analysis. Conventional denoising methods often modify image intensities through filtering or smoothing, potentially altering biologically relevant fluorescence signals. We introduce Sparse Pixel Cluster Cleaning (SPC-Clean), a topology-aware method that removes poorly supported foreground pixels through iterative neighborhood analysis of a thresholded mask. SPC-Clean is deterministic, training-free, preserves original fluorescence intensities for practical microscopy workflows.
bioinformatics2026-08-27v1EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-08-26v4UMITIC: An unsupervised framework for the joint characterization of cellular phenotypes and spatial neighborhoods in multiplex and hyperplex immunofluorescence imaging data
Sangüesa Recalde, M.; De Andrea, C. E.; Ariz, M.Abstract
Multiplexed imaging technologies enable the simultaneous measurement of dozens of protein markers while preserving context, providing a high-resolution view of tissue organization schemes. However, extracting meaningful insights from these high-dimensional datasets--particularly in hyperplex settings (>20 markers)--remains a major computational challenge, especially in the absence of annotated data. Here, we present UMITIC (Unsupervised Analysis of Multiplex Images via TIssue Characterization), a modular and unsupervised computational framework for the joint characterization of cell phenotypes and tissue neighborhoods from multiplex imaging data. UMITIC integrates three components: (i) CellCut, a strategy that combines nuclear and cytoplasmic predictions to improve the delineation capabilities of the framework; (ii) CellMap, a contrastive learning approach that generates low-dimensional representations of single-cell image crops that are enriched with morphological features; and (iii) TissueNet, a graph neural network that models spatial cell-cell interactions to identify tissue neighborhoods. We evaluated UMITIC across four datasets of increasing complexity to assess its robustness, scalability and biological relevance. With respect to a 7-plex human tonsil dataset, the framework identified canonical immune cell populations and reconstructed well-established anatomical regions. When applied to a 43-plex tonsil image, UMITIC preserved these tissue-level structures while enabling a finer cell subtype stratification process driven by increased marker dimensionality. We further validated our method on a 58-plex colorectal cancer cohort, where UMITIC was able to recover previously reported immune composition differences and spatial organization variations between patient groups with different prognoses. Finally, when an expert-annotated mass cytometry imaging dataset concerning human lung tissue was used, UMITIC achieved higher agreement with the reference tissue annotations than the existing approaches did, demonstrating improved lung microanatomy reconstruction accuracy. Together, these results show that UMITIC enables consistent and interpretable analyses of both cellular phenotypes and tissue architectures across diverse multiplex and hyperplex imaging datasets without the need for manual annotations.
bioinformatics2026-08-26v3CRISPR-HAWK: Haplotype- and Variant-aware Guide Design Toolkit for CRISPR-Cas
Kumbara, A.; Tognon, M.; Carone, G.; Fontanesi, A.; Bombieri, N.; Giugno, R.; Pinello, L.Abstract
Current CRISPR guide RNA design tools rely on reference genomes, overlooking how genetic variation impacts editing outcomes. As genome editing advances toward clinical applications, incorporating population diversity becomes essential for ensuring therapeutic efficacy across diverse populations. We present CRISPR-HAWK, a framework integrating individual- and population-scale variants and haplotypes into gRNA design. Analyzing therapeutic targets across 79,648 genomes reveals that genetic variants substantially alter guide performance. For the clinically approved sickle cell disease therapeutic guide targeting BCL11A, we identify haplotypes that completely abolish predicted cutting activity. Across seven therapeutic loci, 82.5% of guides contain variants modifying on-target activity. Variants also create novel protospacer adjacent motif sites generating individual-specific guides invisible to reference-based design. These findings demonstrate that variant-aware selection is critical for equitable genome editing. CRISPR-HAWK is available at https://github.com/pinellolab/CRISPR-HAWK and https://github.com/InfOmics/CRISPR-HAWK
bioinformatics2026-08-26v3Enrichment-free glycoproteomics harnessing real-time mass defect-driven glycopeptide classification reveals sex differences in murine fucosylation
Zhang, B.; Chau, T. H.; Bienes, K. M.; Arakawa, H.; Hane, M.; Sato, C.; Yokoi, A.; Kaji, H.; Ashwood, C.; Matsui, Y.; Kawahara, R.; Thaysen-Andersen, M.Abstract
Glycopeptide enrichment remains a cornerstone in glycoproteomics, but bias and reproducibility issues continue to hinder biological insight and clinical translation. Employing curated glycoproteomics datasets and machine learning, we trained a glycopeptide classifier to recognize N-glycopeptide precursors through mass defect signatures. Integration of the classifier into a data-dependent acquisition framework facilitated real-time prediction of N-glycopeptides from human serum and revealed sex differences in murine plasma fucosylation opening avenues for enrichment-free glycoproteomics.
bioinformatics2026-08-26v3MONTE enables unified pan-cancer tumor purity estimation andmethylation correction from bulk DNA methylation arrays
Kim, M.; Lee, W.-H.; Yao, V.Abstract
Bulk DNA methylation profiling is widely used to study cancer epigenomics in clinical settings, but these measurements aggregate signals from malignant and non-malignant cells, introducing composition-dependent confounding that complicates tumor-intrinsic interpretation and cross-cohort analyses. While existing methods can estimate tumor purity and, in some cases, correct methylation measurements, they typically require cancer-specific reference models, matched normal samples, or predefined probe sets, limiting their applicability to rare cancers, different clinical cohorts, and cross-dataset comparisons. We present MONTE (Methylation-based Observation Normalization and Tumor purity Estimation), a unified, cancer label-free framework for tumor purity inference and CpG-resolved methylation correction from bulk DNA methylation data. MONTE learns probe-wise relationships between methylation and tumor purity using an empirical Bayes-moderated linear model and infers purity in new samples via signal-to-noise weighted aggregation, without requiring matched normals, cancer labels, or predefined probe sets. A single pan-cancer MONTE model outperforms existing cancer-specific methods for purity estimation across 21 cancer types, generalizes across purity references, and runs orders of magnitude faster on full-dataset analyses. MONTE also introduces Bayesian transfer learning, which enables efficient recalibration to alternative purity definitions, validated on three independent external cohorts. Methylation correction with MONTE further amplifies tumor-relevant regulatory signal and improves the reproducibility of differential methylation analyses. By unifying purity estimation and correction in a single flexible, scalable, and interpretable framework, MONTE broadens the accessibility of tumor-intrinsic methylation analysis across cancer types and datasets.
bioinformatics2026-08-26v2