Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
FlashDeconv reveals resolution horizons in atlas-scale spatial transcriptomics
Yang, C.; Chen, J.; Zhang, X.Abstract
Coarsening Visium HD resolution from 8 to 64 m can flip cell-type co-localization from negative to positive (r = -0.12 [->] +0.80), yet many widely used compositional deconvolution workflows require coarsening or subsampling at million-bin scale. Here we introduce FlashDeconv, which combines leverage-score importance sampling with sparse spatial regularization to achieve competitive benchmark accuracy while processing 1.6 million bins in 153 seconds on commodity hardware. Systematic multi-resolution analysis of Visium HD mouse intestine reveals a tissue-specific resolution horizon (8-16 m), the scale at which this sign inversion occurs, validated by Xenium ground truth. Below this horizon, FlashDeconv provides, to our knowledge, the first sequencing-based quantification of Tuft cell chemosensory niches (15.3-fold stem cell enrichment). In a 1.6-million-bin human colorectal cancer cohort, FlashDeconv uncovers neutrophil inflammatory microdomains co-localized with immunoregulatory dendritic cells (mRegDC) at the tumor-stroma interface, spatial niches largely missed by discrete-label summaries, with RCTD doublet mode labeling only 2.3% of hotspot bins as neutrophil singlets.
bioinformatics2026-09-11v5Fast and accurate taxonomic domain assignment of short metagenomic reads using BBERT
Alekhin, D.; Alon, M.; Sidi, T.; Perez Mazeh, S.; Carmi, G.; Finkel, O. M.; Erez, A.Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses.
bioinformatics2026-09-11v4Histology-Aware Graph for Modeling Intercellular Communication in Spatial Transcriptomics
Wang, X.; Tao, C.; Jiang, Y.; Jiang, Y.; Liu, H.; Jiang, Z.; Zhu, P.; Que, N.; Xi, J.; Price, S.; Mou, Y.; Xu, J.; Li, C.Abstract
Cell-cell communication (CCC) is essential to how life forms and functions. Recent tools achieve single-cell-resolved CCC inference utilizing spatial transcriptomics (ST). However, most ignore the modeling of tissue contexts surrounding cells, causing high false-positive/negative rates. Here, we propose HARMONIC, a CCC inference method integrating multimodal ST and hematoxylin and eosin (H&E)-stained images. HARMONIC causally modeling the transcriptomic-to-contextual relationships for CCC inference. The state-of-the-art performance was verified across ST platforms, species and healthy/diseased status, on both synthetic and biological samples. HARMONIC was applied in various real-world scenarios, especially on tissues with clear morphological boundaries, including cortical layers in mouse brain, medullary-cortex structures in mouse kidney, as well as tumor-stromal/immune interface. Significant refinement of false-positive/negative predictions was observed compared to ST-only CCC tools.
bioinformatics2026-09-11v2scPyviewer: a Python-native interactive viewer from AnnData single-cell data
Xuan, H.; Huang, Y.; Bian, J.; Liu, X.Abstract
Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.
bioinformatics2026-09-11v2SenSASP: A Unified, Multi-Layer Database of Senescence and SASP Genes
Xuan, H.; Huang, Y.; Bian, J.Abstract
Research on cellular senescence and the senescence-associated secretory phenotype (SASP) draws on independently curated gene resources that differ in scope, identifiers, and update cycles, making cross-resource integration error-prone. We unified four widely used resources, CellAge, GenAge, the SenMayo signature, and the Reactome Cellular Senescence pathway, onto a single canonical identifier (the Ensembl gene ID) and enriched every gene with three annotation layers absent from all four inputs: cross-species conservation, tissue and cell-type expression, and high-confidence protein-protein interactions. Unification collapsed 1,460 summed source entries into 1,250 unique genes (210 redundant entries removed, 14.4%) while preserving full source provenance: 173 genes are corroborated by two or more resources and two (IL6, JUN) by all four. The three annotation layers reach 95.8%, 97.9%, and 93.0% of genes, with 89.4% annotated across all three. A 500-gene random sample of identifier mappings was validated against HGNC and Ensembl (98.0% exact match). The result, SenSASP, is a single, machine-readable, provenance-tracked database of harmonized identifiers and net-new functional context, illustrated here with a gene-prioritization score and a tissue-expression atlas. SenSASP is freely available at https://xuan13hao.github.io/sensasp/
bioinformatics2026-09-11v2Distinct geometries, comparable interfaces: binding modes and thermodynamic implications in conventional and single domain antibodies
Hauser, A.; Dangla-Pelissier, G.; Cazals, F.Abstract
Heavy-chain only antibodies, produced by the adaptive immune systems of camelids and cartilaginous fish, complement canonical antibodies comprising both heavy and light chain variable domains. Using an integrated interface model that unifies interfacial atoms--including solvent molecules, contacts, buried surface areas, and interface curvature measures, we shed light on two aspects of antibody binding that have so far remained elusive when comparing single domain (SdAb) and double domain (DdAb) antibodies. First, contrary to previous reports of smaller SdAb interfaces, we show that SdAb achieve an interface size comparable to that of DdAb despite using a single variable domain and fewer interface residues, a consequence of a geometric pattern driven by convexity and curvature effects. Second, we show that SdAb exhibit a broader diversity of binding modes than previously reported, with a prominent role played by FR regions. Finally, we discuss the thermodynamic implications of these findings for the design of high-affinity single domain antibodies, with particular relevance to protein engineering and design.
bioinformatics2026-09-11v2Kintsugi decides, gene by gene, where spatial transcriptomics borrows information
Yang, C.; Zhang, X.; Chen, J.Abstract
Subcellular spatial transcriptomics captures where RNA is in tissue, but a single location holds too few molecules of any one gene to estimate composition alone. Every current method fixes in advance where to borrow -- a smoothing scale, a cell outline or a factor model -- and the fixed choice shapes what is visible. Kintsugi removes the fixed choice and lets held-out molecules decide, gene by gene, how much to borrow from spatial neighbours and from other genes at the same location. On a lung section measured by both Xenium and Visium HD, the data-chosen allocation placed an epithelial programme where the Xenium molecules were, ahead of smoothing, cell segmentation and a factor model; the result replicated across tissues and against protein. Across a 45-core pulmonary fibrosis cohort, separating composition from captured amount shows that a fibroblastic focus is not a place with more RNA but a place with different RNA: 2.8-fold higher in activated-fibroblast composition while segmented nuclear density is at most 1.08-fold higher.
bioinformatics2026-09-11v2scOLAR: Ontology-Anchored Open-Set Annotation of Single-Cell RNA-seq Data
Liu, Y.; Yi, S.; Yin, H.; Ju, W.Abstract
Single-cell RNA sequencing profiles cellular heterogeneity at atlas scale, making automated annotation essential. However, target datasets often contain novel cell types missing from incomplete references. We present scOLAR, an ontology-guided open-set framework that learns prototypes over the Cell Ontology and uses both reference and target expression to annotate known classes while detecting unfamiliar populations. Guided by ontology hierarchies and decision-boundary regularization, scOLAR penalizes coarse-lineage misclassification and groups novel cells without requiring predefined cluster counts. Across benchmarks, scOLAR achieves a novelty-detection AUROC of 0.9726 and an average precision of 0.9871, enabling structured post-hoc lineage-level interpretation of populations absent from the reference.
bioinformatics2026-09-11v1Geomosaic: a flexible bioinformatics platform integrating complementary metagenomic analyses from sequencing reads to genomes
Corso, D.; Taccaliti, E.; Barosa, B.; Giovannelli, D.Abstract
Metagenomic analyses can be performed at multiple analytical levels, including read-based, assembly-based, and genome-resolved approaches, each capturing complementary biological information while introducing distinct analytical biases and trade-offs. However, existing workflows are commonly optimized for a single analytical strategy, making it difficult to integrate these complementary representations within a unified, reproducible framework. Here we present Geomosaic, a modular framework that integrates complementary analytical representations of metagenomic data, from reads to genomes, within a single scalable, customizable, and reproducible workflow. Built on a graph-based architecture implemented in Snakemake, Geomosaic enables users to construct complete end-to-end workflows or execute individual analytical modules while selecting among interchangeable software packages. The framework supports read preprocessing, quality control, taxonomic and functional profiling, assembly, genome reconstruction, genome-resolved annotation, custom HMM-based analyses, and automated downstream result aggregation. Automatic generation of execution scripts, modular workflows, and multiple analysis entry points make Geomosaic accessible to researchers approaching metagenomic analyses for the first time, while providing the flexibility and control required by expert users. Native support for HPC environments enables efficient analysis of datasets ranging from individual projects to large-scale metagenomic surveys. Rather than treating read-, assembly-, and genome-resolved metagenomics as alternative analytical strategies, Geomosaic integrates them as complementary representations of the same biological system, allowing users to move seamlessly between community-wide patterns and organism-resolved functional interpretation. By combining workflow flexibility, computational reproducibility, standardized analysis-ready outputs, and extensive documentation, Geomosaic provides a unified platform for environmental metagenomic analyses and facilitates reproducible downstream ecological and evolutionary investigations.
bioinformatics2026-09-11v1SpectroVQ: Noise-Aware Compression of Proteomics Data via Vector-Quantized Deep Learning improves MS/MS data storage and Peptide Identification
Lam, H.; Li, J. H. W.; Hoque, A.Abstract
The amount of proteomics data generated has dramatically grown for the past decade due to the wider accessibility to mass spectrometers and technological advances. Current data storage and compression techniques largely treat mass spectra as meaningless series of numbers, wasting storage on useless noise and limiting the compression ratio. Here, we present SpectroVQ, a noise-aware vector-quantized autoencoder to compress and denoise peptide tandem mass spectra without any prior annotation by exploiting peptide fragmentation pattern using deep-learning Evaluation results showed that SpectroVQ can preferentially retain useful signals in spectra from diverse peptide ions, including those in unseen datasets. SpectroVQ achieved over 3-fold increase in compression ratio over mzMLb while maintaining over 0.9 in average cosine similarity and 90% agreement in peptide identifications. In addition, we develop a novel strategy to increase peptide identifications by ~15% via ordinary library searching, by leveraging the tuneable denoising capability of SpectroVQ.
bioinformatics2026-09-11v1CIDER: detecting changes in gene regulatory networks that are associated with changes in phenotype
Jung, W. J.; Ding, M.; Liao, S.; Erdenebaatar, Z.; Brent, M.Abstract
Changes in gene regulatory networks may drive quantitative traits, or may transmit the effects of one trait, such as blood lipid level, on another, such as cardiovascular health. Yet the standard tools, differential correlation and differential network analysis, compare two discrete groups, while the contexts of interest - circulating lipids, inflammation, and blood glucose - vary continuously; applying them forces dichotomization, discarding within-trait variation. We introduce Continuous Interaction-based Differential Edge Regulation (CIDER), which tests whether a gene regulatory network edge, the relationship between a transcription factor and its target gene, varies with a continuous trait: the target gene's expression is modeled as a function of the TF's expression level, the trait, and their interaction, with the interaction coefficient measuring the trait dependence. To limit multiple testing, CIDER tests only the edges of a reference regulatory network. A generalized additive extension detects interactions that change the shape of the relationship, not only its slope, including forms that cannot be expressed as a difference between two correlations. In simulations it outperformed four two-group methods across sample sizes, effect sizes, and noise levels, with most of its advantage from keeping the trait continuous. In whole-blood transcriptomes from four independent human cohorts across ten quantitative health traits, CIDER identified 63 replicated cases in which a TF's regulation of its target varies with the trait, including coupling of the glucocorticoid-receptor (NR3C1) to the granulocyte colony-stimulating-factor receptor (CSF3R) that strengthens as triglycerides rise, and a pair whose regulation reverses direction across the observed range of C-reactive protein.
bioinformatics2026-09-11v1Early terminated transcripts and missing proteins reflect artifacts in bacterial proteomes
Insana, G.; Martin, M. J.; Pearson, W. R.Abstract
The high redundancy of many bacterial proteomes can be used to evaluate proteome quality and identify sequence errors. We have used MMseqs2 clustering with subsequent filtering to identify clusters that contain sequences from at least 50% of the clustered proteomes to build sets of core proteins that include proteins from 95% of the clustered bacteria. These clusters typically capture more than 80% of proteins in the bacteria. Because these clusters have highly uniform length (the median cluster has more than 99% of its proteins at the mode length), short (<75% of mode length) or long (>133%) proteins are likely artifacts. Most "outlier" proteins are found in fewer than 10% of clusters, and "high-outlier" clusters are over-represented in a small fraction of proteomes, which often have poor proteome BUSCO fragment scores. Short-outlier proteins are artifacts; at least 80% of short-outlier genomes contain mode-length copies of the protein, which were missed because of frame-shifts, termination codons, or initiation codon choice. MMseqs2 clustering with 50% participation provides robust sets of core bacterial proteins and can be used to identify lower-quality proteomes and proteins.
bioinformatics2026-09-09v5Pareto optimization of masked superstrings improves compression of pan-genome k-mer sets
Plachy, J.; Sladky, O.; Brinda, K.; Vesely, P.Abstract
The growing interest in k-mer-based methods across bioinformatics calls for compact k-mer set representations that can be optimized for specific downstream applications. Recently, masked superstrings have provided such flexibility by moving beyond de Bruijn graph paths to general k-mer superstrings equipped with a binary mask, thereby subsuming Spectrum-Preserving String Sets and achieving compactness on arbitrary k-mer sets. However, existing methods optimize superstring length and mask properties in two separate steps, possibly missing solutions where a small increase in superstring length yields a substantial reduction in mask complexity. Here, we introduce the first method for Pareto optimization of k-mer superstrings and masks, and apply it to the problem of compressing pan-genome k-mer sets. We model the compressibility of masked superstrings using an objective that combines superstring length and the number of runs in the mask. We prove that the resulting optimization problem is NP-hard and develop a heuristic based on iterative deepening search in the Aho-Corasick automaton. Using microbial pan-genome datasets, we characterize the Pareto front in the superstring-length/mask-run space and show that the front contains points that Pareto-dominate simplitigs and matchtigs. Finally, we demonstrate that Pareto-optimized masked superstrings improve pan-genome k-mer set compressibility by 12-19% when combined with neural-network compressors, achieving less than 1.2 bits per k-mer in common scenarios.
bioinformatics2026-09-09v3Sampling in structure-token space enables accurate prediction of multiple protein conformations
Wang, Z.; Yu, Y.; Zheng, W.-M.; Yu, C.; Bu, D.Abstract
Protein function is fundamentally mediated by ensembles of distinct metastable states. However, existing methods, such as AlphaFold 3, typically exhibit a bias toward predicting a single dominant state, failing to capture alternative conformations or provide robust metrics for identifying high-quality multi-state conformations. Here, we present MultiStateFold (MSFold), a framework that integrates Parallel Tempering into the discrete structure token space of the ESM3 protein language model. By conceptualizing the model's latent space as an implicit energy landscape, MSFold enables global exploration and barrier crossing, thereby overcoming the local sampling limitations inherent in base generative models. Across a benchmark of 313 multi-conformation pairs, MSFold sets a new performance standard: it achieves the highest success rate in modeling native states and substantially outperforms leading methods, including AlphaFold 3, on challenging alternative conformations, while maintaining competitive accuracy for primary structures. Furthermore, we propose Sequence Log-Likelihood (SLL), a novel confidence metric derived from sequence-structure consistency. Our results demonstrate that SLL offers a modest improvement over standard metrics such as pTM and pLDDT. This work establishes a new paradigm for conformational sampling, bridging classical statistical physics with protein language models.
bioinformatics2026-09-09v3Vizitig a pangenome and pantranscriptome explorer
Degardins, B.; Paperman, C.; MARCHET, C.Abstract
Vizitig is the first platform for real-time exploration and querying of DNA and RNA sequence de Bruijn graphs across many samples, unifying visualization, metadata, and flexible search. It constructs compacted colored de Bruijn graphs from raw sequencing reads and reference sequences, then provides an interactive web interface for graph exploration. It integrates raw and reference-based data, handles complex variation, and provides a human-readable feature-based (also referred to as metadata) query language with scalable graph loading. Its domain-specific query language supports composable searches combining sequences of arbitrary size, genomic features such as gene or exon identifiers, and experimental factors such as sample identifier or abundance thresholds. On-demand subgraph loading retrieves only regions of interest, enabling interactive exploration of large datasets in the graphical user interface without loading the entire graph into memory. We demonstrate Vizitig's capabilities through case studies in pantranscriptomics and pangenomics. In pantranscriptomics, we recover fusion transcript breakpoints on long and short reads. In pangenomics, we explore sequence variations across yeast, rice, nematode, and human pangenomes. Vizitig scales from small virus genomes to human-scale pangenomes (our largest experiments comprises up to 20 assembled human haplotypes, on a laptop). Vizitig enables fast, reproducible analysis in both pangenomics and pantranscriptomics while providing a deployable and user-friendly working environment.
bioinformatics2026-09-09v3Foldseek-Interface reveals a protein interface universe far from complete
Strom, J. M.; Cha, S.; Kim, R. S.; Sajal, H.; Gilchrist, C. L.; Steinegger, M.; Luck, K.Abstract
Protein-protein interactions mediate a vast range of cellular functions, requiring diverse modes of binding. While recent years have seen major efforts to chart and classify the protein structure universe, we lack comparable methods to assess and cluster that diversity in interface structure at interactome scale. Here, we present Foldseek-Interface, a method that converts 3D interface structures into searchable sequences to enable fast alignment and clustering of protein interaction interfaces. It matches the accuracy of state-of-the-art tools while running up to 230 times faster. Applying it to all biological assemblies in the PDB, we cluster 3.1 million dimers into 77{,}167 interface clusters and use this resource to characterise interface diversity, evolution, and pathogen mimicry. Application of Foldseek-Interface to resources of predicted protein complex structures rapidly revealed putatively novel interface types worth further experimental interrogation. Foldseek-Interface and the interface cluster resource are freely available as webservers for search https://search.foldseek.com/interface and exploration https://interface.foldseek.com.
bioinformatics2026-09-09v2FUSED: A Functional Representation for Joint Structural and Elemental Analysis of Protein Ligand Binding Sites
Priyankara, T. M. S.; Ellingson, L.Abstract
Ligand binding site representations are central to the analysis of protein-ligand interactions, with applications in functional characterization, binding-site comparison, and ligand recognition. Many descriptor-based approaches characterize ligand binding sites using a fixed distance threshold from the ligand, despite substantial variability in how such thresholds are defined and the possibility that relevant structural and compositional information changes across spatial scales. We propose Functional Unification of Structural and Elemental Descriptors (FUSED), a multivariate functional representation that jointly models structural and elemental compositional information of ligand binding sites as functions of distance from the ligand. Structural information is captured through covariance-based descriptors derived from the CDPA framework, while chemical composition is represented through isometric log-ratio coordinates to account appropriately for compositional geometry. Treating distance from the ligand as a functional domain allows structural and compositional characteristics to be examined across distance thresholds rather than at a single prespecified value. The resulting representation can be used directly or combined with dimension-reduction, statistical-learning, and/or machine-learning procedures, with the distance interval tailored to the dataset, analytical task, and procedure. We evaluate FUSED on three benchmark datasets spanning complementary ligand binding-site analysis tasks: the Extended Kahraman dataset for multiclass ligand discrimination, TOUGH-C1 for binary binding-site classification, and TOUGH-M1 for pairwise matching of pockets associated with drug-like ligands. FUSED supports strong discrimination in the EK and TOUGH-C1 tasks using standard statistical-learning procedures, while a supervised Siamese neural network applied to the full FUSED representation achieves a mean ROC-AUC of 0.9375 on TOUGH-M1 under repeated sequence-cluster-disjoint evaluation, approaching the strongest reported benchmark performance. These results demonstrate that FUSED provides a flexible, alignment-free representation that supports both direct examination of threshold-dependent binding-site characteristics and competitive downstream classification and pocket matching while remaining computationally practical.
bioinformatics2026-09-09v2EVd3x: A Bioinformatics Platform for Evidence-aware Interpretation of EV Cargo
Ait Ouares, K.; Weerakkody, J. S.Abstract
Extracellular vesicle (EV) profiling can identify hundreds to thousands of RNAs and proteins, but converting those lists into a coherent biological interpretation remains slow because the relevant evidence is scattered across cargo repositories, publications, pathway databases, cell atlases, and interaction resources. We introduce EVd3x.com, a source-linked bioinformatics platform that brings these evidence layers into one continuous workflow for EV research. EVd3x harmonizes more than 22 million records and interactions from 28 public resources into 17 analysis tables, including 234,090 reported EV-cargo records linked to 3,350 publications, 2.6 million miRNA-target relationships, 404,758 pathway memberships, more than 501,000 disease associations, 1.6 million cell-context records, 25,779 ligand-receptor pairs, and 13.7 million protein-interaction records. Each query generates a fixed, exportable evidence packet that follows the same molecules from reported EV detection through disease, pathway, cellular, regulatory, and interaction analyses. Starting with the complex disease query [early-onset Alzheimers disease with behavioral disturbance], EVd3x arrived at a focused, PSEN1-centered evidence packet and organized thousands of linked records into a source-traceable interpretation. The workflow recovered expected Alzheimer biology while keeping reported EV detection separate from broader contextual evidence, thereby exposing evidence gaps and context-dependent relationships. EVd3x provides EV researchers with a practical route from a cargo list, or from a disease, pathway, or cell-context question to a prioritized set of molecules, relationships, evidence gaps, and unresolved questions that researchers can use to formulate hypotheses and design subsequent experimental studies.
bioinformatics2026-09-09v2Hub Facade: View Track Hubs in Integrated Genome Browser
Freese, N. H.; Raveendran, K.; Sirigineedi, J. S.; Chinta, U. L.; Badzuh, P.; Marne, O.; Shetty, C.; Naylor, I.; Jagarapu, S.; Loraine, A. E.Abstract
Summary: Genome browsers are essential for understanding genomic data in detail but use incompatible data repository formats and access protocols, limiting their use. We present a new web application Hub Facade that translates the Track Hub format used by the UCSC Genome Browser into the Quickload format used by the Integrated Genome Browser (IGB), and vice versa. We used the Facade's translation capability to write a single-page web application for IGB users to search and add UCSC-curated Hubs to IGB for exploration and visual analysis. We created a new IGB App (GenArk Genomes) that uses the Facade to automate importing genome assemblies from GenArk, a Hub-based repository with nearly 50,000 genomes hosted by the UCSC Genome Browser team. An example use case investigating alternative splicing of human gene MEOX1 shows how using both browsers to view the same data via the Hub Facade promotes understanding and discovery. Availability and Implementation: Hub Facade is free, open-source software deployed at translate.bioviz.org. Code is available from git repositories at [bitbucket.org|github.com]/lorainelab/hub-facade. The use case is available as Supplemental File 1.
bioinformatics2026-09-09v2Evaluating Few-Shot Meta-Learning using STUNT for Microbiome-Based Disease Classification
Peng, C.; Abeel, T.Abstract
The human gut microbiome is increasingly explored as a diagnostic indicator for disease, yet machine learning models trained on metagenomic data are often constrained by limited sample sizes and poor cross-cohort generalizability. Meta-learning, a machine learning paradigm that optimizes models for rapid adaptation to new tasks with limited examples, offers a promising strategy to address this by leveraging the potential shared microbial structure across publicly available metagenomic datasets. Here, we evaluated STUNT, a framework combining self-supervised pretraining with metric-based meta-learning (Prototypical Networks), for few-shot microbiome-based disease classification. Using over 5,000 species-level gut metagenomic profiles from 57 cohorts in GMrepo v2, we meta-trained STUNT on 52 cohorts and evaluated the pretrained embedding on five held-out disease cohorts covering rheumatoid arthritis (RA), gestational diabetes mellitus during pregnancy (GDM), non-alcoholic fatty liver disease (NAFLD), diabetes mellitus, type 1 (T1D), and inflammatory bowel disease (IBD). We compared Prototypical Networks, Logistic Regression, and Random Forest with and without STUNT-derived embeddings across shot sizes of 1 to 10 samples per class. We found that STUNT-derived embeddings provided a modest benefit only under extreme data scarcity (one labeled sample per class) and this advantage rapidly diminished and reversed with additional samples, indicating that the meta-learned representations impose an information bottleneck limiting access to task-specific signals. Classification performance varied substantially across cohorts, consistent with PERMANOVA-estimated microbiome-disease separability. These results highlight the need for representation learning approaches that preserve disease- and cohort-specific variation and suggest that intrinsic biological signal strength is the primary determinant of classification success.
bioinformatics2026-09-09v2ProteinSage: From implicit learning to explicit structural constraints for efficient protein language modeling
Shen, L.; Chao, L.; Liu, T.; Liu, Q.; Zhou, G.; Wang, H.; Dong, X.; Li, T.; Zhang, X.; Ni, J.Abstract
While protein language models typically rely on sequence-only pretraining objectives, this approach often fails to capture structural regularities and demands large datasets. To address this, we introduce ProteinSage, a pretraining framework that learns protein representations under explicit structural constraints. ProteinSage incorporates structural signals via structure-guided masking and a causal objective designed to model longrange dependencies. This structure-constrained pretraining equips ProteinSage with transferable representations using less data and computation, yet achieves competitive or superior performance across diverse structure-aware and general protein modeling benchmarks. To determine whether these gains stem from genuine structural generalization rather than task-specific fitting, we applied ProteinSage to a structure-driven protein discovery task, focusing on proteins with multi-pass transmembrane helical architectures such as distantly related microbial rhodopsins. The model successfully identified six previously unannotated microbial rhodopsin homologs. Together, our work establishes structure-constrained pretraining as an effective pathway toward data-efficient and structurally faithful protein representation learning.
bioinformatics2026-09-09v2Building an open infrastructure for molecular neuroimaging: standards and tools from the OpenNeuroPET initiative
Ganz, M.; Norgaard, M.; Pernet, C.; Matheson, G. J.; Galassi, A.; Ceballos, E. G.; Wighton, P.; Bilgel, M.; Eierud, C.; Gonzalez-Escamilla, G.; Buckholtz, J.; Blair, R.; Markiewicz, C. J.; Hardcastle, N.; Greve, D. N.; Thomas, A. G.; Poldrack, R. A.; Calhoun, V. D.; Innis, R. B.; Knudsen, G. M.Abstract
Molecular neuroimaging with positron emission tomography (PET) and single-photon emission computed tomography (SPECT) enables quantification of specific molecular targets in the living brain. Despite its scientific impact, molecular neuroimaging research has historically faced challenges due to high costs, small sample sizes, laboratory-specific analysis pipelines, and limited large-scale data sharing. These factors have hindered reproducibility and the broader reuse of valuable PET datasets. The OpenNeuroPET initiative was established to address these barriers by developing standards, infrastructure, and open-source tools currently focused on PET. Through collaborations across Europe and North America, OpenNeuroPET has supported the PET extension of the Brain Imaging Data Structure (PET-BIDS), providing a standardized framework for PET datasets and metadata. Building on PET-BIDS, tools such as PET2BIDS, ezBIDS, and BIDSCoin facilitate data conversion and curation. In parallel, OpenNeuro now hosts PET-BIDS datasets for open sharing, while complementary platforms such as PublicnEUro provide controlled-access pathways designed to support compliance with the European Union General Data Protection Regulation (GDPR). Emerging open-source workflows further support automated, reproducible PET analysis, promoting harmonization across centers. Together, these developments mark an important step toward an open molecular neuroimaging community in which datasets, software, and workflows can be transparently shared, reused, and scaled for collaborative research.
bioinformatics2026-09-09v2GENETHOFF: a flexible workflow for genome wide profiling of CRISPR/Cas off targets
Corre, G.; Rouillon, M.; Mombled, M.; Amendola, M.Abstract
We developed GENETHOFF, a flexible single-command versatile Snakemake workflow designed for the comprehensive analysis of CRISPR/Cas9 related OFF-targets genomic positions from GUIDE-Seq derived protocols. It efficiently processes multiplexed libraries from different organisms, PCR orientations, and Cas nucleases with varying PAM specificities in a single run, all based on a simple user-specified datasheet containing sample metadata.
bioinformatics2026-09-09v2Cell Painting-Based Tool for the Risk Assessment of Mammary Carcinogens and Endocrine Disruptors
ACHEBOUCHE, R.; Taboureau, O.Abstract
Breast cancer is the most common cancer in women worldwide and chemicals disrupting estrogen or progesterone signaling are recognized as potential risk factors. However, chemicals that alter the mammary gland (MG) development and function remain understudied, highlighting the need for additional research in this area. To address this gap, we investigate the relevance of using high-content imaging assays and more specifically, Cell Painting technology to measure cell morphology perturbation caused by chemical exposure and identify morphological features that could characterize mammary carcinogens (MC) risk factors. Using a dataset of MC and non-mammary carcinogens (Non-MC) with Cell Painting profiles from the JUMP-CP dataset, we retrieved 51 compounds: 28 MC, 23 non-genotoxic Non-MC. We characterized the morphological data by non-linear dimensionality reduction (UMAP) and hierarchical clustering. We, then, developed a Guilt-By-Association (GBA) framework comparing multiple configurations of similarity metrics, risk-score aggregation approaches and features representations. Morphological profiles clustered by mechanism of action rather than carcinogenicity label: genotoxic MC produced strong perturbations in endoplasmic reticulum, mitochondria, nucleus, and RNA compartments, whereas hormonally active compounds were indistinguishable from controls, reflecting the lack of functional steroid hormone receptors in the cell line used. Our best GBA configuration achieved an AUC-ROC of 0.630 and an AUC-PR of 0.696. Applied prospectively to endocrine disruptors chemicals, it prioritized clofentezine, 3-methylpyrazole, resorcinol, 2-tert-butyl-4-methoxyphenol, and thiabendazole as candidates for confirmatory testing. This study provides a transparent, interpretable tool for prioritizing chemicals in mammary carcinogenicity assessment, highlights limitations and clarifies where Cell Painting datasets must be complemented by hormone-sensitive models.
bioinformatics2026-09-09v1Predicted and experimental protein-ligand coordinates with distance labels and evaluation splits
Klamt, T.Abstract
Structure prediction tools now provide protein-ligand geometry for pairs no laboratory experiment has resolved thus far. This data is typically provided without calibrated, practically usable per-system measures of whether a given complex is correct and how much to trust it. Agreement between independently constructed methods is an established signal for this type of problem. What has been missing is a way to derive what a given level of agreement is worth. We present PLI-Parallax, which deposits predicted geometry together with experimental data needed to calibrate it. Chai-1, two Boltz-2 configurations, and the docking engine smina were run over shared inputs across an experimentally resolved crystal tier of 19,350 complexes and a corpus tier of 31,746 cross-docked pairs without experimental ground truth on predicted receptors, yielding 307,314,646 residue-to-ligand-atom distance records, which the stored coordinates allow a consumer to recompute at a cutoff of their own choosing. Their mutual agreement is fitted against observed accuracy where experimental data permits it and carried to where it does not. This way 30,567 systems carry predicted label accuracy together with a split-conformal interval. On the crystal tier that interval covers observed accuracy at the stated rate. On the corpus tier it ranks systems by expected label quality, since both distributions differ. The deposit is accompanied by 646 evaluation configurations across seven data split families, and 631 of them report how far their training and test entities separate under a two-sample test. The 906 protein accessions were partitioned to control sequence leakage, so a protein-cold split here tests generalisation across sequence space. The two tiers support evaluation against experimental data and, where this is absent, supervision weighted by how far the configurations agree.
bioinformatics2026-09-09v1Novel Target Combinations in Lung Squamous Cell Carcinoma proposed by the Emet AI Research Environment and supported by discovery stage experimentation
Soman, J.; Kundu, A.; Newington, J.; Wong, S.; Desai, N.; Leung, S.; Cudini, J.; Suarez, F.; Grandsard, P.Abstract
Early drug discovery is frequently bottlenecked by target identification, a challenge that becomes particularly difficult in complex diseases driven by overlapping, redundant pathways rather than a single dominant driver. Lung squamous cell carcinoma (LUSC) is one such disease: mutationally complex, lacking a dominant actionable target, and marked by a long series of failed single-agent targeted trials. We hypothesized a research environment, capable of reasoning across interconnected datasets, would be well-suited to generate novel, mechanistically grounded therapeutic hypotheses for LUSC. Emet, an AI research environment utilizing a biomedical knowledge graph of more than 1.5 billion triples connected to over 100 specialised biological databases, was paired with an Agentic Research Director that pursues each hypothesis through a graph-of-thoughts search, invoking scoped link-prediction and retrieval tools and committing every round of findings to persistent memory. The top-ranked hypotheses were two-target combinations rather than a single target. The four highest-ranked combinations were advanced to dose-matrix testing in NCI-H520 and SK-MES-1 cells, and two produced combination effects exceeding either single agent; to our knowledge neither of these two pairings had previously been evaluated in LUSC. Dual inhibition of CDC7 (TAK-931) and PKMYT1 (RP-6306) produced statistically significant synergy in both lines that strengthened from day 5 to day 7 (HSA 19.24 to 24.19 in NCI-H520; 11.07 to 13.73 in SK-MES-1). Dual inhibition of USP13 (spautin-1) and PI3K (alpelisib) produced an additive, cell-line-dependent effect, and immunoblotting confirmed the predicted mechanism: time-dependent depletion of MCL-1, c-Myc and SOX-2 driven by the USP13 arm. Agentic reasoning over multi-domain biomedical evidence can therefore nominate testable, mechanism-bearing combination hypotheses that survive experimental scrutiny.
bioinformatics2026-09-09v1Scaling Quantum Optimisation Beyond Hardware Limits for Real-World Scientific Workloads: Genome Assembly on Current Quantum Hardware
G Sankar, N.; Miliotis, G.; Caton, S.Abstract
Genome assembly is important in infectious disease surveillance, antimicrobial resistance monitoring, and cancer genomics. The task of reconstructing full genomic sequences from fragmented reads, can be framed as a large scale combinatorial optimisation problem. Recent advances in quantum computing have introduced new optimisation algorithms with potential advantages for navigating complex combinatorial search spaces. However, practical deployment is limited by noisy intermediate-scale quantum (NISQ) hardware, including restricted qubit counts, limited connectivity, and high error rates. In this research, we employ the Hamiltonian Auto Decomposition Optimisation Framework (HADOF), an algorithm agnostic framework that enables scalable quantum optimisation through federated solving across small subproblems. HADOF enabled the quantum-assisted genome assembly of a 7.1 Million base pairs Pseudomonas aeruginosa genome, to our knowledge, representing the largest genome assembly graph studied on real quantum hardware to date. The results achieved a 99.348% genome fraction and 1.0 duplication ratio, demonstrating that biologically plausible genome reconstructions can be obtained despite current hardware limitations.
bioinformatics2026-09-09v1Independent benchmark of H&E-based gene expression prediction in skin
Shaikhutdinova, R.; Gansberger, S.; Staller, J.; Singh, N.; Oyarzun, I.; Simon, M.; Sterniczky, B.; Tschandl, P.; Griss, J.Abstract
Given the widespread availability of H&E slides, there is considerable interest in determining whether molecular information can be inferred directly from tissue morphology, potentially reducing the need for costly spatial transcriptomic profiling. We assessed three state-of-the-art methods for predicting single-cell gene expression from H&E images across three skin disease contexts and two Xenium panels. As controls, we included simple linear regression models trained on embeddings from multiple foundation models, totalling 16 models evaluated in this study. We show that all models performed poorly: for most genes, prediction accuracy was near zero, and reliable predictions were largely restricted to keratinocyte-associated genes. Predicted expression failed to preserve cell-type identity and spatial organisation, with only keratinocytes forming coherent clusters, while immune, fibroblast, and other dermal populations were extensively mixed. Notably, simple ridge regression on pretrained embeddings matched or outperformed the more complex published architectures, indicating that the predictive signal originates primarily from image representations rather than model design. Our results demonstrate that current H&E-based gene expression prediction methods are not yet suitable for single-cell-level interpretation of spatial transcriptomics in skin tissue.
bioinformatics2026-09-09v1A geometry-over-coevolution principle governs protein complex assembly in AlphaFold
Li, S.; Mu, Z.; Yan, C.Abstract
AlphaFold has revolutionized protein complex structure prediction, yet how it assembles intermolecular interfaces remains poorly understood. Contrary to the prevailing view that inter-protein coevolution drives complex prediction, we uncover a geometry-over-coevolution principle governing assembly in AlphaFold-Multimer and AlphaFold3. Through systematic perturbation of evolutionary and structural inputs and development of residue-level constraint propagation mapping to trace the emergence and propagation of geometric information within the network, we show that prediction accuracy is governed primarily by monomer-derived structural geometry and interface-specific sequence-geometry compatibility, rather than direct inter-protein coevolutionary signals. Our mapping reveals a hierarchical assembly mechanism in which monomer-level geometric representations are established first and progressively propagated to constrain cross-chain interfaces. This mechanism further explains why antigen-antibody complexes are predicted less accurately, as their intrinsic interface plasticity and non-canonical architectures limit the propagation of geometric constraints across interfaces. Together, these findings establish a mechanistic framework for understanding how artificial intelligence models assemble protein complexes and provide principles for improving next-generation structure prediction.
bioinformatics2026-09-08v4ProMaya: a hierarchical universal Deep Learning framework for accurate and interpretable Protein-Protein interaction identification
Bhati, U.; Gupta, S.; kesarwani, V.; Shankar, R.Abstract
Protein-protein interactions (PPIs) are molecular lego which define the physical states of cells. Accurately identifying PPIs remains challenging due to the interplay of several factors ranging from electrostatic to molecular geometry, topology, and physics. Existing computational approaches capture only fragments of this orchestra, limiting their generalizability across protein families and interaction types. Here, we present ProMaya, a hierarchical multi-scale Graph-transformer framework that integrates 3D atomic geometry, electronic distribution, residue-level structure and disorder, surface mass-density signatures, and large protein language-model embeddings of interacting proteins. Highly comprehensively benchmarked across nine species and 47 GB experimentally validated data, ProMaya achieved consistently >95% average accuracy, outperforming state-of-the-art tools by >12%. As driven by its explainability, the first time introduced atomic and protein language information dramatically boosted it to an outstanding level for PPI discovery in any species, potent to even bypass costly experiments. ProMaya system is freely accessible at https://scbb.ihbt.res.in/ProMaya/
bioinformatics2026-09-08v2SPARKLE: evidence-constrained correction of local RNA leakage in high-resolution spatial transcriptomics
Wang, S.; Zhu, B.; Li, S.; Wei, X.Abstract
High-resolution sequencing-based spatial transcriptomics, including Stereo-seq and Visium HD, aggregates dense capture units into cell-resolved expression matrices. During tissue processing and permeabilization, RNA released from source cells can spread to neighbouring capture locations, reducing cell-type specificity and biasing downstream analyses. Here we developed SPARKLE (Spatial Ambient RNA Kernel-based Leakage Estimator), a cell-level correction method that uses capture locations outside cell-segmentation masks as within-sample spatial evidence of leakage. SPARKLE fits sparse spatial kernels to out-of-mask observations to estimate a sample-level propagation scale and gene-specific leakage coefficients. It corrects only genes supported by out-of-mask goodness of fit and uses expression-dependent conservative shrinkage to protect highly expressing source cells. In ten simulated scenarios, SPARKLE achieved the highest cell-wise concordance in eight and the lowest RMSE in nine. In axolotl brain, mouse brain and human ovarian cancer, SPARKLE removed ectopic marker signal from neighbouring cells while retaining source-cell expression, improved agreement with independent single-cell and single-nucleus references, and recovered an inferred fibroblast-to-tumor COL1A2-SDC4 communication route that was obscured by ectopic COL1A2 expression. Conclusions remained stable across plausible spatial scales and background-bin sizes. Runtime scaled near-linearly with tissue-window area and was further accelerated on GPU. SPARKLE is therefore a reference-free, fast and scalable method for correcting local RNA leakage from evidence contained within each sample, improving the reliability of cell-type localization, tissue-compartment identification and cell-cell communication inference.
bioinformatics2026-09-08v2A shared functional organisation underlies vascular disease remodelling
Bradford, A.; Bidula, S.; Fabian, L.; Warren, D.Abstract
Cardiovascular disease involves coordinated remodelling across multiple biological processes. Gene-level signatures vary substantially between studies because different combinations of genes can support similar biological processes. Pathway-level analyses can provide more stable representations of these processes but typically consider pathways independently. We hypothesised that grouping related pathways into conserved biological functions and quantifying their relative weighting would reveal a higher-order, transferable property of vascular tissue that we term functional organisation. We quantified the relative weighting of six conserved biological functions across independent transcriptomic datasets spanning human vascular disease, experimental models and therapeutic interventions. Vascular tissues exhibited a reproducible functional organisation defined by the balance of these functions. A dominant remodelling trajectory captured coordinated, nonlinear rebalancing of structural, immune, signalling and metabolic programmes, while a second dimension distinguished contractile/ECM organisation from immune activity. This organisation was preserved across independently reconstructed reference cohorts and remained robust to analytical sensitivity testing. When independent datasets were projected into this fixed framework, biological and clinical phenotypes occupied coherent positions along the remodelling landscape. Plaque-derived vascular, stromal and immune cell populations also occupied ordered functional positions, linking cellular heterogeneity to tissue-level organisation. The same organisational structure generalised across atherosclerosis, peripheral vascular disease and abdominal aortic aneurysm and revealed continuous biological heterogeneity within conventional clinical classifications. Genetic, pharmacological and dietary perturbations reproducibly shifted functional organisation, demonstrating that organisational state is dynamic and biologically responsive. Together, these findings identify functional organisation as a reproducible tissue-level property of vascular remodelling and provide a framework for understanding cardiovascular disease as coordinated rebalancing of biological functions rather than alteration of individual pathways in isolation.
bioinformatics2026-09-08v1AnnoAudit: a marker-based protocol for auditing single-cell atlas annotations reveals annotation-driven artifacts in a widely used traumatic brain injury resource
Zhang, L.; Rao, H.; Li, M.; Qian, X.; Zhang, Y.; Yan, Q.; Gao, R.Abstract
Single-cell atlas annotations are routinely treated as ground truth but rarely validated before use. We present AnnoAudit, a marker-based audit protocol that combines four convergent checks - marker scoring, unsupervised clustering, margin-gated module scoring, and applicability-gated pretrained models - into a composite Annotation Contamination Score (ACS), plus a trajectory-correlation fingerprint tracing suspicious signals to their cell type of origin. Applied to CEREBRI (GSE269748), the most widely used single-cell TBI atlas (73 citations), the official "glutamatergic neuron" label is systematically contaminated: of 45,051 labeled cells, only 963 (2.1%) are marker-confirmed glutamatergic neurons; the remainder are microglia (31.4%), astrocytes (18.6%), oligodendrocytes (15.9%), and other types. The contamination generates coherent false signals - a biphasic trajectory for 40 of 307 ion-channel genes, a KCNC3-specific OXPHOS signature, and an inversion of KCNC3 regulation at 7 days - whereas the corrected response is a sustained acute KCNC3 up-regulation conserved across four datasets and three injury models, and the fingerprint traces 96.3% of 244 informative ion-channel trajectories to non-neuronal populations. Simulation-calibrated ACS is 82.6% for CEREBRI and 70.4% for a human ALS atlas (GSE330130); three marker-defined controls pass. AnnoAudit needs only the deposited count matrix and a canonical panel; we propose it as a routine quality step.
bioinformatics2026-09-08v1Microhomology-Driven Genomic Alterations in Cancer Genomes: Patterns, Prevalence, and Clinical Implications.
Kostka, D.; Sztromwasser, P.; Kimmel, M.; Jaksik, R.Abstract
Homologous recombination deficiency (HRD) can force cancer cells to rely on alternative DNA repair pathways, including microhomology-mediated end joining (MMEJ), an error-prone mechanism that can generate deletions with microhomology at repair junctions. Because such patterns may reflect DNA repair defects with clinical relevance, this study aimed to assess the biological and prognostic significance of microhomology-associated deletions and optimize their detection parameters to improve patient stratification and inform targeted therapeutic strategies, including PARP inhibition. We optimized the microhomology length threshold (Mlt parameter) using signal-to-noise ratios and survival models. We evaluated the utility of whole-exome (WES) versus whole-genome sequencing (WGS). We then investigated how Loss of Function (LoF) alterations affect microhomology-associated deletion burden and assessed their prognostic significance across an ovarian cancer (OV) cohort and a TCGA Pan-Cancer dataset. Applying an Mlt = 2 threshold maximized the precision and statistical reliability of microhomology-associated deletion identification. WGS yielded higher reproducibility, whereas WES proved insufficient for accurate estimation. Inactivation of tumor suppressor genes, including RB1 and CCDC122, was associated with increased microhomology-associated deletion burden. Importantly, higher burden correlated with extended overall survival in ovarian cancer. Furthermore, while baseline microhomology-associated deletion burden varies by tumor type, Pan-Cancer analysis identified candidate gene-level alterations, including FDX1 and PDE8B, whose LoF increased microhomology-associated deletion burden across tumors. Precise parameterization combined with WGS resolution highlights the extent to which specific gene losses affect deletion burden consistent with MMEJ-mediated repair. Identifying these genetic vulnerabilities and the resulting deletion burden may provide a candidate prognostic biomarker and foundation for personalized therapies.
bioinformatics2026-09-08v1NANOCUTSIGHT: A NANOPORE-SEQUENCING APPROACH AND ANALYSIS PIPELINE TO ASSESS GENOME EDITING EFFICACY IN VARIOUS CELL POPULATIONS
Bergeron, D.; Gaudreault, V.; Duval, M.; Nassari, S.; Boudreau, F.; Durand, M.; Choquet, K.; Jean, S.Abstract
Genome editing has revolutionized biomedical sciences and is now an essential tool to define molecular pathways through genetic interaction and loss-of-function studies. Through its diverse variations, it allows for the generation of specific knockout cell lines or organisms, as well as the creation of endogenously edited gene regions. While high-throughput methodologies exist to map CRISPR/Cas9 genetic modifications, the validation of guide efficiencies in cell populations is often performed through analysis of the targeted gene product by western blotting or by deconvolution of Sanger sequencing chromatograms using TIDE or ICE assays. Here, we highlight a rapid nanopore sequencing pipeline, which we have named NanoCutSight, to quantify the percentage of indels at a specific genomic locus and to identify the types of modifications generated. We also benchmarked the methodology on various guide RNAs and in both cultured cell and organoid models. We believe that NanoCutSight will simplify the analysis of complex sample editing and enable the rapid screening of edited samples.
bioinformatics2026-09-08v1EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
Ding, H.; Wu, N.; Qiu, T.Abstract
DNA foundation models such as Evo2~7B adopt hybrid Hyena/attention architectures (StripedHyena2) whose single-stream autoregressive decoding is bounded by weight bandwidth at ${\sim}45$~\tokps. Speculative decoding on such hybrids faces a systems problem that prior SSM work solves only partially: after a draft is verified, rolling back model state cannot be reduced to truncating a KV cache, because the inference state mixes fixed-length FIR sliding windows, infinite-memory IIR recurrences, and append-only KV caches. We present \textbf{EvSpark}. (i)~A block-verification forward pass together with a per-position \emph{state-slicing} rollback protocol that handles all three state classes jointly with zero recomputation, reducing speculative overhead to 1.05--1.16$\times$ under an identity drafter and eliminating the second per-round forward pass of snapshot-and-replay schemes. (ii) A DSpark-style distilled neural drafter ported to a hybrid architecture ($\gam$-parallel trunk $+$ target hidden-state prefix $+$ full-matrix Markov head); controlled injection-layer comparisons at matched budget and two seeds find no layer-choice effect beyond seed noise at realistic budgets, while smoke-scale budgets systematically misrank layer types; the ${\sim}10^{11}$ residual magnitudes of late blocks require an fp32-scale distillation codec. (iii)~Under a 24-prompt $\times$ 1024-token $\times$ 2-seed protocol, EvSpark delivers \textbf{\speed{3.16}} end-to-end at its deployment configuration ($\gam{=}12$, 80M supervised positions) and \speed{2.88} on the reference model used for our deep-dive analyses ($\gam{=}7$, 150M), holding from 1k to 262k context and across 32k tokens of generation depth. In exact arithmetic the scheme preserves the target distribution; empirically, greedy decoding reproduces native outputs token-for-token (zero non-tie divergences across 24 prompts $\times$ 4 checkpoints), and the sampling path sits at the native bootstrap noise floor for 7/10 prompts with bounded bf16-level deviation elsewhere (${\le}1.7\times$ floor; worst-case unigram TVD 5.6\%). The budget-sweet-spot configuration adds \textbf{1.06 GPU-hours} of training on RTX~4090-class hardware and reaches \speed{3.03} (incremental to a one-off teacher-signal dump of ${\approx}19$ GPU-h shared by all configurations).
bioinformatics2026-09-08v1FFPERescuer: deep unsupervised domain adaptation for the reconstruction of gene expression profiles derived from formalin-fixed paraffin-embedded samples
He, l.; Song, K.; Li, Y.; Dong, Y.; Wong, C. Y. N.; Qi, L.; Zhang, X.; Lenos, K.; Back, T. d.; Elbers, C.; Xu, C.; Leung, R. M. H.; Deng, R.; Zhang, Y.; Qiao, S.; Gao, F.; Chen, Y.; Ng, S. S.-M.; Zhou, S.; Vermeulen, L.; Wang, X.Abstract
Formalin-fixed paraffin-embedded (FFPE) tumor tissues often suffer from RNA degradation, posing a long-standing challenge for reliable transcriptomic profiling. Here, we propose FFPERescuer, a deep learning framework employing unsupervised domain adaptation, to rectify distorted gene expression data. FFPERescuer comprises a partial encoder that maps a small subset of genes to high-level representations and a decoder to reconstruct full gene expression profiles. On simulated data with varying noise levels, FFPERescuer faithfully recovered gene expression profiles, achieving high Pearson correlation coefficients (PCCs > 0.85) with the ground truth. In FF-FFPE-matched cohorts, FFPERescuer significantly enhanced expression profile concordance, with average PCCs increased by 23% (P < 0.05). Applying to cancer subtyping, FFPERescuer improved classification accuracy from 67% to 92%, recapitulated subtype-specific biological properties lost in the FFPE-derived data, and enhanced survival associations. Our studies provide a powerful framework for reliable transcriptomic profiling from FFPE-archived tumor samples that are widely available in the clinic.
bioinformatics2026-09-08v1TCRdenoise - an unsupervised similarity-based approach for denoising of TCR-pMHC specificity data
Lund, J. M.; Deleuran, S. N.; Nielsen, M.Abstract
Public repositories of T cell receptor (TCR)-peptide-MHC (pMHC) interactions constitute a critical resource for studying adaptive immunity and developing predictive models of TCR specificity. However, recent evidence suggests that a substantial fraction of reported TCR-pMHC interactions may be incorrectly annotated, limiting the quality of downstream analyses and machine learning applications. Here, we present an unsupervised sequence similarity-based framework for denoising peptide-specific TCR repertoires. The method combines pairwise TCR similarity metrics derived from TCRbase and TCRdist3 with hierarchical clustering and a novel adaptation of the silhouette score designed to address the prevalence of singleton clusters and highly imbalanced cluster structures. By incorporating a pseudo-cluster containing singleton and background TCRs, and by optimising both clustering distance thresholds and minimum cluster-size criteria, the proposed approach identifies TCRs likely to represent true antigen-specific binders while filtering putative noise. Using experimentally validated repertoires from TCRvdb, we demonstrate that the modified silhouette score closely tracks clustering solutions that maximise separation between binding and non-binding TCRs, achieving strong agreement with independent validation based on the Matthews correlation coefficient. Extension to a large collection of peptide-specific TCR data revealed a strong negative correlation between the percentage of TCRs classified as noise and the predictive performance of peptide-specific binding models. Further, the denoising classification labels on this data set were corroborated using structural modeling confidence scores of the peptide-TCR interface extracted from a refined AlphaFold 3 modeling pipeline. Additionally, retraining NetTCR on denoised data improved internal cross-validated performance compared with models trained on the full data set, whereas models trained exclusively on TCRs classified as noise performed close to random. Together, these results demonstrate that sequence similarity-based denoising can effectively enrich for biologically meaningful TCR-pMHC interactions and improve the quality of training data for predictive immunological models. The proposed framework provides a scalable strategy for improving the reliability of public TCR databases and facilitating the development of more accurate TCR specificity prediction methods.
bioinformatics2026-09-08v1Atlas-scale single-cell analysis beyond in-memory paradigm with scAtlasPy
Xu, H.; Ye, Y.; Zhang, S.; Xie, R.; Li, J.; Lin, J.; Hu, Y.; Gao, L.Abstract
Single-cell atlases are rapidly outgrowing the memory capacity of standard workstations, challenging the in-memory paradigm underlying mainstream computational ecosystems. Here, scAtlasPy decouples scale of atlas from memory capacity by leveraging the disk-resident computing. It enables full-resolution analysis of a 100-million-cell atlas with only 42.9 GB peak memory, whereas state-of-the-art platforms are limited at 3 million cells with 512 GB memory. scAtlasPy achieves 137,745 cells/s, 10.4x faster than scDataset with 82.6% lower memory usage for random minibatch retrieval. Its extensible architecture offers a flexible platform for diverse atlas-scale analytical tasks, facilitating the discovery of complex cellular heterogeneity and functions in massive cell atlases.
bioinformatics2026-09-08v1CellART: a unified framework for extracting single-cell information from high-resolution spatial transcriptomics
Chen, Y.; Liu, Y.; Wang, Z.; Zeng, Y.; Chao, Z.; Jiang, P.; Chen, H.; Wang, J.; Xiao, J.; Yang, C.Abstract
Understanding how different cell types assemble into tissues and organs, as well as how they interact to transmit and receive biological signals, is essential for advancing biomedical and biological research. Recent advancements in spatial transcriptomics (ST) technologies have opened new avenues for investigating biological systems by achieving subcellular spatial resolution. Since cells are the fundamental units of life, extracting single-cell information from high-resolution ST data is crucial. However, existing ST platforms often capture sparse transcript counts per spot or measure only a limited number of genes, complicating the extraction of comprehensive single-cell information. In this study, we introduce CellART, a unified framework designed to extract single-cell information across diverse high-resolution ST platforms, including VisiumHD, Xenium, MERFISH, and Stereo-seq. By leveraging multimodal data, such as staining images, spatial transcriptomics data, and single-cell RNA sequencing references, CellART simultaneously performs cell segmentation and cell type annotation through a seamless integration of deep learning and probabilistic modeling. We demonstrate the efficiency, generalizability, and robustness of CellART across various high-resolution spatial transcriptomics platforms, capable of processing datasets containing millions of spots. Comprehensive experiments validate the biological relevance and accuracy of the recovered cellular information within spatial configurations. Notably, we highlight the utility of CellART in breast and colorectal cancer datasets, showcasing its ability to fully leverage high-resolution ST data. By enhancing cellular resolution, CellART facilitates the identification of transient cancer cell states and immune cell subtypes. Furthermore, CellART enables investigations into cancer-immune cell communication, uncovering both established interactions and novel ligand-receptor pairs. The outputs of CellART are compatible with widely used community tools, facilitating a variety of downstream analyses.
bioinformatics2026-09-08v1Novel Pipeline for Large-Scale Comparative Population Genetics
Majoros, S. E.; Cottenie, K.; Adamowicz, S. J.Abstract
As scientists continue to ask complex questions about biodiversity and deal with increasingly large amounts of data, there is a demand for new methods and computational developments to perform scientific analyses. Analytical pipelines and modules can provide a way to meet these demands and ensure reproducibility in scientific methods and analyses. The goal of this study was to create efficient, reproducible, reusable programming modules that are publicly available for future research. These modules were used to determine population genetic structure measures and compare these measures across species with different biological traits. The functionality of the modules is shown through a case study on Diptera (true fly) species from Canada and Greenland. We leveraged high-throughput DNA sequencing data from Northern areas, as it is a valuable resource and provides new opportunities to study the Arctic. Data were pulled from public databases (Barcode of Life Data System and Global Biodiversity Information Facility), as well as taxon-specific literature. The pipeline we developed in R includes fifteen modules, including modules to prepare and filter the data, calculate population genetic structure measures (e.g., FST), and run a multiple regression. These modules can be easily adapted and applied to a diverse set of animal groups, geographic regions, and biological traits. Best practices were followed for pipeline development, and the modules were designed and tested to work for datasets of different sizes by providing multiple different analyses and filtering options. Biological results were also obtained for Diptera species. Habitat and larval diet were both significantly related to population genetic structure. Evidence of isolation by distance and a relationship between population genetic structure and both latitude and longitude were also found. Overall, this study has created efficient, reusable bioinformatics modules, and provided insight into the factors affecting population genetic structure in Northern fly communities.
bioinformatics2026-09-07v4Learning Universal Representations of Intermolecular Interactions with ATOMICA
Fang, A.; Desgagne, M.; Zhang, Z.; Zhou, A.; Loscalzo, J.; Pentelute, B. L.; Zitnik, M.Abstract
Molecular interactions underlie nearly all biological processes, yet most representation models describe isolated entities or specialize in a single molecular setting. Here, we introduce ATOMICA, an interaction-centered geometric deep learning model designed to learn transferable representations of intermolecular interfaces across proteins, small molecules, metal ions, and nucleic acids. Self-supervised pretraining on 2,037,972 interaction complexes yields representations spanning atoms, molecular building blocks, and complete interfaces. The latent space captures molecular identity and interaction context, supporting sequence recovery and zero-shot prioritization of residues involved in non-covalent interactions. ATOMICA provides structural information complementary to sequence representations on RNA and protein-pocket ligand classification. Across protein-pocket analyses, ATOMICA distinguishes ATP- and ADP-associated pocket states and retrieves ligand-matched pockets across proteins without detectable structural alignment. The latent space also enables cross-modal comparison, with orthosteric inhibitor embeddings retrieving regions proximal to native peptide and protein interfaces. Applied to the dark proteome, ATOMICA-Ligand predicts candidate ions or cofactors for 2,646 pockets, and five heme candidates show Soret-band shifts consistent with heme association. Together, these results show how interaction-centered molecular representations can transfer structural information across molecular interaction types and generate experimentally testable hypotheses.
bioinformatics2026-09-07v4EISCA and EISTA: Full-Spectrum Pipelines for Single-Cell and Spatial Transcriptomics Analysis
Wu, H.; Lister, A.; Macaulay, I. C.; Long, K.; Uauy, C.; Lan, Y.; Wickham, G. J.; Swarbreck, D.; Videm, P.; Stubbs, A.; Soranzo, N.; de Waard-van Baardwijk, M.; Nilchi, A. N.; Papatheodorou, I.Abstract
Single-cell and spatial transcriptomics are transforming our understanding of cellular heterogeneity and tissue organization, yet their analytical complexity remains a major bottleneck. Here, we present EISCA and EISTA, two standardized, end-to-end pipelines for single-cell RNA-seq and imaging-based spatial transcriptomics analysis. Built on the Nextflow nf-core framework, both pipelines implement modular, scalable, and reproducible workflows spanning primary, secondary, and tertiary analyses, from raw data processing to advanced downstream analyses. EISCA supports droplet- and plate-based scRNA-seq technologies, while EISTA is tailored for high-resolution spatial platforms including Vizgen MERFISH and 10x Xenium. Together, they integrate state-of-the-art methods for quality control, normalization, clustering, integration, cell-type annotation, differential expression, and cell-cell communication, with EISTA further enabling spatial statistical analyses. A central design principle is to balance standardization with flexibility: workflows can be executed end-to-end or modularly, enabling iterative, exploratory analyses with minimal overhead. Both pipelines deliver rapid preliminary results alongside an out-of-the-box report, facilitating immediate data assessment and accelerating downstream discovery. Case studies in plant immunity and human sepsis demonstrate that EISTA and EISCA reproducibly can be used to recover biologically meaningful insights. Collectively, these pipelines provide efficient, flexible, and scalable solutions for comprehensive single-cell and spatial transcriptomics analyses.
bioinformatics2026-09-07v1Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics
Nieuwoudt, M.; Reverenna, M.; Patel, D.; Catzel, R.; Houngue, I. H. J.; Daniel, J.; Eloff, K.; Santos, A.; Lopez Carranza, N.; Jenkins, T. P.; Van Goey, J.; Kalogeropoulos, K.Abstract
Mass spectrometry-based proteomics increasingly relies on machine learning, yet existing models are trained for defined supervised tasks such as peptide identification, de novo sequencing or fragment intensity prediction, limiting transfer across datasets, instruments and acquisition methods. Here we present InstaNovo-FM, a self-supervised foundation model for bottom-up proteomics trained to reconstruct masked regions of tandem mass spectra. We assemble a diverse training corpus spanning 1.47 billion MS/MS spectra and 184.6 million high-confidence annotations. We train an encoder-only transformer on the annotated tier using a physics-aware masked reconstruction objective. We demonstrate that the InstaNovo-FM embeddings encode fundamental experimental and biological properties, including fragmentation method, sequence properties and post-translational modifications, without requiring peptide labels. Furthermore, this foundation model directly enables diverse downstream applications, including de novo peptide sequencing, database-free identification and analytical run classification. InstaNovo-FM establishes a unified representation space for peptide fragmentation spectra, enabling robust transferability across the proteomics ecosystem.
bioinformatics2026-09-07v1Predicting Endometriosis Status and Menstrual Cycle Phase Using DNA Methylation
Nagasuri, A.; Khan, U.; Grosjean, P.; Siddharth, A.; Kosti, I.; Mortlock, S.; Houshdaran, S.; Rahmioglu, N.; Missmer, S. A.; Zondervan, K. T.; Montgomery, G.; Becker, C. M.; Rogers, P.; Irwin, J.; Oskotsky, T.; Lindquist, K.; Seaman, C.; Giudice, L. C.; Sirota, M.Abstract
Endometriosis is a chronic inflammatory disease associated with pelvic pain, infertility, and delayed diagnosis. Growing evidence suggests that altered DNA methylation contributes to disease development and could serve as a biomarker for disease. We developed a leakage-safe machine learning pipeline to classify endometriosis case-control status and menstrual cycle phase using genome-wide DNA methylation data from eutopic endometrial tissue. The dataset consisted of 984 samples profiled using the Illumina Infinium MethylationEPIC array, with measurements across approximately 759,000 CpG sites. Technical variation was corrected using SmartSVA batch correction. Ridge logistic regression models were trained using stratified 80/20 train-test splits, with regularization strength selected via stratified cross-validation. Feature selection approaches included ridge coefficient ranking, per-CpG t-tests, and univariate logistic regression with FDR correction. Model validity was evaluated using label-shuffling analyses. Menstrual cycle phase classification showed strong performance (mean cross-validation AUROC: 0.971, held-out test AUROC: 0.989), reflecting genome-wide hormonally driven methylation. Ridge regression produced lower but meaningful performance for endometriosis classification (mean cross-validation AUROC: 0.854, held-out test AUROC: 0.875). Ridge coefficient-based feature selection identified compact predictive CpG sets, supporting the hypothesis that endometriosis-associated methylation signal is distributed across many loci rather than a few highly predictive CpGs. Pathway enrichment analyses identified substantial enrichment for menstrual cycle phase but limited enrichment for disease status following FDR correction, consistent with a diffuse endometriosis-associated signal. These findings demonstrate that ridge regression can detect methylation patterns associated with both endometriosis and menstrual cycle phase, highlighting the importance of accounting for cycle-related epigenetic variation in endometrial DNA methylation studies.
bioinformatics2026-09-07v1Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. II. Antibody Properties, Lipid-RNA Interactions
Bhutada, P.; Goyal, N.; Lakhankiya, T. K.; Narahari, S. D.; Thangaraju, S. S. N.; Nayak, T. S.; Peng, Y.; Thota, G. S. A.; Thota, R. S. D.; Wang, Z.; Lee, K.; Sinitskiy, A.Abstract
Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their utility for real-world industrial research remains insufficiently characterized. Extending the analysis presented in the first paper of this series, we evaluated the same five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and the AI Scientist-v2 from Sakana AI) on two more projects of high practical importance for biopharmaceutical development: predicting antibody developability properties with the use of pretrained protein language model embeddings, and modeling non-covalent lipid-RNA interactions in lipid nanoparticles with all-atom molecular dynamics (MD) simulations. The AI frameworks again showed genuine strengths, including unprompted identification of subtle methodological issues, successful use of pretrained protein embeddings, and consistent reporting of p-values and confidence intervals often absent from the original papers. However, no framework approached the scope of the original studies, and severe failures and hallucinations were observed. Our results confirm and extend the conclusion of the first paper that real published research from pharmaceutical companies that we tried to reproduce proved to be considerably harder for current AI frameworks than standard benchmarks suggest.
bioinformatics2026-09-07v1WGCNA+: AI-powered WGCNA for Integration of Multi-Omics Data
Zito, A.; Escriba' Montagut, X.; Cano-Muniz, S.; Martinelli, A.; Akhmedov, M.; Kwee, I. W.Abstract
Background: Weighted Gene Co-expression Network Analysis (WGCNA) is a widely adopted systems biology method to discover gene modules and module-trait associations, mostly from transcriptomics. Designed for a single layer, it cannot jointly analyze multi-omics layers, a consequential limitation in modern biomedical research. WGCNA modules are often hard to interpret, requiring vast follow-up for contextualization. Moreover, no integrated framework exists to visualize condition-specific, cross-omics relationships at module or feature level. Results: To address these limitations, we developed WGCNA+, a novel R package extending WGCNA to multi-omics. WGCNA+ offers key innovations: (i) a unified multi-omics pipeline for per-layer network inference and cross-layer module enrichment; (ii) SVD-accelerated topological overlap matrix calculation that greatly reduces computation time; (iii) a consensus framework identifying modules reproducible across independent datasets/conditions; (iv) LASAGNA, a companion R package for phenotype-conditioned, multi-partite graph visualization of cross-omics relationships; (v) AI-powered annotation and infographics offering immediate biological insight. We tested WGCNA+ across public transcriptomics, proteomics, and miRNA datasets. WGCNA+ detects biologically meaningful modules, cross-omics feature and phenotype correlations, and provides AI-powered interpretation that accelerates research. Conclusions: WGCNA+ addresses existing gaps with a principled, efficient framework for co-expression network analysis across omics. It detects cross-omics regulatory modules and their phenotype association to support basic research, biomarker discovery and pathway analysis. It uniquely offers AI-assisted interpretation and infographics, aiding hypothesis generation. Complementing WGCNA+, LASAGNA is a phenotype-aware multi-partite visualization framework to explore cross-omics relationships. Altogether, these features make WGCNA+ an innovative, powerful tool for clinical and translational research. Availability and implementation: WGCNA+ and LASAGNA are implemented in R language for statistical computing, version[≥]3.5. WGCNA+ and LASAGNA are fully and freely available with no restrictions (https://github.com/bigomics/WGCNAplus; https://github.com/bigomics/lasagna)
bioinformatics2026-09-07v1A systematic evaluation of SIRT6 as transcriptomic biomarker of aging
Kashuk, E.; Tarakanova, A.; Malygina, A.; Kuzovkina, N.; Ponomareva, A.; Toiber, D.; Khrameeva, E.; Smirnov, D.Abstract
SIRT6 is a NAD+-dependent sirtuin that plays central roles in chromatin regulation, DNA repair, telomere maintenance, metabolic homeostasis, and inflammatory control. Although SIRT6 have long been implicated in aging because of its association with several hallmarks of aging, the available evidence remains largely context-dependent and mechanistic, limiting the interpretation of SIRT6 as a robust and evolutionarily conserved biomarker of aging. To comprehensively investigate the role of SIRT6 as an aging biomarker, we established SIRT6.db, a multi-species transcriptomic resource that integrates SIRT6-targeted perturbation experiments across diverse biological systems and organisms form all available SIRT6-related publications collected via a large-scale textual analysis of the SIRT6 literature using topic modeling. Based on this database, we identified both species-specific and evolutionarily conserved transcriptional and functional signatures associated with SIRT6 perturbation and established their relevance to hallmarks of aging. We showed that SIRT6 expression is generally stable during normal aging, but becomes dysregulated in Alzheimer's disease in cell type and stage-specific manner, highlighting the context dependence of its potential as a transcriptomic biomarker of aging.
bioinformatics2026-09-07v1Keloid transcriptomics reveal heterogeneity in fibroblast subtype enrichment, gene expression, and immune cell responses
Panzer, J. J.; Pan, M.; Nair, M.; Loveless, I. M.; Adrianto, I.; Huang, L.; Chitale, D.; Francescone, R.; Vendramini-Costa, D. B.; de Guzman Strong, C.; Levin, A. M.; Jones, L. R.Abstract
Keloid disease (KD) is a fibroproliferative skin disorder resulting from abnormal scar formation that causes pain, itching, and decreased quality of life. While multiple KD transcriptomic studies exist, the influence of cell type composition on bulk tissue gene expression is unknown. We characterized fibroblast subtype and immune cell enrichment using bulk RNA-Seq of head and neck keloid and matched adjacent normal skin tissue (MANST) from 14 patients (10 African American and 4 European American). Cell type enrichment was calculated by single sample gene set enrichment analysis. Linear mixed-effects models were employed for 1) differential cell type enrichment across tissue, 2) tissue type-specific associations between fibroblast subtypes and immune cells, and 3) differentially expressed genes (DEGs) across tissue. Validation was conducted in an independent cohort of 8 African Americans. Three fibroblast subtypes and 14 immune cell types were differentially enriched across tissue type. Further, 17 tissue type-specific fibroblast subtype-immune cell enrichment associations were identified, with 14 exhibiting decreased association in keloid tissue relative to MANST. After adjustment for cell type enrichment, MIR31HG and NR4A2 were significant DEGs with the largest positive and negative fold-changes, respectively. By considering cell type enrichment, underlying keloid tissue-specific cell type and gene expression associations were revealed.
bioinformatics2026-09-07v1PharmCast: rapid generation of three-dimensional pharmacophore fingerprints from two-dimensional structure without conformer generation
Muskal, S. M.; McGregor, M. J.Abstract
A three-dimensional pharmacophore fingerprint records the binding features a molecule can present. It is a description of a hand in search of a glove. Because it is defined by presented features instead of two-dimensional structure, it can identify pharmacophoric similarity between structurally distinct compounds and support scaffold hopping and the identification of structurally distinct compounds with comparable binding features. The descriptor has remained a niche tool because its cost is dominated by conformer generation. In the reference pipeline, generating 100 conformers requires 2.82 s of the 2.86 s needed to fingerprint one screening collection compound; the bit calculation requires 0.039 s. We therefore removed the conformational stage. PharmCast is a feedforward neural network that predicts all 10,549 bits of a PharmPrint ensemble fingerprint directly from a SMILES string. On the same machine, PharmCast generated pharmacophore fingerprints for two molecules and compared them in 0.584 ms, whereas the conventional conformer-based pipeline took 5.71 s. PharmCast version 10 was trained on 5,887,229 molecules drawn from a screening collection, activity-backed ChEMBL compounds from 142 to 1000 Da, and peptide loops excised from crystal structures. We evaluated 155,648 purchasable catalog compounds excluded from every training set, 139,700 activity-backed ChEMBL compounds not present in the version 10 training set, and 13,500 peptide loops reserved for testing. Median fingerprint error, Pearson r, and pairwise ranking accuracy were 0.008, 0.980, and 0.936 for screening collection chemistry; 0.016, 0.984, and 0.952 for loop peptides; and 0.027, 0.936, and 0.889 for activity-backed ChEMBL compounds. The reference calculation reproduces itself at an error of 0.006 and r of 0.995. Ensemble pharmacophore fingerprints can therefore be predicted from constitution alone at a cost suitable for screening the collection and optimization. Keywords: pharmacophore, fingerprint, scaffold hopping, surrogate model, virtual screening, applicability domain
bioinformatics2026-09-07v1