Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
The RdRp Thumb-1 Pocket is a Conserved Target for Broad-Spectrum Antiviral Development
Woods, V.; Umansky, T.; Russell, S. M.; Gallay, P.; Smith, D.; Haders, D.Abstract
RNA viruses cause human diseases ranging from mild colds to deadly pandemics. Direct-acting, broad-spectrum, non-nucleoside antivirals have been characterized as impossible to develop because allosteric binding sites are poorly conserved. The HCV NS5B RNA-dependent RNA polymerase (RdRp) Thumb-1 allosteric site and its interaction with the HCV NS5B {Lambda}1-loop governs an essential conformational change required for polymerase initiation. The only approved NS5B Thumb-1 inhibitor, beclabuvir, has been shown to be inactive against a broad panel of non-HCV viruses, including poliovirus, rhinovirus, coronavirus, coxsackievirus, influenzavirus, and HIV. A conserved, homologous allosteric site on RdRp that spans multiple viral families has not been reported. Here, we report GALILEO's discovery that the Thumb-1 pocket, its associated {Lambda}1-loop and their interaction are conserved across RNA viral families for the first time. The discovery is validated through comparative structural analysis of Protein Data Bank (PDB) deposited viral polymerases utilizing a method that allows independent, public validation by any researcher. We further demonstrate that beclabuvir's dependence on its indole C6 carbonyl to interact with the HCV-specific residue R503 restricts its activity to HCV. We validate the target discovery with MDL-001, which does not contain a C6 carbonyl substituent. MDL-001 directly blocks viral RNA synthesis in isolated replication complexes and selects for the canonical Thumb-1 resistance mutation P495S in HCV NS5B. MDL-001 demonstrates broad-spectrum in vitro inhibition of both HCV and SARS-CoV-2. Preclinical proof of concept and development of MDL-001 across HCV, HBV, HDV, influenza, SARS-CoV-2, and RSV have been previously reported. These findings establish RdRp Thumb-1 as a conserved allosteric pocket and a druggable target for broad-spectrum direct-acting antiviral development.
bioinformatics2026-09-18v7DeepSpaceDB 2.0: an interactive spatial transcriptomics database for large-scale Xenium data exploration
Honcharuk, V.; Takemoto, K.; Masalunga, M. C.; Zhao, H.; Diez, D.; Kawaoka, S.; Vandenbon, A.Abstract
The 10x Genomics Xenium platform enables high-resolution spatial transcriptomics at single-cell and subcellular scales, but effective reuse of public Xenium datasets is hindered by large data sizes and heterogeneous file formats. We previously developed DeepSpaceDB, a spatial transcriptomics database designed for interactive, in-depth analysis of tissues and tissue microenvironments. Here, we present a major expansion of DeepSpaceDB that integrates large-scale single-cell spatial transcriptomics data generated by the Xenium platform. In this update, we systematically collected 1,539 public Xenium datasets from multiple repositories and processed them through a robust, standardized pipeline that validates, repairs, and harmonizes heterogeneous inputs into a unified representation. To support efficient exploration of these data, we introduced a redesigned DeepSpaceDB interface and complementary Zarr-based storage formats optimized for gene-centric visualization and spatially localized queries, enabling sub-second response times for common interactive operations. The updated platform supports real-time visualization of spatial data and analysis of regions of interest directly in the web browser. Together, this expansion establishes DeepSpaceDB as a unified resource for single-cell spatial transcriptomics, substantially lowering the barrier to accessing, exploring, and reusing large-scale public Xenium datasets.
bioinformatics2026-09-18v2LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design
Waththe Liyanage, W. W.; Rigoni, D.; Bove, F.; Righelli, D.; Romano, S.; Visone, R.; Iorio, M. V.; Grassia, M.; Mangioni, G.; Lio, P.; Taccioli, C.Abstract
The discovery of novel small molecules is challenging because of the vastness of chemical space and the complexity of protein-ligand interactions, leading to low success rates and time-consuming workflows. Here, we present LLMsFold, a computational framework that combines Large Language Models (LLMs) and biophysical foundation tools to design and validate new small molecules targeting pathogenic proteins. The pipeline starts by identifying viable binding pockets on a target protein through geometry-based pocket detection. A 70-billion-parameter transformer model from the LlaMA family then generates candidate molecules as SMILES strings under prompt constraints that enforce drug-likeness. Each molecule is evaluated by Boltz-2, a diffusion-based model for protein-ligand co-folding that predicts bound 3D structure and binding affinity. Promising candidates are iteratively optimized through a reinforcement learning loop that prioritizes high predicted affinity and synthetic accessibility. We demonstrate the approach on two challenging targets: ACVR1 (Activin A Receptor Type 1), implicated in fibrodysplasia ossificans progressiva (FOP), and CD19, a surface antigen expressed on most B-cell lymphoma and leukemia cells. Top candidates show strong in silico binding predictions and favorable drug-like profiles. All code and models are made available to support reproducibility and further development.
bioinformatics2026-09-18v2WITHDRAWN: A Comprehensive Analysis of the Electrolytic Hydrogen Water Mechanism via a Feedforward Loop and its Functional Role in Intestinal Cells In Vitro
LI, J.Abstract
This manuscript has been withdrawn following a formal investigation by the Graduate School of Human Sciences, Waseda University
bioinformatics2026-09-18v2Incomplete references leave bulk deconvolution targets non-identifiable, but identification has a computable precision price
Jiang, H.; Gao, F.; Liu, P.; Wu, Y.; Jie, Y.; Li, Y.; Jiang, Y.Abstract
Background. Reference-based deconvolution estimates cell-type proportions from bulk profiles. Incomplete references compromise these estimates, yet many remedies return point estimates. We separate the operator provenance that reproduces an estimate from the information and precision that identify its target. Results. Two operator histories at one reduced reference give different estimators: 28.2% of 196,420 sample-deletion pairs differed by >0.1 total variation; locking every learned component made paths identical. Across nine cohorts, operator history reversed 160 of 952 association signs, 31 with a significant path, and changed significance for 126. Observationally equivalent completions can fill the open simplex and reverse retained-type rankings. Under a joint zero-exposure condition, shared structure leaves inherited bounds unchanged; one to six uncalibrated views also gave identical bounds. A profile library contracted estimator-output envelopes by 98.93% yet covered the full-reference effect for 42.77%, with 19 wrong-sign certificates; conditional sharp bounds stayed at [-1, 1]. Treating RNA yields as exact collapsed intervals to points covering none of seven flow-measured targets: contraction without coverage is false certainty. Calibrated cross-modal anchors contract width to 0.51 at six types but certify no valid sign. Exact RNA yields identify cell fractions from RNA contributions; a decisive sign in blood requires {+/-}2.6% proxy accuracy with near-total contraction of donor-heterogeneity and dynamic-range envelopes. Conclusions. Incomplete-reference deconvolution is an identification problem, not only an estimation problem. Remedies must be scored on shrinkage, coverage, certification and false certification against held-out targets. fitdrop implements this scoring and a precision frontier for planning decisive measurements.
bioinformatics2026-09-18v2OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome
Teixeira, A. A. R.; Zhu, H.; Kothiwal, D.; Cao, R.; Mills, A.Abstract
Choosing which region of a protein to express remains poorly standardized in antibody discovery, recombinant reagent generation, structural biology and computational binder design. For human cell-surface and secreted proteins, this requires reconciling topology, processing, predicted and experimental structure, modifications, interaction partners, orthologs, paralogs and cross-reactivity risk before ordering DNA. OpenAntigens is a free, no-login database of construct-design reports for 5328 human secreted, GPI-anchored, single-pass and multipass proteins. It integrates UniProt topology, AlphaFold pLDDT and PAE, PDB precedent, InterPro and Pfam domains, mouse and cynomolgus orthologs, paralog and family context, Open Targets disease associations, partner and assembly context, and BLAST searches. It provides 55 305 construct suggestions spanning full design regions, PDB-backed boundaries, annotated domains, pLDDT/PAE-derived regions and membrane-expression options, plus 148 722 sequence-similarity hits to help choose constructs and assess cross-reactivity. For targets with compatible AlphaFold models, the interactive designer links sequence, structure, pLDDT and PAE, allowing users to revise boundaries and export species-equivalent sequences with real-time cysteine and modification warnings. OpenAntigens places reproducible construct suggestions, comparative context and browser editing in one workflow, reducing manual reconciliation across resources. OpenAntigens is available at openantigens.org.
bioinformatics2026-09-18v2Novel two-stage deep learning-based approach applied to gene expression data pertaining to esophageal adenocarcinoma boosting biological knowledge discovery
Jamie, F.; Turki, T.; Alsolami, F.; Taguchi, Y.-h.Abstract
Esophageal cancer (EC) is characterized by complex transcriptional alterations and therapeutic resistance, posing challenges for traditional computational methods. In this study, we propose a deep learning (DL)-based computational framework to identify important genes and biologically relevant pathways in bulk cell RNA-seq data (GSE234304 and GSE273848), which comprise tumor and non-tumor esophageal tissue samples. A fully connected feedforward neural network was trained for binary classification, and two feature selection strategies were implemented: Neural Network followed by Support Vector Regression (NN+SVR) and Integrated Gradients combined with SVR (IG+SVR). The genes were then ranked according to their weights in SVR deriving the importance scores, and the top 100 genes were subjected to enrichment analysis using Enrichr and Metascape. The proposed DL-based approaches identified a greater number of expressed genes across established esophageal cancer cell lines than LIMMA, SAM, and the t-test did. Specifically, in the GSE234304 dataset, IG + SVR, our best method, identified a total of 9 expressed genes while the best baseline method, LIMMA, identified a total of 3 expressed genes. In terms of GSE273848 dataset, IG + SVR was also the best identifying a total of 11 expressed genes while the best baseline method, t-test, had a total of 7 expressed genes. The key genes identified included CEBPB, SUMO1, RORA, STAT1, GATA, OCT1, RUNX1, and NR3C1, as well as pathways related to nucleoprotein maturation, collagen fibril organization, the immunoglobulin-mediated immune response, immune regulation, and insulin signaling. These results show that combining neural networks and attribution-based regression creates an effective and interpretable framework for selecting genes in esophageal cancer research.
bioinformatics2026-09-18v1Pep-PU-GAN: Positive-Unlabeled Adversarial Learning for Peptide Function Prediction
Midjani, F.; Hashemi, S.; Keshtkar, F. Z.; Malekpour, M.; Saberzadeh Ardestani, B.; Khosravi, B.Abstract
Peptide classification remains challenging in bioinformatics because of limited labeled data, particularly the scarcity of verified negative examples, and the complex relationship between amino acid sequences and biological functions. This study introduces Pep-PU-GAN, a deep learning framework that combines positive-unlabeled (PU) learning, generative adversarial networks (GANs), and graph neural networks (GNNs) for peptide classification. Peptides are represented as sequence-derived residue graphs, with amino acids as nodes and edges connecting adjacent residues, enabling attention-based message passing over local neighborhoods. The architecture includes a generator that produces synthetic peptide embeddings in encoder space and a dual-function discriminator that distinguishes real from synthetic embeddings while performing PU classification. Training uses a custom loss integrating non-negative PU (nnPU) risk estimation with adversarial objectives. A self-training mechanism further incorporates high-confidence synthetic positive embeddings to augment the training set and improve performance. Evaluated on neuropeptide classification using 4,049 positive neuropeptides and 8,558 unlabeled peptides, Pep-PU-GAN outperformed baseline models, achieving an F1 score of 0.93 and an AUROC of 0.98 on an independent held-out benchmark. Pep-PU-GAN provides a promising approach for peptide classification tasks with scarce labeled and abundant unlabeled data, with potential applications in computational biology and drug discovery.
bioinformatics2026-09-18v1TRACEDD: A Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design
Vangala, S. R.; Kasturi, V. V.; Bung, N.; Roy, A.Abstract
Drug discovery depends on coordinated decisions across target validation, structure analysis, molecular design, developability assessment and synthetic feasibility, but current computational methods often operate as disconnected tools. Here, we introduce TRACEDD (Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design), a framework that makes three primary contributions: (1) It establishes a 'tool-first' multi agentic architecture where LLMs orchestrate validated computational tools rather than replace them, ensuring scientific rigor. (2) It implements a multi-agent system that mirrors expert discovery teams, enabling transparent and traceable decision-making through a Reason-Act-Observe loop. (3) It demonstrates an end-to-end workflow, from target validation to synthesis planning, that adaptively handles real-world data variability, such as the absence of experimental structures. The framework decomposes discovery into specialized agents for target validation, druggability assessment, molecular generation, lead optimization, ADMET evaluation, literature evidence integration and retrosynthesis, all operating through a Reason Act Observe workflow. Using JAK2 as a representative case, we show that the system can retrieve experimental protein structures, invoke AlphaFold when structures are unavailable, identify druggable pockets and perform de novo molecular generation. Known JAK2 inhibitors are used to define design hypotheses and guide reinforcement learning-based molecular generation, with docking scores/predicted pIC50 and other physicochemical/ADMET properties serving as reward and prioritization signals. The framework demonstrates a tool-first, reasoning-driven approach in which each major decision is linked to explicit tool invocation, intermediate evidence. By combining agentic orchestration with domain-specific computational tools, the system supports transparent, adaptable and human-verifiable molecular design workflows, providing a foundation for more reliable AI-assisted drug discovery.
bioinformatics2026-09-18v1Sparse Machine Learning Pipeline with Stabl Identifies Cord Blood Multi-Omic Signatures of Bronchopulmonary Dysplasia
Mestan, K.; Newar, J.; Zhao, J.; Chakraborty, A.; Reiss, J.; Funk, W.; Stelzer, I.; Waked, B.; Bellan, G.; Durand, X.; Hedou, J.Abstract
Background: Several omics studies have been completed in recent years, with the goal of identifying biomarkers of complex multifactorial diseases, such as bronchopulmonary dysplasia (BPD). Objective: To evaluate the performance of 3 distinct omics platforms, using a machine learning pipeline with integration of sparse, reliable and adaptive biomarker identification (Stabl). Methods: Using a well-characterized birth cohort, cord blood metabolomics, proteomics and adductomics data were integrated with Least Absolute Shrinkage and Selection Operator (LASSO) regression and Stabl, to evaluate predictive performance for BPD. Results: Sparse multivariable modeling of 45,000 features measured in 217 infants (52 term, 165 extremely preterm <28 weeks; 82 with BPD and 35 with severe BPD/death) identified a perfect signature for preterm birth with both LASSO and Stabl (AUROC=1.0; p<0.001). Analysis of the preterm group yielded excellent predictive power for severe BPD (AUROC=0.83; p=0.005). Stabl identified a set of 12 biomarkers (2 adducts, 3 proteins and 7 metabolites) with good performance for predicting grade III BPD (AUROC=0.76; P=0.03). Biomarkers across the 3 omics platforms revealed dysregulated pathways of innate/adaptive immune responses, metabolic programming and oxidative stress. Conclusions: The sparse machine learning pipeline is a complementary approach for identifying novel pathways and biomarkers of multifactorial BPD and its endotypes.
bioinformatics2026-09-18v1Setting the SCENE for Interpretable Cell-Gene Embeddings in Single-Cell RNA-seq
Moberg, O. L.; Petersen, M. B.; Herlau, T.; Kristensen, L. E.; Jessen, L. E.; Morup, M.Abstract
Single-cell RNA sequencing measures cellular states at high resolution, but sparse high-dimensional count data remain difficult to model interpretably. We introduce the Single-Cell Euclidean Network Embedding (SCENE), a probabilistic latent-distance model that jointly embeds cells and genes from Unique Molecular Identifier (UMI) counts. SCENE treats the count matrix as a weighted bipartite cell-gene graph, where Euclidean distances represent transcriptional affinity, and combines this geometry with a zero-inflated count likelihood that separates gene detection from expression magnitude. Across real and simulated scRNA-seq datasets, SCENE recovers biologically structured cell and gene embeddings with state-of-the-art performance. Surprisingly, major biological structure is preserved in native two- and three-dimensional latent spaces, enabling directly interpretable visualization. Perturbation analyses show that SCENE organizes glucocorticoid-response genes and T-cell receptor regulatory programs coherently in gene space, capturing biology beyond cell-type separation. SCENE provides a transparent representation learning framework in which low-dimensional Euclidean geometry supports accurate modeling and biological interpretation.
bioinformatics2026-09-18v1Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics
Chen, T.; Hicks, S. C.Abstract
Models that combine tissue images with pretrained gene representations aim to predict spatial expression for genes not used to fit the downstream predictor. Yet success on held-out genes can reflect two capabilities: estimating a gene's mean expression across tissue locations and recovering its spatial variation. Across four cohorts spanning three human brain regions and HER2-positive breast cancer, we evaluated held-out genes in held-out individuals and separated these components. For spatial predictors using fixed gene representations from Decima or scGPT, reductions in gene-mean error accounted for more than 91% of the reduction in mean squared error relative to matched random vectors. Independently fitted mean-only models using the same representations but no tissue images retained 90-99% of the corresponding gain in full-matrix correlation. Spatial gains were smaller on average, increased with expression variation in training tissue and differed across cohorts and representations. Across these settings, pretrained gene representations broadly transferred mean expression but selectively improved spatial recovery, showing that cross-gene generalization in virtual spatial transcriptomics is not a single capability.
bioinformatics2026-09-18v1OpenLipid: a large language model workflow for targeted analysis of DIA mass spectrometry data in lipidomics
Li, J.; Rost, H.Abstract
Mass spectrometry has become a central technology for lipidomics, with data-independent acquisition (DIA) enabling broad and reproducible sampling of lipid signals. However, the multiplexed fragment-ion spectra in DIA data complicate lipid identification. Here, we introduce OpenLipid, a large language model (LLM)-based workflow for targeted DIA lipidomics. Using assay libraries built from data-dependent acquisition (DDA) results, OpenLipid directly evaluates extracted ion chromatograms (XICs) from DIA data in a zero-shot setting to identify target lipid peaks and generate human-readable rationales for individual lipid identification decisions. We benchmarked OpenLipid against manual annotations across four datasets comprising human plasma and mouse feces analyzed in positive and negative ionization modes. The plasma assay libraries contained 199 target lipids in positive mode and 147 in negative mode. The fecal assay libraries contained 181 target lipids in positive mode and 264 in negative mode. At a 5% false discovery rate (FDR) threshold, OpenLipid identified 110 (55.3%) and 28 (19.0%) library targets in plasma and 130 (71.8%) and 84 (31.8%) in feces, in positive and negative ionization modes, respectively. OpenLipid achieved an overall identification rate comparable to that of DIAMetAlyzer (57.8%, 19.7%, 89.0%, and 17.8% across the corresponding datasets) and substantially higher than that of untargeted MS-DIAL DIA analysis (12.6%, 0.0%, 47.5%, and 0.0%) on the same assay-library targets. LLM-derived chromatographic features also enabled supervised discrimination between correct and incorrect candidate peak groups for target lipids, with median cross-validation average precision values of 0.838-0.912. Together, these results demonstrate that OpenLipid is an effective LLM-based workflow for FDR-controlled targeted analysis of DIA lipidomics data.
bioinformatics2026-09-18v1Mapping Gene Expression to an Interpretable Semantic Space
Duan, X.; Aggarwal, M.; Periwal, V.Abstract
Cell embeddings organize single-cell expression data, but their dimensions have no biological meaning, so clusters are interpreted afterward. We present MESIC (Mapping Expression to Semantic space with Interpretable Components), which builds the written knowledge about genes held in curated databases into the dimensions themselves. A biomedical language model converts each gene's summary into a semantic embedding. MESIC compresses these embeddings into a small number of components, each concentrated on a small set of genes and explained by their annotations. The components are computed once from the summaries, so any expression dataset can be mapped onto them, and every cluster, outlier, or cell-type assignment is then characterized by named genes. In cardiomyocytes, outliers in the component space were enriched for hypertrophic cardiomyopathy. In a lung atlas, unsupervised clusters in that space matched the broad cell types that experts had annotated. In both, the components that separated the cells matched their known biology. For about half of the cells that the atlas itself had left unannotated, the same space gave a confident cluster assignment, and with it an interpretation through component-associated genes. Gene summaries thus give single-cell analysis a coordinate system in which every result is traced to genes and what is written about them.
bioinformatics2026-09-18v1A Unified 3D Generative Model for Synthesizable Structure-Based Drug Design
Igashov, I.; Schneuing, A.; Dobbelstein, A. W.; Morozova, I.; Neeser, R. M.; Zielinski, K.; Abriata, L. A.; Petruzzella, A. S.; Pavel Iosub, D. R.; Gampp, O.; Lyubimov, A. Y.; Elizarova, E.; Ferrara, I.; Sousa, P. M. F.; Lemos, A. R.; Testori, F.; Miranda Herrera, P. A.; Kanis, L.; Schmidt, J.; Braza, M. K. E.; Amaro, R. E.; Thoma, N.; Ferraris, D. M.; Riek, R.; Fraser, J. S.; Schwaller, P.; Bronstein, M.; Correia, B.Abstract
Traditional screening-based drug discovery is inherently limited by the astronomical scale of the chemical space. Generative modelling offers a compelling alternative to the classical search paradigm and enables rational, bottom-up design of novel and target-specific small molecules. However, its impact has been hampered by challenges in synthetic accessibility of the designed compounds and lack of large-scale experimental validation. Here, we introduce LDDM (Large Drug Discovery Model), a generative framework that supports a range of drug discovery tasks, including constrained and unconstrained docking, fragment linking and growing, and de novo design. We further introduce a programmable design algorithm that enables accurate design of synthetically accessible compounds satisfying various fine-grained objectives. We experimentally validated the designed or optimised ligands for five therapeutically relevant protein targets. In all cases, LDDM achieved high success rates, allowing us to identify molecules with confirmed binding affinity while synthesizing only a small number of generated compounds. The best designs were structurally characterised through NMR spectroscopy and X-ray crystallography, demonstrating high prediction accuracy. Overall, LDDM provides a scalable and flexible platform for the rapid and tailored design of small molecules and non-natural peptides for therapeutic applications.
bioinformatics2026-09-18v1Physical priors improve performance of structure-based binding affinity models
Kaminow, B.; Payne, A. M.; MacDermott-Opeskin, H. I.; Chodera, J. D.; Singh, S.Abstract
Structure-based drug discovery is a widely used paradigm for the rational design of novel small molecule therapeutics. However, the benefits conferred by the use of structural information has seen limited adoption in machine learning, where ligand-only ("2D") models are still the industry standard for molecular property or binding affinity prediction. Structure-based ("3D") ML models for binding-affinity prediction promise to present a clear advantage, but have not yet overtaken existing 2D models. Here, we show that physics-based priors can improve predictive performance of structure-based models by comparing different model architectures with varying physical priors on several prediction tasks. We present the Modular Training and Evaluation of Neural Networks (mtenn) package, where we decompose affinity prediction into separate steps of embedding structure into learned representations and combining those embeddings into a predicted binding affinity. We consider both E(3)-invariant and E(3)-equivariant architectures to determine the importance of encoding roto-translational inductive biases, as well as different methods for combining learned embeddings. By first optimizing several aspects of model construction using the general purpose PDBBind dataset, we are able to improve the performance and data efficiency of structure-based models. When subsequently trained and evaluated on the COVID Moonshot small molecule drug discovery dataset, our tuned models perform on par with industry standard ligand-only models. Our decomposed model framework highlights that encoding some physical priors improves model performance, while more complex biases such as equivariance offer limited benefit. Additionally, structure-based models generalize better to an unseen target and display higher training efficiency. Overall, these results emphasize that structure-based models benefit from their ability to incorporate physics-informed constraints, giving promising directions for model architecture development. These results also suggest that the strength of these models may be in tasks specifically aimed at generalizability, providing guidelines for their use in early-stage drug discovery campaigns.
bioinformatics2026-09-18v1Rapid Shift Toward Pulsed Field Ablation and Precision Risk Stratification in High-Impact Atrial Fibrillation Research
Su, Z.; Li, T.Abstract
Conventional bibliometrics rely on lifetime citations, obscuring immediate shifts in cardiovascular research paradigms. To track emerging trends in atrial fibrillation management, we performed a comparative bibliometric analysis of the 50 highest-cited original research articles per year in OpenAlex topic T10065 across consecutive 2023 (Class of 2025) and 2024 (Class of 2026) publication cohorts. Articles were ranked using a fixed 18-month post-publication citation window, and extracted concepts were normalized into 719 canonical topics and 32 parent themes using large language model curation. Concept frequency tracking demonstrated a swift technological shift, with pulsed field ablation showing the largest topic frequency increase (+0.08, from 0.30 to 0.38) to become the leading canonical topic in 2024, displacing conventional thermal pulmonary vein isolation (-0.18, 0.40 to 0.22). Simultaneously, stroke prevention focus shifted toward refined predictive modeling, with increases in Risk Stratification and Predictive Models (+0.08, 0.30 to 0.38) and CHA2DS2-VASc scoring (+0.08, from 0.10 to 0.18). High-impact atrial fibrillation research is rapidly pivoting toward non-thermal ablation safety profiling and precision risk stratification, highlighting the utility of fixed-window concept mining for capturing real-time scientific evolution. Online explorer of the result is available at https://pri.pepkio.com.
bioinformatics2026-09-18v1Comparative Mitogenomics and Molecular Phylogeny of Agriculturally Significant Tephritid Fruit Fly Pests in Bangladesh
Rahman, S.; Shormi, F. A.Abstract
Tephritid fruit flies of the tribe Dacini rank among the most economically destructive agricultural pests globally, with several Bactrocera Macquart and Zeugodacus Hendel species causing severe losses to fruit and vegetable production across South and Southeast Asia. Bangladesh harbors five dacine species of primary agricultural significance: Bactrocera dorsalis (Hendel), B. carambolae Drew & Hancock, B. zonata (Saunders), B. correcta (Bezzi), and Zeugodacus cucurbitae (Coquillett). Here we present a comparative mitogenomic and molecular phylogenetic framework based on 21 unique mitogenome records retrieved from NCBI GenBank, representing 19 dacine ingroup taxa and two outgroups. Direct parsing of the GenBank sequences confirmed the canonical complement of 13 protein-coding genes (PCGs) and two rRNA genes in all 21 records. Among the 14 Bactrocera ingroup taxa, genome size ranged from 15,273 to 15,977 bp. Whole-genome AT content ranged from 66.6% in Bactrocera tsuneonis to 82.2% in Drosophila melanogaster; the five Bangladesh-relevant dacine pest species showed tightly clustered AT content of 72.9-73.6%. A concatenated alignment of 13,600 nucleotide sites (13 PCGs + 12S + 16S rRNA) was analyzed by maximum-likelihood inference under the GTR+FO model (IQ-TREE v1.6.11; 1,000 ultrafast bootstrap replicates). The ML tree recovers the B. dorsalis complex (UFBoot = 79-100) and places B. correcta and B. zonata as a maximally supported sister pair (UFBoot = 100). This study provides a sequence-verified mitogenomic reference framework for molecular identification and pest surveillance in Bangladesh, and should be interpreted as a curated comparative baseline rather than a population-genomic analysis, as no newly collected Bangladeshi specimens were sequenced.
bioinformatics2026-09-18v1One-dimensional CNNs for Near-Infrared Prediction of Protein and Moisture in Cereal Grains: The Effects of Architecture and Input Preparation
Deng, G.; Chi, W.; Yu, P.; WU, F.Abstract
Near-infrared (NIR) spectroscopy is widely used for the rapid, non-destructive determination of constituents such as protein and moisture in cereal grains, but model comparisons in this field are often confounded by differences in the inputs received by each model. We benchmark a compact one-dimensional CNN derived by one-factor-at-a-time ablation (CNN-Baseline) and a randomly searched CNN (CNN-RS) against PLSR, SVR, XGBoost on six cereal-grain datasets (n=500-5046) under three input conditions: raw spectra, optimal preprocessing, and preprocessing plus wavelength selection. A single-filter with a kernel size of 11 of convolution and batch normalization proved sufficient. Given a unified raw input, CNN-RS was most accurate on five of the six datasets, while CNN-Baseline came within 0.003-0.017 in test R2 on the four larger datasets. PLSR led only on the smallest protein dataset (XDS, n=500) and was the most stable model. Preprocessing had almost no effect on CNN accuracy on the four larger datasets, but improved accuracy by 0.034-0.047 (CNN-Baseline) and 0.017-0.055 (CNN-RS) on the two smallest, and brought the networks to their optimum in a median of 49% fewer epochs. Therefore, replacing explicit preprocessing with learned filters was only justified when sufficient training data were available. Once each model received its own preprocessing and wavelength subset, SVR became most accurate on four of six datasets and XGBoost rose from the weakest model to within 0.01-0.03 of the best model on four of the six datasets, suggesting that the principal advantage of CNNs lied in robustness to raw input data rather than in attainable accuracy. Overall, model selection for NIR calibration of cereal grains depends more on sample size and the input preparation than on network architecture.
bioinformatics2026-09-18v1Virtual experiments bridge sequence and microscopy with generative models
Zheng, D.; Hong, K.; Huang, B.Abstract
Large-scale screening and mapping efforts have produced vast libraries of perturbation-readout data. Converting these measurements into mechanistic insights requires models that link perturbation and genetic input to phenotypes, i.e., labels from experimental readouts, which are usually task specific. We propose a different, virtual experiment modeling approach: train generative models to recreate readouts conditioned on the experimental context, and then let established downstream models extract phenotypes from the synthetic data. As an illustrative case, we develop a bidirectional sequence-image generative framework, CELL-FM, that maps protein sequence and cellular context to fluorescence microscopy images and back, enabling in silico localization prediction, image-conditioned functional motif analysis and generation, and large-scale virtual mutagenesis revealing the amino acid features controlling condensate formation of intrinsically disordered peptides. This approach decouples representation learning from task-specific annotation, reuses rich experimental modalities across many downstream tasks, and preserves the spatial and organizational detail that hand-crafted labels often discard.
bioinformatics2026-09-18v1Facilitating genome annotation using ANNEXA and long-read RNA sequencing
Hoffmann, N.; Besson, A.; Cadieu, E.; Lorthiois, M.; Le Bars, V.; Houel, A.; Hitte, C.; Andre, C.; Hedan, B.; Derrien, T.Abstract
With the advent of complete genome assemblies, genome annotation has become essential for the functional interpretation of genomic data. Long-read RNA sequencing (LR-RNAseq) technologies have significantly improved transcriptome annotation by enabling full-length transcript reconstruction for both coding and non-coding RNAs. However, challenges such as transcript fragmentation and incomplete isoform representation persist, highlighting the need for robust quality control (QC) strategies. This study presents ANNEXA, a pipeline designed to enhance genome annotation using LR-RNAseq data while also providing QC for reconstructed genes and transcripts. ANNEXA integrates two transcriptome reconstruction tools, StringTie2 and Bambu, applying stringent filtering criteria to improve annotation accuracy. It also incorporates deep learning models to evaluate transcription start sites (TSSs) and employs the tool FEELnc for the systematic annotation of long non-coding RNAs (lncRNAs). Additionally, the pipeline offers intuitive visualisations for comparative analyses of coding and non-coding repertoires. Benchmarking against multiple reference annotations revealed distinct patterns of sensitivity and precision for both known and novel genes and transcripts and mRNAs and lncRNAs. To demonstrate its utility, ANNEXA was applied in a comparative oncology study involving LR-RNAseq of two human and eight canine cancer cell lines. The pipeline successfully identified novel genes and transcripts across species, expanding the catalog of protein-coding and lncRNA annotations in both species. Implemented in Nextflow for scalability and reproducibility, ANNEXA is available as an open-source tool: https://github.com/IGDRion/ANNEXA.
bioinformatics2026-09-17v5Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme
Adamov, A.; Mueller, C. L.; Bokulich, N.Abstract
Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured, and predictive modeling from them typically relies on ad hoc feature representation choices that obscure their impact on performance and interpretation. We present ritme, an open-source Python package that jointly optimizes microbiome-specific feature representation and model selection - combined algorithm selection and hyperparameter optimization - tailored to these data. ritme systematically searches taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment together with model class and hyperparameters, using state-of-the-art optimizers that scale from a laptop to a compute cluster. Across three real-world use cases, ritme outperformed the original study pipelines by 7-29% on the primary task metric and surpassed three AutoML baselines in six of seven comparisons, while selecting substantially fewer features and exposing how feature and model choices drive performance. Open-source and modular, ritme supports reproducible, parsimonious predictive modeling, downstream biological investigation, and extension to other multi-omics modalities.
bioinformatics2026-09-17v4Locat: Joint enrichment and depletion testing identifies localized marker genes in single-cell transcriptomics
Lewis, W. R.; Aizenbud, Y.; Strino, F.; Kluger, Y.; Parisi, F.Abstract
Several methods identify marker genes that delineate cell populations in single-cell transcriptomic data, yet most emphasize enrichment within candidate populations without testing whether expression is significantly reduced elsewhere. We present Locat, a framework for identifying highly specific localized genes by testing whether expression is concentrated within compact regions of a cellular embedding and depleted outside them. For each gene, Locat fits weighted Gaussian mixture models to gene-specific and background densities, computes concentration and depletion statistics, and integrates them into a unified localization score. Across synthetic benchmarks with controlled ground truth, Locat detects uni-modal, multi-modal, and sparse localized patterns and loses significance when expression becomes indistinguishable from background structure. In developmental, perturbation, and differentiation datasets, Locat identifies compact marker sets that capture lineage organization, condition-specific programs, and temporal dynamics. These sets are often smaller than highly variable gene selections, while embeddings built from them preserve major cell populations and developmental programs in several cases. In murine dermis, interferon-treated PBMCs, and retinoic acid-induced embryonic stem cell differentiation, localized genes recover differentiation trajectories, stimulus-responsive programs, and reproducible stage-specific patterns. Together, these results show that jointly assessing concentration and depletion yields specific, interpretable marker genes.
bioinformatics2026-09-17v3FORGE audits residue-level information encoded in RNA tertiary structure geometry
Gow, L.; Li, J.; Tan, X.; Liang, K.; Gui, N.; Luo, B.Abstract
Coarse RNA coordinate representations are widely used, yet the biological information they encode remains unquantified. We introduce FORGE, which converts a seven-atom RNA geometry representation into 935 interpretable descriptors and reports which residue-level annotations this geometry supports. On 4,135 post-2025 RNA chains, FORGE recovered 64.6% of native nucleotides; a six-atom control lacking the glycosidic nitrogen retained 58.5%, locating most of this signal in phosphate-sugar geometry. Confidence was sharply graded: abstaining from the least-confident half of positions raised accuracy to 94.4%, yet many chains remained only partially identifiable. The same descriptors predicted base-pair state far better than a DMS-like proxy or protein-proximal context. Native-decoy, OpenKnot and solved-pseudoknot analyses showed that nucleotide identifiability, foldability and experimental design score are separable: AlphaFold3 reproduced the experimental fold for one of four AI-designed constructs and none of the sequences FORGE read from their geometry. FORGE provides a reproducible audit layer for RNA structural interpretation.
bioinformatics2026-09-17v3Unlocking Sensitive Data with SPHERE in the Age of AI
He, Z.; Park, J.; Pulgrossi, R. C.; Lee, J.; Butler, R. R.; Weber, A.; Tian, L.; Zhang, X.; Wang, J.; Sha, S.; Mormino, E. C.; Wyss-Coray, T.; Henderson, V. W.; Longo, F. M.; Zou, J.; Desai, M.; Altman, R.Abstract
Sensitive human data underpin discoveries across medicine, biology and the social sciences, yet privacy regulation often prevents sharing them with collaborators or artificial intelligence (AI) systems. We introduce SPHERE, a model-free method that makes sensitive datasets directly usable by AI and shareable for open science as a synthetic twin, while the original records never leave the local environment. Across 33 datasets spanning five scientific domains, SPHERE protects individual privacy against adversarial re-identification attacks while preserving the data's statistical structure: means, variances and correlations are reproduced exactly, effect size and P value in linear statistical analysis is numerically identical, nonlinear machine-learning utility is retained, and each twin is generated in seconds on a laptop. Frontier AI agents running on the twin reach the same scientific conclusions as on the original records. Analyses of the twin reproduce genome- and proteome-wide results at UK Biobank scale and recover the findings of landmark studies across three independent cohorts and consortia. The approach also extends to deep-learning embeddings across language, vision and time-series, with minimal utility loss. We make the Stanford Alzheimer's Disease Research Center cohort openly available for the first time, as a SPHERE twin spanning nine modalities that any registered researcher can analyze without an approval process. We release SPHERE with certification of each twin's privacy and fidelity, and an AI agent that autonomously executes research tasks on sensitive data without ever accessing it. Sensitive datasets that are currently closed to research could thus become routine inputs to open science and AI to enable key discoveries.
bioinformatics2026-09-17v2Lacuna: Cryptic Binding Pocket Discovery via Conformational Ensemble Analysis
Moore, C.Abstract
Lacuna, an open-source Python tool for discovering cryptic binding pockets: sites that are absent or too small to detect in a protein's unbound structure and open only during conformational fluctuation. Most binding-site predictors score a single static structure, which is precisely the structure in which a cryptic site is invisible. Lacuna instead generates a conformational ensemble from any input structure, detects pockets independently in every conformer, clusters the detections into persistent sites across the ensemble, and ranks those sites with a model fitted on within-structure pairs. Ensemble generation is pluggable: normal mode analysis by default, with implicit-solvent molecular dynamics, Boltz-2 diffusion sampling, or a user-supplied ensemble as alternatives. On the designated test fold of CryptoBench, Lacuna recovers 55.6% of cryptic sites in its top five predictions with the zero-dependency default and 66.1% with an optional PLM-assisted ranker; pooling the geometric detector with an optional learned surface detector recovers 73.9% while raising the fraction of sites found from 68.5% to 86.4%, measured on the held-out fold at five conformers. It recovers 73%, 45% and 87% on the PocketMiner set, a curated set of literature apo/holo pairs, and COACH420 respectively. The default backend completes in a median of 2.6 seconds per chain on one CPU core, so ensemble-based pocket finding does not require a simulation budget. Every site carries a continuous crypticity score, and outputs are emitted as docking-ready Boltz YAML constraints, AutoDock Vina boxes, pseudoatom PDB files, and the generated conformational ensemble as a multi-model PDB. Lacuna is MIT licensed and available at https://github.com/mooreneural/lacuna and on PyPI as lacuna-pockets.
bioinformatics2026-09-17v2Tractography from Serial Optical Coherence Tomography: How and Why?
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.Abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles, invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Due to its high resolution and its 3D nature, serial optical coherence tomography (S-OCT) offers promise for studying WM at the microscale. However, whether the reflectivity contrast from S-OCT supports tractography at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that multiscale Frangi filters outperforms structure tensor analysis for estimating ODF. We also show that anatomically-constrained particle filtering tractography enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing thalamocortical WM projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles that are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
bioinformatics2026-09-17v2iDriver: A patient-centric framework for genome-wide cancer driver discovery
Bahari, F.; Montazeri, H.Abstract
Tumor genomes harbor a mixture of neutral and positively selected mutations, yet distinguishing true cancer drivers remains a major challenge. Several factors can obscure the detection of selection signals, among which patient-specific variation in mutational burden plays a significant role. Current approaches often fail to account for the heterogeneity in mutation burden across different patients; in particular, no existing method explicitly accounts for it when integrating both mutation recurrence and functional impact. Here we present iDriver, a probabilistic graphical model that integrates both mutation recurrence and functional impact at the individual-patient level, enabling an enhanced estimation of positive selection across functional genomic elements. Applying iDriver to 29 cancer types, we identify both known and previously unrecognized drivers spanning coding and noncoding regions, and provide evidence for their clinical and biological relevance. In comprehensive benchmarks against 12 established driver discovery methods, iDriver consistently outperformed all competitors, achieving the highest rankings for known cancer drivers across both coding and noncoding elements.
bioinformatics2026-09-17v2LRP2: A proteogenomics pipeline for long-read informed protein isoform analysis and discovery
Schertzer, M. D.; Lewandowski, J. T.; Watts, E. F.; Rosenow, W.; Mehlferber, M. M.; Jeffery, E. D.; Adamson, S. I.; Bruand, J.; Tseng, E.; Neelamraju, Y.; Garrett-Bakelman, F. E.; Dolzhenko, E.; Knowles, D. A.; Sheynkman, G.Abstract
Most human genes produce multiple RNA isoforms, yet it remains unclear which isoforms are translated into stable, functional proteins. Long-read RNA sequencing resolves full-length transcript structures and, when paired with mass spectrometry, can provide empirical evidence of isoform translation. Despite this opportunity, comprehensive workflows integrating isoform discovery, open reading frame prediction, peptide identification, and protein inference remain limited, leaving users to handle these steps piecemeal. Here, we present LRP2, a modular, end-to-end long-read proteogenomics pipeline built in Nextflow. LRP2 scales transcript discovery to hundreds of samples via PacBio's latest Isocall tool, removes technical artifacts with SQANTI QC, generates and classifies predicted proteomes via CPAT and SQANTI Protein, performs multi-group differential expression and usage analysis via edgeR, DRIMSeq, and a long-read adaptation of LeafCutter, and integrates protein-level evidence from DDA and DIA MS data through FragPipe. For cross-dataset comparison of novel isoforms, LRP2 employs deterministic splice-junction, coordinate-based isoform identifiers. Used as an integrated pipeline, LRP2 enables the detection of novel peptides and improves the protein isoform inference to confirm protein isoform translation.
bioinformatics2026-09-17v2Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models
De Santis, F.; Malpetti, D.; Gualdi, F.; Mangili, F.Abstract
Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)
bioinformatics2026-09-17v1Targeting Outer Membrane β-barrel Proteins of Burkholderia mallei to design a Multi-epitope Vaccine Against Human Glanders
Kapoor, J.; Panda, A.; Kumar, S.; Bandyopadhyay, A.Abstract
Burkholderia mallei, a facultative intracellular Gram-negative pathogen, is the causative agent of glanders that primarily affects solipeds and is sporadically transmitted to humans. Current interventions mainly rely on antibiotics; however, increasing antimicrobial resistance and the lack of a licensed vaccine further complicate disease management. Surface exposed outer membrane {beta}-barrel (OMBB) proteins serve as excellent targets for vaccine development. In the present study, a consensus-based computational framework was employed on the B. mallei turkey2 proteome that identified 59 OMBB proteins - including porins, TonB receptors, autotransporters, and efflux components. These OMBB proteins were leveraged to predict B- and T-cell epitopes which were manually curated, and mapped onto the corresponding protein models to identify surface-exposed epitopes with direct accessibility to the host immune cells. These epitopes were linked together to construct a multi-epitope vaccine (MEV) that was predicted to be antigenic, and soluble upon overexpression. The tertiary structure of the MEV was generated which was used for molecular docking with TLR4 and TLR2. Molecular dynamics simulation and flexibility analysis confirmed the structural stability of the MEV-TLR4/TLR2 complexes. In-silico immune simulation showed the capability of MEV to induce a strong immune response. Codon optimization and in-silico cloning were performed to evaluate its efficient expression in the E. coli host. The findings suggest that surface exposed OMBB proteins can serve as promising antigenic candidates for designing an MEV construct.
bioinformatics2026-09-15v3SEAHORSE: A Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments
Quackenbush, A.; Kolluri, J.; Biju, R.; Nhong, S.; DeConti, D.; Wu, H.; Quackenbush, J.; Saha, E.; Eicher, T. D.Abstract
Large public molecular atlases such as the Genotype-Tissue Expression (GTEx) project and The Cancer Genome Atlas (TCGA) invite systematic discovery, yet most analyses remain hypothesis-driven and interrogate a tiny fraction of possible relationships among phenotypic, clinical, and molecular variables. We developed SEAHORSE (Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments), a discovery engine and accompanying R package that exhaustively precomputes all pairwise associations across heterogeneous data types and presents them as a searchable association landscape. Using GTEx (948 donors, 43 tissues, 154 phenotypes), SEAHORSE generated 341,008 phenotype-phenotype associations, 125,246,938 phenotype-gene associations, and 10,269,910,410 gene-gene correlations. In parallel analyses spanning 33 tumor types in TCGA, SEAHORSE generated 625,042 phenotype-phenotype associations, 183,080,369 phenotype-gene associations, and 12,096,948,950 gene-gene correlations. Across GTEx, height was repeatedly associated with enrichment of transcriptional programs, most strikingly the Kyoto Encyclopedia of Genes and Genomes (KEGG) term "Pathways in Cancer," significant in 17 tissues, offering a molecular entry point into long-reported links between stature and cancer risk. Height was also associated with immune, cardiovascular, and neurologic programs. In TCGA, age was consistently associated with WNT signaling, translation, cell differentiation, and cell cycle programs across tumors. These findings illustrate a new paradigm: large cohorts should be treated not merely as repositories for testing preconceived hypotheses but as association landscapes that can generate unexpected biological hypotheses.
bioinformatics2026-09-15v2Ryder: Epigenome normalization using a two-tier model and internal reference regions
Cao, Y.; Ge, G.; Zhao, K.Abstract
Motivation: Sequencing-based epigenomic profiling methods are powerful but suffer from technical variability that complicates cross-sample comparisons and can obscure true biological signals. While existing normalization methods using spike-in controls or computational approaches have been proposed, they often rely on assumptions that may not hold across diverse experimental conditions or require additional data types. Results: We present Ryder, a flexible and robust Python package for the normalization of epigenomic signal tracks. Ryder leverages stable internal reference regions, such as invariant CTCF binding sites, to correct for technical artifacts genome-wide. Our results show that it effectively adjusts both background noise and signal intensity, ensuring accurate signal alignment across samples while preserving genuine biological differences. We demonstrate that Ryder performs robustly across diverse assays including DNase-seq, CUT&RUN, ATAC-seq, MNase-seq, and ChIP-seq, with or without spike-in controls. By reducing technical noise, Ryder improves the detection of genuine biological changes, such as quantitative reduction of chromatin accessibility at key enhancer elements by depletion of BRG1, a key subunit of the chromatin remodeling BAF complexes. Availability and Implementation: The Ryder source code, documentation and test data are freely available at: https://github.com/YaqiangCao/ryder . The software version used in this study is archived at Zenodo: https://zenodo.org/records/21267457 .
bioinformatics2026-09-15v2Characterization of Outer Membrane β-barrel Proteins of Pasteurella multocida for Multi-Epitope Vaccine Design against Human Pasteurellosis
Panda, A.; Kapoor, J.; Kumar, S.; Bandyopadhyay, A.Abstract
Pasteurella multocida is a facultative anaerobic, Gram-negative coccobacillus that causes pasteurellosis in companion animals, livestock, and poultry and poses a significant zoonotic risk to humans through bite wounds, scratches, licking, and transfer of bodily fluids. Although vaccines are available for livestock and poultry, no vaccine is currently licensed for human use. In this study, we systematically identified and characterized 29 outer membrane {beta}-barrel (OMBB) proteins in P. multocida Past9 proteome and classified them into functional categories, including TonB-dependent receptors, porins, autotransporters, adhesins, and efflux pumps. B-cell, cytotoxic T-lymphocyte (CTL), and helper T-lymphocyte (HTL) epitopes were predicted from the identified proteins and screened based on antigenicity, non-allergenicity, and non-toxicity. Epitopes conserved across eight human-infecting P. multocida strains and located within the extracellular loop (ECL) region were incorporated into a multi-epitope vaccine (MEV) construct. The designed MEV was predicted to be antigenic, non-allergenic, and soluble. Its tertiary structural model was iteratively refined and validated. Molecular docking with human toll-like receptors 4/2 (TLR4/TLR2) predicted stable interactions, further supported by 100 ns molecular dynamics simulations. Immune simulation of the MEV construct predicted a strong simulated immune response. Furthermore, codon optimization and in silico cloning supported the feasibility of recombinant MEV expression in E. coli. The construct was benchmarked against OmpH, a well-known antigenic protein and exhibited broadly comparable predicted immune responses and receptor-binding energetics. This study proposes a designed MEV candidate against human pasteurellosis and highlights OMBB proteins as potential immunogenic targets for vaccine development.
bioinformatics2026-09-15v2Effects of UniProtKB restructuring and taxonomic database restrictions on downstream Unipept peptide-centric profiling
Vande Moortele, T.; Van de Vyver, S.; Binke, B.-B.; Van Den Bossche, T.; Dawyndt, P.; Martens, L.; Verschaffelt, P.; Mesuere, B.Abstract
Metaproteomics identifies the proteins present in a microbial community and, from them, which organisms are active and what functions they carry out. Tools such as Unipept assign peptides to taxa and functions by matching them to UniProtKB proteins and taking the lowest common ancestor of the matching taxa, deriving functional annotations from the same matches. This analysis therefore depends directly on the content of UniProtKB, which was substantially restructured in 2025-2026, reducing it by more than 100 million protein entries. We investigated how these changes affect taxonomic interpretation and how restricting the database to sample-specific taxa identified by rRNA profiling modifies that effect. Fixed, non-redundant peptide lists from human-gut and marine-hatchery studies were reanalysed with Unipept against UniProtKB releases 2025_03, 2025_04, and 2026_02. Peptide mapping coverage fell from 85.9% to 73.3% (gut) and from 82.3% to 68.7% (marine) between release 2025_03 and 2026_02, yet dominant taxonomic profiles remained robust: family- and genus-level distributions changed little and all 15 dominant gut species were retained. Root-level assignments dropped from 18.7% to 7.0% (gut) and 25.8% to 14.3% (marine). Restricting UniProtKB 2026_02 to taxa detected by large-subunit ribosomal RNA (LSU rRNA) profiling reduced coverage to 64.4% (gut) and 41.8% (marine), while genus- and species-level proportions changed by less than one percentage point. The major UniProtKB restructuring of 2025-2026 therefore narrowed mapping breadth without destabilizing the dominant taxonomic profiles in these two datasets.
bioinformatics2026-09-15v2TCRdenoise - an unsupervised similarity-based approach for denoising of TCR-pMHC specificity data
Lund, J. M.; Deleuran, S. N.; Nielsen, M.Abstract
Public repositories of T cell receptor (TCR)-peptide-MHC (pMHC) interactions constitute a critical resource for studying adaptive immunity and developing predictive models of TCR specificity. However, recent evidence suggests that a substantial fraction of reported TCR-pMHC interactions may be incorrectly annotated, limiting the quality of downstream analyses and machine learning applications. Here, we present an unsupervised sequence similarity-based framework for denoising peptide-specific TCR repertoires. The method combines pairwise TCR similarity metrics derived from TCRbase and TCRdist3 with hierarchical clustering and a novel adaptation of the silhouette score designed to address the prevalence of singleton clusters and highly imbalanced cluster structures. By incorporating a pseudo-cluster containing singleton and background TCRs, and by optimising both clustering distance thresholds and minimum cluster-size criteria, the proposed approach identifies TCRs likely to represent true antigen-specific binders while filtering putative noise. Using experimentally validated repertoires from TCRvdb, we demonstrate that the modified silhouette score closely tracks clustering solutions that maximise separation between binding and non-binding TCRs, achieving strong agreement with independent validation based on the Matthews correlation coefficient. Extension to a large collection of peptide-specific TCR data revealed a strong negative correlation between the percentage of TCRs classified as noise and the predictive performance of peptide-specific binding models. Further, the denoising classification labels on this data set were corroborated using structural modeling confidence scores of the peptide-TCR interface extracted from a refined AlphaFold 3 modeling pipeline. Additionally, retraining NetTCR on denoised data improved internal cross-validated performance compared with models trained on the full data set, whereas models trained exclusively on TCRs classified as noise performed close to random. Together, these results demonstrate that sequence similarity-based denoising can effectively enrich for biologically meaningful TCR-pMHC interactions and improve the quality of training data for predictive immunological models. The proposed framework provides a scalable strategy for improving the reliability of public TCR databases and facilitating the development of more accurate TCR specificity prediction methods.
bioinformatics2026-09-15v2Cell-level random splits leak group-owned answers in single-cell benchmarks
Sun, S.; Cang, H.Abstract
Machine learning models in single-cell biology increasingly forecast differentiation, reprogramming and therapeutic response from early transcriptomic profiles. Testing whether a model has learned real biology requires held-out cells. Single-cell data, however, are grouped: cells from the same clone, patient or batch share the same label. A random split therefore places relatives of each test cell, carrying its label, in the training set, and a model can score well by memorizing a relative instead of learning a transferable rule. Grouped validation removes this leakage but leaves far fewer independent units behind each error bar. Here we show how to estimate this leakage before training any model, from two properties of the data: exposure, the fraction of test cells with relatives in training, and retrievability, how often a nearest-neighbor search returns such a relative rather than an unrelated cell. Across lineage-barcoded and patient data, exposure determines whether a random split opens a leakage channel, and retrievability determines how much it can inflate the score. The inflation is negligible where cell state has decoupled from ancestry, much larger where clonal sisters remain close in expression space, and in a patient cohort large enough to overturn a clinical conclusion. We also provide eakcheck, which computes both properties in seconds, before the outcome model is fitted.
bioinformatics2026-09-15v1Accurate and scalable decontamination of imaging-based spatial transcriptomics via optimal transport
Chen, Y.; Liu, Y.; Chao, Z.; Han, S.; Zeng, Y.; Yu, B.; Zhang, F.; Wu, A.; Wang, J.; Chen, H.; Xiao, J.; Yang, C.Abstract
Imaging-based spatial transcriptomics enables molecule-resolved profiling of gene expression and tissue organization in situ. However, segmentation errors, transcript spillover and three-dimensional cell overlap can introduce misassigned transcripts into cell-level expression profiles, compromising biological interpretation and obscuring genuine signals. Existing methods either remove suspect expression at the cost of signal loss or lack a biologically grounded criterion for transcript assignment. Here we present CellDot, an optimal-transport framework that determines the fate of each transcript by retaining it in its host cell, reassigning it to a plausible neighboring cell or removing it as background. By integrating reference-guided expression compatibility with spatial information and data-adaptive constraints, CellDot enables accurate and traceable molecule-level correction while preserving biologically meaningful variation. In evaluations across multiple human tumor datasets, CellDot exhibited superior performance compared to existing decontamination methods, successfully restoring spatial expression patterns that matched independent cross-platform measurements. Moreover, it significantly enhanced the recovery of cellular states, intercellular communication, and spatial niche programs. Our experiments using real data demonstrated CellDot's scalability and established it as the only method applicable to a whole-transcriptome Atera dataset, underscoring its distinct advantages in the field of spatial transcriptomics.
bioinformatics2026-09-15v1Interpretable Machine Learning Reveals Complementary Age-Related Signatures in the Oral and Gut Microbiome
Ruthbah, C. A.; Sadi, T. H.; Jahan, N. E. S.; Adib, A. N. M. T.Abstract
Whether combining microbiome data from multiple body sites improves prediction, and whether different sites carry complementary or redundant information, are distinct questions that most studies conflate into a single accuracy metric. This work makes two contributions, one methodological and one biological, using paired stool and oral cavity microbiome samples from 44 subjects across two age groups, healthy adults and newborns (Ferretti et al., 2018). Methodologically, we show that a subject-matched fusion design combined with SHAP-based (SHapley Additive exPlanations) site attribution can detect complementary information between body sites even when no measurable accuracy gain results. This is a pattern that conventional model comparison would misread as a null result. Gut (stool) composition alone achieved near-perfect classification (area under the receiver operating characteristic curve, AUC = 1.00), and combined stool-oral models never exceeded this ceiling. A null baseline, bootstrap confidence intervals, and preprocessing sensitivity checks confirmed that this ceiling reflects genuine biological signal rather than an artifact. Despite the flat accuracy curve, SHAP analysis of the fused model showed that oral cavity features carried more total feature importance than stool features (58.1% versus 41.9%), indicating that the model draws on real, non-redundant information from both sites. Biologically, the taxa driving this pattern include Malassezia restricta, Staphylococcus epidermidis, and Prevotella melaninogenica. These taxa behave in a manner consistent with their established roles as early colonizers of the neonatal gut, skin, and oral cavity, once their model-specific behavior is verified directly against abundance data rather than inferred from the literature alone. An independent, substantially larger paired-cohort study using a different analytical method reports a compatible pattern. Together, these results support a model of oral-gut microbiome maturation as two distinct, complementary processes, and demonstrate that detecting this kind of relationship requires examining a model's internal reasoning rather than its accuracy alone.
bioinformatics2026-09-15v1Accounting for pseudo-replication of Linkage Disequilibrium for contemporary Ne estimation
Zhou, L.; Hui, T.-Y. J.; Burt, A.Abstract
The Linkage Disequilibrium (LD) of unlinked loci can be used to estimate contemporary effective population size (Ne) of one to a few generations ago. In genomic datasets loci on different chromosomes are considered unlinked, but there are many more pairs of unlinked loci than there are independent pairs of chromosomes, resulting to confidence intervals (C.I.) being too narrow if the non-independence is not taken into account. Simulations were run to investigate the correlation structure among LD of unlinked loci, which can be expressed by the LD of loci along the same chromosomes, based on a discovery of a novel Random Probe LD estimator. We classify the correlation into two categories: overlapping of loci and disjoint pairs. The former is induced from the same locus being considered twice and is the stronger form of correlation. These correlations feed into {rho}, a parameter to quantify the degree of pseudo-replication in a dataset, and further a correction formula from which C.I. can be properly inferred. We demonstrate the use of our method via an analysis of genomic data from the malaria-transmitting Anopheles gambiae s.s mosquitoes. Apart from the point and C.I. estimates, we find that Var((r^2 ) ) is inflated by about 550 times due to pseudo-replication, highlighting the danger of not handling genetic correlation properly.
bioinformatics2026-09-15v1Hierarchical temporal transformer for cancer grade prediction and cross cancer transfer learning from pathology reports
Brimo, N.; Anand, R.; Harb, H.; Serdaroglu, D. C.Abstract
Language models for cancer clinical reports carry two blind spots. They read each report in isolation, ignoring how a patients disease changes across visits, and they are evaluated only on cancer types present in their training data. We present the Hierarchical Temporal Transformer (HTT), a two-level architecture that addresses both. Level 1 encodes each report with BiomedBERT adapted by low-rank adaptation (LoRA). Level 2 is a temporal transformer that reads a patients full report sequence using a continuous-time positional encoding built from the measured number of days between visits, with learnable cancer-type embeddings supplying per-family conditioning. Two experiments test the two capabilities separately, since no fully open corpus contains longitudinal reports for many cancer types. On a controlled synthetic corpus of sequential radiology reports, in which progression phrases are inserted from templated trajectories, HTT reaches a validation AUROC of 0.942 against 0.881 for a single-report baseline and transfers to held-out pancreatic cancer at 0.995 against 0.949 while the two models are indistinguishable on a 60-patient test set. On 4,786 real pathology reports from the TCGA-Reports corpus spanning 14 cancer types HTT predicts tumor grade for three types withheld entirely from training, reaching AUROC 1.000 on thyroid carcinoma, 0.960 on sarcoma and 0.808 on lung squamous cell carcinoma. The mean held-out AUROC of 0.923 equals the in-distribution test AUROC of 0.923, so transfer to unseen cancer families incurred no measurable penalty. Ablation on the real corpus shows that the transfer is carried by the pre-trained encoder rather than by the temporal components, which, with one report per patient, contribute 0.39 AUROC points. Grade-related pathological language therefore appears to be learnable in a cancer-type agnostic way, which points toward unified cancer NLP systems that require no per-type retraining.
bioinformatics2026-09-15v1Codon language model scores provide information beyond protein language models for missense variant interpretation
Chen, R.; Palpant, N.; Foley, G.; Boden, M.Abstract
Predicting variant effects remains a central challenge in genomics. Protein language models (PLMs) capture amino-acid-level sequence constraints, whereas codon language models operate on coding sequences and may retain information that is lost upon translation. Here, we tested whether scores from the codon language model CaLM provide predictive information beyond protein-level representations for missense-variant interpretation. Across 71,436 ClinVar missense variants from 11,554 genes, adding CaLM to PLM baselines produced modest but reproducible improvements under gene-held-out cross-validation. PLM-only ensemble controls and explicit mutational-context analyses indicated that this improvement could not be explained solely by generic ensembling or simple codon-substitution features. Aggregating CaLM probabilities across synonymous codons attenuated codon-degeneracy-associated discordance while preserving most of the broader differences between CaLM and PLM scores. Gene-level analyses further showed that CaLM contribution varied continuously across genes and depended partly on the protein-model background. Across ClinMAVE functional assays, however, improvements were less consistent, indicating that codon-protein complementarity is context dependent rather than universal. Together, these results identify a modest but reproducible component of variant-effect information in CaLM-derived codon-level scores that is not fully captured by protein-level language-model representations.
bioinformatics2026-09-14v4Differential co-localisation analysis of multi-sample and multi-condition experiments with spatialFDA
Emons, M.; Scheipl, F.; Gunz, S.; Purdom, E.; Robinson, M. D.Abstract
Advances in spatial omics data generation have led to an explosion in new datasets that record the spatial location of transcripts and proteins. However, challenges remain in the analysis of spatial omics data. One important analysis is differential cellular co-localisation (CCoL): the quantification of the clustering, or spacing, of one or more cell types across multiple conditions. Our framework spatialFDA combines methodology from spatial statistics with functional data analysis to accurately quantify and test for differences between conditions in CCoL across spatial scales. Using two simulation studies, we show that spatialFDA performs well in controlled settings. Furthermore, spatialFDA recovers known biological processes in type-1 diabetes and adds insights about the CCoL strength in space. spatialFDA is readily available as an open-source Bioconductor R package.
bioinformatics2026-09-14v2Benchmarking niche identification via domain segmentation for spatial transcriptomics data
Wang, Y.; Chen, Y.; Yang, L.; Wang, C.; Cai, J.; Xin, H.Abstract
Tissue niches are spatially organized microenvironments in which coordinated multicellular interactions shape cellular states and biological functions. Currently, niche identification is routinely performed using domain segmentation frameworks. While interrelated, spatial domains and niches are not fundamentally equivalent. The former emphasizes intra-domain compositional consistency and transcriptomic homogeneity, whereas the latter is defined by the emergent properties of localized signaling gradients and the functional reciprocity between key cell lineages. Here, we present a high-resolution reference by thoroughly annotating single-cell resolution CosMx ST data of a human follicular lymphoid hyperplasia lymph node, a dynamic, non-compartmentalized tissue containing several critical immune niches defined by specific lineage architectures. We systematically benchmarked 16 contemporary domain segmentation algorithms, demonstrating that most methods in their default configurations fail to recapitulate biologically defined niche boundaries. Our analysis reveals that the definitive, disjoint spatial distributions of key functional lineages are frequently obscured by the stochastic infiltration of peripheral cell types. Such reduction in the spatial signal-to-noise ratio represents a primary bottleneck for existing algorithms, which prioritize local transcriptomic variance over global architectural logic. Following this observation, we demonstrate that strategic weighting of core functional lineages can restore the resolution of spatial niches in select domain segmentation frameworks. Cross-comparison against compartmentalized tissues further underscores the unique challenges of niche identification in non-mechanically separated environments and clarifies the fundamental divergence between structural domain segmentation and functional niche discovery. Our work delineates the limitations of current paradigms and advocates for the development of specialized computational approaches tailored specifically to the complexity of functional microenvironments.
bioinformatics2026-09-14v2SpaCoEx: Sparse Gene Selection for Spatially Varying Co-expression in Spatial Transcriptomics
Jian, M.; Pei, S.; Alterovitz, G.Abstract
Spatial transcriptomics enables gene expression to be measured while preserving tissue location, but most existing analyses focus on spatial variation in individual genes or expression-defined domains. Here, we introduce SpaCoEx, a sparse spatial representation framework that integrates gene-expression levels with spatially varying gene-gene co-expression. SpaCoEx first estimates local co-expression matrices from neighboring spatial spots, maps them into a log-Euclidean representation, and performs structured gene selection by retaining or removing the full row and column associated with each gene. The selected genes are then used to construct both expression-level features and local co-expression features, which are combined through an -weighted joint representation for downstream spatial analysis. We applied SpaCoEx to human cutaneous squamous cell carcinoma and annotated human breast cancer spatial transcriptomics datasets. In the cutaneous squamous cell carcinoma dataset, SpaCoEx selected 29 of 45 keratinocyte-related genes while preserving 96.74% of the spatial co-expression variation. In the breast cancer dataset, SpaCoEx identified spatially varying co-expression between B2M and HLA-C, a biologically meaningful major histocompatibility complex (MHC) class I antigen-presentation gene pair. Their local correlation was significantly higher in cancer-associated regions than in non-cancer regions (mean difference = 0.30, spatially adjusted SE = 0.045, P<1.0x10^(-10)), whereas B2M and HLA-C expression individually did not differ significantly between cancer and non-cancer. In benchmarking against manual tissue annotations, co-expression-only SpaCoEx achieved the strongest spatial coherence (percentage of abnormal spots [PAS] = 0.076), while the joint expression/co-expression representation achieved the highest annotation agreement, with an adjusted Rand Index (ARI) of 0.584 at = 0.60 and and normalized mutual information (NMI) of 0.663 at = 0.40. By integrating marginal gene-expression information with local gene-gene co-expression structure, SpaCoEx provides a sparse, low-dimensional, and interpretable representation of spatial transcriptomics data that captures complementary aspects of tissue organization beyond expression-based variation alone.
bioinformatics2026-09-14v2ForceFlowAb: physics-aware mixture-of-experts flow matching model for antibody CDRs design
Li, Z.; Lv, Z.; Zhang, G.Abstract
Abstract Motivation: Antibodies are a major class of therapeutic molecules, and their recognition of target antigens is largely mediated by complementarity-determining regions (CDRs), making antigen-conditioned CDR design a central problem in antibody engineering. Recent generative methods have enabled antigen-conditioned co-design of CDR sequences and structures, but their limited capacity to capture local interface heterogeneity and lack explicit energy-based guidance during sampling, which may result in unfavorable antibody-antigen interaction energies. Overcoming these limitations requires methods that better represent diverse interface environments via adaptive routing and incorporate physical guidance to steer sampling toward energetically favorable conformations. Results: We present ForceFlowAb, a physics-aware mixture-of-experts flow-matching framework for antigen-conditioned CDR sequence-structure co-design. The framework models heterogeneous interface environments through specialized expert routing and applies differentiable force-field guidance during sampling to guide CDR generation toward energetically favorable conformations. For CDR-H3 design, ForceFlowAb achieved more favorable antibody-antigen interaction energies than FlowDesign and Diffab, with improvement rates (IMP) of 46.5% versus 35.0% and 35.5%, respectively. For simultaneous six-CDR design, ForceFlowAb also outperformed Diffab, with IMP values of 16% versus 9%. These results suggest complementary roles for interface-adaptive modeling and energy-based guidance, with the former capturing binding-mode diversity and the latter leveraging physical constraints to ensure biophysical feasibility. Availability and implementation: The web server is freely available at http://zhanglab-bioinf.com/ForceFlowAb. The source code and implementation are available at https://github.com/iobio-zjut/ForceFlowAb. Contact: zgj@zjut.edu.cn Supplementary information: Supplementary data are available at Bioinformatics online.
bioinformatics2026-09-14v1LIGER2: Scalable Single-Cell Integration with On-Disk Datasets
Wang, Y.; Robbins, A.; Gadhvi, G.; Welch, J. D.Abstract
Correcting batch effects and integrating single-cell sequencing datasets has been a crucial step in large-scale biological studies. Many methods have been published for this task, and excel in various scenarios. Our previous work, LIGER, leveraging integrative non-negative matrix factorization (iNMF), stands out in providing an interpretable low-dimensional representation. To adapt to the modern need for integrating millions of cells, we developed a highly-optimized parallel factorization solution with on-demand loading from disk. The upgraded LIGER algorithm shows significant improvements in time and memory efficiency for single-cell data integration. We also developed a new downstream embedding alignment method significantly improved performance in conserving biological variation while still aligning corresponding cell types across datasets.
bioinformatics2026-09-14v1TransBind2: Improving Transcription Factor-DNA Binding Prediction with Multimodal Data and Bidirectional Cross Attention
Basnet, S.; Cheng, J.Abstract
Accurate genome-wide prediction of transcription factor (TF)-DNA binding remains challenging because many models focus mainly on DNA sequence and overlook chromatin context and TF structure. We previously developed TransBind, a protein-aware model that combines TF and DNA representations through cross-attention. Here, we introduce TransBind2, which improves on TransBind in several ways. It incorporates DNase-seq accessibility and genome mappability tracks as additional input, uses a biomodal protein language model (ProstT5) to capture both TF sequence and structure, and applies bidirectional cross-attention so DNA and protein features can refine each other. We also frame prediction as binary classification of individual <DNA bin, TF, cell type> triplets, allowing the model to generalize to new TFs and cell types. Across 690 human ChIP-seq experiments covering 161 TFs and 91 cell types, TransBind2 achieves a macro AUROC of 0.9648 and AUPR of 0.4215, outperforming TransBind and other baselines, with a [≥]12.67% relative AUPR gain. The model trained on human data also performs well in cross-species zero-shot prediction on mouse data. Saliency analysis shows that it can identify TF-binding peaks with a median error of 12-38 base pairs (bps) despite being trained on window-level labels. Ablation studies further show that TF structure, chromatin accessibility, and bidirectional attention each improve performance. Overall, these results show that combining TF structure with chromatin context leads to more accurate and generalizable TF-DNA binding predictions.
bioinformatics2026-09-14v1PHACTn enables training-free, context-independent inference of nucleotide variant tolerance across the genome
Yildirim, C.; Kuru, N.; Adebali, O.Abstract
Accurate prediction of single-nucleotide variant (SNV) tolerability across the entire human genome remains a fundamental challenge in computational genomics, particularly for non-coding regions where the regulatory landscape is vast and poorly understood. Machine learning classifiers suffer from data circularity and demographic bias, while genomic language models demand massive computational resources and offer little biological interpretability. Here, we present PHACTn (Phylogeny-Aware Computing of Tolerance for nucleotide variants), a training-free, parameter-minimal method that infers nucleotide variant tolerability by traversing the mammalian phylogenetic tree and explicitly modeling the evolutionary independence of observed substitutions and their distance from the query species. With only 4 interpretable parameters, no training and no GPU requirement, PHACTn outperforms all evaluated tools on non-coding variants curated from both the ClinVar, and on non-coding variants potentially responsible for selected Mendelian diseases curated from OMIM. Additionally, it achieves state-of-the-art performance on variants within the informative range of alignment-based inference. These results establish that principled probabilistic phylogenetic modeling captures evolutionary constraint signals that large-scale sequence models fail to recover, offering a powerful, accessible, and mechanistically transparent alternative for genome-wide variant effect prediction.
bioinformatics2026-09-14v1ImmuneLens: linking transcriptional states and TCR clonotypes through disentangled multimodal learning
Duan, Z.; Wang, Y.; Li, C.; Li, G.; Cao, Y.; Bai, X.; Yang, F.; Song, S.Abstract
Single-cell multi-omics technologies simultaneously capture the transcriptome and TCR sequence of T cells, providing an opportunity to study the relationship between transcriptional states and clonal architectures. However, jointly modeling the relationships between transcriptional states and TCR sequences while preserving modality-specific information remains challenging. Here, we present ImmuneLens, an interpretable multimodal representation learning framework designed for paired single-cell transcriptome and TCR sequence data. ImmuneLens supports the construction of a transferable multi-cohort immune reference atlas and enables unsupervised mapping of external query data. The complementarity between GEX and TCR information improves the stability of antigen-specificity prediction. In neoadjuvant immunotherapy cohorts, ImmuneLens resolves response-associated T cell heterogeneity and reveals links between clonal expansion and CD8 T cell functional states. Overall, ImmuneLens provides a
bioinformatics2026-09-14v1