Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
A critical evaluation of Gene Ontology priors in biologically-informed neural networks
Verlaan, T.; Lieftinck, M. A.; Mwine, W.; Reinders, M. J. T.Abstract
Biologically-informed neural networks (BINNs) embed prior knowledge such as the Gene Ontology (GO) into their architecture to produce structurally interpretable representations, yet whether and how this prior improves performance or interpretation remains unclear. Here, we introduce GONNECT, a BINN incorporating GO into an autoencoder. We evaluate GO constraints in the encoder, decoder, or both on RNA-seq tumour samples from The Cancer Genome Atlas (TCGA), comparing against published BINNs (OntoVAE and VEGA), randomized-prior controls, and an unconstrained baseline. Across metrics, GO structure adds little to reconstruction or latent-space organization, frequently matched by randomized or unconstrained models. Its value lies in node activations, particularly in the encoder, where they correlate with a gene set enrichment analysis (GSEA)-derived reference. GONNECT-SL introduces regularized connections outside GO, but these soft links are unstable across seeds and concentrate where the ontology is sparse, appearing to compensate for the priors constraints rather than reveal new biology. They recover near-unconstrained reconstruction, keeping encoder activations interpretable. We identify the soft-link encoder as most promising. Our results clarify what biological priors contribute: their value lies not in the identity of the imposed connections or in improved performance, but in organizing activations into biologically meaningful units that can be interrogated directly.
bioinformatics2026-09-01v3A Visually Interpretable Histopathology-Based Immune Model Predicts T-effector Biology and Response to Immune checkpoint inhibition in Clear Cell Renal Cell Carcinoma Clinical Trial and Contemporary Real-World Datasets
Perny, A.; Jarmale, V.; Jasti, J.; Zhong, H.; Christie, A. L.; Miyata, J.; Nielsen, A. W.; Kontoyiannis, P.; Rakheja, D.; Modrusan, Z.; Huseni, M.; Kadel, W.; Brugarolas, J.; Kapur, P.; Rajaram, S.Abstract
Immune checkpoint inhibitors (ICI) are central to the treatment of metastatic clear cell renal cell carcinoma (ccRCC), yet only a subset of patients derive durable benefit, and clinically deployable predictive biomarkers remain an unmet need. RNA-based T-effector signatures capture cytotoxic immune biology and have been associated with ICI response in clinical trial cohorts; however, their clinical implementation is limited by the marked spatial heterogeneity of ccRCC, as well as cost, long turnaround time, sample quality requirements, and limited accessibility. Here, we developed a visually interpretable deep learning (DL) model that predicts a T-cell-enriched immune score directly from hematoxylin and eosin (H&E)-stained whole-slide images. To overcome the inability of H&E morphology alone to distinguish lymphocyte subsets, we trained the model using multimodal spatial supervision from CD8, PAX8, and ERG IHC, which respectively identified cytotoxic T-cell-rich regions, tumor cells, and endothelial cells, thereby constraining immune predictions to relevant tumor microenvironmental niches. The resulting H&E DL Immune score was validated by pathologist review, comparison with held-out CD8 IHC annotations, and independent datasets. The H&E DL Immune score correlated with T-effector RNA scores across independent institutional and IMmotion150 clinical trial cohorts (spearman correlations of 0.726; p=5.90x10-15 and 0.706; p=4.04x10-19). As a proof of principle, the score was used to characterize associations with key biological features across large cohorts, including sarcomatoid differentiation, BAP1 and PBRM1 mutation status, and additional transcriptomic signatures. In IMmotion150 clinical trial cohort, a median-dichotomized H&E DL Immune score, similar to RNA-based T-effector score, was significantly associated with clinical benefit from atezulumab therapy. In contemporary institutional cohorts of patients treated with frontline ipilimumab plus nivolumab or in initial 3 lines of nivolumab monotherapy, patients in the top quartile of H&E DL Immune score had significantly longer progression-free survival. Collectively, these findings support a scalable and interpretable H&E-based biomarker that captures T-effector biology and can help identify patients with ccRCC more likely to benefit from ICIs.
bioinformatics2026-09-01v2Large Language Models for Accessible Reporting of Bioinformatics Analyses in Interdisciplinary Contexts
Yu, L.; Kim, D.; Cao, Y.; Shu, M. W. S.; Shen, M.; Liang, X.; Gu, J.; Jayakumar, R.; Ding, W.; Yang, F.; Zhang, X.; Kim, J.; Yang, P.; Yang, J. Y. H.Abstract
Health and life scientists frequently rely on quantitative experts to perform complex data analyses, yet interpretation of these results often becomes a bottleneck due to communication barriers across disciplines. Large Language Models (LLMs) have the potential to act as intermediaries that translate analytical outputs into accessible scientific narratives, but their ability to reliably interpret outputs from real-world biomedical data analyses across disciplines remains unclear. In this study, we investigate how state-of-the-art LLMs behave when integrated into real bioinformatics analysis workflows as report-generation assistants for interdisciplinary teams. We evaluated LLM-generated summaries using both automated and human evaluation frameworks to ensure holistic evaluation. Automated assessment employed multiple choice questions designed using Bloom's taxonomy to assess multiple levels of understanding, while human evaluation tasked scientists to score summaries for factual consistency, lack of harmfulness, comprehensiveness, and coherence. All generally produced readable and largely safe summaries, confirming their value for first-pass translation of technical analyses, however frequently misinterpreted visualisations, produced verbose summaries and rarely offered novel insights beyond what was already contained in the analytics. Our findings suggest that LLMs are best suited for easing interdisciplinary communication rather than replacing domain expertise and human oversight remains essential to guarantee accuracy, interpretative depth, and the generation of genuinely novel scientific insights.
bioinformatics2026-09-01v2An Integrated Transcriptomic Landscape of Lung Cancer Identifies Tumor Clusters of Biological Significance
Arora, S.; Suresh, L.; Thirmanne, H. N.; Jensen, M.; Glatzer, G.; Fatherree, J.; Konnick, E.; Levine, K.; Brooks, A. N.; Houghton, A. M.; Pritchard, C.; MacPherson, D.; Berger, A.; Holland, E. C.Abstract
Lung cancer encompasses multiple histological entities with substantial molecular heterogeneity that remain incompletely resolved at population scale. Here, we constructed a unified reference landscape of lung cancer by analyzing raw RNA sequencing data from 1,824 tumors spanning adenocarcinoma (n=966), squamous cell carcinoma (n=628), small cell lung cancer (n=150), and unclassified non small cell lung cancer (n=80). Following batch correction, samples were analyzed using consensus clustering and visualized with PaCMAP to generate a molecular atlas annotated with clinical and biological metadata. Rather than segregating by pathological diagnosis, tumors organized along conserved transcriptional axes defined by tumor-intrinsic biology including proliferative or metabolic programs and immune-infiltrated states. Consensus clustering resolved nine robust molecular clusters, including an adenocarcinoma-associated subgroup, a neuroendocrine-like adenocarcinoma marked by ASCL1 activation, immune-associated regions, and bifurcation of both small cell and squamous carcinomas into biologically distinct states. Spatially restricted expression of selected clinically relevant transcripts nominated state-specific therapeutic hypotheses requiring future functional and clinical validation. Projection of patient tumors and patient-derived xenografts onto the atlas demonstrated preservation of transcriptional identity and enabled quantitative assessment of model fidelity. This integrated framework organizes lung cancer as a structured continuum of transcriptional states and provides a reference resource for biological interpretation and future translational studies.
bioinformatics2026-09-01v2Spatial Transcriptomics As Rasterized Image Tensors (STARIT) characterizes cell states with subcellular molecular heterogeneity
Velazquez, D.; Hallinan, C.; An, R.; Clifton, K.; Fan, J.Abstract
Abstract Imaging-based spatially resolved transcriptomics (imSRT) technologies provide high-throughput molecular-resolution spatial characterization of genes within cells. Conventional analysis methods to identify cell-types and states in imSRT data rely on gene count matrices derived from tallying the number of mRNA molecules detected for each gene per segmented cell, thereby overlooking subcellular heterogeneity that can be useful in defining cell states. To take advantage of the molecular-resolution information in imSRT data and potentially identify cell-states based on subcellular heterogeneity, we developed STARIT (Spatial Transcriptomics As Rasterized Image Tensors). STARIT converts transcripts within segmented cells in imSRT data into an image-based tensor representation that can be combined with deep learning computer vision models for downstream analysis. Using simulated and real imSRT data, we demonstrate that STARIT distinguishes transcriptionally distinct cell-types and further separates cell states based on subcellular transcript localization, which conventional gene count analysis fails to capture. By providing a standardized framework to encode subcellular molecular information in imSRT data, STARIT will enable deeper insights into subcellular heterogeneity and enhance the identification and characterization of cell-types and states that are overlooked by gene count representations.
bioinformatics2026-09-01v2PathFold: Predicting the Entire Protein Folding Pathway from Protein Sequence Alone
Zhang, Z.; Ibtehaz, N.; Kagaya, Y.; Xu, Z.; Punuru, P.; Kihara, D.Abstract
Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured {Phi}-values.
bioinformatics2026-09-01v1CyChat: a conversational Cytoscape app for no-code, reproducible network analysis
Liebold, J.; Stahl, M.; Schulze, J.-O.; Razavi, M. M.; Bader, G. B.; Kurtz, S.; Baumbach, J.Abstract
Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.
bioinformatics2026-09-01v1The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI
Nelen, J.; Khan, O.; Adams, E.; Aschenbrenner, J. C.; Thompson, W.; Ebrahim, A.; Capkin, E.; Vallee, C.; OpenBind, ; Shotton, E. J.; Griffen, E. J.; Chodera, J. D.; Deane, C. M.; von Delft, F.; AlQuraishi, M.; Imrie, F.Abstract
High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.
bioinformatics2026-09-01v1GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler
Uzum, A. S.; Haliloglu, T.Abstract
Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.
bioinformatics2026-09-01v1Sphingolipid metabolism-related genes as key regulatory hubs in white smoke inhalation induced lung injury
Meng, F.; Xin, H.; Li, R. R.Abstract
Objective White smoke inhalation injury (WSI) causes severe acute lung damage with no specific therapy currently available. Sphingolipid metabolism is implicated in pulmonary inflammation, but its transcriptional regulatory landscape in WSI remains unexplored. This study aimed to identify key sphingolipid metabolism related genes and evaluate their regulatory roles and therapeutic potential in WSI. Methods We established a rat model of WSI and performed integrated bulk RNA sequencing, weighted gene coexpression network analysis (WGCNA), and single-cell RNA sequencing (scRNAseq) to screen for differentially expressed sphingolipid metabolism-related genes (DESRGs). Protein-protein interaction (PPI) network with four centrality algorithms was used to prioritize hub genes. In silico gene knockout and molecular docking were conducted to assess regulatory functions and identify potential drug candidates. Results We identified 22 DESRGs that were predominantly enriched in DNA replication and cell cycle pathways rather than canonical sphingolipid metabolic processes. PPI consensus prioritized three hub genes--Top2a, Ttk, and Ccna2--with Top2a exhibiting the highest expression in epithelial cells and significant downregulation after smoke exposure. ScRNAseq revealed immune cell infiltration and epithelial differentiation trajectories. Virtual knockout showed that Top2a depletion affected the largest transcriptomic fraction (~0.4%) and was enriched in lysosome biogenesis, innate immunity, phagocytosis, and lipid catabolism. Molecular docking identified thalidomide as a high affinity ligand for Top2a (Vina score: -8.5 kcal/mol). Conclusion Our multiomics integrative framework identifies Top2a as a central regulatory hub linking sphingolipid associated inflammation to epithelial responses in WSI, and nominates thalidomide as a potential drug repurposing candidate. These findings provide prioritized targets for future translational investigation.
bioinformatics2026-09-01v1RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution
Yang, X.; Hao, N.; Zhao, R.; Angel, S.; Tan, Y.; Lian, C. G.; Zhou, L.; Olson, D.; Yu, K.-H.; Ruiz de Luzuriaga, A.; Wan, G.Abstract
Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.
bioinformatics2026-09-01v1The DYNAM-O Toolbox: Characterizing Individualized Neural Signatures in Sleep EEG
He, M.; Saremsky, S. R.; Noamany, H.; Chen, S.; Prerau, M. J.Abstract
Conventional sleep electroencephalography (EEG) measures often rely on predefined bands, thresholds, and averages that incompletely capture transient oscillatory dynamics across an entire night. Here, we introduce the Dynamic Oscillation (DYNAM-O) Toolbox, an open-source, cross-platform (MATLAB, Python, and Rust) software package for data-driven characterization of individualized neural dynamics in sleep EEG. DYNAM-O identifies transient oscillations as time-frequency peaks on multitaper spectrograms using a novel multi-resolution procedure, computes intrinsic and sleep-state-dependent extrinsic features for each event, and represents the overnight distributions of tens of thousands of TF-peaks as feature histograms spanning oscillation frequency, slow oscillation power, and slow oscillation phase. This distributional representation preserves continuous brain-state variation that could be obscured by averaging within conventional sleep stages. The toolbox further provides Gaussian and spline basis-based dimensionality reduction, visualization, and whole-histogram statistical testing tools to support both exploratory and hypothesis-driven analyses. To demonstrate its use for group-level inference, we analyzed overnight C3-channel EEG from 133 adults (71 females, 72 males; ages 20-35 years) in the Cleveland Family Study. Whole-histogram and parameterized-mode analyses reproduced the established higher center frequency of fast-spindle activity in females and additionally revealed greater low-alpha transient oscillatory activity in females, a pattern outside the conventional sleep spindle range. By completing the analysis cycle from TF-peak extraction to statistical inference, DYNAM-O provides an accessible and interpretable framework for studying individualized sleep physiology and identifying subtle, reproducible electrophysiological patterns.
bioinformatics2026-09-01v1Calibration-free compression brings Evo 2 to its full million-token context on a single GPU
Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.Abstract
Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.
bioinformatics2026-09-01v1Intelligent differential ion mobility spectrometry (iDMS): A deep neural network that predicts optimal space-resolved ion mobility parameters for isomeric monoglycosphingolipids
Nguyen-Tran, T.; Shi, X. X.; Hashimoto-Roth, E.; Organ, M. G.; Lavallee-Adam, M.; Perkins, T. J.; Bennett, S. A. L.Abstract
Simultaneous quantification of monoglycosphingolipid stereoisomers is required to monitor changes in defective enzymatic pathways linked to diseases such as Gaucher Disease, Parkinson's Disease, and Krabbe Disease. Resolution of beta-glucosyl and beta-galactosyl epimers cannot be achieved by standard liquid chromatography, electrospray ionization, tandem mass spectrometry (LC-ESI-MS/MS). Separation becomes possible when field asymmetric ion mobility spectrometry (FAIMS), also known as differential mobility mass spectrometry (DMS), is added as an orthogonal separation technique to LC. FAIMS/DMS separates epimeric ion clusters in a high versus low electric field (separation voltage, SV) then redirects the target epimeric ions to the mass spectrometer through the application of a direct current (compensation voltage, CoV). Resolving SVs and CoVs must be manually determined for each lipid. Manual derivation is a labour-intensive process that requires pure synthetic standards, limiting the number of stereoisomers a user can include in an assay. To address this problem, we introduce here intelligent DMS (iDMS). iDMS is an in silico supervised neural network model that learns the ion mobility relationships between SV and CoV and the monoglycosphingolipid structural features of sugar headgroup, N-acyl chain length, and N-acyl degree of unsaturation. iDMS predicts the SV and CoV combinations capable of resolving any stereoisomer pair from a training dataset of composed of measured signal intensities across a range of SVs and CoVs of 12 lipids. This machine learning alternative to manual DMS optimization promises to accelerate the deployment of multiple-reaction-monitoring mode (MRM) RPLC-ESI-DMS-MS/MS assays for the routine and rapid quantification of biologically relevant monoglycosphingolipid stereoisomers.
bioinformatics2026-09-01v1Mural-VISTA: A tool for mural cell-vessel interaction assessment and multiscale single-cell topo-morphological analysis
Zeng, H.; Hu, M.; Phng, L.-K.; Matsunaga, Y. T.Abstract
Three-dimensional (3D) mural cell morphology is heterogeneous and coupled to vessel geometry, however, measurements from two-dimensional (2D) maximum intensity projections (MIP) obscure overlapping processes and cell-vessel contacts. Accordingly, we developed Mural-VISTA, a semi-automated Python workflow for mural cell-vessel interaction and single-cell topo-morphology analysis of reconstructed surface meshes. This workflow integrates mesh pretreatment, interactive centerline extraction, hierarchical segmentation of cell soma, main axis and secondary processes (branches), and extraction of 36 multiscale (cell process segment level, process level, and whole cell level) topo-morphological and vessel-referenced metrics. Mural-VISTA identified morphological changes in pericytes and vascular smooth muscle cells (vSMCs) with altered RhoA activity. Constitutive active RhoA (RhoA CA) over-expression reduced branch complexity and increased process alignment in both cell types, while increased whole-cell and branch solidity only in vSMCs. Dominant negative RhoA (RhoA DN) over-expression increased branch abundance and reduced branch solidity in pericytes but not vSMCs, suggesting cell-type specific effect of reduced RhoA activity. In conclusion, Mural-VISTA enables quantitative 3D profiling of mural cell architecture and its spatial relationship with the vessel.
bioinformatics2026-09-01v1Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R
Zeng, Z.; Wang, Y.Abstract
Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.
bioinformatics2026-09-01v1Automatic bioinformatic software named entity recognition from literature
Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.Abstract
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.
bioinformatics2026-09-01v1From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology
qin, y.; Pang, J.; Zhang, X.Abstract
Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.
bioinformatics2026-09-01v1Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach
Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.Abstract
Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.
bioinformatics2026-09-01v1An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators
Li, X.; Wei, P.Abstract
Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.
bioinformatics2026-09-01v1AmPair: automating housekeeping-gene primer design for species-level metataxonomics
Xu, X.; Yang, X.Abstract
Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.
bioinformatics2026-09-01v1Spatial Transcriptomics Reveals Compartment-Specific Immune Activation Signatures in Ileal and Lymph Node Tissue in Treated HIV Infection
Barrett, M.; Anderson, J.; Escandon, K.; Wieking, G.; Schroeder, T.; Swanson, E.; Graham, M. L.; Gale, M.; Schacker, T. W.; Klatt, N. R.; Basting, C. M.Abstract
People with HIV (PWH) on long-term antiretroviral therapy (ART) continue to experience elevated rates of morbidities and mortality driven by persistent immune activation despite viral suppression. Known contributors include low-level HIV provirus activity, microbial translocation in part from epithelial barrier dysfunction, microbiome dysfunction, and co-infections. However, how these interact and where they predominate across tissue compartments remains incompletely defined. Here, we applied spatial transcriptomics to characterize compartment-specific transcriptional programs in ileum (epithelium, Peyer's patches, lamina propria) and inguinal lymph nodes (B Cell follicles and T cell zone) from ten PWH on long-term ART, stratified by CD4/CD8 ratio into low-ratio and high-ratio groups, with low-ratio as a proxy for immune activation and increased risk for non-AIDS related serious event. Comparison of global expression found significant differences between groups in four of five compartments. Differential expression analysis identified 483 differentially expressed genes across four of five compartments, with the greatest burden in the T-cell zone and none in the lamina propria. Gene set enrichment analysis identified 116 enriched pathways predominantly in the low-ratio group, spanning immune activation, infection-associated, and metabolic programs, with Peyer's patches showing the broadest transcriptional divergence of any compartment. Cross-compartment signals included higher expression of ORMDL3 and ARL17B in the low-ratio group implicating mitochondrial stress and inflammasome activation, lower expression of CCL3L3 and FCMR in the low-ratio group suggesting impaired immune execution, and divergent ribosomal protein programs between B-cell follicles and the T-cell zone. Cell deconvolution identified compartment-specific differences in estimated immune cell proportions, and T-cell zone gene expression showed significant associations with HIV reservoir measures and plasma markers of microbial translocation and immune activation. Together these findings support spatially heterogeneous immune activation as a feature of persistent immune dysregulation in treated HIV infection and provide compartment-resolved, hypothesis-generating evidence for the tissue-specific mechanisms driving inflammation in this population.
bioinformatics2026-09-01v1GlyComboCLI enables command line-based FAIR workflows for glycan composition assignment in mass spectrometry data
Kelly, M. I.; Thang, W. C. M.; Pang, C. N. I.; Gustafsson, O. J. R.; Ashwood, C.Abstract
Glycans are integral biomolecules whose presence cannot be predicted from genomic data alone, necessitating experimental characterisation through approaches including mass spectrometry. Assignment of glycan compositions to observed mass to charge ratios is computationally challenging due to the potential monosaccharide diversity and existing tools lack the required flexibility for integration into automated bioinformatic workflows. Here, we present GlyComboCLI, an open-source command-line application for the assignment of glycan compositions to mass spectrometry data which expands upon the previous GUI application, GlyCombo. GlyComboCLI accepts mass lists and vendor-neutral mzML files, supports diverse monosaccharides, derivatisation states, reducing-end modifications and adducts, and is validated through automated tests spanning supported search parameters, input formats, and instrument vendors. Outputs are compatible with downstream tools including Skyline and GlycoWorkBench while GlyTouCan accessions support persistent identification of registered base compositions. Deployment as a standalone executable, a Docker container, and a Galaxy tool, supports FAIR and reproducible workflows. Applied to published mouse and human glycomics datasets, GlyComboCLI reproduced major qualitative glycomic motifs and relative abundance trends, while Galaxy workflow reproduced expected effects of sialidase treatment. These results demonstrate a flexible, scalable, and reproducible approach for glycan composition assignment within automated glycomics workflows.
bioinformatics2026-08-31v3OmniSplice: detection of non-canonical splicing events from RNA-seq
Lannes, R.; Li, R. Y.; Fingerhut, J. M.; Cummings, R. A.; Salagean, A. D.; Yamashita, Y. M. M.Abstract
Splicing generates mature mRNA by removing introns from nascent transcripts and is widely studied using RNA sequencing. However, most RNA-seq analysis pipelines classify RNA-seq reads according to predefined splice-junction structures and discard those that do not conform to such predefined models, potentially obscuring biologically meaningful splicing events. In this study, we developed OmniSplice, a computational framework that captures and analyzes RNA-seq reads that overlap annotated exon ends without assuming predefined splicing architectures. This approach enables systematic detection of non-canonical splicing events that are often overlooked by conventional analyses. Applying OmniSplice to Drosophila splicing factor mutants and mouse TDP-43 mutant datasets, we found widespread splicing defects with non-canonical junctions that were not previously recognized, including back-splicing and trans-splicing. Together, these results demonstrate that RNA-seq datasets may contain a substantial reservoir of overlooked splicing information, warranting more comprehensive approaches for analyzing RNA-seq data for splicing events.
bioinformatics2026-08-31v2JMod: Joint modeling of mass spectra for empowering multiplexed DIA proteomics
McDonnell, K.; Geiszler, D. J.; Wamsley, N.; Derks, J.; Sipe, S.; Cohen, Z. A.; Warinner, L. K.; Yeh, M.; Koo, E.; Leduc, A.; Zwang, T. J.; Specht, H.; Slavov, N.Abstract
Parallelization of data acquisition substantially increases the throughput of mass spectrometry-based proteomics. However, parallelization also increases the density of mass spectra and consequently the overlap between ions, frustrating their analysis. To improve sequence identification and quantification from such spectra, we developed an open-source software for Joint Modeling of mass spectra (JMod). JMod models overlapping peaks as linear superpositions of their components in both MS1 and MS2 space, which permits multiplexed DIA with smaller mass offsets to increase the multiplexing capacity and thus proteomics throughput for a given plexDIA tag. This enables 9-plexDIA using 2 Da offset PSMtags, increasing throughput 9-fold while preserving quantitative accuracy and coverage depth. Furthermore, we use JMod to deconvolve simultaneous labeling by mass tags and heavy amino acids, thus increasing the throughput of metabolic pulse experiments measuring protein synthesis and degradation rates in single cells from mouse liver. By supporting enhanced decoding of highly multiplexed DIA spectra, JMod provides an open and flexible software that increases the throughput of sensitive proteomics.
bioinformatics2026-08-31v2Decoding heterogeneous aging clocks and disease risk stratification using MetAgeFormer
Xu, Y.; Zou, B.; Xie, G.; Chen, T.; Jia, W.; Zhang, L.Abstract
Metabolomic aging clocks estimate biological age by modeling metabolite concentrations, thereby capturing aging signals from healthspan and adverse outcomes. However, existing clocks generally assume homogeneous aging trajectories and yield only a single age acceleration metric, limiting their capacity to capture inter-individual metabolic heterogeneity and characterize nuanced individual-level representations. To address these limitations, we proposed MetAgeFormer, a transformer-based metabolomic model pre-trained on nuclear magnetic resonance (NMR) metabolomic profiles from over 430,000 participants in UK Biobank via self-supervised learning. This large-scale pre-training enables MetAgeFormer to learn a metabolomic representation space that captures the complex, nonlinear structure of systemic metabolism as reflected in NMR data. Building on MetAgeFormer, we developed a mortality-informed metabolomic aging clock by fine-tuning an attached survival module, deriving age acceleration that demonstrates significant associations with multiple age-related diseases and factors. We further validated zero-shot transfer in the independent Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort. More importantly, we utilized embeddings generated by MetAgeFormer to identify 13 distinct metabolic subtypes and consolidated them into four meta-subtypes with markedly divergent susceptibility profiles for major age-related diseases, particularly type 2 diabetes and neurodegenerative disorders. This finding empirically demonstrated substantial metabolic heterogeneity across populations, persisting even at comparable levels of age acceleration. To enhance clinical applicability, we further employed contrastive learning to distill a lightweight model that approximates the learned metabolomic representation space using only 14 routine clinical blood test measurements as inputs. Both hold-out testing within UK Biobank and external validation in the China Health and Retirement Longitudinal Study replicated similar disease onset patterns across the identified subtypes, underscoring the robust generalizability of MetAgeFormer and supporting its translational potential as a scalable framework for metabolomic aging assessment and early disease risk stratification.
bioinformatics2026-08-31v2seqproc: An efficient, flexible, and concise tool for sequence geometry description and transformation
Cape, N.; Fisher, E.; Liu, D.; Patro, R.Abstract
Complex sequencing protocols encode technical information in read structure and require accurate, e[ff]icient preprocessing. We introduce seqproc, which compiles concise sequence-geometry descriptions into execution graphs. Across four single-cell RNA-sequencing protocols, seqproc has the lowest mean runtime at every tested thread count and uses substantially less memory than the next-fastest tool. It has the highest F1 agreement with conservative structural references on all three discriminative chemistries and ties both alternatives on the 10x length-filter control. By separating protocol description from execution, seqproc makes complex read transformations compact, reusable, and efficient.
bioinformatics2026-08-31v2Prioritizing peptides for targeted mass spectrometry experiments using deep learning
Sonthalia, S.; Wen, B.; Dasgupta, P.; Hsu, C.; MacCoss, M. J.; Noble, W. S.Abstract
One critical step in any targeted mass spectrometry experiment is selecting, from each protein of interest, a small number of peptides that respond well in the mass spectrometer and can serve as reliable proxies for protein quantification. Existing methods select target peptides either by relying on prior empirical measurements, limiting their applicability to previously observed peptides, or using machine learning to predict peptide behavior from sequence alone. However, current machine learning tools suffer from various limitations, including using detectability as an indirect proxy for intensity, relying on small training sets, or ignoring the precursor charge state. In this study, we introduce Bromo, a transformer-based deep learning model that ranks peptide precursors from a given protein by their relative response, taking charge state into account. Trained on millions of annotated peptide pairs derived from large-scale, publicly available data-independent acquisition mass spectrometry data, Bromo consistently outperforms existing sequence-based methods across diverse, independent datasets. Furthermore, we show that fine-tuning Bromo on experiment-specific data can account for differences in sample preparation, sample matrix, and instrument platform, all of which influence which peptides serve as optimal targets. This adaptability makes Bromo a practical tool for selecting target peptides for selected reaction monitoring and parallel reaction monitoring assay development across a wide range of experimental conditions.
bioinformatics2026-08-31v2A self-supervised DNA foundation model with collapse-resistant multimodal fusion
Chen, Y.Abstract
Genomic foundation models pretrained on DNA sequence have achieved strong performance across many tasks, but sequence-only representations cannot fully capture regulatory information from additional DNA-centric modalities. Existing multimodal genomic models are optimized for specific prediction tasks rather than reusable embeddings. Directly fusing heterogeneous modalities is challenging because sparse, peak-shaped regulatory signals and dense sequence embeddings have markedly different statistical structures, making naive alignment prone to near-zero solutions. We present a self-supervised DNA-centric multimodal foundation model integrating DNA sequence embeddings with local and global chromatin accessibility in a shared encoder to produce reusable window-level embeddings. We show that global normalization alleviates this collapse, enabling effective joint learning. The resulting embeddings improve regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection, achieving a 4.6-fold AUPRC improvement over the DNA-only baseline, with further gains on external ClinVar, GTEx eQTL and PBMC caQTL datasets.
bioinformatics2026-08-31v2Universal physical principles of protein structural genesis emerge in language-model representation space
Chuanyang, L.; Liu, J.; Qiu, X.; Wu, X.; Li, W.; Min, L.; Zhang, G.; Zhang, S.; Zhu, L.Abstract
Protein structure is usually treated as the endpoint of sequence by most AI models, yet its true biological emergence is a process of ordered change, whose logic remains hidden in opaque black box. ProtGenesis creates a bidirectional mirror world: it maps amino acid assembly, elongation and mutation in biological space into quantitative trajectories and ensembles in protein language model representation space, then translates spatial geometry into testable biophysical hypotheses. Across peptides, reporter proteins and protein families, three universal general principles emerged: hierarchical directional assembly (Principle I), quantitative and reproducible structural-emergence trajectories (Principle II) and discrete topological transitions (Principle III) from short- to long-range order. Three novel metrics of spatial features [D,{rho} ,{delta} ], complemented by squared Gaussian Wasserstein-2 measures, located structural anchors, sensitive regions and state transitions, enabling split-protein engineering and programmable protein design. ProtGenesis makes latent representations mechanistically interpretable, offering AI a route to discovering, rather than merely predicting, scientific principles. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=173 HEIGHT=200 SRC="FIGDIR/small/706798v2_ufig1.gif" ALT="Figure 1"> View larger version (60K): org.highwire.dtl.DTLVardef@18f690org.highwire.dtl.DTLVardef@e36903org.highwire.dtl.DTLVardef@35f15org.highwire.dtl.DTLVardef@1577b81_HPS_FORMAT_FIGEXP M_FIG C_FIG
bioinformatics2026-08-31v2The Metabarcoding Analysis Pipeline (MAP): Simple, accurate, and flexible metabarcoding
Prosser, S. W.; Bard, N. W.; Thompson, K. A.; Floyd, R. A.; Padhye, S.; Ozsahin, E.; Jafarpour, S.; Hebert, P. D. N.Abstract
Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP.
bioinformatics2026-08-31v2S2F-Agent: Harnessing sequence-to-function models for verifiable genome interpretation
Li, J.; Qin, T.; Li, J. G.; Bao, Z.Abstract
Sequence-to-function (S2F) models offer a revolutionary paradigm for genotype-phenotype mapping, yet their broader application is bottlenecked by the need for reliable orchestration and interpretation across a fragmented model ecosystem. While general-purpose language models can automate scientific workflows, they are not inherently grounded in the model-specific execution constraints required for robust S2F analysis. Here, we present S2F-Agent, a human-in-the-loop framework designed for the verifiable orchestration of the heterogeneous S2F ecosystems. The framework employs a contract-based harness to bridge model-specific capabilities (Skills) and model-agnostic biological objectives (Playbooks), seamlessly translating free-form biological requests into reliable execution and rigorous downstream interpretation. Evaluated on a benchmark of 54 query cases derived from published S2F workflows, S2F-Agent systematically outperformed general-purpose LLMs, demonstrating superior reliability accuracy in routing, groundedness, and end-to-end task execution success. We further demonstrate the robustness and scalability of S2F-Agent across model adaptation, variant interpretation, genome-scale functional profiling and personal-genome analysis. First, the agent autonomously adapts a genomic foundation model to quantitative chromatin profiles, resolving sequence features associated with primed and active regulatory states. Second, integrating multi-perspective variant effect predictions prioritized 42 high-priority candidate variants among CAD-associated variants (>16,000), and identified tissue-resolved regulatory mechanisms including the hepatic SORT1 axis. Third, genome-scale profiling of multiple traits GWAS atlas variants (>250,000) revealed pervasive context dependence in molecular consequences and regulatory architecture, highlighting the analytical focus toward fine-grained, tissue-specific regulatory variants. Finally, evidence-gated analysis of personal genomes expanded functional hypothesis generation beyond clinically annotated variants to thousands of prioritized candidates per individual while imposing explicit evidence-dependent boundaries on clinical claims. Collectively, these results establish S2F-Agent as a general framework for converting heterogeneous sequence-to-function capabilities into verifiable, scalable, and evidence-aware genomic analyses. By bridging the chasm between LLMs, specialized S2F ecosystems and rigorous genomic science, this framework democratizes the S2F paradigm for unlocking the full potential of these advanced models in real-world discoveries.
bioinformatics2026-08-31v2HESTIA: Scalable Multimodal Integration of Histology and High-Resolution Spatial Transcriptomics for Robust Spatial Domain Identification
Zhong, Z.; Zhu, X.; Guo, J.; Liao, S.; Chen, A.Abstract
Spatial omics has revolutionized molecular biology by providing invaluable insights into how native tissue microenvironments regulate cellular functions and disease mechanisms. Accurately capturing this structural complexity and decoding the underlying biological processes requires effectively integrating data from multiple modalities. However, transitioning to subcellular resolutions introduces massive data scales and severe transcriptomic sparsity, which challenge current analytical frameworks. To address this, we present HESTIA (Histology-Enhanced Scalable cross-Resolution inTegration for spatial trAnscriptomics), a highly efficient multimodal algorithm designed for identifying spatial domains in large-scale, high-resolution spatial omics data. By circumventing memory-intensive computations, HESTIA efficiently processes massive datasets on which existing algorithms fail due to memory constraints. HESTIA outperforms current multimodal methods in clustering accuracy and spatial continuity, accurately delineating fine structural boundaries. Furthermore, applying HESTIA to large-scale pathological samples successfully dissects clinically relevant intratumoral heterogeneity and maps distinct immune microenvironments in lung and colorectal cancers.
bioinformatics2026-08-31v2Lineage-specific X chromosome inactivation escape and skew underlie sex-biased immune gene dosage and deleterious variant exposure
Kavanagh, D.; Steel, A.; King, H. E.; Vieira, H. G. S.; Kumar, K. R.; Masle-Farquhar, E.; King, C.; Skvortsova, K.; Weatheritt, R. J.Abstract
The X chromosome carries an unusually high density of immune genes and is a major contributor to sex differences in immune function and autoimmune diseases. In females, X-chromosome inactivation (XCI) has two major functional consequences: it shapes X-linked gene dosage through XCI escape and determines the cellular exposure of heterozygous X-linked variants through XCI skew. Yet because XCI creates a mosaic of cells expressing different parental X chromosomes, these properties have remained largely inaccessible in individual women, becoming measurable only where XCI is non-random or after aggregation across large cohorts. Consequently, how X-linked variation contributes to sex-biased immunity and differs between individual women has remained unresolved. Here we present scDaisyChain, a graph-based framework that reconstructs chromosome-scale X haplotypes directly from heterozygous SNPs and single-cell long-read transcriptomes. scDaisyChain achieves near-ground-truth accuracy in highly polymorphic mouse hybrids and shows strong concordance with orthogonal long-read whole-genome phasing in human samples. Applied to peripheral blood immune cells from healthy women, it reveals a lineage-specific escape program in which lymphoid cells escape XCI more broadly than monocytes, with corresponding gains in the inactive X chromatin accessibility and female-biased expression. Lineage-specific skew further alters the proportion of cells expressing each heterozygous X-linked variant, a property we term variant exposure. Predicted deleterious variants are preferentially found in low-exposure states, exemplified by a splice-altering TLR8 variant expressed in few cytotoxic T cells. In rheumatoid arthritis (RA), the monocyte compartment - which has the lowest escape in health - shows reproducible inactive X dysregulation converging on a trained-immunity programme linked to disease flare and synovial macrophage activation, with elevated escape of IL13RA1 and HDAC8. These findings establish lineage-specific escape, skew and variant exposure as quantifiable, patient-resolved determinants of sex-biased immune gene dosage and X-linked variant penetrance in health and autoimmune disease, resolving a dimension of female biology that has been previously inaccessible in individual donors.
bioinformatics2026-08-31v1LRSPAT: A low-rank framework for spatial omics statistics
Frost, H. R.Abstract
We describe LRSPAT (low-rank spatial toolkit), a fast and memory-efficient framework for approximating measures of spatial association for high-dimensional data. While LRSPAT can be applied to any multivariate spatial dataset, development was motivated by the computational challenge of identifying spatially variable genes in high-resolution spatial transcriptomics (ST) data generated by technologies such as 10x Visium HD, Xenium and Atera. LRSPAT leverages a truncated SVD of the expression data and a thresholded spatial weights matrix to perform reduced-rank reconstruction of spatial statistics in the quadratic form family, including global and local versions of Moran's I, Geary's C, and Getis-Ord G. A regularization approach is leveraged to account for the inflated null distribution of spatial statistics computed on latent variables. By performing key operations on the low-dimensional embeddings, LRSPAT is orders of magnitude faster than standard implementations with significantly lower memory requirements. Because the low-rank approach denoises and desparsifies ST data, LRSPAT is also more accurate than standard techniques at identifying genes with true spatial expression patterns. The dramatic improvements in execution time and memory consumption enable the genome-wide analysis of spatially variable genes (SVGs) and exploration of the full range of hyperparameters including spatial scale, distance metric, and embedding rank. This preprint outlines the background and mathematical details of the approach with limited preliminary results and a short conclusion.
bioinformatics2026-08-31v1scPyviewer: a Python-native interactive viewer from AnnData single-cell data
Xuan, H.; Huang, Y.; Bian, J.; Liu, X.Abstract
Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.
bioinformatics2026-08-31v1Survey of the human proteostasis network: the ubiquitin-proteasome system
Elsasser, S.; Powers, E.; Stoeger, T.; Sui, X.; Kurtzbard, R. D.; Martinez-Botia, P.; Wangaline, M. A.; Gama, A. R.; Huttlin, E. L.; Elia, L. P.; Kelly, J. W.; Gestwicki, J. E.; Frydman, J. E.; Finkbeiner, S.; Clerico, E. M.; Morimoto, R.; Prado, M. A.; Vertegaal, A. C. O.; Hofmann, K.; Finley, D.Abstract
Modification by ubiquitination governs the half-lives of thousands of proteins that are fated for elimination by either the proteasome or autophagy pathways, depending on the intricate architectures of ubiquitin modification. This system mediates quality control for individual proteins, protein complexes, and organelles, as well as myriad purely regulatory functions. Here we provide a comprehensive survey of the ubiquitin-proteasome system (UPS), the scope of which is at present poorly defined. The UPS, with the inclusion of pathways involving ubiquitin-like modifiers, comprises in our estimate over 1430 distinct proteins in humans, a vast set of activities whose collective impact on the biology of the cell is pervasive. The UPS is an integral component of the proteostasis network (PN), the remainder of which we have also surveyed in recent studies. With the addition of molecular chaperones, proteins from autophagy-lysosome pathway, and related activities, the PN includes in total over 3150 components by our estimates. Comprehensive and systematic definition of these pathways should support a range of ongoing investigations in the areas of genomics, proteomics, biochemistry, cell biology, and disease research.
bioinformatics2026-08-30v3Trends in Machine Learning and Feature Selection Stability for Human Gut Microbiome (Shotgun Metagenomics) and Metabolomics Matched Datasets
Palmer, S. N.; Mishra, A. A.; Zarek, C. M.; Gan, S.; Wang, R.; Kim, J.; Liu, D.; Koh, A.; Zhan, X.Abstract
Microbiome research is often limited by methodological inconsistencies that reduce reproducibility and functional insight. Traditional taxonomy-based profiling is limited by sparse data, variable resolution, and reliance on gDNA sequencing which provides only indirect links to microbial function. Multi-omics integration offers a framework for linking community composition to functional outputs, but progress has been hindered by the lack of standardized frameworks and inconsistent use of machine learning. In particular, feature selection stability, which is central to biomarker discovery and experimental validation, remains underexplored. Here, we systematically benchmarked three widely used algorithms (Elastic Net, Random Forest, XGBoost) across seven multi-omics integration strategies and single-omics models. Additionally, we evaluated the impact of transforming metabolomics and taxonomic abundance data. Using human gut microbiome datasets that integrate metagenomic taxonomic profiles with metabolomics, we evaluated models for 9 binary and 8 continuous outcomes across 20 train/test splits per dataset. We further assessed the effect of feature reduction on both predictive accuracy and feature selection stability. Nonlinear learners were most consistently competitive: continuous outcomes favored metabolomics-dominant models, whereas binary outcomes favored stacked multi-omics models. Random Forest and XGBoost also yielded greater feature selection stability, particularly for full dimensional metabolomics data. Together, these findings demonstrate how integration strategy, algorithm choice, and data preprocessing jointly shape predictive performance and feature selection reproducibility in multi-omics microbiome modeling.
bioinformatics2026-08-29v4CIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations
Reisle, C.; Grisdale, C. J.; Krysiak, K.; Danos, A. M.; Khanfar, M.; Pleasance, E.; Saliba, J.; Hanos, M.; Patel, N. V.; Jain, A.; Seifi, M.; McMichael, J. F.; Venigalla, A. C.; Griffith, M.; Griffith, O. L.; Jones, S. J. M.Abstract
Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. Large language models (LLMs) offer a potential mechanism for accelerating biomedical claim verification, but their rapid turnover, variable availability, and known risks of unsupported reasoning require standardized and reproducible evaluation before integration into curation workflows. To address this, we developed CIViC-Fact, an expert-curated, full-text benchmark and evaluation framework. Domain experts linked structured cancer-variant claims to sentence-level evidence from source publications, including evidence from full-text articles, tables, and non-abstract sections that are commonly omitted from existing biomedical question-answering and scientific fact-checking datasets. Claim-verification reference labels were derived from CIViC records, revision histories and controlled data augmentation. A major finding of CIViC-Fact is that abstracts are insufficient for realistic biomedical claim verification. In the evaluated development subset of text-verifiable entries with full-text access, fewer than 30% could be fully validated from the abstract alone, highlighting the importance of full-text evaluation for biomedical curation. Upon the application of our fact-checking pipeline to newly submitted CIViC entries, after excluding entries requiring supplementary material or images for validation, automated retrieval successfully identified appropriate evidence for most cases (93%), supporting low-incremental-effort evaluation of future systems. Fine-tuning improved agreement with CIViC-Fact reference labels on the static benchmark, but larger general-purpose models performed better on a heterogeneous post-cutoff cohort. These findings support CIViC-Fact primarily as a reproducible framework for comparing evolving retrieval and verification systems rather than as validation of a single deployment-ready model. These findings suggest that, in a rapidly changing model landscape, the durable contribution is not a single optimized model but a reproducible benchmark framework that enables continual testing, model substitution, and lightweight updating through small high-quality few-shot exemplar sets.
bioinformatics2026-08-29v3Genome-Wide in silico analysis reveals activation of a silent resistome driving imipenem resistance in Pseudomonas aeruginosa
Anwar, S.; Aromal, A. R.; Anurag Anand, A.; Samanta, S. K.Abstract
Resistance to imipenem in Pseudomonas aeruginosa relies on multiple factors that remain poorly understood. In our work, we performed a systemic analysis of genome-wide changes involved in resistance in a set of 95 clinically unrelated strains, including 41 resistant (MIC [≥] 64 mg/L) and 54 susceptible (MIC [≤] 2 mg/L) isolates. Our approach is based on the pan-genomics analysis, combining the use of core-genome phylogenetic analysis, MLST (Multilocus Sequence Typing), GWAS (Genome-Wide Association Studies) and variant level profiling of the blaOXA genes. Higher-order structure within the set was studied using methods of the co-occurrence networks and WGCNA (weighted gene co-expression network analysis) specifically adjusted to handle presence/absence data. Despite having a broader and more diverse resistome, no clonal grouping of the resistant isolates was observed indicating independent evolutionary origins. The LASSO model using a lineage-aware approach showed robust predictive capability (AUC = 0.836) that validates the polygenic characteristic of resistance. Twelve accessory genes were found to be significant determinants of resistance; however, only four genes (group_10880, group_10887, group_4947, and phzB) were identified using both GWAS and gene network analysis, showing involvement in protein folding, metal stress response, genome plasticity, and metabolic adaptation. Interestingly, some carbapenemase-active variants of blaOXA were also found in imipenem-susceptible strains, showing that gene presence alone does not ensure resistance. We therefore propose the Silent Resistome Activation Model, where resistance genes become functional only with support from identified accessory genes and coordinated interactions at both the genomic and network levels.
bioinformatics2026-08-29v2FuncSeek: Multi-PLM contrastive learning for protein functional similarity search
Cloete, L. J.; Patterton, H. G.Abstract
Below 30% pairwise sequence identity, alignment-based methods struggle to reliably distinguish true homologs from chance, and enzyme function prediction degrades accordingly: on proteins in this regime, even advanced methods (CLEAN) achieves only 55.1% accuracy at full EC specificity on the CARE benchmark. To this end Protein Language Models (PLMs) have gained favor as alternatives. However, PLMs often encode only a subset of the biology, whereas the understanding of enzyme function requires among other things a combination of sequence, structure and functional-context simultaneously. In this work, we describe FuncSeek, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities). Using SwissProt data, these 2816-D embeddings are labelled with Enzyme Commission numbers (EC) and are trained through a supervised contrastive head into a 256-D space. FuncSeek attains 64.6% nearest-neighbour EC4 accuracy on the CARE out-of-distribution benchmark set (ood30; proteins below 30% identity to training set), outperforming CLEAN (55.1%) and Diamond BLASTp (51.4%), and obtains 93.7% nearest-neighbour EC4 accuracy on the promiscuous, multi-functional enzymes benchmark (CLEAN, 69.4%). We also show that the learned representations transfer without retraining to the TrEMBL database, achieving 97.3% nearest-neighbour EC4 accuracy on an 8,031 BRENDA-validated enzyme set, never seen during training. Because only projected embeddings are stored in the target index, and function is inferred from an annotated reference set, we propose this paradigm for rapidly searching extremely large metagenomic databases, bypassing costly sequence alignment and annotation pipelines.
bioinformatics2026-08-29v2HANSEN: An Integrated Structural and Functional Proteome Resource for Structure-Guided Drug Discovery in Mycobacterium leprae
Vedithi, S. C.; Rees, R.; Malhotra, S.; Munir, A.; Matusevicius, M.; Alsulami, A. F.; Beaudoin, C. A.; Sunkara, K. S.; Das, M.; Blundell, T. L.; Floto, R. A.Abstract
Leprosy remains a leading infectious cause of preventable disability, yet its causative agent, Mycobacterium leprae (M. leprae), is structurally under-characterised. Only ten Protein Data Bank (PDB) entries represent seven of its 1,603 protein-coding genes. We present HANSEN, a proteome-wide structural and functional resource for M. leprae. Monomeric and oligomeric models were generated with AlphaFold 3, Boltz-1, Boltz-2 and Chai-1, and annotated with per-residue confidence, predicted aligned error and, for assemblies, interface confidence. Ligand-binding pockets were predicted with AF2BIND, P2Rank and fpocket, template-derived ligands were modelled within oligomeric complexes, residue-level B-cell epitope propensity was estimated with DiscoTope-3.0, and gene essentiality was transferred from Mycobacterium tuberculosis transposon-sequencing labels. These features are integrated in a relational database with interactive visualisation and combined into a calibrated Target Priority Score that ranks all 1,603 proteins into four tiers and recovers established antimycobacterial targets. HANSEN (https://hansen-leprosy.medschl.cam.ac.uk/home) provides a practical basis for target prioritisation and structure-guided drug discovery in leprosy.
bioinformatics2026-08-29v2Bacterial metagenome in plaque, saliva, and tumor samples from individuals with and without OSCC by next-generation sequencing
ERIRA, A.; ROBAYO, D. A. G.; GAMBOA, F.; CHALA, A.; MORENO, A.; ARREGUI, A. C.; MUNOZ, E.; NOGUERA, J.; TOBAR-TOSSE, F.Abstract
Background: Oral dysbiosis has been associated with oral squamous cell carcinoma (OSCC); however, most microbiome studies rely on 16S ribosomal RNA (rRNA) gene sequencing, limiting species-level taxonomic resolution. Methods: Dental plaque, saliva, and tumor tissue samples from 10 patients with OSCC and dental plaque and saliva samples from 10 healthy controls were analyzed in this exploratory cross-sectional study. DNA was extracted and subjected to shotgun metagenomic sequencing using the Illumina MiSeq platform. Sequence reads were quality filtered with fastp, taxonomically classified using Kraken2 v2.1.3, and species-level abundances were re-estimated with Bracken v2.9 following the removal of human reads and low abundance taxa. Relative abundances were compared using the Mann Whitney U test with the Benjamini Hochberg false discovery rate correction, while the Bray Curtis principal coordinate analysis was used as an exploratory approach to visualize microbial community patterns. Results: Shotgun metagenomic sequencing revealed distinct bacterial community profiles across the oral microenvironment. Dental plaque exhibited the highest taxonomic diversity and relative abundance. The control plaque was enriched in Streptococcus koreensis, Capnocytophaga sp. oral taxon 878, Treponema sp. Marseille Q4132, and Leptotrichia sp. oral taxon 498, whereas the plaque from patients with OSCC showed a higher relative abundance of Pyramidobacter piscolens, Parvimonas parva, and Gemella sanguinis. Salivary samples displayed lower diversity and a more homogeneous composition, predominantly comprising Capnocytophaga endodontalis, Prevotella jejuni, Aggregatibacter aphrophilus, and Gemella sanguinis. The tumor tissue showed relatively higher abundance of Sellimonas catena, Escherichia coli, Solobacterium moorei, and Lacrimispora sp. HJ 01. Conclusions: This exploratory study provides species-level characterization of the oral microbiome across multiple oral microenvironments in OSCC and generates hypotheses for future integrative metagenomic and functional studies investigating the potential contribution of oral bacterial communities to OSCC pathogenesis.
bioinformatics2026-08-29v1Quantifying the Rearrangement Complexity of Pangenomes
Bohnenkaemper, L.; Stoye, J.Abstract
The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.
bioinformatics2026-08-29v1When AI encounters natural history: Morphological OTUs reshape our understanding of Earth's life
Zhan, Z.; Ye, M.; Orr, M. C.; Chen, W.; Liu, X.; Yue, L.; Sun, X.; Zhang, F.Abstract
Biodiversity can be quantified only after organisms are assigned to reproducible units, yet most individuals encountered in nature lack reliable species-level identifications. Molecular operational taxonomic units can organize unnamed diversity, whereas image-based approaches generally depend on predefined species levels. Here, we show that operational biodiversity units can be derived directly from phenotypes. We developed morphOTU, a framework combining self-supervised representation learning, operational metric supervision, and adaptive hierarchical clustering to organize specimen images in continuous phenotypic space. Across five benchmark datasets comprising flowers, wood anatomy, and beetle habitus, morphOTUs recovered coherent species-level structure and produced -diversity estimates close to those obtained from expert identifications. This structure remained informative when species were excluded from representation learning, when labeled data were sparse, and per-species sampling was limited. In a heterogeneous field-survey dataset of 4,717 insects representing 269 species across 12 orders, fine-tuning on only 28 common species produced diversity estimates close to expert labels (Shannon index, 3.75 versus 3.53). Visual explanations localized variation to biologically meaningful structures, including body outliers, surface sculpture, floral symmetry, and wood vessels. Morphological units therefore provide an operational layer for organizing and quantifying biodiversity before, alongside, or in the absence of formal species names.
bioinformatics2026-08-28v3Can SMILES be fragmented into a concatenable ordered sequence of retrosynthetically interesting string block ?
Reboul, E.; Prabakaran, H.; Waldispuhl, J.; Taly, A.Abstract
Molecules generated by deep learning models are often difficult to synthesize. Their synthetic accessibility can be improved with automated retrosynthetic analysis, which allows for identifying synthons. However, synthons in a SMILES can be scattered throughout the string depending on the path taken through the molecular graph used to generate the SMILES. We tested whether the ensemble of possible SMILES for a molecule can be used to generate a concatenable ordered sequence of string fragments (blocks) from SMILES that match potential synthons obtained through automated retrosynthetic analysis. We found that exhaustively sampling the SMILES space of a molecule improves the coverage of retrosynthetic breaks. We achieved full coverage of retrosynthetic bonds in string form for 85\% of the 1.9 million molecules in the MOSES dataset. Doing so allowed us to test our block SMILES in an unconditional de novo drug design test case with MolGPT and Monte Carlo Tree Search (MCTS). We found that using blocks as an LLM's token did degrade MolGPT performance due to the curse of dimensionality. However, using the SMILES selected by our blocking algorithm with the default SMILES tokenizer improved the reproduction of physico-chemical properties of samples and also improved uniqueness, novelty, and validity. The MCTS outperforms our MolGPT models in terms of validity and novelty. However, samples generated by the MCTS had physico-chemical properties that were further away from the MOSES baseline than the samples produced by molGPT, with an improved distribution of quantitative estimation of drug-likeness (QED).
bioinformatics2026-08-28v2Evaluating Aggregated Gene Level eQTL Scores
Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.Abstract
Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.
bioinformatics2026-08-28v2Gene expression inference from cell-free DNA using uncertainty-aware deep learning
Patton, R. D.; McDeed, A. P.; Netzley, A.; Pawar, A.; Persse, T. W.; Nair, A.; Galipeau, P. C.; Coleman, I. M.; Itagi, P.; Chandra, P.; Sayar, E.; Adil, M.; Vashisth, M.; Hiatt, J. B.; Dumpit, R.; Kollath, L.; Demirci, R. A.; Ghodsi, A.; Lam, H.-M.; Morrissey, C.; Chen, D. L.; Schweizer, M. T.; Iravani, A.; Hsieh, A. C.; MacPherson, D.; Haffner, M. C.; Nelson, P. S.; Ha, G.Abstract
Tumor gene expression profiling provides crucial diagnostic information for guiding therapy, but standard tissue biopsies are invasive, spatially biased, and may inadequately sample metastatic disease. Cell-free DNA (cfDNA) provides a minimally invasive alternative for tumor genotyping, yet reconstructing robust, transcriptome-wide expression from standard-depth cfDNA whole-genome sequencing (WGS) remains a major challenge. We developed a deep learning framework comprising Triton, for comprehensive cfDNA feature extraction, and Proteus, a probabilistic model that infers single-gene expression from standard-depth cfDNA WGS. Proteus outperformed prior cfDNA approaches in reconstructing molecular phenotypes from matched tumor transcriptomes across multiple cancer types, including prostate, lung, and bladder cancer cohorts, with uncertainty-guided withholding improving model reliability. Proteus further enabled assessment of therapeutic target activity, prognostic transcriptional programs, and candidate treatment-emergent resistance states, establishing a generalizable framework for minimally invasive functional genomics in precision oncology.
bioinformatics2026-08-28v2OmicsFM brings proteomics into the foundation model era
Heyndrickx, S.; Gabriels, R.; Ramadasan, H.; Martens, L.; Claeys, T.Abstract
While foundation models have been shown to learn biological representations from large transcriptomic atlases, it remained unknown whether proteomics data allow the same. We here therefore introduce OmicsFM, a modality-agnostic transformer pretrained through masked abundance reconstruction on an unprecedented proteomics data corpus of 48,837 quality-filtered proteomics profiles from 1,397 reprocessed PRIDE projects. Interestingly, despite training on 14- to 93-fold fewer profiles than matched bulk- and single-cell transcriptomic models, respectively, our proteomics model rivals both. On held-out projects, OmicsFM attention networks recovered more molecular relationships than co-expression methods and existing single-cell foundation models across nine reference databases that reveal pathway-level organization. Sample-level embeddings preserved biological structure across independent studies, and its representations transferred successfully to cell-type classification, gene-essentiality prediction, and perturbation-response prediction, while consistently outperforming task-specific models. Moreover, our results show that proteomics and transcriptomics representations capture complementary biology. OmicsFM thus firmly establishes the possibility of training highly performant proteomics-based foundation models, and their importance in modelling and uncovering fundamental biology.
bioinformatics2026-08-28v1A pretrained unified model enables cellular functional profile prediction and multi-objective virtual drug screening
Chen, R.; Huang, L.; Qiao, Y.; Mandal, S.; Mo, L.; Li, L.; Leshchiner, D.; Zhang, X.; Pu, J.; Xie, Y.; Girgis, R.; Ellsworth, E.; Huang, L.; Chen, X.; Li, X.; Zhou, J.; Chen, B.Abstract
Cells are characterized by molecular states, coordinated molecular interactions, regulatory programs, and responses to perturbations. Systematic mapping of these cellular functional profiles across biological contexts remains experimentally costly and fragmented. Here we present InsilicoCell, a pretrained multi-modal, multi-task model that unifies prediction of cellular functional profiles spanning molecular states, molecular interactions, and perturbation-induced responses. Built on a supervised transformer architecture and pretrained on more than 88 million measurements across seven tasks, including drug sensitivity, drug-induced gene expression, and drug-protein binding, InsilicoCell learns a shared representation that links molecular profiles to cellular phenotypes, improves performance over task-specific models, and generalizes to unseen entities, contexts, and conditions. InsilicoCell extends beyond cell line systems to patient, spatial and single-cell settings, and enables multi-objective virtual drug screening. It identifies novel candidate compounds with experimental validation, including c-Myc activity inhibitors, antifibrotic agents and stemness-inducing compounds. Together, InsilicoCell provides a scalable framework for predictive cellular biology and therapeutic discovery.
bioinformatics2026-08-28v1