Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
Ding, H.; Nie, W.; Wu, N.; Qiu, T.Abstract
Hybrid DNA foundation models combine convolutional, recurrent, and attention layers, making speculative decoding more difficult than truncating a KV cache. We present EvSpark, a speculative decoding system for Evo2 that verifies draft blocks in parallel and restores all three classes of inference state by selecting retained intermediate states, without replay. A compact, hidden-state-conditioned drafter proposes each block in one parallel forward pass. On Evo2 7B, a 48-prompt benchmark with three training seeds yields 2.96x on 43 real-sequence prompts and 3.27x including the five synthetic controls. Acceleration persists at 262k-token context (1.84x - 2.43x on two bacterial genomes) and over 32k generated tokens. Retraining the same drafter architecture for Evo2 20B and 40B yields 2.18x - 2.46x on real sequences and 2.51x - 2.78x on the full suite. Autoregressive drafter comparisons and batch measurements show why low draft latency, rather than acceptance alone, determines the gain. The method preserves the target distribution in exact arithmetic. In bf16, greedy tests find no non-tie divergences across 48 prompts and six checkpoints; sampling tests expose residual numerical sensitivity, especially in repetitive sequences. A cost-efficient 7B drafter requires 1.06 incremental GPU-hours of training, excluding teacher-data collection, and achieves 2.82x on real sequences. In regulatory-DNA design, EvSpark achieves a median complete-workflow speedup of 1.57x over a calibrated batched native baseline.
bioinformatics2026-09-21v3Product-stabilized filamentation by human glutamine synthetase allosterically tunes metabolic activity
Greene, E.; Muniz, R.; Yamamura, H.; Hoff, S. E.; Bajaj, P.; Tecson, M. C. B.; Geluz, C.; Lee, D. J.; Thompson, E. M.; Arada, A.; Lee, G. M.; Bonomi, M.; Kollman, J. M.; Fraser, J. S.Abstract
To maintain metabolic homeostasis, enzymes must adapt to fluctuating nutrient levels through mechanisms beyond gene expression. Here, we demonstrate that human glutamine synthetase (GS) can reversibly polymerize into filaments aided by a composite binding site formed at the filament interface by the product, glutamine. Time-resolved cryo-electron microscopy (cryo-EM) confirms that glutamine binding stabilizes these filaments, which in turn exhibit reduced catalytic specificity for ammonia at physiological concentrations. This inhibition appears induced by a conformational change that remodulates the active site loop ensemble gating substrate entry. Metadynamics ensemble refinement revealed >10 [A] conformational range for the active site loop and that the loop is stabilized by transient contacts. This disorder is significant, as we show that the transient contacts which stabilize this loop in a closed conformation are essential for catalysis both in vitro and in cells. We propose that GS filament formation constitutes a negative-feedback mechanism, directly linking product concentration to the structural and functional remodeling of the enzyme.
bioinformatics2026-09-21v3Covariance Nonstationarity is Evident in Spatial Transcriptomics and Provides a New Categorization of Spatially Varying Genes
Velidi, P.; Wei, Z.; Nathoo, F.Abstract
Gaussian process models underlie many spatial transcriptomics tools but typically assume stationary covariance. Covariance non-stationarity has long been recognized in spatial statistics as an important feature of spatial data, yet it has received little attention in spatial transcriptomics. We show that this omission is consequential: covariance non-stationarity is substantially evident across spatial transcriptomic datasets and alters the characterization of spatially varying genes. While typically ignored, non-stationarity of spatial covariance in gene expression may correspond to tissue heterogeneity or cell aggregates. Across 12 Visium datasets, we use approximate Bayes factors from R-INLA to compare stationary and non-stationary Mat'ern covariance functions. Evidence for covariance non-stationarity appears in 3% to 50% of genes across tissue samples. We find that gene sets associated with immune, cytokine, and other effector functions are enriched among genes favoring non-stationary spatial covariance. Covariance stationarity is therefore not a benign technical simplification in spatial transcriptomics; it is frequently violated, the violation is biologically structured, and it changes the definition and classification of spatially varying genes.
bioinformatics2026-09-21v3Integrating structural homology with deep learning to achieve highly accurate protein-protein interface prediction for the human interactome
Xiong, D.; Torres, M.; Murray, D.; Zhang, Z.; Li, L.; Naravane, A. C.; Fragoza, R.; Honig, B.; Yu, H.Abstract
A significant portion of disease-causing mutations occur at protein-protein interfaces however, the number of structurally resolved multi-protein complexes is extremely small. Here we present a computational pipeline, PIONEER2, that integrates 3D structural similarity with geometric deep learning to accurately predict protein binding partner-specific interfacial residues. We compare the performance of PIONEER2 to that of AlphaFold3 and found, using a test set of PDB structures, that their performance is quite similar. However, about 20% of AlphaFold3 predictions for protein-protein complexes in the PDB have AlphaFold3 ranking scores below 0.5, which indicates an uncertain model. For these structures, PIONEER2 outperforms AlphaFold3 at discriminating interfacial from non-interfacial residues. Further, about half of the AlphaFold3 ranking scores on high confidence protein-protein interactions (PPIs) not associated with a PDB structure are below 0.5 indicating that PIONEER2 offers superior interface prediction for a large number of PPIs for which structures are not available. We created a comprehensive 3D structurally informed interactome encompassing all 352,124 experimentally detected binary human PPIs in the current literature and made PIONEER2 interface predictions for each. We experimentally validated these predictions by generating 1,866 mutations and testing their disruptive impact on 5,010 mutation-interaction pairs. PIONEER2-predicted interfaces are found to be comparable to PDB structures in their ability to predict disruptive mutations while AlphaFold3 performance is reduced. Similarly, PIONEER2-predicted interfaces outperform AlphaFold3 in accounting for the depletion of non-deleterious common population variants and the enrichment of disease-related mutations on protein surfaces. Overall, our results suggest that PIONEER2-predicted interfaces provide a valuable tool for studying disease etiology, advancing personalized medicine and for fundamental research. We further implemented PIONEER2 as a user-friendly web server (https://pioneer2.yulab.org) platform for users to explore our 3D interactome models and conduct genome-wide functional genomics studies.
bioinformatics2026-09-21v2Benchmarking confidence estimation and rescoring for cyclic peptide-protein complex predictions
Li, Z.; Yuan, Y.; Hu, K.; Pan, P.; He, F.Abstract
Cyclic peptides are a rapidly expanding class of therapeutics, but the reliability of deep-learning structure prediction for cyclic peptide-protein complexes has not been systematically evaluated. We assembled a curated benchmark of 111 nonredundant complexes spanning five cyclization chemistries and assessed two co-folding models, Boltz and Protenix, each generating 100 poses per target (22,200 total). Stratifying all poses by complex attributes, we found that disulfidecyclized peptides and small protein targets (200 or fewer target residues) were predicted significantly worse by both tools, with target size the largest and most consistent effect; overall accuracy nevertheless remained high (median top-pose DockQ of about 0.89, 96-98% of targets Acceptable or better), indicating that pose generation is rarely the bottleneck. Conversely, native model ranking scores correlated only moderately with pose quality (Spearman rank correlations of 0.53-0.66): approximately 12% of poses showed high model ranking score/confidence despite poor pose DockQ quality, and the highest-quality pose was not ranked first for nearly every target. We therefore augmented the native score with externally computed interface descriptors normalized by chain length, principally the per-residue density of inter-chain hydrogen bonds, in a gradient-boosted rescoring model evaluated under target-grouped cross-validation that prevents leakage, improving out-of-fold ROC-AUC for both tools, significantly so for Protenix. Together, these findings identify pose ranking, rather than pose generation, as the major limitation of current cyclic peptide-protein complex prediction and demonstrate that complementary structural features can improve confidence-based pose selection.
bioinformatics2026-09-21v2Is level-1 blob reconstruction under the network multispecies coalescent easy?
Dai, J.; Molloy, E.Abstract
Hybridization is an important evolutionary process, commonly modeled by the network multispecies coalescent. Reconstructing evolutionary histories under this model is notoriously costly, even for level-1 networks where hybridization events are isolated from each other. The widely used methods that combine speed with statistical guarantees rely on quartet concordance factors computed for all subsets of four species, resulting in an O(n^4k) bottleneck that severely limits scalability to large numbers of species (n) and genes (k). Among quartet-based methods, NANUQ+ is notable because it decomposes the problem into two steps: first reconstructing a tree of blobs, which compresses each non-treelike part of the network, called a blob, into a single vertex, and second reconstructing the internal structure of each level-1 blob, specifically its circular order and hybrid vertex. Here, we investigate whether level-1 blob reconstruction is difficult once the tree of blobs is known. We present a fast and statistically consistent algorithm, called NetCS, based on two simple primitives: majority voting and merge sort, circumventing the bottleneck of computing all quartet concordance factors. In simulations, NetCS achieved comparable accuracy to NANUQ+ and was dramatically faster, enabling analyses of 200 taxa and 1000 genes in only a few minutes. Both methods attained near-perfect accuracy when given the true tree of blobs; however, their performance degraded in end-to-end pipelines due to errors in tree of blobs reconstruction. Strikingly, even methods that reconstruct level-1 networks directly struggled to accurately predict hybrid ancestry. Our results suggest that reconstructing level-1 blobs is unexpectedly easy once the tree of blobs is known, and that a major challenge for phylogenetic network inference lies in accurate tree of blobs reconstruction.
bioinformatics2026-09-21v2Optimized Multiple Circular Sequence Alignment for Cyclic Peptide Motif Discovery
Yuan, Y.; Li, Z.; Hu, K.; Pan, P.; He, F.Abstract
Head-to-tail (H2T) cyclized peptides are an increasingly important modality in drug discovery, combining high target affinity and selectivity with metabolic sta-bility. Because their underlying chemistry is still that of a linear amino-acid chain, their linear sequence representation is the native input format of main- stream sequence generative models now driving de novo peptide design (Slough et al., 2018; Rettie et al., 2025a;b). Discovering the conserved motifs responsible for a family's function across a library of such candidates requires a multiple sequence alignment (MSA). Because a cyclic peptide can be linearised at any residue, the alignment must additionally solve for the unknown rotation of each sequence, which is the multiple circular sequence alignment (MCSA) problem. However, leading MCSA heuristics (e.g. Ayad & Pissis, 2017) were tuned for the genomic regime (a few tens of long sequences) and become prohibitively slow on the cyclic peptide library regime (hundreds to thousands of shorter sequences). We close this gap by identifying quality-preserving optimisation opportunities, notably the library-scale preset tailored to short-sequence inputs (algorithmic details in Appendix A), and by adding an orthogonal multi-core and SIMD backend for further performance tuning, which gives near-linear thread scaling on the pairwise-comparison stage. We validate the pipeline on a library of 1,000 H2T cyclized peptides of length 18 targeting the oncoprotein Mouse double minute 2 human homolog (MDM2) produced by an internal peptide-design engine. In this practical setup, the optimised MCSA recovers the underlying positional motif of MDM2 binders at the same fidelity as the original MCSA implementation while running over 650x faster. Our optimised MCSA tool thus enables library-scale cyclic peptide sequence alignment and is publicly available at https://github.com/IVB-Generative-Biology/mars-turbo.
bioinformatics2026-09-21v2Pitfalls in understanding PhIP-Seq data: technical variability, replicability, and best practices for interpretation
Trgovec-Greif, L.; Vogl, T.Abstract
Phage Display Immunoprecipitation sequencing (PhIP-Seq) is a high throughput method allowing to measure antibody binding against hundreds of thousands of potential antigens. Typically, blood samples of hundreds to thousands of individuals are measured in 96-well plates, necessitating distribution of samples on multiple plates and processing in batches. To correctly interpret PhIP-Seq results, it is crucial to understand its nature and technical limitations. Here, we analyzed 195 technical replicates of controls (anchor samples) from 49 plates measured in 13 immunoprecipitation runs to gauge technical variability and replicability. While there were few false-positive enriched peptides, we observed a substantial fraction of false negatives. The total number of enriched peptides can be affected by batch effects. Additionally, the type of biological material used (serum or plasma) influences enrichment profiles with an antigen library containing bacterial proteins due to the interaction of phages with plasma proteins. Finally, replicating measurements in different labs produces comparable enrichment profiles, but slight differences in a subset of peptides are noticeable. These results suggest best practices for PhIP-Seq experiments. Most importantly, due to batch effects, samples of cases and controls should be evenly distributed between 96-well plates to avoid confounding with biological effects. Also, for highly complex libraries, the total number of enriched peptides can be a technical artifact and potentially not a feature useful in comparisons.
bioinformatics2026-09-21v1Discriminating rare disease cases from their controls based on observed and excluded phenotypes
Guo, Y.; Yao, Y.; Zhang, Y.; Duan, G.; Zhang, N.; shan, g.Abstract
Rare diseases are individually uncommon but collectively prevalent. Their primary clinical challenge lies not in treatment but in diagnosis. In the early stages of clinical management, it is frequently unclear whether the observed phenotypes are associated with a rare disease. Leveraging machine learning methods to mine latent associations between these phenotypes and rare diseases for early diagnosis offers a viable strategy to alleviate this diagnostic dilemma. In the present study, we demonstrated that machine learning can effectively discriminate rare disease cases from their controls using observed and excluded phenotypes as features. Among them, the Random Forest model achieved the best classification performance with a certain degree of generalizability. Based on further analysis of the feature selection results, we conclude that the two factors, specificity and occurrence count, are important for phenotype selection in rare disease discrimination, and comparable importance should be attached to both observed and excluded phenotypes, during feature construction.
bioinformatics2026-09-21v1POAnoise: A Graph-based Denoising Pipeline for Amplicon Sequencing Data
Ardiyansyah, M.; Furneaux, B.; Ovaskainen, O.Abstract
High-throughput DNA metabarcoding enables large-scale biodiversity assessment by identifying taxa from environmental samples, but its accuracy critically depends on denoising methods that separate true biological variation from PCR and sequencing errors. A persistent challenge is robust reconstruction of sequence diversity across abundance distributions, where low-abundance variants are particularly difficult to recover. We introduce POAnoise, a graph-based denoising framework that uses Partial Order Alignment (POA) to model relationships among noisy sequencing reads. POAnoise incrementally constructs sequence graphs that represent substitutions and indels, and derives consensus sequences from graph-supported paths using a weighted consensus strategy. By combining graph-based alignment with abundance-aware clustering, the method provides a structured way to reconstruct sequence variants from noisy amplicon data across heterogeneous abundance regimes. We evaluated POAnoise on simulated ITS and 16S datasets and compared its performance with established denoising methods, DADA2 and UNOISE3, across multiple parameter settings. Across the benchmark datasets, POAnoise generally achieved higher F1-scores and exhibited more stable performance across parameter configurations. In abundance-aware analyses, POAnoise showed reconstruction ratios closer to unity and reduced abundance-dependent deviation compared with DADA2, while remaining broadly comparable to UNOISE3 across most abundance classes. Overall, these results indicate that POAnoise can provide a robust alternative for amplicon denoising.
bioinformatics2026-09-21v1A single-cell RNA-seq catalog of ground truth gene coregulation
Svensson, V.Abstract
Single-cell RNA-sequencing measurements are uniquely well-poised to identify coregulated gene transcription. Coregulation should be apparent as correlation of expression levels, but at which unit to quantify expression for calculation of correlations is not clear. To enable evaluation of normalization methods for identification of coregulation we have curated and preprocessed 35 scRNA-seq samples with known regulature relationships into a catalog. These datasets contain promoter-reporter data representing known positive coregulation and B cells where isotypic mutual exclusion of {kappa} and {lambda} light chains represent known negative coregulation. As an example, we use the catalog to evaluate seven normalization methods. Different normalization methods perform better for positive or negative coregulation, and results indicate implementation choices for advanced normalization methods The catalog may serve as a resource for evaluating novel proposed normalization methods and is available at https://huggingface.co/datasets/valsv/scrna-coregulation-benchmark
bioinformatics2026-09-21v1Detecting cell segmentation errors using doublet methods
Sarwar, A.; Gillis, J.Abstract
In spatial transcriptomics, cell segmentation is used to draw boundaries around cells. Molecules located within a cell's boundary are assigned to it, making its gene expression profile dependent on segmentation accuracy. To identify potentially problematic cells, studies increasingly use doublet detection methods, a class of algorithms developed to recognize molecular admixture from two cells in scRNA-seq. To evaluate their suitability in spatial data, we model cell segmentation errors as a continuum of partial molecular admixture between neighboring cells, generated by varying the loss of a cell's own transcripts and the gain of transcripts from its neighbor. Evaluating 8 doublet methods across 16 spatial datasets, we characterize the conditions under which they perform well and identify their failure modes. Detection improves with increasing molecular admixture and transcriptional dissimilarity between neighboring cells, with cxds2, scDblFinder.score, and a simple baseline (unique_genes) performing best. These patterns are consistent across datasets spanning different gene panels, technological platforms, staining techniques, segmentation algorithms, and error frequencies. In unperturbed spatial data, elevated doublet scores localize to regions consistent with segmentation problems, suggesting that they can help prioritize cells or regions for inspection, segmentation refinement, or transcript reassignment. Our results establish when doublet methods can provide useful quality control signals for cell segmentation, supporting their use in spatial transcriptomics.
bioinformatics2026-09-21v1Serum metabolomics reveals signatures associated with physical resilience trajectories from middle to older age
Seo, J. I.; Scheurink, T. A. W.; Wang, C. X.; Kvitne, K. E.; Nunes, W. D. G.; Zemlin, J.; Bergstrom, J.; Dorrestein, P. C.; Mohanty, I.; Molina, A. J. A.Abstract
Lifecourse physical resilience is defined by the ability to maintain abilities across multiple domains of physical performance. While the importance of physical resilience in functional independence and mobility disability is clear, studies investigating metabolomic signatures of physical resilience are lacking. Here, we performed untargeted metabolomics on serum samples from a community-based cohort of 237 individuals followed over 28 years, and applied spectral data mining tools to map identified metabolites to health phenotypes from public repositories. We identified metabolites across multiple chemical classes, including acylcarnitines, glutamine conjugates, and phosphocholines, that were differentially associated with physical resilience status. Notably, medium-chain acylcarnitines negatively associated with physical resilience were more frequently observed in disease phenotypes than in healthy individuals. Kynurenine, a tryptophan metabolite linked to age-related functional decline, increased more steeply with age in individuals with low physical resilience. We also found that metabolites of the antihypertensive drug verapamil were associated with physical resilience in a metabolism-dependent manner, differing between oxidative and glucuronidated forms. Together, these metabolic signatures offer a resource for identifying biochemical pathways and biomarkers relevant to physical resilience for healthy aging.
bioinformatics2026-09-21v1MaSkel and napari-MaSkel: fast morphological skeleton feature extraction from segmentation masks
Wittmann, S.; Pysch, D.; Uderhardt, S.; Blumenthal, D. B.; Moeller, A.Abstract
Summary: Existing workflows for extracting morphological features from network-like biomedical structures often require several heterogeneous tools, which can limit integration and reproducibility. We present MaSkel and napari-MaSkel, open-source Python tools that provide GUI- and CLI-based workflows for 2D and 3D skeletonization and graph-based feature extraction, using a substantially accelerated implementation of the widely used Lee94 thinning algorithm. Availability and Implementation: MaSkel and napari-MaSkel are open-source Python packages available under the MIT license on PyPI and GitHub: https://github.com/bionetslab/maskel/, https://github.com/bionetslab/napari-maskel/. Documentation: https://bionetslab.github.io/maskel/, https://bionetslab.github.io/napari-maskel/. Benchmarking and validation: https://github.com/bionetslab/maskel-evaluations.
bioinformatics2026-09-21v1ARCHER-LD: Rapid Long-Range Linkage Disequilibrium Calculations at Biobank Scale using GPU Acceleration
Kumar, R.; Singhal, P.; Zhang, D.; Carson, C.; Conery, M.; Rodriguez, A. A.; Nandi, T.; Ritchie, M.; Thavappiragasam, M.; Voight, B. F.; Madduri, R.; Verma, A.Abstract
Linkage disequilibrium (LD) information from one's own dataset is considered optimal for downstream analyses such as statistical fine-mapping. However, the computational complexity ({approx}N2 / 2 computations for N variants) lead most studies to use external reference panels, such as from 1000 Genomes. To capture LD across all variant pairs in biobank-scale whole-genome sequencing (WGS) datasets with hundreds of millions of variants, new computational strategies are essential. We present a novel approach that uses multi-GPU distributed computing to compute R2 for every variant pair in a dataset. On chromosome 22 (1.8 million variants) of the 30x WGS 1000 Genomes dataset, our method using eight consumer-level 12GB GPUs (NVIDIA RTX 2080TIs) takes <10 minutes, while the same calculation with PLINK using a high-end 64-threaded CPU (Intel Xeon Gold 6338) takes >80 minutes, corresponding to a {approx}8x speedup. We further demonstrate true biobank-scale performance in the Penn Medicine Biobank (PMBB; 57,170 samples), computing chromosome 1 LD ({approx}1.38 million variants) in 49 minutes versus 22.7 hours for PLINK, a {approx}28x speedup. With this method, we successfully computed on the aforementioned 30x WGS 1000 Genomes dataset ({approx}120 million variants and {approx}2500 samples) the entire LD matrix (>1e16, or 10 quadrillion elements) in under 6 hours using 512 NVIDIA 40GB A100 GPUs on the Department of Energy Argonne Leadership Computing Facility Polaris Supercomputer. We make this tool, coded in Python using CuPy, publicly available. Using this tool, researchers can leverage the full extent of their genomic data without relying on external LD reference panels and acquire more accurate, population-specific findings, particularly for groups underrepresented in existing databases.
bioinformatics2026-09-21v1Uniformly processed transcriptome-wide alternative splicing profiles for pediatric cancer research
Liang, C. E.; Shapiro, J. A.; Beale, H. C.; Taroni, J. N.; Vaske, O. M.Abstract
Alterations in regulatory processes like alternative splicing contribute to pediatric cancer development. Although splicing aberrations have been observed in pediatric leukemias, alternative splicing has yet to be studied in pediatric cancers at scale, due to a lack of uniformly processed, sample-level pediatric cancer splicing profiles with non-diseased tissue comparators. We address this need by quantifying splice event usage for a curated set of bulk RNA-seq datasets from the NCI's Therapeutically Applicable Research to Generate Effective Treatments (TARGET, n = 1152) and Genotype-Tissue Expression (GTEx, n = 1098) as a comparator. This Treehouse Splice Compendium is accompanied by a reproducible workflow that was used to generate the data in the compendium and reflects the largest known RNA-seq dataset processed by the splice quantification tool Shiba. The compendium is part of a suite of large, uniformly processed datasets aggregated by the UCSC Treehouse Childhood Cancer Initiative and Alex's Lemonade Stand Foundation's Childhood Cancer Data Lab, which include the Treehouse Expression Compendia, refine.bio, and the Single-cell Pediatric Cancer Atlas.
bioinformatics2026-09-21v1Two Comparators May Be All We Need
Muskal, S. M.Abstract
A compound in a cell meets a spectrum of proteins drawn from many families at once, while screening most often interrogates one target at a time. Two questions asked many times in a rank ordering workflow ultimately guide decisions on what gets made and what gets counter-screened, and both are comparative: which of two targets does a compound prefer, and which of two compounds does a target prefer. We built one model for each, over a roster of 1,879 human proteins covering 34 protein families. Each model is given two chemical structures and a sequence, or two sequences and a chemical structure, and returns which member of the pair is preferred together with how firmly it holds that view. No conformational analysis, protein structure, binding site or docked pose is used. Across families, asked which of two targets a compound prefers, the model is correct 0.75 of the time over 8,689 held-out comparisons, and 0.93 of the time on the third of them it holds most confidently. Asked which of two compounds a single target prefers, it is correct 0.71 of the time over 65,725 held-out comparisons, rising to 0.96 on the most confidently held. Neither compound in any of those comparisons appeared anywhere in training. Accuracy in both tracks the size of the real difference between the two measurements, from near chance where they fall within half a log unit to about 0.90 where they differ by more than two logs, and it holds across 31 protein families, not only the best-measured one. Within the chemistry and the targets they were built on, these models rank compounds and rank targets well. Both models can be explored and downloaded at familyfoundationmodel.com. Keywords: target preference; compound preference; pairwise comparison; polypharmacology; off-target triage; ESM2; random forest; structure-free prediction; ChEMBL
bioinformatics2026-09-21v1celltypeEnrich: a consensus-based scRNA-seq cluster annotation tool
Rutledge, S.; Tuteja, G.Abstract
Motivation Single-cell RNA sequencing (scRNA-seq) cluster annotation is a critical step in data analysis. Current methods are time-consuming, difficult to reproduce, or limited in tissue or species coverage. Results We developed celltypeEnrich, a cluster-level annotation tool that uses a hypergeometric test to identify enrichment of cell-type-specific genes from input gene lists. Enrichment results from up to 26 reference datasets are used to determine a consensus annotation. Benchmarking using scRNA-seq datasets from three tissues spanning two species showed 62-72% annotation accuracy for celltypeEnrich, generally outperforming other tools, which had either lower accuracy, incomplete tissue coverage, or the need for parameter optimization. The performance of celltypeEnrich remained stable when input gene lists were down-sampled to 25% of their original size. Availability and Implementation celltypeEnrich is freely available at (https://celltypeenrich.gdcb.iastate.edu) as an R Shiny web application under the MIT license for non-profit academic use.
bioinformatics2026-09-21v1Individual-level expression deconvolution and assessment of cross-sample variation
Kang, K.; Xie, K.Abstract
Recovering cell-type-specific gene expression from bulk RNA sequencing would facilitate the study of transcriptional variation among individuals. However, accuracy can differ substantially among genes and cell types. We describe a reference-informed Bayesian deconvolution framework and a score that identifies gene--cell-type pairs likely to have more accurate estimates of cross-sample variation. The score uses bulk counts, reference expression profiles, and estimated RNA proportions. Known component expression is used to train and evaluate the score, but is not needed to calculate predictions from a trained model. We evaluated the approach in a ROSMAP-derived simulation with 40 target donors, 2,000 genes, and seven cell types. Median gene-wise correlation was 0.801 for raw allocated counts and 0.296 after normalization within each donor and cell type. To evaluate the score, we divided genes into five sets, kept linked genes together, and scored each set using a model trained on the other four. Retaining approximately 20\% of pairs within each cell type increased the median normalized correlation to 0.622. Ranking pairs only by the estimated share of a gene's bulk RNA contributed by the cell type yielded 0.570 at the same retained count. These results show that observable information can help prioritize pairs with more accurately recovered cross-sample variation.
bioinformatics2026-09-21v1Correlation-aware discovery of co-occurring mutational signatures in cancer
Jin, H.; Geiger, B.; Glodzik, D.; Gulhan, D. C.; Park, P. J.Abstract
Somatic mutations in cancer genomes record the activities of diverse mutational processes. Mutational signature analysis has advanced mechanistic understanding of mutagenesis and informed clinical decision-making, yet existing methods assume independence among signatures--an unrealistic assumption that can produce composite or contaminated signatures, reduce detection power, and yield inconsistent results. Here we present Cornet (CORrelated NMF ExTraction), a framework for mutational signature discovery that explicitly models co-occurring processes and jointly infers signatures and their correlation structure. Benchmarking on simulated data shows Cornet more accurately recovers distinct signatures under strong correlations. Applied to cancer genomes, Cornet enables unsupervised discovery of the colibactin-associated signature SBS88 in oral cancers and identifies the tobacco smoking signature SBS4 in bladder cancer, where it was previously thought absent. Cornet also uncovers a novel mutational process implicated in early-onset colorectal cancer and a signature arising from the interplay between tobacco smoking and ERCC2-mutation-driven nucleotide-excision repair deficiency. Together, these results demonstrate that modeling correlations among mutational processes is essential for high-resolution signature discovery and dissecting the mutational etiology of human cancer.
bioinformatics2026-09-21v1Gravlax: an annotation-independent molecular evidence archive for single-cell RNA-seq
Patro, R.Abstract
A cell-by-gene count matrix is the artifact of a single-cell RNA-seq experiment that is most often stored, shared, and reanalyzed. It is the output of a computation whose inputs are the sequenced molecules and a gene annotation, and while the molecules never change, the annotation is revised continually. Once the matrix has been produced, the evidence behind it can no longer be reinterpreted. Recovering that evidence means returning to raw reads or alignments that are large, costly to process, and frequently unavailable. We ask whether a concise representation of the molecules themselves can be extracted once and reused indefinitely, to quantify under any future annotation, to query and discover features that no annotation yet describes, and to pool evidence across cells and samples. Our starting observation is that the procedures that turn alignments into counts, gene assignment and UMI collapse, never read most of what an alignment file contains. They consume relations among molecules such as shared genomic geometry, shared placements, barcode identity, and the equality or near-equality of UMIs. We show that these relations form a statistic that is sufficient for such consumers, and we design a compact, seekable, content-authenticated archive that stores them while deferring every annotation-dependent decision to analysis time. Archives compose into content-addressed collections that route cohort queries to the molecules that can answer them without copying molecules. We implement these ideas in a tool called gravlax. Across four human 10x 3' datasets, gravlax archives require 11--18 bits per read and are 9.0--12.7x smaller than tag-preserving CRAM. Count matrices replayed from an archive deviate from direct STARsolo quantification by 0.24--0.75% of normalized UMI mass, whereas changing GENCODE v32 to v49 moves 2.12--4.64%, and quantification replay is 34--82x faster than STARsolo at matched thread budgets. A federated index over eight archives occupies 2.96% of their size, answers a 96-query junction panel 2.59x faster than the archives alone, and screens the cohort genome-wide for unannotated splice events that recur across donors in just 9 seconds. Because the molecules are retained, the archives also answer questions the matrix has discarded. An analysis of four peripheral-blood archives recovers a validated FYB1 immune-cell splicing switch, a cross-fitted fragment model appropriate for 3' chemistry reveals an eight-donor shift in NTRK2 terminal-isoform usage from astrocyte and neural-stem-cell populations to mature neurons, and pooling evidence across cells within the context of an expectation-maximization algorithm recovers 75--98% of withheld multi-gene molecule labels. Gravlax is open source, implemented in Rust, licensed under the BSD 3-clause license, and available at https://github.com/COMBINE-lab/gravlax.
bioinformatics2026-09-21v1A transcriptional continuum from clinically normal to lesional skin: Single-cell trajectory analysis of psoriasis
Uzun, G.; Konak, D.; Kazan, H.; Pir, P.Abstract
Psoriasis is a chronic inflammatory skin disease characterized by immune dysregulation and complex cellular interactions involving keratinocytes (KCs), T cells, and endothelial cells. Single-cell studies have mapped the cellular diversity of psoriatic plaques, yet most analyses treat skin as either healthy or diseased and therefore overlook the graded changes that precede a visible lesion. We asked whether psoriasis instead progresses along a transcriptional continuum and whether clinically normal skin from patients already carries a measurable disease signature. Using single-cell RNA sequencing (scRNA-seq) of healthy skin (NS), clinically normal skin adjacent to lesions (PN), and lesional skin (PP), we combined two complementary designs: a complete set spanning all three states and a within-patient paired set of matched PN and PP skin biopsies that removes inter-individual variability. Trajectory analysis using a novel pseudotime approach reveals transcriptional changes within cell populations and dynamic cellular states underlying psoriatic pathology. To distinguish lesion-driven effects from disease-associated transcriptional changes, we applied a differential expression strategy comparing PP and PN psoriatic tissue with NS and integrating these contrasts. This approach enabled the identification of disease-specific gene signatures while minimizing lesion-specific bias. As a result, we identified a coherent, cell-type-resolved continuum of transcriptional states in which endothelial cells, IL20+ keratinocytes (KCs: IL20), and KGF+ keratinocytes (KCs: KGF) undergo trajectory-dependent state transitions. These trajectories converge on features including epidermal thickening, barrier dysfunction, and plaque maintenance. Endothelial cells transition from an inflamed, angiogenic state to a mature, barrier-stabilized phenotype, while KCs: IL20 and KCs: KGF shift from PN states toward hyperproliferative, stress-adapted lesional phenotypes. These findings reframe psoriasis as a dynamic, multicellular continuum and identify candidate temporal windows for early intervention.
bioinformatics2026-09-21v1A functional landscape of human microproteins
Slivak, M.; Choteau, S. A.; Spinelli, L.; Souville, L.; Pierre, P.; Carvunis, A.-R.; Bornberg-Bauer, E.; Zanzoni, A.; Brun, C.Abstract
Short open reading frames (sORFs) are widespread yet often poorly annotated, and an increasing number are known to encode functional sORF-encoded peptides (sPEPs, or microproteins). Here, we use a system approach to determine the functions of sPEPs in three tissues: monocytes, skeletal muscle and brain. We first predicted interactions between sPEPs and canonical proteins and characterized their interaction interfaces. sPEPs preferentially target specific canonical partners rather than interacting promiscuously, and two functional groups emerge from their interaction features. A small subset (~2%) exhibits canonical protein-like features, with structured domains interacting with diverse proteins involved in tissue-specific functions. In contrast, most sPEPs contain short linear motifs with extensive phosphorylation potential, suggesting they may be involved in signaling. We then constructed the first sPEP-containing interactome networks and infer sPEP functions from network topology. Most sPEPs appear to act as regulatory elements modulating diverse cellular processes rather than acting directly as functional components. Knowledge-based validations using characterized sPEPs, show that our framework can capture context-specific functions beyond current annotations. Together with their limited evolutionary constraint, the pervasive regulatory roles of sPEPs suggest that they constitute a reservoir of functional novelty and an evolutionarily flexible layer of cellular regulation.
bioinformatics2026-09-19v2The RdRp Thumb-1 Pocket is a Conserved Target for Broad-Spectrum Antiviral Development
Woods, V.; Umansky, T.; Russell, S. M.; Gallay, P.; Smith, D.; Haders, D.Abstract
RNA viruses cause human diseases ranging from mild colds to deadly pandemics. Direct-acting, broad-spectrum, non-nucleoside antivirals have been characterized as impossible to develop because allosteric binding sites are poorly conserved. The HCV NS5B RNA-dependent RNA polymerase (RdRp) Thumb-1 allosteric site and its interaction with the HCV NS5B {Lambda}1-loop governs an essential conformational change required for polymerase initiation. The only approved NS5B Thumb-1 inhibitor, beclabuvir, has been shown to be inactive against a broad panel of non-HCV viruses, including poliovirus, rhinovirus, coronavirus, coxsackievirus, influenzavirus, and HIV. A conserved, homologous allosteric site on RdRp that spans multiple viral families has not been reported. Here, we report GALILEO's discovery that the Thumb-1 pocket, its associated {Lambda}1-loop and their interaction are conserved across RNA viral families for the first time. The discovery is validated through comparative structural analysis of Protein Data Bank (PDB) deposited viral polymerases utilizing a method that allows independent, public validation by any researcher. We further demonstrate that beclabuvir's dependence on its indole C6 carbonyl to interact with the HCV-specific residue R503 restricts its activity to HCV. We validate the target discovery with MDL-001, which does not contain a C6 carbonyl substituent. MDL-001 directly blocks viral RNA synthesis in isolated replication complexes and selects for the canonical Thumb-1 resistance mutation P495S in HCV NS5B. MDL-001 demonstrates broad-spectrum in vitro inhibition of both HCV and SARS-CoV-2. Preclinical proof of concept and development of MDL-001 across HCV, HBV, HDV, influenza, SARS-CoV-2, and RSV have been previously reported. These findings establish RdRp Thumb-1 as a conserved allosteric pocket and a druggable target for broad-spectrum direct-acting antiviral development.
bioinformatics2026-09-18v7OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome
Teixeira, A. A. R.; Zhu, H.; Kothiwal, D.; Cao, R.; Mills, A.Abstract
Choosing which region of a protein to express remains poorly standardized in antibody discovery, recombinant reagent generation, structural biology and computational binder design. For human cell-surface and secreted proteins, this requires reconciling topology, processing, predicted and experimental structure, modifications, interaction partners, orthologs, paralogs and cross-reactivity risk before ordering DNA. OpenAntigens is a free, no-login database of construct-design reports for 5328 human secreted, GPI-anchored, single-pass and multipass proteins. It integrates UniProt topology, AlphaFold pLDDT and PAE, PDB precedent, InterPro and Pfam domains, mouse and cynomolgus orthologs, paralog and family context, Open Targets disease associations, partner and assembly context, and BLAST searches. It provides 55 305 construct suggestions spanning full design regions, PDB-backed boundaries, annotated domains, pLDDT/PAE-derived regions and membrane-expression options, plus 148 722 sequence-similarity hits to help choose constructs and assess cross-reactivity. For targets with compatible AlphaFold models, the interactive designer links sequence, structure, pLDDT and PAE, allowing users to revise boundaries and export species-equivalent sequences with real-time cysteine and modification warnings. OpenAntigens places reproducible construct suggestions, comparative context and browser editing in one workflow, reducing manual reconciliation across resources. OpenAntigens is available at openantigens.org.
bioinformatics2026-09-18v2Incomplete references leave bulk deconvolution targets non-identifiable, but identification has a computable precision price
Jiang, H.; Gao, F.; Liu, P.; Wu, Y.; Jie, Y.; Li, Y.; Jiang, Y.Abstract
Background. Reference-based deconvolution estimates cell-type proportions from bulk profiles. Incomplete references compromise these estimates, yet many remedies return point estimates. We separate the operator provenance that reproduces an estimate from the information and precision that identify its target. Results. Two operator histories at one reduced reference give different estimators: 28.2% of 196,420 sample-deletion pairs differed by >0.1 total variation; locking every learned component made paths identical. Across nine cohorts, operator history reversed 160 of 952 association signs, 31 with a significant path, and changed significance for 126. Observationally equivalent completions can fill the open simplex and reverse retained-type rankings. Under a joint zero-exposure condition, shared structure leaves inherited bounds unchanged; one to six uncalibrated views also gave identical bounds. A profile library contracted estimator-output envelopes by 98.93% yet covered the full-reference effect for 42.77%, with 19 wrong-sign certificates; conditional sharp bounds stayed at [-1, 1]. Treating RNA yields as exact collapsed intervals to points covering none of seven flow-measured targets: contraction without coverage is false certainty. Calibrated cross-modal anchors contract width to 0.51 at six types but certify no valid sign. Exact RNA yields identify cell fractions from RNA contributions; a decisive sign in blood requires {+/-}2.6% proxy accuracy with near-total contraction of donor-heterogeneity and dynamic-range envelopes. Conclusions. Incomplete-reference deconvolution is an identification problem, not only an estimation problem. Remedies must be scored on shrinkage, coverage, certification and false certification against held-out targets. fitdrop implements this scoring and a precision frontier for planning decisive measurements.
bioinformatics2026-09-18v2DeepSpaceDB 2.0: an interactive spatial transcriptomics database for large-scale Xenium data exploration
Honcharuk, V.; Takemoto, K.; Masalunga, M. C.; Zhao, H.; Diez, D.; Kawaoka, S.; Vandenbon, A.Abstract
The 10x Genomics Xenium platform enables high-resolution spatial transcriptomics at single-cell and subcellular scales, but effective reuse of public Xenium datasets is hindered by large data sizes and heterogeneous file formats. We previously developed DeepSpaceDB, a spatial transcriptomics database designed for interactive, in-depth analysis of tissues and tissue microenvironments. Here, we present a major expansion of DeepSpaceDB that integrates large-scale single-cell spatial transcriptomics data generated by the Xenium platform. In this update, we systematically collected 1,539 public Xenium datasets from multiple repositories and processed them through a robust, standardized pipeline that validates, repairs, and harmonizes heterogeneous inputs into a unified representation. To support efficient exploration of these data, we introduced a redesigned DeepSpaceDB interface and complementary Zarr-based storage formats optimized for gene-centric visualization and spatially localized queries, enabling sub-second response times for common interactive operations. The updated platform supports real-time visualization of spatial data and analysis of regions of interest directly in the web browser. Together, this expansion establishes DeepSpaceDB as a unified resource for single-cell spatial transcriptomics, substantially lowering the barrier to accessing, exploring, and reusing large-scale public Xenium datasets.
bioinformatics2026-09-18v2LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design
Waththe Liyanage, W. W.; Rigoni, D.; Bove, F.; Righelli, D.; Romano, S.; Visone, R.; Iorio, M. V.; Grassia, M.; Mangioni, G.; Lio, P.; Taccioli, C.Abstract
The discovery of novel small molecules is challenging because of the vastness of chemical space and the complexity of protein-ligand interactions, leading to low success rates and time-consuming workflows. Here, we present LLMsFold, a computational framework that combines Large Language Models (LLMs) and biophysical foundation tools to design and validate new small molecules targeting pathogenic proteins. The pipeline starts by identifying viable binding pockets on a target protein through geometry-based pocket detection. A 70-billion-parameter transformer model from the LlaMA family then generates candidate molecules as SMILES strings under prompt constraints that enforce drug-likeness. Each molecule is evaluated by Boltz-2, a diffusion-based model for protein-ligand co-folding that predicts bound 3D structure and binding affinity. Promising candidates are iteratively optimized through a reinforcement learning loop that prioritizes high predicted affinity and synthetic accessibility. We demonstrate the approach on two challenging targets: ACVR1 (Activin A Receptor Type 1), implicated in fibrodysplasia ossificans progressiva (FOP), and CD19, a surface antigen expressed on most B-cell lymphoma and leukemia cells. Top candidates show strong in silico binding predictions and favorable drug-like profiles. All code and models are made available to support reproducibility and further development.
bioinformatics2026-09-18v2WITHDRAWN: A Comprehensive Analysis of the Electrolytic Hydrogen Water Mechanism via a Feedforward Loop and its Functional Role in Intestinal Cells In Vitro
LI, J.Abstract
This manuscript has been withdrawn following a formal investigation by the Graduate School of Human Sciences, Waseda University
bioinformatics2026-09-18v2Pep-PU-GAN: Positive-Unlabeled Adversarial Learning for Peptide Function Prediction
Midjani, F.; Hashemi, S.; Keshtkar, F. Z.; Malekpour, M.; Saberzadeh Ardestani, B.; Khosravi, B.Abstract
Peptide classification remains challenging in bioinformatics because of limited labeled data, particularly the scarcity of verified negative examples, and the complex relationship between amino acid sequences and biological functions. This study introduces Pep-PU-GAN, a deep learning framework that combines positive-unlabeled (PU) learning, generative adversarial networks (GANs), and graph neural networks (GNNs) for peptide classification. Peptides are represented as sequence-derived residue graphs, with amino acids as nodes and edges connecting adjacent residues, enabling attention-based message passing over local neighborhoods. The architecture includes a generator that produces synthetic peptide embeddings in encoder space and a dual-function discriminator that distinguishes real from synthetic embeddings while performing PU classification. Training uses a custom loss integrating non-negative PU (nnPU) risk estimation with adversarial objectives. A self-training mechanism further incorporates high-confidence synthetic positive embeddings to augment the training set and improve performance. Evaluated on neuropeptide classification using 4,049 positive neuropeptides and 8,558 unlabeled peptides, Pep-PU-GAN outperformed baseline models, achieving an F1 score of 0.93 and an AUROC of 0.98 on an independent held-out benchmark. Pep-PU-GAN provides a promising approach for peptide classification tasks with scarce labeled and abundant unlabeled data, with potential applications in computational biology and drug discovery.
bioinformatics2026-09-18v1TRACEDD: A Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design
Vangala, S. R.; Kasturi, V. V.; Bung, N.; Roy, A.Abstract
Drug discovery depends on coordinated decisions across target validation, structure analysis, molecular design, developability assessment and synthetic feasibility, but current computational methods often operate as disconnected tools. Here, we introduce TRACEDD (Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design), a framework that makes three primary contributions: (1) It establishes a 'tool-first' multi agentic architecture where LLMs orchestrate validated computational tools rather than replace them, ensuring scientific rigor. (2) It implements a multi-agent system that mirrors expert discovery teams, enabling transparent and traceable decision-making through a Reason-Act-Observe loop. (3) It demonstrates an end-to-end workflow, from target validation to synthesis planning, that adaptively handles real-world data variability, such as the absence of experimental structures. The framework decomposes discovery into specialized agents for target validation, druggability assessment, molecular generation, lead optimization, ADMET evaluation, literature evidence integration and retrosynthesis, all operating through a Reason Act Observe workflow. Using JAK2 as a representative case, we show that the system can retrieve experimental protein structures, invoke AlphaFold when structures are unavailable, identify druggable pockets and perform de novo molecular generation. Known JAK2 inhibitors are used to define design hypotheses and guide reinforcement learning-based molecular generation, with docking scores/predicted pIC50 and other physicochemical/ADMET properties serving as reward and prioritization signals. The framework demonstrates a tool-first, reasoning-driven approach in which each major decision is linked to explicit tool invocation, intermediate evidence. By combining agentic orchestration with domain-specific computational tools, the system supports transparent, adaptable and human-verifiable molecular design workflows, providing a foundation for more reliable AI-assisted drug discovery.
bioinformatics2026-09-18v1Sparse Machine Learning Pipeline with Stabl Identifies Cord Blood Multi-Omic Signatures of Bronchopulmonary Dysplasia
Mestan, K.; Newar, J.; Zhao, J.; Chakraborty, A.; Reiss, J.; Funk, W.; Stelzer, I.; Waked, B.; Bellan, G.; Durand, X.; Hedou, J.Abstract
Background: Several omics studies have been completed in recent years, with the goal of identifying biomarkers of complex multifactorial diseases, such as bronchopulmonary dysplasia (BPD). Objective: To evaluate the performance of 3 distinct omics platforms, using a machine learning pipeline with integration of sparse, reliable and adaptive biomarker identification (Stabl). Methods: Using a well-characterized birth cohort, cord blood metabolomics, proteomics and adductomics data were integrated with Least Absolute Shrinkage and Selection Operator (LASSO) regression and Stabl, to evaluate predictive performance for BPD. Results: Sparse multivariable modeling of 45,000 features measured in 217 infants (52 term, 165 extremely preterm <28 weeks; 82 with BPD and 35 with severe BPD/death) identified a perfect signature for preterm birth with both LASSO and Stabl (AUROC=1.0; p<0.001). Analysis of the preterm group yielded excellent predictive power for severe BPD (AUROC=0.83; p=0.005). Stabl identified a set of 12 biomarkers (2 adducts, 3 proteins and 7 metabolites) with good performance for predicting grade III BPD (AUROC=0.76; P=0.03). Biomarkers across the 3 omics platforms revealed dysregulated pathways of innate/adaptive immune responses, metabolic programming and oxidative stress. Conclusions: The sparse machine learning pipeline is a complementary approach for identifying novel pathways and biomarkers of multifactorial BPD and its endotypes.
bioinformatics2026-09-18v1Setting the SCENE for Interpretable Cell-Gene Embeddings in Single-Cell RNA-seq
Moberg, O. L.; Petersen, M. B.; Herlau, T.; Kristensen, L. E.; Jessen, L. E.; Morup, M.Abstract
Single-cell RNA sequencing measures cellular states at high resolution, but sparse high-dimensional count data remain difficult to model interpretably. We introduce the Single-Cell Euclidean Network Embedding (SCENE), a probabilistic latent-distance model that jointly embeds cells and genes from Unique Molecular Identifier (UMI) counts. SCENE treats the count matrix as a weighted bipartite cell-gene graph, where Euclidean distances represent transcriptional affinity, and combines this geometry with a zero-inflated count likelihood that separates gene detection from expression magnitude. Across real and simulated scRNA-seq datasets, SCENE recovers biologically structured cell and gene embeddings with state-of-the-art performance. Surprisingly, major biological structure is preserved in native two- and three-dimensional latent spaces, enabling directly interpretable visualization. Perturbation analyses show that SCENE organizes glucocorticoid-response genes and T-cell receptor regulatory programs coherently in gene space, capturing biology beyond cell-type separation. SCENE provides a transparent representation learning framework in which low-dimensional Euclidean geometry supports accurate modeling and biological interpretation.
bioinformatics2026-09-18v1Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics
Chen, T.; Hicks, S. C.Abstract
Models that combine tissue images with pretrained gene representations aim to predict spatial expression for genes not used to fit the downstream predictor. Yet success on held-out genes can reflect two capabilities: estimating a gene's mean expression across tissue locations and recovering its spatial variation. Across four cohorts spanning three human brain regions and HER2-positive breast cancer, we evaluated held-out genes in held-out individuals and separated these components. For spatial predictors using fixed gene representations from Decima or scGPT, reductions in gene-mean error accounted for more than 91% of the reduction in mean squared error relative to matched random vectors. Independently fitted mean-only models using the same representations but no tissue images retained 90-99% of the corresponding gain in full-matrix correlation. Spatial gains were smaller on average, increased with expression variation in training tissue and differed across cohorts and representations. Across these settings, pretrained gene representations broadly transferred mean expression but selectively improved spatial recovery, showing that cross-gene generalization in virtual spatial transcriptomics is not a single capability.
bioinformatics2026-09-18v1OpenLipid: a large language model workflow for targeted analysis of DIA mass spectrometry data in lipidomics
Li, J.; Rost, H.Abstract
Mass spectrometry has become a central technology for lipidomics, with data-independent acquisition (DIA) enabling broad and reproducible sampling of lipid signals. However, the multiplexed fragment-ion spectra in DIA data complicate lipid identification. Here, we introduce OpenLipid, a large language model (LLM)-based workflow for targeted DIA lipidomics. Using assay libraries built from data-dependent acquisition (DDA) results, OpenLipid directly evaluates extracted ion chromatograms (XICs) from DIA data in a zero-shot setting to identify target lipid peaks and generate human-readable rationales for individual lipid identification decisions. We benchmarked OpenLipid against manual annotations across four datasets comprising human plasma and mouse feces analyzed in positive and negative ionization modes. The plasma assay libraries contained 199 target lipids in positive mode and 147 in negative mode. The fecal assay libraries contained 181 target lipids in positive mode and 264 in negative mode. At a 5% false discovery rate (FDR) threshold, OpenLipid identified 110 (55.3%) and 28 (19.0%) library targets in plasma and 130 (71.8%) and 84 (31.8%) in feces, in positive and negative ionization modes, respectively. OpenLipid achieved an overall identification rate comparable to that of DIAMetAlyzer (57.8%, 19.7%, 89.0%, and 17.8% across the corresponding datasets) and substantially higher than that of untargeted MS-DIAL DIA analysis (12.6%, 0.0%, 47.5%, and 0.0%) on the same assay-library targets. LLM-derived chromatographic features also enabled supervised discrimination between correct and incorrect candidate peak groups for target lipids, with median cross-validation average precision values of 0.838-0.912. Together, these results demonstrate that OpenLipid is an effective LLM-based workflow for FDR-controlled targeted analysis of DIA lipidomics data.
bioinformatics2026-09-18v1Mapping Gene Expression to an Interpretable Semantic Space
Duan, X.; Aggarwal, M.; Periwal, V.Abstract
Cell embeddings organize single-cell expression data, but their dimensions have no biological meaning, so clusters are interpreted afterward. We present MESIC (Mapping Expression to Semantic space with Interpretable Components), which builds the written knowledge about genes held in curated databases into the dimensions themselves. A biomedical language model converts each gene's summary into a semantic embedding. MESIC compresses these embeddings into a small number of components, each concentrated on a small set of genes and explained by their annotations. The components are computed once from the summaries, so any expression dataset can be mapped onto them, and every cluster, outlier, or cell-type assignment is then characterized by named genes. In cardiomyocytes, outliers in the component space were enriched for hypertrophic cardiomyopathy. In a lung atlas, unsupervised clusters in that space matched the broad cell types that experts had annotated. In both, the components that separated the cells matched their known biology. For about half of the cells that the atlas itself had left unannotated, the same space gave a confident cluster assignment, and with it an interpretation through component-associated genes. Gene summaries thus give single-cell analysis a coordinate system in which every result is traced to genes and what is written about them.
bioinformatics2026-09-18v1A Unified 3D Generative Model for Synthesizable Structure-Based Drug Design
Igashov, I.; Schneuing, A.; Dobbelstein, A. W.; Morozova, I.; Neeser, R. M.; Zielinski, K.; Abriata, L. A.; Petruzzella, A. S.; Pavel Iosub, D. R.; Gampp, O.; Lyubimov, A. Y.; Elizarova, E.; Ferrara, I.; Sousa, P. M. F.; Lemos, A. R.; Testori, F.; Miranda Herrera, P. A.; Kanis, L.; Schmidt, J.; Braza, M. K. E.; Amaro, R. E.; Thoma, N.; Ferraris, D. M.; Riek, R.; Fraser, J. S.; Schwaller, P.; Bronstein, M.; Correia, B.Abstract
Traditional screening-based drug discovery is inherently limited by the astronomical scale of the chemical space. Generative modelling offers a compelling alternative to the classical search paradigm and enables rational, bottom-up design of novel and target-specific small molecules. However, its impact has been hampered by challenges in synthetic accessibility of the designed compounds and lack of large-scale experimental validation. Here, we introduce LDDM (Large Drug Discovery Model), a generative framework that supports a range of drug discovery tasks, including constrained and unconstrained docking, fragment linking and growing, and de novo design. We further introduce a programmable design algorithm that enables accurate design of synthetically accessible compounds satisfying various fine-grained objectives. We experimentally validated the designed or optimised ligands for five therapeutically relevant protein targets. In all cases, LDDM achieved high success rates, allowing us to identify molecules with confirmed binding affinity while synthesizing only a small number of generated compounds. The best designs were structurally characterised through NMR spectroscopy and X-ray crystallography, demonstrating high prediction accuracy. Overall, LDDM provides a scalable and flexible platform for the rapid and tailored design of small molecules and non-natural peptides for therapeutic applications.
bioinformatics2026-09-18v1Physical priors improve performance of structure-based binding affinity models
Kaminow, B.; Payne, A. M.; MacDermott-Opeskin, H. I.; Chodera, J. D.; Singh, S.Abstract
Structure-based drug discovery is a widely used paradigm for the rational design of novel small molecule therapeutics. However, the benefits conferred by the use of structural information has seen limited adoption in machine learning, where ligand-only ("2D") models are still the industry standard for molecular property or binding affinity prediction. Structure-based ("3D") ML models for binding-affinity prediction promise to present a clear advantage, but have not yet overtaken existing 2D models. Here, we show that physics-based priors can improve predictive performance of structure-based models by comparing different model architectures with varying physical priors on several prediction tasks. We present the Modular Training and Evaluation of Neural Networks (mtenn) package, where we decompose affinity prediction into separate steps of embedding structure into learned representations and combining those embeddings into a predicted binding affinity. We consider both E(3)-invariant and E(3)-equivariant architectures to determine the importance of encoding roto-translational inductive biases, as well as different methods for combining learned embeddings. By first optimizing several aspects of model construction using the general purpose PDBBind dataset, we are able to improve the performance and data efficiency of structure-based models. When subsequently trained and evaluated on the COVID Moonshot small molecule drug discovery dataset, our tuned models perform on par with industry standard ligand-only models. Our decomposed model framework highlights that encoding some physical priors improves model performance, while more complex biases such as equivariance offer limited benefit. Additionally, structure-based models generalize better to an unseen target and display higher training efficiency. Overall, these results emphasize that structure-based models benefit from their ability to incorporate physics-informed constraints, giving promising directions for model architecture development. These results also suggest that the strength of these models may be in tasks specifically aimed at generalizability, providing guidelines for their use in early-stage drug discovery campaigns.
bioinformatics2026-09-18v1Rapid Shift Toward Pulsed Field Ablation and Precision Risk Stratification in High-Impact Atrial Fibrillation Research
Su, Z.; Li, T.Abstract
Conventional bibliometrics rely on lifetime citations, obscuring immediate shifts in cardiovascular research paradigms. To track emerging trends in atrial fibrillation management, we performed a comparative bibliometric analysis of the 50 highest-cited original research articles per year in OpenAlex topic T10065 across consecutive 2023 (Class of 2025) and 2024 (Class of 2026) publication cohorts. Articles were ranked using a fixed 18-month post-publication citation window, and extracted concepts were normalized into 719 canonical topics and 32 parent themes using large language model curation. Concept frequency tracking demonstrated a swift technological shift, with pulsed field ablation showing the largest topic frequency increase (+0.08, from 0.30 to 0.38) to become the leading canonical topic in 2024, displacing conventional thermal pulmonary vein isolation (-0.18, 0.40 to 0.22). Simultaneously, stroke prevention focus shifted toward refined predictive modeling, with increases in Risk Stratification and Predictive Models (+0.08, 0.30 to 0.38) and CHA2DS2-VASc scoring (+0.08, from 0.10 to 0.18). High-impact atrial fibrillation research is rapidly pivoting toward non-thermal ablation safety profiling and precision risk stratification, highlighting the utility of fixed-window concept mining for capturing real-time scientific evolution. Online explorer of the result is available at https://pri.pepkio.com.
bioinformatics2026-09-18v1Comparative Mitogenomics and Molecular Phylogeny of Agriculturally Significant Tephritid Fruit Fly Pests in Bangladesh
Rahman, S.; Shormi, F. A.Abstract
Tephritid fruit flies of the tribe Dacini rank among the most economically destructive agricultural pests globally, with several Bactrocera Macquart and Zeugodacus Hendel species causing severe losses to fruit and vegetable production across South and Southeast Asia. Bangladesh harbors five dacine species of primary agricultural significance: Bactrocera dorsalis (Hendel), B. carambolae Drew & Hancock, B. zonata (Saunders), B. correcta (Bezzi), and Zeugodacus cucurbitae (Coquillett). Here we present a comparative mitogenomic and molecular phylogenetic framework based on 21 unique mitogenome records retrieved from NCBI GenBank, representing 19 dacine ingroup taxa and two outgroups. Direct parsing of the GenBank sequences confirmed the canonical complement of 13 protein-coding genes (PCGs) and two rRNA genes in all 21 records. Among the 14 Bactrocera ingroup taxa, genome size ranged from 15,273 to 15,977 bp. Whole-genome AT content ranged from 66.6% in Bactrocera tsuneonis to 82.2% in Drosophila melanogaster; the five Bangladesh-relevant dacine pest species showed tightly clustered AT content of 72.9-73.6%. A concatenated alignment of 13,600 nucleotide sites (13 PCGs + 12S + 16S rRNA) was analyzed by maximum-likelihood inference under the GTR+FO model (IQ-TREE v1.6.11; 1,000 ultrafast bootstrap replicates). The ML tree recovers the B. dorsalis complex (UFBoot = 79-100) and places B. correcta and B. zonata as a maximally supported sister pair (UFBoot = 100). This study provides a sequence-verified mitogenomic reference framework for molecular identification and pest surveillance in Bangladesh, and should be interpreted as a curated comparative baseline rather than a population-genomic analysis, as no newly collected Bangladeshi specimens were sequenced.
bioinformatics2026-09-18v1One-dimensional CNNs for Near-Infrared Prediction of Protein and Moisture in Cereal Grains: The Effects of Architecture and Input Preparation
Deng, G.; Chi, W.; Yu, P.; WU, F.Abstract
Near-infrared (NIR) spectroscopy is widely used for the rapid, non-destructive determination of constituents such as protein and moisture in cereal grains, but model comparisons in this field are often confounded by differences in the inputs received by each model. We benchmark a compact one-dimensional CNN derived by one-factor-at-a-time ablation (CNN-Baseline) and a randomly searched CNN (CNN-RS) against PLSR, SVR, XGBoost on six cereal-grain datasets (n=500-5046) under three input conditions: raw spectra, optimal preprocessing, and preprocessing plus wavelength selection. A single-filter with a kernel size of 11 of convolution and batch normalization proved sufficient. Given a unified raw input, CNN-RS was most accurate on five of the six datasets, while CNN-Baseline came within 0.003-0.017 in test R2 on the four larger datasets. PLSR led only on the smallest protein dataset (XDS, n=500) and was the most stable model. Preprocessing had almost no effect on CNN accuracy on the four larger datasets, but improved accuracy by 0.034-0.047 (CNN-Baseline) and 0.017-0.055 (CNN-RS) on the two smallest, and brought the networks to their optimum in a median of 49% fewer epochs. Therefore, replacing explicit preprocessing with learned filters was only justified when sufficient training data were available. Once each model received its own preprocessing and wavelength subset, SVR became most accurate on four of six datasets and XGBoost rose from the weakest model to within 0.01-0.03 of the best model on four of the six datasets, suggesting that the principal advantage of CNNs lied in robustness to raw input data rather than in attainable accuracy. Overall, model selection for NIR calibration of cereal grains depends more on sample size and the input preparation than on network architecture.
bioinformatics2026-09-18v1Virtual experiments bridge sequence and microscopy with generative models
Zheng, D.; Hong, K.; Huang, B.Abstract
Large-scale screening and mapping efforts have produced vast libraries of perturbation-readout data. Converting these measurements into mechanistic insights requires models that link perturbation and genetic input to phenotypes, i.e., labels from experimental readouts, which are usually task specific. We propose a different, virtual experiment modeling approach: train generative models to recreate readouts conditioned on the experimental context, and then let established downstream models extract phenotypes from the synthetic data. As an illustrative case, we develop a bidirectional sequence-image generative framework, CELL-FM, that maps protein sequence and cellular context to fluorescence microscopy images and back, enabling in silico localization prediction, image-conditioned functional motif analysis and generation, and large-scale virtual mutagenesis revealing the amino acid features controlling condensate formation of intrinsically disordered peptides. This approach decouples representation learning from task-specific annotation, reuses rich experimental modalities across many downstream tasks, and preserves the spatial and organizational detail that hand-crafted labels often discard.
bioinformatics2026-09-18v1Novel two-stage deep learning-based approach applied to gene expression data pertaining to esophageal adenocarcinoma boosting biological knowledge discovery
Jamie, F.; Turki, T.; Alsolami, F.; Taguchi, Y.-h.Abstract
Esophageal cancer (EC) is characterized by complex transcriptional alterations and therapeutic resistance, posing challenges for traditional computational methods. In this study, we propose a deep learning (DL)-based computational framework to identify important genes and biologically relevant pathways in bulk cell RNA-seq data (GSE234304 and GSE273848), which comprise tumor and non-tumor esophageal tissue samples. A fully connected feedforward neural network was trained for binary classification, and two feature selection strategies were implemented: Neural Network followed by Support Vector Regression (NN+SVR) and Integrated Gradients combined with SVR (IG+SVR). The genes were then ranked according to their weights in SVR deriving the importance scores, and the top 100 genes were subjected to enrichment analysis using Enrichr and Metascape. The proposed DL-based approaches identified a greater number of expressed genes across established esophageal cancer cell lines than LIMMA, SAM, and the t-test did. Specifically, in the GSE234304 dataset, IG + SVR, our best method, identified a total of 9 expressed genes while the best baseline method, LIMMA, identified a total of 3 expressed genes. In terms of GSE273848 dataset, IG + SVR was also the best identifying a total of 11 expressed genes while the best baseline method, t-test, had a total of 7 expressed genes. The key genes identified included CEBPB, SUMO1, RORA, STAT1, GATA, OCT1, RUNX1, and NR3C1, as well as pathways related to nucleoprotein maturation, collagen fibril organization, the immunoglobulin-mediated immune response, immune regulation, and insulin signaling. These results show that combining neural networks and attribution-based regression creates an effective and interpretable framework for selecting genes in esophageal cancer research.
bioinformatics2026-09-18v1Facilitating genome annotation using ANNEXA and long-read RNA sequencing
Hoffmann, N.; Besson, A.; Cadieu, E.; Lorthiois, M.; Le Bars, V.; Houel, A.; Hitte, C.; Andre, C.; Hedan, B.; Derrien, T.Abstract
With the advent of complete genome assemblies, genome annotation has become essential for the functional interpretation of genomic data. Long-read RNA sequencing (LR-RNAseq) technologies have significantly improved transcriptome annotation by enabling full-length transcript reconstruction for both coding and non-coding RNAs. However, challenges such as transcript fragmentation and incomplete isoform representation persist, highlighting the need for robust quality control (QC) strategies. This study presents ANNEXA, a pipeline designed to enhance genome annotation using LR-RNAseq data while also providing QC for reconstructed genes and transcripts. ANNEXA integrates two transcriptome reconstruction tools, StringTie2 and Bambu, applying stringent filtering criteria to improve annotation accuracy. It also incorporates deep learning models to evaluate transcription start sites (TSSs) and employs the tool FEELnc for the systematic annotation of long non-coding RNAs (lncRNAs). Additionally, the pipeline offers intuitive visualisations for comparative analyses of coding and non-coding repertoires. Benchmarking against multiple reference annotations revealed distinct patterns of sensitivity and precision for both known and novel genes and transcripts and mRNAs and lncRNAs. To demonstrate its utility, ANNEXA was applied in a comparative oncology study involving LR-RNAseq of two human and eight canine cancer cell lines. The pipeline successfully identified novel genes and transcripts across species, expanding the catalog of protein-coding and lncRNA annotations in both species. Implemented in Nextflow for scalability and reproducibility, ANNEXA is available as an open-source tool: https://github.com/IGDRion/ANNEXA.
bioinformatics2026-09-17v5Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme
Adamov, A.; Mueller, C. L.; Bokulich, N.Abstract
Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured, and predictive modeling from them typically relies on ad hoc feature representation choices that obscure their impact on performance and interpretation. We present ritme, an open-source Python package that jointly optimizes microbiome-specific feature representation and model selection - combined algorithm selection and hyperparameter optimization - tailored to these data. ritme systematically searches taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment together with model class and hyperparameters, using state-of-the-art optimizers that scale from a laptop to a compute cluster. Across three real-world use cases, ritme outperformed the original study pipelines by 7-29% on the primary task metric and surpassed three AutoML baselines in six of seven comparisons, while selecting substantially fewer features and exposing how feature and model choices drive performance. Open-source and modular, ritme supports reproducible, parsimonious predictive modeling, downstream biological investigation, and extension to other multi-omics modalities.
bioinformatics2026-09-17v4FORGE audits residue-level information encoded in RNA tertiary structure geometry
Gow, L.; Li, J.; Tan, X.; Liang, K.; Gui, N.; Luo, B.Abstract
Coarse RNA coordinate representations are widely used, yet the biological information they encode remains unquantified. We introduce FORGE, which converts a seven-atom RNA geometry representation into 935 interpretable descriptors and reports which residue-level annotations this geometry supports. On 4,135 post-2025 RNA chains, FORGE recovered 64.6% of native nucleotides; a six-atom control lacking the glycosidic nitrogen retained 58.5%, locating most of this signal in phosphate-sugar geometry. Confidence was sharply graded: abstaining from the least-confident half of positions raised accuracy to 94.4%, yet many chains remained only partially identifiable. The same descriptors predicted base-pair state far better than a DMS-like proxy or protein-proximal context. Native-decoy, OpenKnot and solved-pseudoknot analyses showed that nucleotide identifiability, foldability and experimental design score are separable: AlphaFold3 reproduced the experimental fold for one of four AI-designed constructs and none of the sequences FORGE read from their geometry. FORGE provides a reproducible audit layer for RNA structural interpretation.
bioinformatics2026-09-17v3Locat: Joint enrichment and depletion testing identifies localized marker genes in single-cell transcriptomics
Lewis, W. R.; Aizenbud, Y.; Strino, F.; Kluger, Y.; Parisi, F.Abstract
Several methods identify marker genes that delineate cell populations in single-cell transcriptomic data, yet most emphasize enrichment within candidate populations without testing whether expression is significantly reduced elsewhere. We present Locat, a framework for identifying highly specific localized genes by testing whether expression is concentrated within compact regions of a cellular embedding and depleted outside them. For each gene, Locat fits weighted Gaussian mixture models to gene-specific and background densities, computes concentration and depletion statistics, and integrates them into a unified localization score. Across synthetic benchmarks with controlled ground truth, Locat detects uni-modal, multi-modal, and sparse localized patterns and loses significance when expression becomes indistinguishable from background structure. In developmental, perturbation, and differentiation datasets, Locat identifies compact marker sets that capture lineage organization, condition-specific programs, and temporal dynamics. These sets are often smaller than highly variable gene selections, while embeddings built from them preserve major cell populations and developmental programs in several cases. In murine dermis, interferon-treated PBMCs, and retinoic acid-induced embryonic stem cell differentiation, localized genes recover differentiation trajectories, stimulus-responsive programs, and reproducible stage-specific patterns. Together, these results show that jointly assessing concentration and depletion yields specific, interpretable marker genes.
bioinformatics2026-09-17v3LRP2: A proteogenomics pipeline for long-read informed protein isoform analysis and discovery
Schertzer, M. D.; Lewandowski, J. T.; Watts, E. F.; Rosenow, W.; Mehlferber, M. M.; Jeffery, E. D.; Adamson, S. I.; Bruand, J.; Tseng, E.; Neelamraju, Y.; Garrett-Bakelman, F. E.; Dolzhenko, E.; Knowles, D. A.; Sheynkman, G.Abstract
Most human genes produce multiple RNA isoforms, yet it remains unclear which isoforms are translated into stable, functional proteins. Long-read RNA sequencing resolves full-length transcript structures and, when paired with mass spectrometry, can provide empirical evidence of isoform translation. Despite this opportunity, comprehensive workflows integrating isoform discovery, open reading frame prediction, peptide identification, and protein inference remain limited, leaving users to handle these steps piecemeal. Here, we present LRP2, a modular, end-to-end long-read proteogenomics pipeline built in Nextflow. LRP2 scales transcript discovery to hundreds of samples via PacBio's latest Isocall tool, removes technical artifacts with SQANTI QC, generates and classifies predicted proteomes via CPAT and SQANTI Protein, performs multi-group differential expression and usage analysis via edgeR, DRIMSeq, and a long-read adaptation of LeafCutter, and integrates protein-level evidence from DDA and DIA MS data through FragPipe. For cross-dataset comparison of novel isoforms, LRP2 employs deterministic splice-junction, coordinate-based isoform identifiers. Used as an integrated pipeline, LRP2 enables the detection of novel peptides and improves the protein isoform inference to confirm protein isoform translation.
bioinformatics2026-09-17v2Lacuna: Cryptic Binding Pocket Discovery via Conformational Ensemble Analysis
Moore, C.Abstract
Lacuna, an open-source Python tool for discovering cryptic binding pockets: sites that are absent or too small to detect in a protein's unbound structure and open only during conformational fluctuation. Most binding-site predictors score a single static structure, which is precisely the structure in which a cryptic site is invisible. Lacuna instead generates a conformational ensemble from any input structure, detects pockets independently in every conformer, clusters the detections into persistent sites across the ensemble, and ranks those sites with a model fitted on within-structure pairs. Ensemble generation is pluggable: normal mode analysis by default, with implicit-solvent molecular dynamics, Boltz-2 diffusion sampling, or a user-supplied ensemble as alternatives. On the designated test fold of CryptoBench, Lacuna recovers 55.6% of cryptic sites in its top five predictions with the zero-dependency default and 66.1% with an optional PLM-assisted ranker; pooling the geometric detector with an optional learned surface detector recovers 73.9% while raising the fraction of sites found from 68.5% to 86.4%, measured on the held-out fold at five conformers. It recovers 73%, 45% and 87% on the PocketMiner set, a curated set of literature apo/holo pairs, and COACH420 respectively. The default backend completes in a median of 2.6 seconds per chain on one CPU core, so ensemble-based pocket finding does not require a simulation budget. Every site carries a continuous crypticity score, and outputs are emitted as docking-ready Boltz YAML constraints, AutoDock Vina boxes, pseudoatom PDB files, and the generated conformational ensemble as a multi-model PDB. Lacuna is MIT licensed and available at https://github.com/mooreneural/lacuna and on PyPI as lacuna-pockets.
bioinformatics2026-09-17v2Unlocking Sensitive Data with SPHERE in the Age of AI
He, Z.; Park, J.; Pulgrossi, R. C.; Lee, J.; Butler, R. R.; Weber, A.; Tian, L.; Zhang, X.; Wang, J.; Sha, S.; Mormino, E. C.; Wyss-Coray, T.; Henderson, V. W.; Longo, F. M.; Zou, J.; Desai, M.; Altman, R.Abstract
Sensitive human data underpin discoveries across medicine, biology and the social sciences, yet privacy regulation often prevents sharing them with collaborators or artificial intelligence (AI) systems. We introduce SPHERE, a model-free method that makes sensitive datasets directly usable by AI and shareable for open science as a synthetic twin, while the original records never leave the local environment. Across 33 datasets spanning five scientific domains, SPHERE protects individual privacy against adversarial re-identification attacks while preserving the data's statistical structure: means, variances and correlations are reproduced exactly, effect size and P value in linear statistical analysis is numerically identical, nonlinear machine-learning utility is retained, and each twin is generated in seconds on a laptop. Frontier AI agents running on the twin reach the same scientific conclusions as on the original records. Analyses of the twin reproduce genome- and proteome-wide results at UK Biobank scale and recover the findings of landmark studies across three independent cohorts and consortia. The approach also extends to deep-learning embeddings across language, vision and time-series, with minimal utility loss. We make the Stanford Alzheimer's Disease Research Center cohort openly available for the first time, as a SPHERE twin spanning nine modalities that any registered researcher can analyze without an approval process. We release SPHERE with certification of each twin's privacy and fidelity, and an AI agent that autonomously executes research tasks on sensitive data without ever accessing it. Sensitive datasets that are currently closed to research could thus become routine inputs to open science and AI to enable key discoveries.
bioinformatics2026-09-17v2