Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Early terminated transcripts and missing proteins reflect artifacts in bacterial proteomes
Insana, G.; Martin, M. J.; Pearson, W. R.Abstract
The high redundancy of many bacterial proteomes can be used to evaluate proteome quality and identify sequence errors. We have used MMseqs2 clustering with subsequent filtering to identify clusters that contain sequences from at least 50% of the clustered proteomes to build sets of core proteins that include proteins from 95% of the clustered bacteria. These clusters typically capture more than 80% of proteins in the bacteria. Because these clusters have highly uniform length (the median cluster has more than 99% of its proteins at the mode length), short (<75% of mode length) or long (>133%) proteins are likely artifacts. Most "outlier" proteins are found in fewer than 10% of clusters, and "high-outlier" clusters are over-represented in a small fraction of proteomes, which often have poor proteome BUSCO fragment scores. Short-outlier proteins are artifacts; at least 80% of short-outlier genomes contain mode-length copies of the protein, which were missed because of frame-shifts, termination codons, or initiation codon choice. MMseqs2 clustering with 50% participation provides robust sets of core bacterial proteins and can be used to identify lower-quality proteomes and proteins.
bioinformatics2026-09-22v6Decoding the Functional Interactome of Non-Model Organisms with PHILHARMONIC
Sledzieski, S.; Ocitti, F.; Versavel, C.; Singh, R.; Devkota, K.; Kumar, L.; Shpilker, P.; Roger, L.; Yang, J.; Lewinski, N.; Putnam, H. M.; Klein-Seetharaman, J.; Berger, B.; Cowen, L.Abstract
Despite the widespread availability of genome sequencing pipelines, many genes remain part of the genome's "dark matter," where existing inference tools cannot even begin to guess the biological function of their proteins from sequence alone. This challenge is especially pronounced in organisms that are highly evolutionarily distant from well-studied models, where homology-based methods break down. Here, we describe PHILHARMONIC, a computational method that combines deep learning-based de novo protein interaction network inference with robust unsupervised spectral clustering and remote homology to illuminate functional organization in any non-model organism. From only a sequenced proteome, we show PHILHARMONIC generates meaningful functional hypothesis for protein functions, functional communities, and higher-order network structure. We validate its performance using experimental gene expression and pathway data in D. melanogaster, and we demonstrate its broad utility by analyzing temperature sensing and stress response pathways in the reef-building coral P. damicornis and its algal symbiont C. goreaui. PHILHARMONIC provides a general-purpose engine for functional discovery and biological hypothesis generation in non-model organisms, enabling systems-level insights across the full diversity of life.
bioinformatics2026-09-22v4A Bayesian modelling framework for inference of latent infection risk patterns from virus neutralisation assay titration data
Alrefae, T. A.; Pons-Salort, M.; Donnelly, C. A.; Lambert, B.; Kamau, E.Abstract
Serological assays remain the standard approach for estimating the cumulative incidence of a pathogen and monitoring population immunity. Titration data from virus neutralisation assays are conventionally analysed using a nearly century-old interpolation-based method that neglects inherent imperfections in the assay and produces estimates with no measure of uncertainty. We introduce a two-part Bayesian framework for inferring latent infection risk directly from raw titration data. First, we develop a mechanistic model for serum antibody titration data that estimates antibody concentrations while quantifying uncertainty. Second, we propagate this uncertainty into an age-structured serocatalytic mixture model by integrating over posterior draws of individual antibody concentrations, allowing joint inference on latent serostate membership, force of infection, and serological waning rate. We applied the framework to three cross-sectional serosurveys of the population of England, for enteroviruses A71 (EV-A71) and D68 (EV-D68) and coxsackievirus A6 (CVA6), which are leading causes of severe respiratory illness and hand, foot, and mouth disease. We estimated consistently higher and more persistent lifetime antibody concentrations for EV-D68 than for EV-A71 and CVA6. The proportion of recently infected individuals peaks around 40% by age 5 years for CVA6 and around 32% by age 7 years for EV-A71, before declining with age. These estimates are conditional on a model structure excluding complete seroreversion, which, where identifiable, implied a half-life of several decades. For EV-D68, the inferred proportion previously infected exceeded 50% by age 13 years and grew, with the force of infection declining more gradually with age. These estimates differ from those previously obtained from binarised versions of the same data. Our framework uncovers the wide-ranging variation in antibody levels that are often obscured by conventional endpoint titre methods and infers infection rates directly from raw titration data, without dichotomising at predetermined seropositivity cut-offs, while making minimal assumptions about virus-specific infection mechanisms.
bioinformatics2026-09-22v3Quantifying the Rearrangement Complexity of Pangenomes
Bohnenkaemper, L.; Stoye, J.Abstract
The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.
bioinformatics2026-09-22v2Two Comparators May Be All We Need
Muskal, S. M.Abstract
A compound in a cell meets a spectrum of proteins drawn from many families at once, while screening most often interrogates one target at a time. Two questions asked many times in a rank ordering workflow ultimately guide decisions on what gets made and what gets counter-screened, and both are comparative: which of two targets does a compound prefer, and which of two compounds does a target prefer. We built one model for each: the target comparison over a roster of 2,279 human proteins in 34 protein families, the compound comparison over 2,079 in 32. Each model is given two chemical structures and a sequence, or two sequences and a chemical structure, and returns which member of the pair is preferred together with how firmly it holds that view. No conformational analysis, protein structure, binding site or docked pose is used. Asked which of two targets a compound prefers, the model is correct 0.75 of the time over 8,689 held-out comparisons setting two families against each other, and 0.78 of the time over 32,738 comparisons between two targets of one family, rising to 0.94 and 0.95 on the most confident third of each. Asked which of two compounds a single target prefers, it is correct 0.71 of the time over 65,725 held-out comparisons, rising to 0.96 on the most confidently held. Neither compound in any of those comparisons appeared anywhere in training. Accuracy follows the gap between the two measurements. Where they sit within half a log unit the models are right 0.57 to 0.64 of the time, and where they differ by more than two logs, 0.90 to 0.92. The compound result holds across 31 protein families, not only the best-measured one. Within the chemistry and the targets they were built on, these models rank compounds and rank targets well. Both models can be explored and downloaded at familyfoundationmodel.com. Keywords: target preference; compound preference; pairwise comparison; polypharmacology; off-target triage; ESM2; random forest; structure-free prediction; ChEMBL
bioinformatics2026-09-22v2MetaboCensoR: A Shiny Application for Data Filtering in Untargeted LC-MS Metabolomics to Enhance Interpretability
Plyushchenko, I. V.; Luzzatto-Knaan, T.Abstract
Untargeted LC-MS metabolomics datasets often contain large numbers of redundant and non-informative features arising from background contaminants, multiple ion forms, poorly integrated peaks, and other low-quality signals. These features complicate downstream analysis by inflating feature space, degrading molecular networks, impeding pathway analysis, and obscuring statistically meaningful changes. Here, we present MetaboCensoR, an input-versatile Shiny application and local R package that bridges the gap between LC-MS peak picking and downstream interpretation through analyte-centric peak table filtering. The workflow integrates four fully interactive modules for blank, ion-species, quality-control, and peak-based filtering, with synchronized processing of associated spectra files. The tool was evaluated across three independent datasets covering plant extracts, human cell lines, and bacterial interactions. Across these case studies, data filtering reduced feature redundancy and improved downstream interpretation in feature-based molecular networking, pathway-level functional analysis, and differential abundance testing, while retaining known analytes. MetaboCensoR was systematically benchmarked against existing tools, demonstrating comparable ion-species annotation coverage and a favorable overall balance of target recovery, adduct precision, and feature reduction. Together, these evaluations support MetaboCensoR as a robust approach for systematic peak table curation, enhancing interpretability and analytical value of untargeted metabolomics data.
bioinformatics2026-09-22v2A Graph-based QSAR Modeling Pipeline for Predicting In vitro PubChem Assays and In vivo Human Hepatotoxicity: Mechanistic Analysis of Caspase-3/7 Activation
Chitikela, Y.; Zhu, c.; Jia, Z.Abstract
Background: Caspase-3 and -7 are key effector caspases in the apoptotic pathway, a form of programmed cell death, and their activities serve as a well-established biomarker for evaluating environmental chemical toxicity and informing chemical risk assessment. Loss of mitochondrial membrane potential is a key event in the activation of Caspase-3/7 signaling and the subsequent induction of apoptosis. Therefore, simultaneous assessment of mitochondrial membrane potential and Caspase-3/7 activity enables elucidation of the mechanisms and pathways through which apoptosis is initiated. Rapid and accurate assessment of the potential toxicity of environmental chemicals and drugs remains a major challenge. Quantitative Structure Activity Relationship (QSAR) modeling have been widely used for toxicity prediction. Graph-based approaches encode compounds directly as molecular graphs, allowing structure-activity relationships to be learnt from molecular topology without the information loss in binary fingerprints. While advanced graph models such as graph transformers (GTs) have shown outstanding performance in many domains, they have not been fully leveraged in QSAR modeling on Caspase and mitochondrial toxicity. Methods: We propose a QSAR modeling pipeline that encompasses assay data preprocessing, feature representations (fingerprints and molecular graphs), and benchmarking machine learning (ML) models, including classic ML models, graph neural networks (GNNs), GTs, and their consensus ensembles. Based on in vitro Caspase and mitochondrial assays in PubChem, we applied the pipeline to predict Caspase-3/7 activation and mitochondrial membrane potential (MMP). Beyond in vitro assays, we also built in vivo QSAR modeling for FDA Drug-Induced Liver Injury (DILI) gold standard on human hepatotoxicity. Moreover, mechanistic analysis on Caspase-3/7 activation was conducted by comparing with MMP disruption to identify chemical substructures that may be responsible for dual activations. We also investigated cell-line-specific responses by identifying structural motifs that selectively induce Caspase-3/7 activation in individual cell lines.Results:Experimental evaluations show that GTs and GNNs outperformed classic ML models when the number of active compounds is large, such as MMP disruption, while classic ML models and GTs performed good for highly imbalance data with limited active compounds, such as Caspase-3/7 activation. For DILI prediction, the full consensus model achieved the highest AUC 0.69 and Graphormer had the highest F1 score 0.79, both surpassing the previous best model with AUC 0.63 and F1 0.65 with a large margin.Our mechanistic analysis shows that phenolic compounds bearing a para-hydroxyphenyl motif, as well as members of the lipophilic chain family with long alkyl chains can trigger the collapse of MMP, leading to the activation of caspases-3 and -7. Human embryonic kidney (HEK293) was the only cell line with a distinct structural motif: 1,1-dichloroethane and chlorobenzene. Human neuroblastoma (SK-N-SH) is uniquely impacted by an epoxide fragment and rat hepatoma (H-4-II-E) is uniquely impacted by a tetramethylcyclohexene motif and an acetaldehyde fragment.Conclusions:The proposed pipeline for QSAR modeling, including data preprocessing, feature representations, and incorporation of advanced graph ML approaches, is highly effective in predicting not only on Caspase-3/7 activation and membrane potential collapse, but also on FDA DILI human hetatotoxicity. As future research directions, we will leverage extra information, e.g., biological activity and findings in existing toxicity literature, and recent advances in large language models and agentic AI to further improve the predictive performance and enable a sensitive and specific framework for assessing human hepatotoxicity of environmental compounds.
bioinformatics2026-09-22v2GO-Term Enrichment of Proteome-Scale Docking Profiles as a Biological Search-Space Reduction Layer for Protein Target Discovery
Harms, C. M.; Klein-Seetharaman, J.Abstract
Identifying protein targets from phenotype-first or mechanism-uncertain compounds remains difficult because proteome-scale docking can generate thousands of structurally plausible interactions per compound. We developed a workflow that converts ranked proteome-scale docking profiles into stable Gene Ontology (GO) Biological Process enrichment signatures and evaluates whether those signatures can reduce the candidate target space while preferentially retaining known drug-target relationships. Docking targets were retained at the PDB-chain level, mapped to unique human gene identities, and analyzed with PANTHER overrepresentation against the screened structural gene universe. GO enrichment was evaluated from the top 25 through 505 ranked proteins in increments of 10, and a stable compound-level GO profile was selected using a Jaccard stability threshold of 0.80 across three consecutive transitions. Benchmarking used Yamanishi drug-target interactions, with 682 mapped compounds assigned to a prespecified development/held-out split (552/130) and evaluated at canonical, mechanistic, and fine-mechanistic biological resolutions. In the full corrected benchmark, 642 compounds with usable canonical GO profiles showed greater within-class than between-class similarity (0.1019 versus 0.0931; delta = 0.0088; 100,000-permutation p = 0.00222), and canonical class explained 1.34% of multivariate GO-profile variation by PERMANOVA (p < 1e-4). Fine-mechanistic labels showed stronger organization in the full dataset (delta = 0.0245; PERMANOVA R-squared = 0.0976; both p < 1e-4). Held-out validation was more modest and metric-dependent: mechanistic labels were significant by PERMANOVA (R-squared = 0.0626, p = 0.0437), whereas the frozen fine-mechanistic analysis showed greater within-class similarity (0.1350 versus 0.1110; p = 0.038) and significant nearest-neighbor recovery (p = 0.0495), but not significant PERMANOVA (p = 0.119). The principal held-out search-space experiment evaluated 107 compounds, 753,492 candidate protein rows, and 424 represented gold-standard targets. A direct GO gate retained 1.26% of candidates while retaining 11.32% of known targets (8.96-fold enrichment); ontology-propagated GO associations retained 5.10% of candidates and 21.46% of known targets (4.21-fold enrichment). At matched candidate-space sizes, GO-associated prioritization retained 27.59% versus 22.41% of known targets at approximately 5% of candidates, 39.39% versus 36.08% at 10%, and 58.73% versus 56.13% at 20%. By contrast, additive protein-level GO reranking was heterogeneous: among 605 evaluable compounds, 21.7% improved their best known-target rank, but mean reciprocal rank decreased from 0.0276 to 0.0189. These results support GO enrichment as an intermediate biological search-space reduction and prioritization layer rather than a universal direct target-scoring function. Keywords: proteome-scale docking; Gene Ontology; target discovery; targetome; PANTHER; target fishing; biological filtering; search-space reduction; Yamanishi benchmark
bioinformatics2026-09-22v1A dependency-free, streamable format and cross-language toolkit for scalable LC-MS feature detection: reading only what you need
Osorio Mosquera, J.; Lawler, N. G.; Cox, M.; Osorio Mosquera, J.; Moreno L, W. A.; Sala, S.; Nambiar, V.; Whiley, L.; Nicholson, J. K.; Holmes, E.; Wist, J.Abstract
Mass spectrometry generates data faster than it can be read, and the exchange standard, mzML, is text-based and must be parsed in full before any spectrum is accessible. Binary alternatives are smaller but depend on storage engines such as HDF5, so access is dictated by the engine, not the file. We present Ionic (.ion), an open-source, compact, streamable binary format, and Quant{middle dot}ion, a processing toolkit built on the former. Ionic stores spectra, chromatograms and metadata as independently compressed, indexed blocks, so a reader retrieves only the bytes it needs, even inside a web browser, and converts losslessly to and from mzML. Ionic was smaller than compressed mzMLb on every acquisition type tested, and extracting one compound took under 40 ms, 35 to 90 times faster than an mzML reader. Quant{middle dot}ion exposes one core to R, Python, and JavaScript with identical results, and recovered 97% of true features on a ground-truth benchmark.
bioinformatics2026-09-22v1Benchmarking Uncertainty and Improving Risk Ranking in Multi-task Bioactivity Prediction
Wang, Z.; Tian, L.; Che, C.-M.Abstract
Multi-task bioactivity prediction transfers information across assays, but experimental prioritization also requires uncertainty estimates that identify unreliable predictions. We benchmarked prediction accuracy and uncertainty calibration for five multi-task predictors on 100 ChEMBL27 test assays and 100 CHEMBL37 assays under the random split and the clustering split. Calibration used native, Deep Ensemble and MC Dropout uncertainty; risk ranking was compared for two Gaussian process (GP) backbones. Adaptive deep kernel fitting (ADKF) achieved the best mean calibration results, while deep kernel transfer (DKT) ranked errors more effectively than ADKF with native uncertainty. We also introduce Influence Calibrated Support Reconstruction (ICSR), a risk score for kernel-based predictors. ICSR combines measured support reconstruction errors with query-specific influence and stabilizes their weighted average toward the full-support mean to adjust Gaussian process uncertainty while preserving predicted activities. With DKT and ADKF, ICSR improved all three mean risk-ranking metrics over native uncertainty and an adapted neighborhood comparator in every panel and split. ICSR achieved a mean half-query MAE reduction (R50) of 14.2-19.6%, where R50 measures the percentage decrease in mean absolute error (MAE) after retaining the lowest-risk half of the queries. A retrospectively selected SARS-CoV-2 main protease case further illustrated its use for selecting more reliable predictions. The benchmark supports joint assessment of calibration and error ranking, while ICSR improves selective use of kernel-based bioactivity predictions.
bioinformatics2026-09-22v1MINT infers the latent single-cell spatial transcriptome from paired histology
Chen, L.; Zhang, Z.; Wang, B.; Ren, W.; Peng, H.; Lu, S.; Xu, M.; Zhou, Y.; Luo, X.; Ye, T.; Zhu, Q.; Yan, F.; Tian, L.Abstract
Sequencing-based spatial transcriptomics captures transcriptome-wide molecular information, yet the corresponding cell-resolved tissue map remains incomplete. Histology records dense cellular morphology, tissue architecture and local context. Here, we present MINT (Morphological Integrative Network for Transcriptomics), which combines this latent cellular organisation with sample-specific spatial measurements to infer tissue-wide, cell-resolved transcriptomes without an external single-cell reference. MINT reconstructs expression from sparse Slide-tags profiles across unmatched adjacent sections; its shared nucleus-centred representation also resolves pooled Visium measurements into cell-level profiles. MINT improved Slide-tags alignment and expression recovery under controlled nuclear loss and cross-section perturbation. It also outperformed reference-free Visium disaggregation methods against Xenium ground truth and was validated with Visium HD sequencing. In human lung adenocarcinoma, MINT reconstructed the tumour immune microenvironment cell by cell and resolved tertiary lymphoid structures and their maturity. These results establish paired histology as a quantitative, sample-intrinsic complement to molecularly rich but cellularly incomplete spatial-transcriptomic assays.
bioinformatics2026-09-22v1StackHPpred: A Stacking-based ensemble learning framework for the identification of peptide hormones using multi-view feature representations
Kirthipati, J.; Chemarthi Ravi, M. K.; Kirthipati, M.Abstract
Peptide hormones are important signaling molecules that regulate diverse physiological processes and have substantial therapeutic relevance. However, their experimental identification can be challenging because of low abundance, limited stability, and complex post-translational processing, highlighting the need for reliable computational approaches. Here, we present StackHPpred, a stacking-based ensemble framework for accurate identification of hormone peptides (HPs) from amino acid sequences. We systematically evaluated 56 feature representations, comprising 35 conventional sequence descriptors and 21 PLM/NLP-based embeddings, across 10 machine-learning classifiers, generating 560 baseline models. Prediction-level correlation analysis was subsequently employed to identify 14 complementary, non-redundant feature representations, which were integrated through feature- and classifier-level stacking strategies using out-of-fold predicted probabilities. The final StackHPpred model, based on the top three feature representations CTDD, KSCTriad, and PTAB, achieved an AUC of 0.9965, MCC of 0.9506, ACC of 0.9752, and F1-score of 0.9749 on the independent dataset. StackHPpred was further evaluated under increasingly imbalanced independent datasets with positive-to-negative ratios ranging from 1:10 to 1:50, demonstrating robust predictive performance under challenging screening conditions. t-SNE visualization and SHAP analysis further showed improved class separation and highlighted the contributions of multiple classifier-feature combinations to the final prediction. Overall, StackHPpred provides a robust and interpretable framework for distinguishing HPs from non-HPs and is freely available as a web server at https://jayasreekirthipati.org/StackHPpred/, with a standalone implementation available at https://github.com/JayasreeKirthipati04/StackHPpred.
bioinformatics2026-09-22v1DeepSCENIC: transfer learning from sequence-to-function models enables causal gene regulatory network inference
Partel, G.; De Winter, S.; Konstantakos, V.; Blaauw, C. H.; Aerts, S.Abstract
Sequence-to-function (S2F) deep learning models have become an important aid to decipher the genomic cis-regulatory code. However, current S2F models do not take the cellular trans-environment of transcription factors (TF) into account. Conversely, methods for gene regulatory network (GRN) inference often rely on heuristics or simple position weight matrices (PWMs), without exploiting the combinatorial grammar of genomic enhancers. Here, we present DeepSCENIC, a deep learning framework that enables causal GRN inference by performing transfer learning from S2F models to single-cell multiome atlases. We first test and validate DeepSCENIC on ENCODE cell lines, demonstrating that the framework accurately predicts single-cell gene expression and chromatin accessibility by leveraging pretrained S2F models like Enformer and Borzoi. We show that DeepSCENIC recovers TF-region (TF-RE) interactions with high fidelity, and captures de novo TF binding motifs without prior PWM knowledge. The model improves enhancer-gene associations (RE-TG) over correlation-based baselines when benchmarked against large-scale CRISPRi screens. After training a DeepSCENIC model, it enables the prediction of perturbation effects during cell state changes by acting as a mechanistic simulator. In a melanoma cell line atlas, the model accurately recapitulates the transcriptional shift from melanocytic to mesenchymal states, and predicted knock-down effects show high concordance with experimental time series data. Finally, we use DeepSCENIC to identify mouse-human cortex conserved GRNs, finding high cross-species concordance in TF activity programs across matched neuronal subclasses and validating top-ranked enhancers against experimental reporter assays. By unifying S2F enhancer representations with single-cell multiomics, DeepSCENIC provides a new paradigm for jointly modeling and simulating cis-sequence and trans-cellular perturbations.
bioinformatics2026-09-22v1Spatial transcriptomics reveals selective vulnerability of cardiac neural crest-derived glial cells in human myocardial infarction
Xia, X.; Ding, Y.; Zhao, X.; He, Q.; Zhang, H.; Gu, Y.; Bai, H.; Pan, D.Abstract
Cardiac neural crest-derived glial cells (CNGs) are recently identified glial cells that support cardiac autonomic innervation and maintain sympathetic-parasympathetic balance. Their fate in myocardial infarction (MI) is unknown. Here, we analyzed 16 human cardiac spatial transcriptomes spanning normal myocardium (n=4), infarct zone (n=5), border zone (inner, n=1; outer, n=3), and remote zone (n=3). CNGs (S100B+/GFAP+, CD68- spots) were selectively lost along a spatial gradient from healthy tissue to the infarct core (Spearman rho = -0.92, p = 4.5 x 10-7). Tissue-area-corrected analysis confirmed that CNG loss (26% retained in infarct zone vs. normal; p = 0.016) vastly exceeded overall tissue loss, indicating selective vulnerability rather than bulk tissue destruction. Mechanistically, oxidative phosphorylation (OXPHOS) activity in CNG-positive spots correlated positively with CNG survival (partial rho = 0.65, p = 0.009 after controlling for spot composition), consistent with a high metabolic demand underlying their ischemic sensitivity. Pyroptosis - but not ferroptosis or apoptosis - signatures correlated with CNG loss (partial rho = -0.67, p = 0.006). These findings identify CNGs as selectively vulnerable cellular components of the cardiac autonomic system in MI and nominate metabolic failure and pyroptosis as candidate mechanisms, providing a spatially resolved foundation for future mechanistic and interventional studies.
bioinformatics2026-09-22v1baltic: the Backronymed Adaptable Lightweight Tree vIsualisation Code
Potter, B. I.; Gangavarapu, K.; Torres Jimenez, M. F.; Bell, S. M.; Dudas, G.Abstract
For ten years, baltic (Backronymed Adaptable Lightweight Tree vIsualization Code) has been used to make annotated phylogeny figures in molecular epidemiology, including during the West African Ebola epidemic, the Zika epidemic in the Americas, and the SARS-CoV-2 pandemic, as well as other fields. At its core, baltic is a Python library used for the efficient parsing, traversal, manipulation, and visualisation of phylogenetic trees. It is a small library with few dependencies that reads tree formats common in phylodynamic analyses and gives the user leverage to interact with a lightweight tree data structure to produce publication-ready figures with matplotlib. We present baltic v1.0, its first formally released and documented version. baltic reads and writes BEAST Nexus, Newick, and Nextstrain/Auspice JSON, and can process large BEAST posterior tree files in parallel to extract user-defined posterior statistics. The same objects are used for tree manipulation and for plotting in a single script. New to this release: a set of rooting methods (midpoint rooting, rerooting on any branch, and root-to-tip regression); support for reticulate evolution, with reassortment and recombination edges; and composite figures that combine a tree with other data, such as Muller plots, skygrid plots, tanglegrams, and plots connecting trees to maps. The release includes a documentation site with an API reference, tutorials, and a matplotlib-style gallery of worked examples.
bioinformatics2026-09-22v1Selective depletion of upper-layer somatostatin interneuron subtypes in schizophrenia
Endresz, N.; Fafouti, M. E.; Arbabi, K.; Zhou, X.; DeLong, T.; Gonzalez-Burgos, G.; Duncan, L.; Sibille, E.; Tripathy, S. J.Abstract
Schizophrenia (SCZ) is associated with cortical GABAergic dysfunction, but whether inhibitory interneurons are lost or persist in an altered molecular state remains unresolved. Here, we harmonized seven prefrontal post-mortem single-nucleus RNA-seq datasets onto a fine-grained taxonomy of cortical cell types and meta-analyzed their gene expression and abundance changes in SCZ (298 controls, 171 SCZ). First, we find a subclass-wide reduction of SST mRNA within somatostatin (Sst) neurons. Second, we find reduced abundance (depletion) of a subset of upper-layer Sst interneurons and increased abundance of L6b excitatory neurons, with both changes confirmed in spatial transcriptomics (12 controls, 12 SCZ). Notably, SCZ genetic risk is enriched in the most depleted Sst cells. Depleted Sst subtypes highly express HCN1, partially correspond to primate-specialized CALB1-expressing double-bouquet cells, and are among the cells lost earliest in Alzheimer's disease. These upper-layer Sst interneurons constitute an intrinsically vulnerable population and a promising target for neuroprotective and compensatory therapies.
bioinformatics2026-09-22v1Covariance Nonstationarity is Evident in Spatial Transcriptomics and Provides a New Categorization of Spatially Varying Genes
Velidi, P.; Wei, Z.; Nathoo, F.Abstract
Gaussian process models underlie many spatial transcriptomics tools but typically assume stationary covariance. Covariance non-stationarity has long been recognized in spatial statistics as an important feature of spatial data, yet it has received little attention in spatial transcriptomics. We show that this omission is consequential: covariance non-stationarity is substantially evident across spatial transcriptomic datasets and alters the characterization of spatially varying genes. While typically ignored, non-stationarity of spatial covariance in gene expression may correspond to tissue heterogeneity or cell aggregates. Across 12 Visium datasets, we use approximate Bayes factors from R-INLA to compare stationary and non-stationary Mat'ern covariance functions. Evidence for covariance non-stationarity appears in 3% to 50% of genes across tissue samples. We find that gene sets associated with immune, cytokine, and other effector functions are enriched among genes favoring non-stationary spatial covariance. Covariance stationarity is therefore not a benign technical simplification in spatial transcriptomics; it is frequently violated, the violation is biologically structured, and it changes the definition and classification of spatially varying genes.
bioinformatics2026-09-21v3Product-stabilized filamentation by human glutamine synthetase allosterically tunes metabolic activity
Greene, E.; Muniz, R.; Yamamura, H.; Hoff, S. E.; Bajaj, P.; Tecson, M. C. B.; Geluz, C.; Lee, D. J.; Thompson, E. M.; Arada, A.; Lee, G. M.; Bonomi, M.; Kollman, J. M.; Fraser, J. S.Abstract
To maintain metabolic homeostasis, enzymes must adapt to fluctuating nutrient levels through mechanisms beyond gene expression. Here, we demonstrate that human glutamine synthetase (GS) can reversibly polymerize into filaments aided by a composite binding site formed at the filament interface by the product, glutamine. Time-resolved cryo-electron microscopy (cryo-EM) confirms that glutamine binding stabilizes these filaments, which in turn exhibit reduced catalytic specificity for ammonia at physiological concentrations. This inhibition appears induced by a conformational change that remodulates the active site loop ensemble gating substrate entry. Metadynamics ensemble refinement revealed >10 [A] conformational range for the active site loop and that the loop is stabilized by transient contacts. This disorder is significant, as we show that the transient contacts which stabilize this loop in a closed conformation are essential for catalysis both in vitro and in cells. We propose that GS filament formation constitutes a negative-feedback mechanism, directly linking product concentration to the structural and functional remodeling of the enzyme.
bioinformatics2026-09-21v3EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
Ding, H.; Nie, W.; Wu, N.; Qiu, T.Abstract
Hybrid DNA foundation models combine convolutional, recurrent, and attention layers, making speculative decoding more difficult than truncating a KV cache. We present EvSpark, a speculative decoding system for Evo2 that verifies draft blocks in parallel and restores all three classes of inference state by selecting retained intermediate states, without replay. A compact, hidden-state-conditioned drafter proposes each block in one parallel forward pass. On Evo2 7B, a 48-prompt benchmark with three training seeds yields 2.96x on 43 real-sequence prompts and 3.27x including the five synthetic controls. Acceleration persists at 262k-token context (1.84x - 2.43x on two bacterial genomes) and over 32k generated tokens. Retraining the same drafter architecture for Evo2 20B and 40B yields 2.18x - 2.46x on real sequences and 2.51x - 2.78x on the full suite. Autoregressive drafter comparisons and batch measurements show why low draft latency, rather than acceptance alone, determines the gain. The method preserves the target distribution in exact arithmetic. In bf16, greedy tests find no non-tie divergences across 48 prompts and six checkpoints; sampling tests expose residual numerical sensitivity, especially in repetitive sequences. A cost-efficient 7B drafter requires 1.06 incremental GPU-hours of training, excluding teacher-data collection, and achieves 2.82x on real sequences. In regulatory-DNA design, EvSpark achieves a median complete-workflow speedup of 1.57x over a calibrated batched native baseline.
bioinformatics2026-09-21v3Is level-1 blob reconstruction under the network multispecies coalescent easy?
Dai, J.; Molloy, E.Abstract
Hybridization is an important evolutionary process, commonly modeled by the network multispecies coalescent. Reconstructing evolutionary histories under this model is notoriously costly, even for level-1 networks where hybridization events are isolated from each other. The widely used methods that combine speed with statistical guarantees rely on quartet concordance factors computed for all subsets of four species, resulting in an O(n^4k) bottleneck that severely limits scalability to large numbers of species (n) and genes (k). Among quartet-based methods, NANUQ+ is notable because it decomposes the problem into two steps: first reconstructing a tree of blobs, which compresses each non-treelike part of the network, called a blob, into a single vertex, and second reconstructing the internal structure of each level-1 blob, specifically its circular order and hybrid vertex. Here, we investigate whether level-1 blob reconstruction is difficult once the tree of blobs is known. We present a fast and statistically consistent algorithm, called NetCS, based on two simple primitives: majority voting and merge sort, circumventing the bottleneck of computing all quartet concordance factors. In simulations, NetCS achieved comparable accuracy to NANUQ+ and was dramatically faster, enabling analyses of 200 taxa and 1000 genes in only a few minutes. Both methods attained near-perfect accuracy when given the true tree of blobs; however, their performance degraded in end-to-end pipelines due to errors in tree of blobs reconstruction. Strikingly, even methods that reconstruct level-1 networks directly struggled to accurately predict hybrid ancestry. Our results suggest that reconstructing level-1 blobs is unexpectedly easy once the tree of blobs is known, and that a major challenge for phylogenetic network inference lies in accurate tree of blobs reconstruction.
bioinformatics2026-09-21v2Integrating structural homology with deep learning to achieve highly accurate protein-protein interface prediction for the human interactome
Xiong, D.; Torres, M.; Murray, D.; Zhang, Z.; Li, L.; Naravane, A. C.; Fragoza, R.; Honig, B.; Yu, H.Abstract
A significant portion of disease-causing mutations occur at protein-protein interfaces however, the number of structurally resolved multi-protein complexes is extremely small. Here we present a computational pipeline, PIONEER2, that integrates 3D structural similarity with geometric deep learning to accurately predict protein binding partner-specific interfacial residues. We compare the performance of PIONEER2 to that of AlphaFold3 and found, using a test set of PDB structures, that their performance is quite similar. However, about 20% of AlphaFold3 predictions for protein-protein complexes in the PDB have AlphaFold3 ranking scores below 0.5, which indicates an uncertain model. For these structures, PIONEER2 outperforms AlphaFold3 at discriminating interfacial from non-interfacial residues. Further, about half of the AlphaFold3 ranking scores on high confidence protein-protein interactions (PPIs) not associated with a PDB structure are below 0.5 indicating that PIONEER2 offers superior interface prediction for a large number of PPIs for which structures are not available. We created a comprehensive 3D structurally informed interactome encompassing all 352,124 experimentally detected binary human PPIs in the current literature and made PIONEER2 interface predictions for each. We experimentally validated these predictions by generating 1,866 mutations and testing their disruptive impact on 5,010 mutation-interaction pairs. PIONEER2-predicted interfaces are found to be comparable to PDB structures in their ability to predict disruptive mutations while AlphaFold3 performance is reduced. Similarly, PIONEER2-predicted interfaces outperform AlphaFold3 in accounting for the depletion of non-deleterious common population variants and the enrichment of disease-related mutations on protein surfaces. Overall, our results suggest that PIONEER2-predicted interfaces provide a valuable tool for studying disease etiology, advancing personalized medicine and for fundamental research. We further implemented PIONEER2 as a user-friendly web server (https://pioneer2.yulab.org) platform for users to explore our 3D interactome models and conduct genome-wide functional genomics studies.
bioinformatics2026-09-21v2Benchmarking confidence estimation and rescoring for cyclic peptide-protein complex predictions
Li, Z.; Yuan, Y.; Hu, K.; Pan, P.; He, F.Abstract
Cyclic peptides are a rapidly expanding class of therapeutics, but the reliability of deep-learning structure prediction for cyclic peptide-protein complexes has not been systematically evaluated. We assembled a curated benchmark of 111 nonredundant complexes spanning five cyclization chemistries and assessed two co-folding models, Boltz and Protenix, each generating 100 poses per target (22,200 total). Stratifying all poses by complex attributes, we found that disulfidecyclized peptides and small protein targets (200 or fewer target residues) were predicted significantly worse by both tools, with target size the largest and most consistent effect; overall accuracy nevertheless remained high (median top-pose DockQ of about 0.89, 96-98% of targets Acceptable or better), indicating that pose generation is rarely the bottleneck. Conversely, native model ranking scores correlated only moderately with pose quality (Spearman rank correlations of 0.53-0.66): approximately 12% of poses showed high model ranking score/confidence despite poor pose DockQ quality, and the highest-quality pose was not ranked first for nearly every target. We therefore augmented the native score with externally computed interface descriptors normalized by chain length, principally the per-residue density of inter-chain hydrogen bonds, in a gradient-boosted rescoring model evaluated under target-grouped cross-validation that prevents leakage, improving out-of-fold ROC-AUC for both tools, significantly so for Protenix. Together, these findings identify pose ranking, rather than pose generation, as the major limitation of current cyclic peptide-protein complex prediction and demonstrate that complementary structural features can improve confidence-based pose selection.
bioinformatics2026-09-21v2Optimized Multiple Circular Sequence Alignment for Cyclic Peptide Motif Discovery
Yuan, Y.; Li, Z.; Hu, K.; Pan, P.; He, F.Abstract
Head-to-tail (H2T) cyclized peptides are an increasingly important modality in drug discovery, combining high target affinity and selectivity with metabolic sta-bility. Because their underlying chemistry is still that of a linear amino-acid chain, their linear sequence representation is the native input format of main- stream sequence generative models now driving de novo peptide design (Slough et al., 2018; Rettie et al., 2025a;b). Discovering the conserved motifs responsible for a family's function across a library of such candidates requires a multiple sequence alignment (MSA). Because a cyclic peptide can be linearised at any residue, the alignment must additionally solve for the unknown rotation of each sequence, which is the multiple circular sequence alignment (MCSA) problem. However, leading MCSA heuristics (e.g. Ayad & Pissis, 2017) were tuned for the genomic regime (a few tens of long sequences) and become prohibitively slow on the cyclic peptide library regime (hundreds to thousands of shorter sequences). We close this gap by identifying quality-preserving optimisation opportunities, notably the library-scale preset tailored to short-sequence inputs (algorithmic details in Appendix A), and by adding an orthogonal multi-core and SIMD backend for further performance tuning, which gives near-linear thread scaling on the pairwise-comparison stage. We validate the pipeline on a library of 1,000 H2T cyclized peptides of length 18 targeting the oncoprotein Mouse double minute 2 human homolog (MDM2) produced by an internal peptide-design engine. In this practical setup, the optimised MCSA recovers the underlying positional motif of MDM2 binders at the same fidelity as the original MCSA implementation while running over 650x faster. Our optimised MCSA tool thus enables library-scale cyclic peptide sequence alignment and is publicly available at https://github.com/IVB-Generative-Biology/mars-turbo.
bioinformatics2026-09-21v2Individual-level expression deconvolution and assessment of cross-sample variation
Kang, K.; Xie, K.Abstract
Recovering cell-type-specific gene expression from bulk RNA sequencing would facilitate the study of transcriptional variation among individuals. However, accuracy can differ substantially among genes and cell types. We describe a reference-informed Bayesian deconvolution framework and a score that identifies gene--cell-type pairs likely to have more accurate estimates of cross-sample variation. The score uses bulk counts, reference expression profiles, and estimated RNA proportions. Known component expression is used to train and evaluate the score, but is not needed to calculate predictions from a trained model. We evaluated the approach in a ROSMAP-derived simulation with 40 target donors, 2,000 genes, and seven cell types. Median gene-wise correlation was 0.801 for raw allocated counts and 0.296 after normalization within each donor and cell type. To evaluate the score, we divided genes into five sets, kept linked genes together, and scored each set using a model trained on the other four. Retaining approximately 20\% of pairs within each cell type increased the median normalized correlation to 0.622. Ranking pairs only by the estimated share of a gene's bulk RNA contributed by the cell type yielded 0.570 at the same retained count. These results show that observable information can help prioritize pairs with more accurately recovered cross-sample variation.
bioinformatics2026-09-21v1Correlation-aware discovery of co-occurring mutational signatures in cancer
Jin, H.; Geiger, B.; Glodzik, D.; Gulhan, D. C.; Park, P. J.Abstract
Somatic mutations in cancer genomes record the activities of diverse mutational processes. Mutational signature analysis has advanced mechanistic understanding of mutagenesis and informed clinical decision-making, yet existing methods assume independence among signatures--an unrealistic assumption that can produce composite or contaminated signatures, reduce detection power, and yield inconsistent results. Here we present Cornet (CORrelated NMF ExTraction), a framework for mutational signature discovery that explicitly models co-occurring processes and jointly infers signatures and their correlation structure. Benchmarking on simulated data shows Cornet more accurately recovers distinct signatures under strong correlations. Applied to cancer genomes, Cornet enables unsupervised discovery of the colibactin-associated signature SBS88 in oral cancers and identifies the tobacco smoking signature SBS4 in bladder cancer, where it was previously thought absent. Cornet also uncovers a novel mutational process implicated in early-onset colorectal cancer and a signature arising from the interplay between tobacco smoking and ERCC2-mutation-driven nucleotide-excision repair deficiency. Together, these results demonstrate that modeling correlations among mutational processes is essential for high-resolution signature discovery and dissecting the mutational etiology of human cancer.
bioinformatics2026-09-21v1Uniformly processed transcriptome-wide alternative splicing profiles for pediatric cancer research
Liang, C. E.; Shapiro, J. A.; Beale, H. C.; Taroni, J. N.; Vaske, O. M.Abstract
Alterations in regulatory processes like alternative splicing contribute to pediatric cancer development. Although splicing aberrations have been observed in pediatric leukemias, alternative splicing has yet to be studied in pediatric cancers at scale, due to a lack of uniformly processed, sample-level pediatric cancer splicing profiles with non-diseased tissue comparators. We address this need by quantifying splice event usage for a curated set of bulk RNA-seq datasets from the NCI's Therapeutically Applicable Research to Generate Effective Treatments (TARGET, n = 1152) and Genotype-Tissue Expression (GTEx, n = 1098) as a comparator. This Treehouse Splice Compendium is accompanied by a reproducible workflow that was used to generate the data in the compendium and reflects the largest known RNA-seq dataset processed by the splice quantification tool Shiba. The compendium is part of a suite of large, uniformly processed datasets aggregated by the UCSC Treehouse Childhood Cancer Initiative and Alex's Lemonade Stand Foundation's Childhood Cancer Data Lab, which include the Treehouse Expression Compendia, refine.bio, and the Single-cell Pediatric Cancer Atlas.
bioinformatics2026-09-21v1Gravlax: an annotation-independent molecular evidence archive for single-cell RNA-seq
Patro, R.Abstract
A cell-by-gene count matrix is the artifact of a single-cell RNA-seq experiment that is most often stored, shared, and reanalyzed. It is the output of a computation whose inputs are the sequenced molecules and a gene annotation, and while the molecules never change, the annotation is revised continually. Once the matrix has been produced, the evidence behind it can no longer be reinterpreted. Recovering that evidence means returning to raw reads or alignments that are large, costly to process, and frequently unavailable. We ask whether a concise representation of the molecules themselves can be extracted once and reused indefinitely, to quantify under any future annotation, to query and discover features that no annotation yet describes, and to pool evidence across cells and samples. Our starting observation is that the procedures that turn alignments into counts, gene assignment and UMI collapse, never read most of what an alignment file contains. They consume relations among molecules such as shared genomic geometry, shared placements, barcode identity, and the equality or near-equality of UMIs. We show that these relations form a statistic that is sufficient for such consumers, and we design a compact, seekable, content-authenticated archive that stores them while deferring every annotation-dependent decision to analysis time. Archives compose into content-addressed collections that route cohort queries to the molecules that can answer them without copying molecules. We implement these ideas in a tool called gravlax. Across four human 10x 3' datasets, gravlax archives require 11--18 bits per read and are 9.0--12.7x smaller than tag-preserving CRAM. Count matrices replayed from an archive deviate from direct STARsolo quantification by 0.24--0.75% of normalized UMI mass, whereas changing GENCODE v32 to v49 moves 2.12--4.64%, and quantification replay is 34--82x faster than STARsolo at matched thread budgets. A federated index over eight archives occupies 2.96% of their size, answers a 96-query junction panel 2.59x faster than the archives alone, and screens the cohort genome-wide for unannotated splice events that recur across donors in just 9 seconds. Because the molecules are retained, the archives also answer questions the matrix has discarded. An analysis of four peripheral-blood archives recovers a validated FYB1 immune-cell splicing switch, a cross-fitted fragment model appropriate for 3' chemistry reveals an eight-donor shift in NTRK2 terminal-isoform usage from astrocyte and neural-stem-cell populations to mature neurons, and pooling evidence across cells within the context of an expectation-maximization algorithm recovers 75--98% of withheld multi-gene molecule labels. Gravlax is open source, implemented in Rust, licensed under the BSD 3-clause license, and available at https://github.com/COMBINE-lab/gravlax.
bioinformatics2026-09-21v1A transcriptional continuum from clinically normal to lesional skin: Single-cell trajectory analysis of psoriasis
Uzun, G.; Konak, D.; Kazan, H.; Pir, P.Abstract
Psoriasis is a chronic inflammatory skin disease characterized by immune dysregulation and complex cellular interactions involving keratinocytes (KCs), T cells, and endothelial cells. Single-cell studies have mapped the cellular diversity of psoriatic plaques, yet most analyses treat skin as either healthy or diseased and therefore overlook the graded changes that precede a visible lesion. We asked whether psoriasis instead progresses along a transcriptional continuum and whether clinically normal skin from patients already carries a measurable disease signature. Using single-cell RNA sequencing (scRNA-seq) of healthy skin (NS), clinically normal skin adjacent to lesions (PN), and lesional skin (PP), we combined two complementary designs: a complete set spanning all three states and a within-patient paired set of matched PN and PP skin biopsies that removes inter-individual variability. Trajectory analysis using a novel pseudotime approach reveals transcriptional changes within cell populations and dynamic cellular states underlying psoriatic pathology. To distinguish lesion-driven effects from disease-associated transcriptional changes, we applied a differential expression strategy comparing PP and PN psoriatic tissue with NS and integrating these contrasts. This approach enabled the identification of disease-specific gene signatures while minimizing lesion-specific bias. As a result, we identified a coherent, cell-type-resolved continuum of transcriptional states in which endothelial cells, IL20+ keratinocytes (KCs: IL20), and KGF+ keratinocytes (KCs: KGF) undergo trajectory-dependent state transitions. These trajectories converge on features including epidermal thickening, barrier dysfunction, and plaque maintenance. Endothelial cells transition from an inflamed, angiogenic state to a mature, barrier-stabilized phenotype, while KCs: IL20 and KCs: KGF shift from PN states toward hyperproliferative, stress-adapted lesional phenotypes. These findings reframe psoriasis as a dynamic, multicellular continuum and identify candidate temporal windows for early intervention.
bioinformatics2026-09-21v1Pitfalls in understanding PhIP-Seq data: technical variability, replicability, and best practices for interpretation
Trgovec-Greif, L.; Vogl, T.Abstract
Phage Display Immunoprecipitation sequencing (PhIP-Seq) is a high throughput method allowing to measure antibody binding against hundreds of thousands of potential antigens. Typically, blood samples of hundreds to thousands of individuals are measured in 96-well plates, necessitating distribution of samples on multiple plates and processing in batches. To correctly interpret PhIP-Seq results, it is crucial to understand its nature and technical limitations. Here, we analyzed 195 technical replicates of controls (anchor samples) from 49 plates measured in 13 immunoprecipitation runs to gauge technical variability and replicability. While there were few false-positive enriched peptides, we observed a substantial fraction of false negatives. The total number of enriched peptides can be affected by batch effects. Additionally, the type of biological material used (serum or plasma) influences enrichment profiles with an antigen library containing bacterial proteins due to the interaction of phages with plasma proteins. Finally, replicating measurements in different labs produces comparable enrichment profiles, but slight differences in a subset of peptides are noticeable. These results suggest best practices for PhIP-Seq experiments. Most importantly, due to batch effects, samples of cases and controls should be evenly distributed between 96-well plates to avoid confounding with biological effects. Also, for highly complex libraries, the total number of enriched peptides can be a technical artifact and potentially not a feature useful in comparisons.
bioinformatics2026-09-21v1Discriminating rare disease cases from their controls based on observed and excluded phenotypes
Guo, Y.; Yao, Y.; Zhang, Y.; Duan, G.; Zhang, N.; shan, g.Abstract
Rare diseases are individually uncommon but collectively prevalent. Their primary clinical challenge lies not in treatment but in diagnosis. In the early stages of clinical management, it is frequently unclear whether the observed phenotypes are associated with a rare disease. Leveraging machine learning methods to mine latent associations between these phenotypes and rare diseases for early diagnosis offers a viable strategy to alleviate this diagnostic dilemma. In the present study, we demonstrated that machine learning can effectively discriminate rare disease cases from their controls using observed and excluded phenotypes as features. Among them, the Random Forest model achieved the best classification performance with a certain degree of generalizability. Based on further analysis of the feature selection results, we conclude that the two factors, specificity and occurrence count, are important for phenotype selection in rare disease discrimination, and comparable importance should be attached to both observed and excluded phenotypes, during feature construction.
bioinformatics2026-09-21v1POAnoise: A Graph-based Denoising Pipeline for Amplicon Sequencing Data
Ardiyansyah, M.; Furneaux, B.; Ovaskainen, O.Abstract
High-throughput DNA metabarcoding enables large-scale biodiversity assessment by identifying taxa from environmental samples, but its accuracy critically depends on denoising methods that separate true biological variation from PCR and sequencing errors. A persistent challenge is robust reconstruction of sequence diversity across abundance distributions, where low-abundance variants are particularly difficult to recover. We introduce POAnoise, a graph-based denoising framework that uses Partial Order Alignment (POA) to model relationships among noisy sequencing reads. POAnoise incrementally constructs sequence graphs that represent substitutions and indels, and derives consensus sequences from graph-supported paths using a weighted consensus strategy. By combining graph-based alignment with abundance-aware clustering, the method provides a structured way to reconstruct sequence variants from noisy amplicon data across heterogeneous abundance regimes. We evaluated POAnoise on simulated ITS and 16S datasets and compared its performance with established denoising methods, DADA2 and UNOISE3, across multiple parameter settings. Across the benchmark datasets, POAnoise generally achieved higher F1-scores and exhibited more stable performance across parameter configurations. In abundance-aware analyses, POAnoise showed reconstruction ratios closer to unity and reduced abundance-dependent deviation compared with DADA2, while remaining broadly comparable to UNOISE3 across most abundance classes. Overall, these results indicate that POAnoise can provide a robust alternative for amplicon denoising.
bioinformatics2026-09-21v1A single-cell RNA-seq catalog of ground truth gene coregulation
Svensson, V.Abstract
Single-cell RNA-sequencing measurements are uniquely well-poised to identify coregulated gene transcription. Coregulation should be apparent as correlation of expression levels, but at which unit to quantify expression for calculation of correlations is not clear. To enable evaluation of normalization methods for identification of coregulation we have curated and preprocessed 35 scRNA-seq samples with known regulature relationships into a catalog. These datasets contain promoter-reporter data representing known positive coregulation and B cells where isotypic mutual exclusion of {kappa} and {lambda} light chains represent known negative coregulation. As an example, we use the catalog to evaluate seven normalization methods. Different normalization methods perform better for positive or negative coregulation, and results indicate implementation choices for advanced normalization methods The catalog may serve as a resource for evaluating novel proposed normalization methods and is available at https://huggingface.co/datasets/valsv/scrna-coregulation-benchmark
bioinformatics2026-09-21v1Detecting cell segmentation errors using doublet methods
Sarwar, A.; Gillis, J.Abstract
In spatial transcriptomics, cell segmentation is used to draw boundaries around cells. Molecules located within a cell's boundary are assigned to it, making its gene expression profile dependent on segmentation accuracy. To identify potentially problematic cells, studies increasingly use doublet detection methods, a class of algorithms developed to recognize molecular admixture from two cells in scRNA-seq. To evaluate their suitability in spatial data, we model cell segmentation errors as a continuum of partial molecular admixture between neighboring cells, generated by varying the loss of a cell's own transcripts and the gain of transcripts from its neighbor. Evaluating 8 doublet methods across 16 spatial datasets, we characterize the conditions under which they perform well and identify their failure modes. Detection improves with increasing molecular admixture and transcriptional dissimilarity between neighboring cells, with cxds2, scDblFinder.score, and a simple baseline (unique_genes) performing best. These patterns are consistent across datasets spanning different gene panels, technological platforms, staining techniques, segmentation algorithms, and error frequencies. In unperturbed spatial data, elevated doublet scores localize to regions consistent with segmentation problems, suggesting that they can help prioritize cells or regions for inspection, segmentation refinement, or transcript reassignment. Our results establish when doublet methods can provide useful quality control signals for cell segmentation, supporting their use in spatial transcriptomics.
bioinformatics2026-09-21v1Serum metabolomics reveals signatures associated with physical resilience trajectories from middle to older age
Seo, J. I.; Scheurink, T. A. W.; Wang, C. X.; Kvitne, K. E.; Nunes, W. D. G.; Zemlin, J.; Bergstrom, J.; Dorrestein, P. C.; Mohanty, I.; Molina, A. J. A.Abstract
Lifecourse physical resilience is defined by the ability to maintain abilities across multiple domains of physical performance. While the importance of physical resilience in functional independence and mobility disability is clear, studies investigating metabolomic signatures of physical resilience are lacking. Here, we performed untargeted metabolomics on serum samples from a community-based cohort of 237 individuals followed over 28 years, and applied spectral data mining tools to map identified metabolites to health phenotypes from public repositories. We identified metabolites across multiple chemical classes, including acylcarnitines, glutamine conjugates, and phosphocholines, that were differentially associated with physical resilience status. Notably, medium-chain acylcarnitines negatively associated with physical resilience were more frequently observed in disease phenotypes than in healthy individuals. Kynurenine, a tryptophan metabolite linked to age-related functional decline, increased more steeply with age in individuals with low physical resilience. We also found that metabolites of the antihypertensive drug verapamil were associated with physical resilience in a metabolism-dependent manner, differing between oxidative and glucuronidated forms. Together, these metabolic signatures offer a resource for identifying biochemical pathways and biomarkers relevant to physical resilience for healthy aging.
bioinformatics2026-09-21v1MaSkel and napari-MaSkel: fast morphological skeleton feature extraction from segmentation masks
Wittmann, S.; Pysch, D.; Uderhardt, S.; Blumenthal, D. B.; Moeller, A.Abstract
Summary: Existing workflows for extracting morphological features from network-like biomedical structures often require several heterogeneous tools, which can limit integration and reproducibility. We present MaSkel and napari-MaSkel, open-source Python tools that provide GUI- and CLI-based workflows for 2D and 3D skeletonization and graph-based feature extraction, using a substantially accelerated implementation of the widely used Lee94 thinning algorithm. Availability and Implementation: MaSkel and napari-MaSkel are open-source Python packages available under the MIT license on PyPI and GitHub: https://github.com/bionetslab/maskel/, https://github.com/bionetslab/napari-maskel/. Documentation: https://bionetslab.github.io/maskel/, https://bionetslab.github.io/napari-maskel/. Benchmarking and validation: https://github.com/bionetslab/maskel-evaluations.
bioinformatics2026-09-21v1ARCHER-LD: Rapid Long-Range Linkage Disequilibrium Calculations at Biobank Scale using GPU Acceleration
Kumar, R.; Singhal, P.; Zhang, D.; Carson, C.; Conery, M.; Rodriguez, A. A.; Nandi, T.; Ritchie, M.; Thavappiragasam, M.; Voight, B. F.; Madduri, R.; Verma, A.Abstract
Linkage disequilibrium (LD) information from one's own dataset is considered optimal for downstream analyses such as statistical fine-mapping. However, the computational complexity ({approx}N2 / 2 computations for N variants) lead most studies to use external reference panels, such as from 1000 Genomes. To capture LD across all variant pairs in biobank-scale whole-genome sequencing (WGS) datasets with hundreds of millions of variants, new computational strategies are essential. We present a novel approach that uses multi-GPU distributed computing to compute R2 for every variant pair in a dataset. On chromosome 22 (1.8 million variants) of the 30x WGS 1000 Genomes dataset, our method using eight consumer-level 12GB GPUs (NVIDIA RTX 2080TIs) takes <10 minutes, while the same calculation with PLINK using a high-end 64-threaded CPU (Intel Xeon Gold 6338) takes >80 minutes, corresponding to a {approx}8x speedup. We further demonstrate true biobank-scale performance in the Penn Medicine Biobank (PMBB; 57,170 samples), computing chromosome 1 LD ({approx}1.38 million variants) in 49 minutes versus 22.7 hours for PLINK, a {approx}28x speedup. With this method, we successfully computed on the aforementioned 30x WGS 1000 Genomes dataset ({approx}120 million variants and {approx}2500 samples) the entire LD matrix (>1e16, or 10 quadrillion elements) in under 6 hours using 512 NVIDIA 40GB A100 GPUs on the Department of Energy Argonne Leadership Computing Facility Polaris Supercomputer. We make this tool, coded in Python using CuPy, publicly available. Using this tool, researchers can leverage the full extent of their genomic data without relying on external LD reference panels and acquire more accurate, population-specific findings, particularly for groups underrepresented in existing databases.
bioinformatics2026-09-21v1Two Comparators May Be All We Need
Muskal, S. M.Abstract
A compound in a cell meets a spectrum of proteins drawn from many families at once, while screening most often interrogates one target at a time. Two questions asked many times in a rank ordering workflow ultimately guide decisions on what gets made and what gets counter-screened, and both are comparative: which of two targets does a compound prefer, and which of two compounds does a target prefer. We built one model for each, over a roster of 1,879 human proteins covering 34 protein families. Each model is given two chemical structures and a sequence, or two sequences and a chemical structure, and returns which member of the pair is preferred together with how firmly it holds that view. No conformational analysis, protein structure, binding site or docked pose is used. Across families, asked which of two targets a compound prefers, the model is correct 0.75 of the time over 8,689 held-out comparisons, and 0.93 of the time on the third of them it holds most confidently. Asked which of two compounds a single target prefers, it is correct 0.71 of the time over 65,725 held-out comparisons, rising to 0.96 on the most confidently held. Neither compound in any of those comparisons appeared anywhere in training. Accuracy in both tracks the size of the real difference between the two measurements, from near chance where they fall within half a log unit to about 0.90 where they differ by more than two logs, and it holds across 31 protein families, not only the best-measured one. Within the chemistry and the targets they were built on, these models rank compounds and rank targets well. Both models can be explored and downloaded at familyfoundationmodel.com. Keywords: target preference; compound preference; pairwise comparison; polypharmacology; off-target triage; ESM2; random forest; structure-free prediction; ChEMBL
bioinformatics2026-09-21v1celltypeEnrich: a consensus-based scRNA-seq cluster annotation tool
Rutledge, S.; Tuteja, G.Abstract
Motivation Single-cell RNA sequencing (scRNA-seq) cluster annotation is a critical step in data analysis. Current methods are time-consuming, difficult to reproduce, or limited in tissue or species coverage. Results We developed celltypeEnrich, a cluster-level annotation tool that uses a hypergeometric test to identify enrichment of cell-type-specific genes from input gene lists. Enrichment results from up to 26 reference datasets are used to determine a consensus annotation. Benchmarking using scRNA-seq datasets from three tissues spanning two species showed 62-72% annotation accuracy for celltypeEnrich, generally outperforming other tools, which had either lower accuracy, incomplete tissue coverage, or the need for parameter optimization. The performance of celltypeEnrich remained stable when input gene lists were down-sampled to 25% of their original size. Availability and Implementation celltypeEnrich is freely available at (https://celltypeenrich.gdcb.iastate.edu) as an R Shiny web application under the MIT license for non-profit academic use.
bioinformatics2026-09-21v1A functional landscape of human microproteins
Slivak, M.; Choteau, S. A.; Spinelli, L.; Souville, L.; Pierre, P.; Carvunis, A.-R.; Bornberg-Bauer, E.; Zanzoni, A.; Brun, C.Abstract
Short open reading frames (sORFs) are widespread yet often poorly annotated, and an increasing number are known to encode functional sORF-encoded peptides (sPEPs, or microproteins). Here, we use a system approach to determine the functions of sPEPs in three tissues: monocytes, skeletal muscle and brain. We first predicted interactions between sPEPs and canonical proteins and characterized their interaction interfaces. sPEPs preferentially target specific canonical partners rather than interacting promiscuously, and two functional groups emerge from their interaction features. A small subset (~2%) exhibits canonical protein-like features, with structured domains interacting with diverse proteins involved in tissue-specific functions. In contrast, most sPEPs contain short linear motifs with extensive phosphorylation potential, suggesting they may be involved in signaling. We then constructed the first sPEP-containing interactome networks and infer sPEP functions from network topology. Most sPEPs appear to act as regulatory elements modulating diverse cellular processes rather than acting directly as functional components. Knowledge-based validations using characterized sPEPs, show that our framework can capture context-specific functions beyond current annotations. Together with their limited evolutionary constraint, the pervasive regulatory roles of sPEPs suggest that they constitute a reservoir of functional novelty and an evolutionarily flexible layer of cellular regulation.
bioinformatics2026-09-19v2The RdRp Thumb-1 Pocket is a Conserved Target for Broad-Spectrum Antiviral Development
Woods, V.; Umansky, T.; Russell, S. M.; Gallay, P.; Smith, D.; Haders, D.Abstract
RNA viruses cause human diseases ranging from mild colds to deadly pandemics. Direct-acting, broad-spectrum, non-nucleoside antivirals have been characterized as impossible to develop because allosteric binding sites are poorly conserved. The HCV NS5B RNA-dependent RNA polymerase (RdRp) Thumb-1 allosteric site and its interaction with the HCV NS5B {Lambda}1-loop governs an essential conformational change required for polymerase initiation. The only approved NS5B Thumb-1 inhibitor, beclabuvir, has been shown to be inactive against a broad panel of non-HCV viruses, including poliovirus, rhinovirus, coronavirus, coxsackievirus, influenzavirus, and HIV. A conserved, homologous allosteric site on RdRp that spans multiple viral families has not been reported. Here, we report GALILEO's discovery that the Thumb-1 pocket, its associated {Lambda}1-loop and their interaction are conserved across RNA viral families for the first time. The discovery is validated through comparative structural analysis of Protein Data Bank (PDB) deposited viral polymerases utilizing a method that allows independent, public validation by any researcher. We further demonstrate that beclabuvir's dependence on its indole C6 carbonyl to interact with the HCV-specific residue R503 restricts its activity to HCV. We validate the target discovery with MDL-001, which does not contain a C6 carbonyl substituent. MDL-001 directly blocks viral RNA synthesis in isolated replication complexes and selects for the canonical Thumb-1 resistance mutation P495S in HCV NS5B. MDL-001 demonstrates broad-spectrum in vitro inhibition of both HCV and SARS-CoV-2. Preclinical proof of concept and development of MDL-001 across HCV, HBV, HDV, influenza, SARS-CoV-2, and RSV have been previously reported. These findings establish RdRp Thumb-1 as a conserved allosteric pocket and a druggable target for broad-spectrum direct-acting antiviral development.
bioinformatics2026-09-18v7OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome
Teixeira, A. A. R.; Zhu, H.; Kothiwal, D.; Cao, R.; Mills, A.Abstract
Choosing which region of a protein to express remains poorly standardized in antibody discovery, recombinant reagent generation, structural biology and computational binder design. For human cell-surface and secreted proteins, this requires reconciling topology, processing, predicted and experimental structure, modifications, interaction partners, orthologs, paralogs and cross-reactivity risk before ordering DNA. OpenAntigens is a free, no-login database of construct-design reports for 5328 human secreted, GPI-anchored, single-pass and multipass proteins. It integrates UniProt topology, AlphaFold pLDDT and PAE, PDB precedent, InterPro and Pfam domains, mouse and cynomolgus orthologs, paralog and family context, Open Targets disease associations, partner and assembly context, and BLAST searches. It provides 55 305 construct suggestions spanning full design regions, PDB-backed boundaries, annotated domains, pLDDT/PAE-derived regions and membrane-expression options, plus 148 722 sequence-similarity hits to help choose constructs and assess cross-reactivity. For targets with compatible AlphaFold models, the interactive designer links sequence, structure, pLDDT and PAE, allowing users to revise boundaries and export species-equivalent sequences with real-time cysteine and modification warnings. OpenAntigens places reproducible construct suggestions, comparative context and browser editing in one workflow, reducing manual reconciliation across resources. OpenAntigens is available at openantigens.org.
bioinformatics2026-09-18v2Incomplete references leave bulk deconvolution targets non-identifiable, but identification has a computable precision price
Jiang, H.; Gao, F.; Liu, P.; Wu, Y.; Jie, Y.; Li, Y.; Jiang, Y.Abstract
Background. Reference-based deconvolution estimates cell-type proportions from bulk profiles. Incomplete references compromise these estimates, yet many remedies return point estimates. We separate the operator provenance that reproduces an estimate from the information and precision that identify its target. Results. Two operator histories at one reduced reference give different estimators: 28.2% of 196,420 sample-deletion pairs differed by >0.1 total variation; locking every learned component made paths identical. Across nine cohorts, operator history reversed 160 of 952 association signs, 31 with a significant path, and changed significance for 126. Observationally equivalent completions can fill the open simplex and reverse retained-type rankings. Under a joint zero-exposure condition, shared structure leaves inherited bounds unchanged; one to six uncalibrated views also gave identical bounds. A profile library contracted estimator-output envelopes by 98.93% yet covered the full-reference effect for 42.77%, with 19 wrong-sign certificates; conditional sharp bounds stayed at [-1, 1]. Treating RNA yields as exact collapsed intervals to points covering none of seven flow-measured targets: contraction without coverage is false certainty. Calibrated cross-modal anchors contract width to 0.51 at six types but certify no valid sign. Exact RNA yields identify cell fractions from RNA contributions; a decisive sign in blood requires {+/-}2.6% proxy accuracy with near-total contraction of donor-heterogeneity and dynamic-range envelopes. Conclusions. Incomplete-reference deconvolution is an identification problem, not only an estimation problem. Remedies must be scored on shrinkage, coverage, certification and false certification against held-out targets. fitdrop implements this scoring and a precision frontier for planning decisive measurements.
bioinformatics2026-09-18v2DeepSpaceDB 2.0: an interactive spatial transcriptomics database for large-scale Xenium data exploration
Honcharuk, V.; Takemoto, K.; Masalunga, M. C.; Zhao, H.; Diez, D.; Kawaoka, S.; Vandenbon, A.Abstract
The 10x Genomics Xenium platform enables high-resolution spatial transcriptomics at single-cell and subcellular scales, but effective reuse of public Xenium datasets is hindered by large data sizes and heterogeneous file formats. We previously developed DeepSpaceDB, a spatial transcriptomics database designed for interactive, in-depth analysis of tissues and tissue microenvironments. Here, we present a major expansion of DeepSpaceDB that integrates large-scale single-cell spatial transcriptomics data generated by the Xenium platform. In this update, we systematically collected 1,539 public Xenium datasets from multiple repositories and processed them through a robust, standardized pipeline that validates, repairs, and harmonizes heterogeneous inputs into a unified representation. To support efficient exploration of these data, we introduced a redesigned DeepSpaceDB interface and complementary Zarr-based storage formats optimized for gene-centric visualization and spatially localized queries, enabling sub-second response times for common interactive operations. The updated platform supports real-time visualization of spatial data and analysis of regions of interest directly in the web browser. Together, this expansion establishes DeepSpaceDB as a unified resource for single-cell spatial transcriptomics, substantially lowering the barrier to accessing, exploring, and reusing large-scale public Xenium datasets.
bioinformatics2026-09-18v2LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design
Waththe Liyanage, W. W.; Rigoni, D.; Bove, F.; Righelli, D.; Romano, S.; Visone, R.; Iorio, M. V.; Grassia, M.; Mangioni, G.; Lio, P.; Taccioli, C.Abstract
The discovery of novel small molecules is challenging because of the vastness of chemical space and the complexity of protein-ligand interactions, leading to low success rates and time-consuming workflows. Here, we present LLMsFold, a computational framework that combines Large Language Models (LLMs) and biophysical foundation tools to design and validate new small molecules targeting pathogenic proteins. The pipeline starts by identifying viable binding pockets on a target protein through geometry-based pocket detection. A 70-billion-parameter transformer model from the LlaMA family then generates candidate molecules as SMILES strings under prompt constraints that enforce drug-likeness. Each molecule is evaluated by Boltz-2, a diffusion-based model for protein-ligand co-folding that predicts bound 3D structure and binding affinity. Promising candidates are iteratively optimized through a reinforcement learning loop that prioritizes high predicted affinity and synthetic accessibility. We demonstrate the approach on two challenging targets: ACVR1 (Activin A Receptor Type 1), implicated in fibrodysplasia ossificans progressiva (FOP), and CD19, a surface antigen expressed on most B-cell lymphoma and leukemia cells. Top candidates show strong in silico binding predictions and favorable drug-like profiles. All code and models are made available to support reproducibility and further development.
bioinformatics2026-09-18v2WITHDRAWN: A Comprehensive Analysis of the Electrolytic Hydrogen Water Mechanism via a Feedforward Loop and its Functional Role in Intestinal Cells In Vitro
LI, J.Abstract
This manuscript has been withdrawn following a formal investigation by the Graduate School of Human Sciences, Waseda University
bioinformatics2026-09-18v2Pep-PU-GAN: Positive-Unlabeled Adversarial Learning for Peptide Function Prediction
Midjani, F.; Hashemi, S.; Keshtkar, F. Z.; Malekpour, M.; Saberzadeh Ardestani, B.; Khosravi, B.Abstract
Peptide classification remains challenging in bioinformatics because of limited labeled data, particularly the scarcity of verified negative examples, and the complex relationship between amino acid sequences and biological functions. This study introduces Pep-PU-GAN, a deep learning framework that combines positive-unlabeled (PU) learning, generative adversarial networks (GANs), and graph neural networks (GNNs) for peptide classification. Peptides are represented as sequence-derived residue graphs, with amino acids as nodes and edges connecting adjacent residues, enabling attention-based message passing over local neighborhoods. The architecture includes a generator that produces synthetic peptide embeddings in encoder space and a dual-function discriminator that distinguishes real from synthetic embeddings while performing PU classification. Training uses a custom loss integrating non-negative PU (nnPU) risk estimation with adversarial objectives. A self-training mechanism further incorporates high-confidence synthetic positive embeddings to augment the training set and improve performance. Evaluated on neuropeptide classification using 4,049 positive neuropeptides and 8,558 unlabeled peptides, Pep-PU-GAN outperformed baseline models, achieving an F1 score of 0.93 and an AUROC of 0.98 on an independent held-out benchmark. Pep-PU-GAN provides a promising approach for peptide classification tasks with scarce labeled and abundant unlabeled data, with potential applications in computational biology and drug discovery.
bioinformatics2026-09-18v1TRACEDD: A Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design
Vangala, S. R.; Kasturi, V. V.; Bung, N.; Roy, A.Abstract
Drug discovery depends on coordinated decisions across target validation, structure analysis, molecular design, developability assessment and synthetic feasibility, but current computational methods often operate as disconnected tools. Here, we introduce TRACEDD (Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design), a framework that makes three primary contributions: (1) It establishes a 'tool-first' multi agentic architecture where LLMs orchestrate validated computational tools rather than replace them, ensuring scientific rigor. (2) It implements a multi-agent system that mirrors expert discovery teams, enabling transparent and traceable decision-making through a Reason-Act-Observe loop. (3) It demonstrates an end-to-end workflow, from target validation to synthesis planning, that adaptively handles real-world data variability, such as the absence of experimental structures. The framework decomposes discovery into specialized agents for target validation, druggability assessment, molecular generation, lead optimization, ADMET evaluation, literature evidence integration and retrosynthesis, all operating through a Reason Act Observe workflow. Using JAK2 as a representative case, we show that the system can retrieve experimental protein structures, invoke AlphaFold when structures are unavailable, identify druggable pockets and perform de novo molecular generation. Known JAK2 inhibitors are used to define design hypotheses and guide reinforcement learning-based molecular generation, with docking scores/predicted pIC50 and other physicochemical/ADMET properties serving as reward and prioritization signals. The framework demonstrates a tool-first, reasoning-driven approach in which each major decision is linked to explicit tool invocation, intermediate evidence. By combining agentic orchestration with domain-specific computational tools, the system supports transparent, adaptable and human-verifiable molecular design workflows, providing a foundation for more reliable AI-assisted drug discovery.
bioinformatics2026-09-18v1Sparse Machine Learning Pipeline with Stabl Identifies Cord Blood Multi-Omic Signatures of Bronchopulmonary Dysplasia
Mestan, K.; Newar, J.; Zhao, J.; Chakraborty, A.; Reiss, J.; Funk, W.; Stelzer, I.; Waked, B.; Bellan, G.; Durand, X.; Hedou, J.Abstract
Background: Several omics studies have been completed in recent years, with the goal of identifying biomarkers of complex multifactorial diseases, such as bronchopulmonary dysplasia (BPD). Objective: To evaluate the performance of 3 distinct omics platforms, using a machine learning pipeline with integration of sparse, reliable and adaptive biomarker identification (Stabl). Methods: Using a well-characterized birth cohort, cord blood metabolomics, proteomics and adductomics data were integrated with Least Absolute Shrinkage and Selection Operator (LASSO) regression and Stabl, to evaluate predictive performance for BPD. Results: Sparse multivariable modeling of 45,000 features measured in 217 infants (52 term, 165 extremely preterm <28 weeks; 82 with BPD and 35 with severe BPD/death) identified a perfect signature for preterm birth with both LASSO and Stabl (AUROC=1.0; p<0.001). Analysis of the preterm group yielded excellent predictive power for severe BPD (AUROC=0.83; p=0.005). Stabl identified a set of 12 biomarkers (2 adducts, 3 proteins and 7 metabolites) with good performance for predicting grade III BPD (AUROC=0.76; P=0.03). Biomarkers across the 3 omics platforms revealed dysregulated pathways of innate/adaptive immune responses, metabolic programming and oxidative stress. Conclusions: The sparse machine learning pipeline is a complementary approach for identifying novel pathways and biomarkers of multifactorial BPD and its endotypes.
bioinformatics2026-09-18v1Setting the SCENE for Interpretable Cell-Gene Embeddings in Single-Cell RNA-seq
Moberg, O. L.; Petersen, M. B.; Herlau, T.; Kristensen, L. E.; Jessen, L. E.; Morup, M.Abstract
Single-cell RNA sequencing measures cellular states at high resolution, but sparse high-dimensional count data remain difficult to model interpretably. We introduce the Single-Cell Euclidean Network Embedding (SCENE), a probabilistic latent-distance model that jointly embeds cells and genes from Unique Molecular Identifier (UMI) counts. SCENE treats the count matrix as a weighted bipartite cell-gene graph, where Euclidean distances represent transcriptional affinity, and combines this geometry with a zero-inflated count likelihood that separates gene detection from expression magnitude. Across real and simulated scRNA-seq datasets, SCENE recovers biologically structured cell and gene embeddings with state-of-the-art performance. Surprisingly, major biological structure is preserved in native two- and three-dimensional latent spaces, enabling directly interpretable visualization. Perturbation analyses show that SCENE organizes glucocorticoid-response genes and T-cell receptor regulatory programs coherently in gene space, capturing biology beyond cell-type separation. SCENE provides a transparent representation learning framework in which low-dimensional Euclidean geometry supports accurate modeling and biological interpretation.
bioinformatics2026-09-18v1Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics
Chen, T.; Hicks, S. C.Abstract
Models that combine tissue images with pretrained gene representations aim to predict spatial expression for genes not used to fit the downstream predictor. Yet success on held-out genes can reflect two capabilities: estimating a gene's mean expression across tissue locations and recovering its spatial variation. Across four cohorts spanning three human brain regions and HER2-positive breast cancer, we evaluated held-out genes in held-out individuals and separated these components. For spatial predictors using fixed gene representations from Decima or scGPT, reductions in gene-mean error accounted for more than 91% of the reduction in mean squared error relative to matched random vectors. Independently fitted mean-only models using the same representations but no tissue images retained 90-99% of the corresponding gain in full-matrix correlation. Spatial gains were smaller on average, increased with expression variation in training tissue and differed across cohorts and representations. Across these settings, pretrained gene representations broadly transferred mean expression but selectively improved spatial recovery, showing that cross-gene generalization in virtual spatial transcriptomics is not a single capability.
bioinformatics2026-09-18v1