Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Transcriptional landscape of direct reprogramming toward the hematopoietic lineage
Cwycyshyn, J.; Stansbury, C.; Golts, S.; Lee, H.; Pickard, J.; Meixner, W.; Rajapakse, I.; Muir, L. A.Abstract
Direct reprogramming of human fibroblasts into hematopoietic stem cells (HSCs) offers a promising strategy for generating autologous cells to treat blood and immune disorders. Current protocols are limited by low efficiency and insufficient tools for evaluating reprogramming outcomes. Although functional assays are the standard for confirming cell identity, they require fully reprogrammed cells, limiting their utility during protocol development. To address this, we assembled a single-cell transcriptomic reference atlas of hematopoietic reprogramming and tested an algorithmically-predicted transcription factor recipe for HSC induction. Long-read single-cell RNA sequencing of CD34+ reprogrammed cells revealed progressive loss of fibroblast identity alongside induction of early hematopoietic and endothelial programs, with reference-atlas benchmarking placing reprogrammed cells in an intermediate transcriptomic state between fibroblasts, endothelial cells, and HSCs. Isoform-level analysis further revealed transcriptional remodeling not captured by gene-level analyses. This experimental-computational framework offers a generalizable strategy for characterizing partially reprogrammed states and guiding optimization of reprogramming protocols.
bioinformatics2026-07-22v5Graph Lens Lite: A browser-based tool for interactive visualization and exploration of biological networks
Ley, M.; Keska-Izworska, K.; Fillinger, L.; Walter, S. M.; Baumgärtel, F.; Bono, E.; Galou, L.; Andorfer, P.; Hauser, P.; Leierer, J.; Kratochwill, K.; Perco, P.Abstract
Biological network visualization together with graph-based analyses are key techniques in systems biology and network medicine to detect patterns and generate hypotheses regarding disease pathobiology, drug target identification, biomarker prioritization, and digital drug discovery. Network representations provide an intuitive way to communicate and share research findings. We have developed Graph Lens Lite, a browser-based tool that combines rich visualization with a streamlined interface for exploring and sharing biological networks. It offers an expressive query language, topological network analysis, interactive filtering, visual grouping, customizable layouts, a data editor, fine-grained property-based styling, animated edge-flow visualization, community detection, and a context-aware locally powered AI assistant particularly suited for exploring molecular models of disease pathobiology or drug mechanism of action. We demonstrate its utility on a curated network model of autosomal dominant polycystic kidney disease. Graph Lens Lite is open source, with a live web version available at https://delta4ai.github.io/GraphLensLite/.
bioinformatics2026-07-22v3Benchmarking the Impact of Data Leakage on the Performance of Knowledge Graph Embedding Models for Biomedical Link Prediction
BRIERE, G.; STOSSKOPF, T.; LOIRE, B.; BAUDOT, A.Abstract
Motivation: Knowledge Graphs (KGs) are increasingly used to organize complex biomedical knowledge into structured representations of entities and relations. Knowledge Graph Embedding (KGE) models facilitate efficient exploration of KGs by learning compact representations, and are widely applied to biomedical link prediction, for instance to uncover new therapeutic uses for existing drugs. Despite extensive work on KGE models, current evaluations often overlook data leakage, which can artificially inflate performance and undermine benchmark validity. Data leakage can arise when (1) there is redundancy between training and test sets, (2) the model leverages illegitimate features, or (3) the test set does not reflect real-world inference scenarios. Results: We assess the impact of data leakage on KGE-based link prediction across three biomedical KGs, using both decoder-only and GNN-based models. We first demonstrate the impact of train-test redundancies and implement a systematic procedure to detect and remove them. Using permutation experiments, we investigate whether node degree acts as an illegitimate predictive feature, and find no evidence that predictions are driven by degree alone. Finally, we evaluate how well common test set sampling strategies reflect real-world inference in drug repurposing. Comparing random and cold-start splits with an independent Orphanet-derived test set, we observe a substantial performance drop on the latter, indicating that current practices may overestimate how well KGE models generalize. Overall, our findings highlight the importance of rigorous benchmark design and careful evaluation of the generalization ability of KGE models for biomedical link prediction. Availability and Implementation: All code and results are openly available on GitHub at https://github.com/galadrielbriere/data_leakage_kge_benchmark.git.
bioinformatics2026-07-22v3CESAR: A R Package for High-Sensitivity Detection of Copy Number Variations in ctDNA Using Segmentation and Anchor Recalibration
Lu, J.; Ni, S.; Wang, L.; Wu, N.; Jiang, X.Abstract
Background: Detecting copy number variations (CNVs) in circulating tumor DNA (ctDNA) is crucial for the companion diagnosis and resistance monitoring of various solid tumors (e.g., NSCLC, Glioblastoma). However, when tumor-derived DNA fractions are extremely low (often <1%), traditional depth-based methods frequently fail due to non-linear sequencing depth fluctuations and probe-specific capture biases inherent to targeted Next-Generation Sequencing (NGS). Methods: We developed CESAR (CNV Estimation with Segmentation and Anchor Recalibration), a novel computational tool optimized for ultra-sensitive, tumor-only CNV detection in targeted NGS panels. CESAR utilizes Circular Binary Segmentation (CBS) to re-partition target regions based on relative capture efficiency. It then introduces a dynamic "anchor" selection algorithm that identifies a personalized set of genomic segments mirroring the non-linear coverage behavior of each target gene. By minimizing the Coefficient of Variation (CV) through iterative anchor selection, CESAR effectively recalibrates the baseline to suppress technical noise. Results: Validation using standard DNA reference materials demonstrated that CESAR successfully identified both amplifications (e.g., MET, ERBB2, EGFR) and relative copy number deletions at ultra-low tumor fractions. Notably, CESAR achieved stable detection of focal alterations as subtle as 2.18 copies (a mere 1.09x fold change relative to the diploid baseline), while maintaining zero false positives in control regions. Evaluation across distinct clinical biofluids, 36 clinical plasma samples and 41 glioma cerebrospinal fluid (CSF) samples, identified critical, previously undetected CNV events, including subtle ERBB2 gains and distinct MET deletions. Furthermore, comprehensive benchmarking revealed that CESAR consistently outperformed the widely used CNVkit, particularly in suppressing technical variance and resolving ultra-low-level copy number gains that CNVkit failed to distinguish from background noise. Conclusions: CESAR provides a highly stable and sensitive algorithmic framework for tumor-only CNV calling in liquid biopsies, facilitating precise therapeutic decision-making in precision oncology.
bioinformatics2026-07-22v2Integrative structure determination of a human mitochondrial contact site and cristae organizing system (MICOS) sub-assembly
Jindal, M.; Mahato, R.; Das, S.; Guha, A.; Majila, K.; Arvindekar, S.; Vaidya, A. T.; Viswanath, S.Abstract
The Mitochondrial contact site and Cristae Organizing System (MICOS) complex is an inner mitochondrial membrane (IMM) assembly present at the cristae junction. It is responsible for regulating cristae formation and remodeling. However, its structure is not known. We applied Bayesian integrative structure determination to characterize the structure of the Mic60, Mic19, Mic10, and Mic13-containing MICOS complex combining AlphaFold predictions with data from crosslinking mass spectrometry, biochemical assays, electron tomography, homology modeling, and sequence alignments. The integrative structure revealed novel mutual interfaces among Mic10N,C, Mic60LBS1,LBS2,mitofilin, and Mic13central,C, which were experimentally validated. Several likely-pathogenic missense mutations also localize to these novel interfaces, highlighting their importance. Our results indicate that Mic13 likely facilitates MICOS assembly by binding Mic10 in the IMM-proximal region and Mic60 in the intermembrane space. Taken together, our integrative approach sheds light on the structure and assembly of the MICOS complex.
bioinformatics2026-07-22v1IntraTalker - Modelling Intracellular Signalling in1 Cellular Crosstalk
Kloker, V.; Nagai, J. S.; Feng, Z.; Mavrommatis, L.; Hermanns, L.; Moscoso, J. M. J.; Ruiz, M.; Kuppe, C.; Costa, I. G.Abstract
Single-cell sequencing has advanced the study of cell-cell communication, yet most methods focus on intercellular ligand-receptor interactions while neglecting downstream intracellular signalling cascades and the possibility that downstream target genes themselves encode ligands, thereby propagating communication across multiple cells. We present IntraTalker+CrossTalkeR that combines intracellular (IntraTalker) and intercellular (CrossTalkeR) signalling from multimodal single-cell data. IntraTalker infers cell-type-specific transcription factor activities and constructs receptomes that link receptors to downstream target genes, which are then integrated with ligand-receptor predictions in CrossTalkeR. To prioritize signalling receptors, the framework performs in silico receptor perturbation. In a murine bone marrow dataset, this recovered the known function of the Il1r1 receptor in driving myeloid progenitor cell-state changes. In a human myocardial infarction dataset, it predicted a novel role for IGF1R signalling in driving the differentiation of fibroblasts towards progenitor fibroblast states.
bioinformatics2026-07-22v1A Standardized Methodology for FAIRness Assessment and Multi-Dimensional Scoring in Agrosystem Research Data Infrastructures
Haleem, A. U.; Arend, D.; Etukala, J. R.; Mazon, E. R.; Schmidt, M.; Jung, J.; Martini, D.; Usadel, B.; Neidiger, C.; Ulrich, R.; Lange, M.Abstract
The NFDI-consortium FAIRagro has established a systematic framework for evaluating the FAIRness of Research Data Infrastructures (RDIs) within the German agrosystem research landscape. While FAIR principles are widely accepted, their practical implementation by research data infrastructures (RDI) remains challenging. By operationalizing the FAIR principles into a reproducible multi-dimentional scoring methodology, this initiative addresses the critical need for a transparent and citable benchmark of RDIs that moves beyond simple compliance. This paper details the underlying assessment criteria, comprising 20 aggregated core metrics, the iterative community-driven validation process, and the integration of these metrics into the FAIRagro Search Hub. This framework evaluates RDIs, like repositories or databases, instead of sampling hosted data sets., across the four distinct categories of FAIR independently, yielding granular, pillar-specific ratings. Unlike aggregate scoring models, which can inadvertently mask technical deficiencies by averaging performance across categories, this multi-dimentional approach ensures that a repositorys distinct strengths and bottlenecks remain fully visible. Our findings demonstrate that standardized scoring not only clarifies data accessibility for users but also highlights specific operational gaps, allowing repository providers to identify precisely where the service implementation can be enhanced. By establishing this data-driven service in the agronomy domain, we provide a scalable template for the broader NFDI and EOSC ecosystems to foster a culture of excellence in research data stewardship.
bioinformatics2026-07-22v1Hobrac: a reference-guided workflow for genome comparison and synteny visualization
Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.Abstract
Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO rthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.
bioinformatics2026-07-22v1scLEMBAS: Context-Aware Modeling of Signaling Pathway Activity at Single-Cell Resolution
Baghdassarian, H. M.; Meimetis, N.; Nordenstorm, O.; Joughin, B.; Nilsson, A.; Lauffenburger, D.Abstract
Cells sense and integrate extracellular cues through intracellular signaling networks that reshape transcription factor activity to dictate cellular responses. Signaling activity is difficult to decipher: it is non-linear, and it contains extensive feedback and crosstalk. Furthermore, the same perturbation can elicit markedly different responses depending on context (e.g., cell type, disease state, and tissue microenvironment) such that identical stimuli produce diverse responses in multicellular populations. Consequently, there is a vast combinatorial space of complex interactions and context-dependent responses that necessitate computational models. Computational models of single-cell perturbation responses are demonstrated to predict cellular responses, but are often limited in mechanistic insight. Prior knowledge networks offer a route to bridge predictive capability and interpretability. Here we present scLEMBAS, a context-aware, gray-box neural network that models signaling pathway activity at single-cell resolution while preserving mechanistic grounding. scLEMBAS encodes a prior-knowledge network of protein-protein interactions as a recurrent neural network whose learnable edge weights correspond to signaling interaction strengths. It also captures context and individual cell variance through compositional bias terms. An adversarial approach allows the model to answer a single-cell counterfactual - what a given cell's TF activity would be under a different perturbation or context - while involving mechanistic rather than simply relational information. Across two scRNA-seq datasets spanning single- and multi-perturbation settings, scLEMBAS accurately predicts out-of-distribution combinations of perturbation and context. Capturing population variance across individual cells enables the model to predict cell subtype specific perturbation responses, despite being agnostic to such labels. Beyond prediction, scLEMBAS learned parameters are biologically interpretable: learned edge weights carry information beyond network topology and "self-prune" spurious interactions, while the categorical bias nominates proteins associated with cell-type-specific perturbation states. Overall, scLEMBAS enables quantitative dissection of how signaling pathway activity is reshaped by perturbation within specific cellular contexts.
bioinformatics2026-07-22v1Learning Minimal Gene Programs for Disease-Aligned Representations
Madduri, A.; Patel, C. J.Abstract
Identifying small, interpretable gene sets that robustly capture disease-associated variation in single-cell transcriptomic data remains a central challenge for biological interpretation and experimental follow-up. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-held-out evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data.
bioinformatics2026-07-22v1Tangerine: A Python framework for dynamic gene regulation analysis from transcriptomic time series
Narendra, T.; Schweikert, G.Abstract
Motivation: Time-series single-cell transcriptomics enables the study of dynamic gene regulation. However, standard computational tools frequently aggregate temporal data into static, dense topologies, obscuring the precise regulatory rewiring that drives developmental transitions. Further, navigating the inherent noise of statistical inference without losing biological interpretability remains an important bottleneck. Results: We present Tangerine, a Python framework for the dynamic reconstruction and interactive exploration of time-varying gene regulatory networks. Tangerine integrates time-constrained metacell aggregation with regularized linear modelling and non-parametric correlation to infer dynamic topologies. To solve the interpretability gap, it features a browser-based visual analytics engine. Tangerine empowers researchers to track macroscopic gene module evolution, interactively filter effect sizes, and link topological rewiring directly to raw transcriptomic evidence. Availability and implementation: Tangerine is implemented in Python and Plotly Dash. The code is available on Github at https://github.com/ntanmayee/tangerine.
bioinformatics2026-07-22v1AI4Life Open Calls and Public Challenges: why, how, and what we have learned.
Galinova, V.; Seifi, M.; Serrano Solano, B.; Lidayova, K.; Dalle Nogare, D.; Corbat, A. A.; Talks, J.; Giacomello, E.; Gomez-de-Mariscal, E.; Ferreira, M. G.; Fuster-Barcelo, C.; Battagliotti, J. M.; Garcia-Lopez-de-Haro, C.; Salmon, B.; Croft, M.; Yie, S. Y.; Rey-Paniagua, G.; Hu, X.; Cho, S.; Sheth, A.; Porwal, C.; Li, X.; AI4Life Consortium, ; Henriques, R.; Li, X.; Krull, A.; Klemm, A.; Munoz Barrutia, A.; Kreshuk, A.; Ouyang, W.; Jug, F.; Deschamps, J.Abstract
Within AI4Life, we ran three Open Calls and three Public Challenges (2023-2025), supporting 22 bioimage analysis projects from 151 applications and engaging 225 challenge participants, with the aim of applying FAIR deep learning in the life sciences. Our experience offers a view of the current state of bioimage analysis, the landscape of available tools, as well as the existing gaps between method developers, tool producers and potential users. It highlights that even after careful selection for AI-ready projects, most still require substantial effort to apply deep learning, and that the field still relies heavily on established, well-rounded methods to solve common problems. We come to the conclusion that for scientific AI in biology, the rate-limiting step is not methods and models but data, annotations, and shared infrastructure underneath them.
bioinformatics2026-07-22v1TrioNsight: Building a meta-predictor to evaluate the clinical impact of TrioN-like Dbl-homology domain variants
Taciroglu, A.; Aydin Son, Y.; Martin, A. C. R.; Orengo, C.Abstract
TRIO is a member of the Rho family of guanine nucleotide exchange factors (Rho GEFs), which promote the exchange of GDP for GTP to activate Rho GTPases and serve as key regulators of cellular signalling pathways. Mutations in TRIO are associated with neurodevelopmental disorders, including intellectual disability and autism spectrum disorders. TRIO contains two GEF units: one N-terminal and one C-terminal, each of which contains a Dbl-homology (DH) domain that drives its GEF activity to activate Rho GTPases. While the human proteome contains 70 highly conserved DH domains, the N-terminal DH domain of TRIO (TrioN) contains one-third of all reported DH domain pathogenic variants and has many variants of unknown significance. Numerous variant impact prediction tools exist, but most lack gene-specific considerations. Here, we describe TrioNsight, a meta-predictor designed to predict mutation impacts for TrioN and 12 highly similar human DH domains, including TrioC. TrioNsight exploits the naive-Bayes algorithm and leverages structural, evolutionary, and physiochemical features of approximately 1500 highly similar DH domains (TrioN-like DH domains) from 294 species. TrioNsight surpasses all available predictors, including AlphaMissense, achieving a Matthews' Correlation Coefficient of 0.890. Additionally, we provide a variant impact map that details the impacts of mutations at each position in the DH domain of these proteins, which can be valuable for clinical assessments. Furthermore, our approach establishes a standardised workflow adaptable for creating domain-specific variant predictors for other protein families, offering a template for improved variant interpretation.
bioinformatics2026-07-22v1IOBRpy enables agentic multi-omics decoding of anti-tumor immunity
Huang, H.; Li, X.; Liu, L.; Gu, W.; Wang, G.; Zeng, D.Abstract
Decoding the tumor immunity is pivotal for cancer immunotherapy, yet transcriptomic pipelines remain bottlenecked by fragmented tools and biased interpretations. Here we present IOBRpy, a Python toolkit driven by an innovative AI dual-agent layer for automated, highly standardized immuno-oncology workflows. Moving beyond conventional expression profiling, IOBRpy enables agentic multi-omics decoding. From raw FASTQ or TPM matrices, it seamlessly integrates upstream quality control, transcript quantification, and downstream TME parsing, encompassing signature scoring, ligand-receptor crosstalk, and cellular deconvolution. Crucially, IOBRpy expands data dimensions by incorporating complementary immunogenomic layers, empowering concurrent high-resolution SpecHLA typing and TRUST4-based TCR/BCR repertoire reconstruction from sequencing data. Deployed across two large-scale cohorts (IMvigor210 and OAKPOPLAR), IOBRpy successfully captured multi-dimensional prognostic insights. While broad HLA-I heterozygosity showed negligible impact, it precisely unmasked treatment-stratified, allele-specific survival associations (e.g., HLA-A*01 and HLA-DPA1*02) tightly coupled with distinct immunosuppressive ligand-receptor networks (such as HMGB1-THBD and EFNB2-EPHB6) and dynamic TCR clonal diversity shifts. Empowering this lifecycle is a paired agent framework: a workflow agent automatically audits project states to execute validated, path-aware commands, while a result agent evaluates tool provenance, handles method-aware adaptive visualizations, and organizes findings into evidence-constrained biological hypotheses. Collectively, IOBRpy provides a reproducible, scalable, and intelligence-augmented Python gateway to transform raw sequencing data into multi-omic, interpretation-ready discoveries for cohort-scale precision immunotherapy (https://iobr.github.io/IOBRpy/).
bioinformatics2026-07-22v1BioPhasor: Decoding Cellular State Tensors from Multi-Omics Phasor Dynamics for Quantum Ready Systems Biology
Sigdel, D.; Panday, N.Abstract
Integrating multi-omics data - transcriptomics, proteomics, metabolomics, single-cell - remains a fundamental challenge in systems biology. We present BioPhasor, a framework that encodes each measurement as a complex phasor z = e^(i{varphi}) on the compact N-torus T^N, modelling the cell as phase-coupled oscillatory programs whose dissipative dynamics generate limit cycles and an attractor landscape. From this geometry we derive the Cell State Tensor (CST), a rank-3 tensor whose axes we root in measured multi-omics quantities: a pathway/module atlas on the regulatory axis and a directional central-dogma modality axis. Across nine scenarios on open public data (GEO, CPTAC), loaded through one unmodified data layer, we report verdicts honestly: four reproduce, three are partial, two do not. A data-driven cell-cycle axis lifts agreement with a reference method from 0.34 to 0.69; an explicit circadian origin cuts peak-time error from 10.6 to 1.4 h; and central-dogma coupling mRNA phase organising protein amplitude clears a surrogate null and is tumour-specific. Grounding the quantum-ready claim, the CST maps to a density-matrix formalism whose coherence and entropy match quantum-information counterparts, and the phasor circuit transpiles gate-for-gate to a variational quantum circuit, though no empirical advantage emerges. One loader regenerates every reported number, and the code is released.
bioinformatics2026-07-22v1Target Preference Maps: A machine learning model generalizing transferable drug-receptor interactions and guiding drug discovery
Menezes, F.; Wahida, A.; Froehlich, T.; Grass, P.; Zaucha, J.; Napolitano, V.; Siebenmorgen, T.; Pustelny, K.; Barzowska-Gogola, A.; Rioton, S.; Didi, K.; Bronstein, M.; Czarna, A.; Hochhaus, A.; Plettenburg, O.; Sattler, M.; Nissen-Meyer, J.; Conrad, M.; Kurzrock, R.; Popowicz, G. M.Abstract
Modern AI models can decode the genomic landscape and protein structure world. Yet, they fail to generalize to one of the most important fields: small-molecule drug discovery. Since the late 1970s, the advent of macromolecular crystallography inspired the notion that structural knowledge alone could enable a lock-and-key approach to drug design. However, drug discovery continues to depend on costly, resource-intensive, and largely serendipitous screening campaigns that probe only an infinitesimal fraction of the drug-like chemical space. Despite some successful cases, our understanding of, and reasoning from, non-bonded interaction chemistry remains limited for general applicability. Furthermore, though structural databases contain hundreds of thousands of entries, a strong historical bias pervades protein-drug structures, hindering reliable advances through AI scaling. Here, we present a machine-learning framework that learns atom-type-specific spatial preference maps from local protein microenvironments in protein-ligand structures. By excluding ligand topology from the model input and learning from local atom-level environments, the framework is designed to reduce dependence on whole-ligand memorization and to capture transferable interaction preferences. The resulting maps recover chemically meaningful interaction patterns, including cases involving bridging waters and metal-dependent environments. The model was validated using retrospective and prospective real-world data in drug optimization when targeting a challenging protein-protein interface. This shows that the method can provide interpretable workflows to guide molecule optimization and provide input for downstream generative or docking workflows.
bioinformatics2026-07-21v11LeafRank: A phylodynamic framework for inferring relative fitness from single-cell phylogenies in chromosomally unstable tumors
Wu, C.; Leder, K.; Wang, Z.; Sun, R.Abstract
Tumors contain cancer cells with diverse growth potentials that shape evolutionary trajectories, yet this fitness diversity remains difficult to quantify in cases of whole-genome duplication (WGD) and chromosomal instability. We present LeafRank, a mathematical framework that leverages single-cell DNA-seq phylogenies to infer the relative fitness of individual cells. Using a multi-type branching process model, LeafRank integrates full tree topology, including branch lengths and bifurcation patterns, to estimate marginal fitness probabilities under punctuated evolutionary regimes driven by rare driver events. To account for elevated aberration rates following WGD, we introduce a tree-rescaling strategy that adjusts for lineage-specific genomic instability. Unlike methods focused on predefined subclones, LeafRank ranks all sampled cells, enabling flexible assessment of growth heterogeneity. Simulations demonstrate high accuracy across spatial and non-spatial virtual tumors. Applied to ovarian cancer, LeafRank reveals directional and parallel selection in WGD tumors and identifies recurrent copy number events enriched in high-fitness lineages. WGD lineages do not show immediate growth advantages but acquire fitness through subsequent alterations.
bioinformatics2026-07-21v4Panomap: Unbiased Nanopore Signal Mapping with Pangenome Variation Graphs
Shih, P. J.; Sanghani, Z.; Guarracino, A.; Gamaarachchi, H.; Batten, C.Abstract
Motivation: Signal-space nanopore mappers enable real-time mapping and filtering decisions directly from raw nanopore signals. However, existing signal-space mappers are built around linear references, and using a single representative reference can introduce reference bias when the sample diverges from that reference. Pangenome reference collections can reduce this bias by representing diversity across related reference sequences, but linear-reference signal mappers must treat each sequence as a separate target, redundantly storing shared sequences. Pangenome variation graphs provide a more compact representation by storing shared sequences once and encoding variants as alternative paths through the graph. Although sequence-to-graph mapping is well established for basecalled reads, existing signal-space methods do not directly use pangenome variation graphs. Results: We present Panomap, the first signal-space mapper that operates on pangenome variation graphs. Panomap maps raw nanopore signals to graph references, allowing signal-space mapping to use pangenome diversity while representing shared sequences once. We evaluate Panomap in three settings. First, when a single reference already maps the sample well, Panomap preserves mapping accuracy as additional reference sequences are added to the reference collection, while state-of-the-art signal-space tools regress. Second, when the exact sample strain is absent from the reference collection, Panomap benefits from adding related assemblies from the same species to the pangenome reference. Third, using a highly polymorphic locus, we show that Panomap can map reads from alleles not represented in the reference collection by using related alleles in the pangenome, with the largest gains for more divergent alleles and for decisions made from short prefixes of the read signal. In addition, Panomap's graph index scales sublinearly with pangenome collection size. Together, these results show that Panomap brings population-aware reference representation into signal-space mapping. Availability and Implementation: Panomap is open source and available at https://github.com/cornell-brg/panomap.
bioinformatics2026-07-21v2Genomic, Transcriptomic, and Regulomic Analyses Do Not Support Profound Autism as a Distinct Biological Category
Eicher, T. D.; Ne'eman, A.; Quackenbush, J. D.Abstract
The Lancet Commission on the Future of Care and Clinical Research in Autism proposed the construct of "profound autism" as a recognizable subtype of autism. Supporters argue that this classification is necessary to ensure that autistic persons with severe impairment receive appropriate research attention and policy support, whereas critics contend that the construct lacks scientific validity and may reflect social or political considerations more than biological distinction. To inform this debate, we evaluate whether the proposed "profound autism" category represents a distinct genetic phenotype using multiple molecular data types collected in a large cohort. Across genomic, transcriptomic, and regulatory analyses, we find no evidence supporting "profound autism" as a biologically distinct phenotypic group. Instead, differences emerge primarily in inferred gene regulatory networks distinguishing nonspeaking from speaking autistic children, suggesting potential regulatory mechanisms contributing to speech ability. These findings suggest that future research into severe impairment may be more productive if focused on specific traits -- such as speech impairment -- rather than attempting to define a distinct biological subtype within the multidimensional phenomenon of autism.
bioinformatics2026-07-21v2Molecular Mimicry in Inflammatory Bowel Disease: Multi-layered Functional and Sequence-level Analysis of Gut Microbial Proteins Mimicking the Human Proteome
Anand, A. A.; Mishra, P.; Srivathsa, V. S.; Yadav, V.; Samanta, S. K.Abstract
Inflammatory bowel disease (IBD), encompassing Crohn's disease (CD) and ulcerative colitis (UC), is a chronic inflammatory disorder whose pathogenesis involves intricate host-microbiome interactions. Molecular mimicry (the structural or functional resemblance between microbial and host proteins) represents a plausible mechanism by which gut microbiota may trigger or perpetuate autoimmune responses. Here we present a comprehensive, multi-layered molecular mimicry in silico pipeline (MMIP) analysis of 39 baseline shotgun metagenomic samples from Human Microbiome Project 2 (HMP2/IBDMDB). Using DIAMOND-based homology search against the SwissProt followed by UniProt Retrieve/ID Mapping (URIM), we characterized microbial protein functional space through three complementary frameworks: (i) normalized GO term frequency comparison across diagnostic groups, (ii) protein family (PFAM) domain enrichment analysis, and (iii) sequence-level mimicry analysis identifying microbial proteins with direct homology to human proteins. Taxonomic profiling using MetaPhlAn3 pre-computed abundance profiles provided further biological context. CD microbiomes exhibited greater enrichment of immune-relevant biological processes and a higher per-sample sequence-level mimicry rate than UC. CD and UC also showed distinct pathobiont profiles, with CD enriched for oral-origin taxa including Haemophilus parainfluenzae and UC enriched for Fusobacterium nucleatum and Prevotella species. Notably, healthy gut microbiome maintains coordinated mimicry of host neuronal, signal-recognition-particle-associated, and antimicrobial peptide machinery, a repertoire dismantled in IBD and replaced by disease-subtype-specific signatures, alongside a candidate mimicry link between CD-enriched bacteria and NOD2. These results represent the first metagenome-wide characterization of molecular mimicry across IBD subtypes using shotgun metagenomic data, offering new mechanistic insight into how microbial dysbiosis may contribute to immune dysregulation in IBD. Keywords: Inflammatory bowel disease (IBD), Crohn's disease (CD), Ulcerative colitis (UC), Gastroenteritis, Molecular mimicry
bioinformatics2026-07-21v2RNAStabFormer: Region-Aware Multi-Task Hybrid Learning for RNA Stability Prediction from Pulse-Chase Transcriptomics
Wang, S.; Zhang, C.Abstract
RNA stability is a major post-transcriptional regulator of gene expression, yet sequence-based prediction from pulse-chase transcriptomics remains difficult because labels depend on the time window, quantification region, and replicate quality. We present RNAStabFormer, a controlled RNA stability framework centered on a Region-Aware Multi-Task Hybrid Transformer (RAMHT). RAMHT encodes the nucleotide context of the 5-prime UTR, coding sequence (CDS), and 3-prime UTR; incorporates an additional CDS codon stream; upgrades engineered sequence features into a tabular interaction branch; and uses gated multi-task regression to predict four ENCODE BrU-seq/BruChase-seq RNA stability proxies, with Exon 6 h/0 h as the primary task. Across 26 outer data splits, including 23 chromosome holdout tests, a heterogeneous three-member RAMHT ensemble achieves a mean Pearson correlation of 0.773 on the primary task, statistically matching an engineered-feature XGBoost baseline, which also achieves 0.773. The mean paired difference is +0.000004, with a bootstrap 95 percent confidence interval from -0.003845 to +0.004077 and a Wilcoxon p-value of 0.8613. The ensemble improves upon the strongest individual RAMHT member, increasing the mean Pearson correlation from 0.768 to 0.773 and outperforming it on 23 of the 26 data splits. A strictly nested XGBoost-RAMHT blended model further increases the correlation to 0.775. Evaluations conducted on identical data splits also show that the ensemble outperforms frozen full-length mRNA language-model embeddings, which achieve a correlation of 0.760, and public LAMAR-DR transfer learning, which achieves a correlation of 0.180. Gate analysis, ablation experiments, and sequence recoding analyses indicate that engineered sequence grammar remains the dominant source of predictive information, whereas the nucleotide and codon branches provide complementary signals localized primarily within the CDS. RNAStabFormer narrows the performance gap between neural RNA sequence models and strong tabular baselines while retaining an extensible architecture for model interpretation and biological data integration.
bioinformatics2026-07-21v2SEEK-VEC: Augmenting topic modeling with spectral ensemble learning
Danning, R.; Ke, Z. T.; Ma, R.; Lin, X.Abstract
Count data are ubiquitous across many applications in which understanding latent patterns is of interest. Topic modeling is a powerful tool for detecting latent structure in count data. However, standard topic modeling methods are often constrained by their restrictive assumptions, susceptible to noise, and sensitive to misspecification of the number of topics. Here, we introduce SEEK-VEC (Spectral Ensembling of topic models with Eigenscore for K-agnostic Vocabulary Embedding and Classification), an ensemble topic modeling framework that integrates insights from multiple candidate topic models through a spectral ensembling procedure. SEEK-VEC produces a meta-structure matrix containing prioritization scores and grouping scores that enable variable classification, interactive pattern discovery, and model diagnostics. Through simulations, we demonstrate that SEEK-VEC augments the performance of standard topic models for identifying important vocabulary words and understanding the relationships among them, particularly when signal strength is weak. We apply SEEK-VEC to the MADStat dataset of statistical abstracts and demonstrate its utility for evaluating the proposed interpretation of a topic model.
bioinformatics2026-07-21v2COSMOS: A FAIR-aligned infrastructure for clinical trial data validation, warehousing, and interactive discovery
Roberts, C. A.; Nilsson-Takeuchi, A.; Stuart, C.; Chivers, M.; Soares, P.; Foure, V.; Griffiths, G.; Niazi, U.Abstract
The Clinical Omics System for Metadata and Outcome Storage (COSMOS) is an open-source, FAIR- and GCP aligned clinical trial unit (CTU) infrastructure designed to streamline the transition of academic clinical and multi-omic trial datasets into curated, analysis-ready repositories within Secure Data Environments (SDEs). By integrating an automated Data Quality and Data Validation (DQ&DV) "Trust Layer" with a relational Structured Query Language (SQL) schema, COSMOS enables programmatic and interactive data access via R Shiny applications. Dynamic integration of omics and clinical data is achieved through analytical data structures (e.g. ExpressionSet objects) linked via relational database identifiers. This lowers technical barriers for researchers and promotes governed data reuse and secondary discovery.
bioinformatics2026-07-21v1ProtSyntax: a protein large language model for decoding post-translational modification syntax and function
Lin, Y.Abstract
Post-translational modifications (PTMs) regulate protein function through dependencies among residue chemistry, sequence context, three-dimensional microenvironments and modification states, yet most predictors model sites independently and do not connect modification propensity to functional consequences. Here we present ProtSyntax, a PTM-centered protein language model trained on 4.25 million examples spanning 40 PTM classes and supervised for kinase specificity, PTM crosstalk and enzyme kinetics. ProtSyntax integrates bidirectional long-range modeling with geometry-gated attention in a sparse mixture-of-experts architecture and uses adaptive multi-objective learning to couple residue-level PTM syntax to protein-level function. Across 40 PTM-site benchmarks, ProtSyntax improved mean MCC and AP by 12.7% and 10.7%, respectively, relative to the best-performing baselines. It also distinguished authentic sites from structurally incompatible motif decoys, transferred to rare PTMs, recovered crosstalk, linked PTM perturbations to enzyme-kinetic changes and identified disease-associated PTM disruptions. Together, ProtSyntax provides an interpretable framework for decoding PTM regulation across the proteome.
bioinformatics2026-07-21v1A new set of DNA methylation variants in the human genome show predominant tissue specificity and sensitivity to reprogramming with a potential for disease susceptibility.
Anne, A.; Kumar, L.; Singh, M.; Choudhury, S.; Das, S.; Zimmer-Bensch, G.; Bandyopadhyay, D.; K, N. M.Abstract
Analyses of 3,370 normal human tissues of ectodermal, endodermal and mesodermal origins identified 12,587 regions averaging ~585 bp with significant differences in DNA methylation levels within identical tissues. These methylation variants (MeVars) occurred in 8,037 genes enriched in neurological disorders and cancers of which, majority were tissue-specific rather than being systemic. This somatic variation was reduced by reprogramming in vitro into iPSCs and in vivo during spermatogenesis. Analysis of prefrontal cortices showed a higher incidence of MeVars in the candidate genes in controls than schizophrenia patients wherein a subset showed significantly altered transcript levels. Similar effects were observed for oral tissues and skin fibroblast cells. MeVars showed significant association with SINE1, simple and low complexity repeats, H3K27me3, H3k9me3 and H3K4me1 modifications and EZH2, SUZ12 and REST binding sites. Collectively, MeVars have postzygotic origins with an ability to reset during reprogramming, adding a new dimension in the form of epigenetic diversity and its relevance to disease susceptibility in humans.
bioinformatics2026-07-21v1Bayesian Factor Analysis for Binary and Ordinal Phenotypes with Missingness
Shashaank, N.; Knowles, D. A.Abstract
Binary and ordinal phenotypes are common in clinical screening and self-reported questionnaires, but many factor analysis and matrix factorization methods are only applicable for quantitative phenotypes with real-valued and/or continuous data distributions. To address this, we propose FABOr (Factor Analysis of Binary and Ordinal data), a Bayesian framework for matrix factorization in which the low-rank matrices are latent variables with continuous priors while the phenotypes are observed variables modeled with appropriate binary/ordinal likelihoods. We also develop missing not at random (MNAR) extensions of FABOr for analyzing data with structured missingness. In experiments with simulated phenotypes, we found that FABOr performs similarly to the best-performing benchmark methods on binary data at imputation and exceeds the performance of all tested benchmark methods on ordinal data. We then applied FABOr to analyze real-world binary and ordinal phenotypes from the Simons Foundation SPARK dataset on autism spectrum disorder (ASD) and found that it improved imputation accuracy by up to 5% on binary data and up to 23% on ordinal data relative to the benchmark methods.
bioinformatics2026-07-21v1PepCL: A replay-based continual learning framework for updating peptide-MHC models
Chati, P. M.; Lashkari, V. D.; Salhotra, A.; Bruno, P. M.; Ntranos, V.Abstract
Understanding peptide-major histocompatibility complex (MHC) class I binding is critical for effective vaccine and immunotherapy design but is a combinatorially complex challenge for which prediction models have become essential. MHC ligands are typically identified at scale via untargeted mass spectrometry (MS), and this has built a strong base for peptide-MHC model training. However, MS incompletely captures the vast peptide-MHC space due to technical, sampling, and biological biases. Although recently developed experimental assays have queried such blind spots yielding complementary information, existing peptide-MHC predictors have not yet incorporated these orthogonal data and are not designed to be updated as new data are generated. Here, we introduce PepCL (Peptide-MHC Continual Learning), a continual learning framework for updating peptide-MHC predictors with new assay data while explicitly preserving prior MS knowledge. To enable PepCL, we also develop MHCPrime, a new state-of-the-art pan-allelic peptide-MHC prediction model, trained on publicly available MS data, that can be effectively updated under our framework. We demonstrate that PepCL allows MHCPrime to learn previously unseen, assay-specific information while preventing catastrophic forgetting that is typically observed with conventional fine-tuning. We evaluate PepCL and MHCPrime in a variety of biological contexts, including infectious disease and cancer, and show improved peptide-MHC prediction that transfers across alleles for broader applicability in clinical settings. Overall, our results establish PepCL as a flexible framework for extending the utility of peptide-MHC models by improving their predictive performance as immunopeptidomics assays continue to evolve and new data become available.
bioinformatics2026-07-21v1An automated platform for spatial functional modeling and fingerprint analysis of tissue molecular landscapes
Hajihosseini, M.; Patino-Martinez, E.; Ghosal, R.; Kaplan, M. J.; Pyne, S.Abstract
Spatial transcriptomics (ST) enables high resolution molecular profiling while preserving tissue architecture, creating new opportunities to investigate how disease-associated pathways are organized within tissues. However, existing analytical approaches largely focus on individual pathways or cell types and do not provide a unified framework for modeling spatially varying pathway interactions across tissue sections and anatomical planes. Here, we present an integrative framework, Spatial Fingerprints Analytics (SFinx), that combines reference-free deconvolution, pathway activity reconstruction, and Spatial Functional Data Analysis (SFDA) to map localized pathway activity and pathway phenotype interactions in complex tissues. Applying SFinx to10x Visium ST datasets from murine lupus nephritis, we reconstructed continuous spatial landscapes of pathway activity and disease-associated phenotypes across kidney sections. This approach identified anatomically restricted inflammatory domains characterized by coordinated activation of immune pathways and revealed substantial spatial heterogeneity in pathway crosstalk across renal compartments. Using generalized additive models with tensor-product splines, we quantified spatially varying associations between lupus nephritis and neutrophil activation pathways across tissue sections, uncovering regions with both positive and negative relationships that would be obscured by conventional bulk analyses. Multi-slice integration further demonstrated reproducible spatial interaction patterns while accounting for section-specific variability. Together, SFinx transforms mixed-spot transcriptomic measurements into interpretable spatial pathway landscapes and interaction maps, providing a general framework for identifying localized disease mechanisms. This approach reveals previously unrecognized spatial organization of inflammatory signaling in lupus nephritis and presents a broadly applicable strategy for studying spatially coordinated biological processes in autoimmune disorders, cancer, and neurodegenerative diseases.
bioinformatics2026-07-21v1Stratified Immune Profiling Uncovers Prognostic Heterogeneity Beyond MYCN Amplification and Age in Neuroblastoma
Magno, J. M.; Muzzi, J. C. D.; Resende, J. S. S.; Querne, L. B. P.; Alvarenga, L. M.; Cavalli, L. R.; Figueiredo, B. C.; Castro, M. A. A.Abstract
Neuroblastoma is the most common extracranial solid tumor in children, presenting remarkable clinical heterogeneity with survival outcomes ranging from spontaneous regression to aggressive progression. MYCN oncogene amplification and age at diagnosis are established prognostic factors that are typically treated as independent covariates in risk stratification, yet their joint influence on the tumor immune microenvironment remains poorly understood. Here we show that stratifying patients by both variables simultaneously reveals six reproducible immune subtypes with distinct transcriptional programs and prognostic significance. Consensus clustering of immunomodulatory gene expression profiles from 149 patients in the TARGET-NBL cohort identified subtypes whose survival trajectories differ significantly within clinical strata defined by MYCN status and age at diagnosis. A linear Support Vector Machine classifier trained on these subtypes, using immunomodulatory gene expression combined with MYCN amplification status and age at diagnosis as predictive features, achieved 97.2% accuracy and a Cohen's Kappa of 0.963 under 10-fold cross-validation, and generalized to an independent cohort of 493 patients (GSE62564). Kaplan-Meier analysis revealed significant survival differences across subtypes in both cohorts (TARGET-NBL: log-rank p = 0.0018; GSE62564: log-rank p < 0.0001). Single-sample gene set enrichment analysis identified differential activation of proliferative and immune response pathways across subtypes, consistent between both cohorts. These findings suggest that integrating MYCN amplification status and age at diagnosis as joint determinants of immune organization may reveal prognostic heterogeneity that is not fully captured when these factors are considered independently.
bioinformatics2026-07-21v1GeneAutomate: A Browser-Based, Integer-Indexed Platform for Dual-Gene-List Functional Annotation and Interactive Network Visualization
Singh, R. P.; Kumar, A.Abstract
Comparative interpretation of two gene lists, for example, two treatment arms, two tissues, or a discovery and a validation cohort, is a routine task in functional genomics. While several tools offer dual-list comparison (e.g., EnrichmentMap, RRHO packages), they typically require local software installation, R/Bioconductor, or manual reconciliation of separate single-list outputs. Most widely used web-based enrichment tools (DAVID, g:Profiler, Enrichr, ShinyGO, WebGestalt) are built around the analysis of a single gene list at a time, and those that support comparison often lack interactive, publication-ready visualization or depend on server-side query latency. Here we present GeneAutomate, a browser-based tool purpose-built for side-by-side comparison of two gene lists. GeneAutomate performs Over-Representation Analysis (ORA) against Gene Ontology (GO) and Reactome using an exact hypergeometric test with Benjamini-Hochberg false discovery rate correction, and Gene Set Enrichment Analysis (GSEA) when ranked (log2 fold-change) input is supplied, alongside Protein-Protein Interaction (PPI) subgraph extraction from BioGRID physical interactions. All reference data (Gene Ontology, Reactome, BioGRID, and NCBI/Ensembl identifier cross-references) are pre-compiled offline into a single integer-indexed database of approximately 32 MB for Homo sapiens, in which every gene identifier Ensembl ID, Entrez ID, official symbol, or alias is resolved to one canonical integer prior to any user query. This design removes live database round-trips from the runtime path, enabling fast, at-your-desk enrichment without installation or a server-side per-query bottleneck. The tool renders thirteen interactive, D3.js- and Cytoscape.js-based comparative visualizations, including a Rank-Rank Hypergeometric Overlap (RRHO) heatmap, a GO-slim "Radar/Spider" functional fingerprint, and chord/edge-bundled cross-talk diagrams that are, to our knowledge, not offered as an integrated set by any existing academic or commercial ORA/GSEA platform. GeneAutomate is an unfunded, individual student project developed with feedback from a professor, and is in its final stage of development. It requires no installation or login. We describe the tool's architecture, statistical methods, and comparative feature set relative to established academic tools (DAVID, ShinyGO, g:Profiler, Enrichr, WebGestalt, STRING, PANTHER, GeneMANIA, Cytoscape, clusterProfiler, GSEA, Metascape) and commercial platforms (IPA, MetaCore, Pathway Studio, iPathwayGuide, Partek Pathway), and we state candidly the current version's limitations, which are planned to be the added in next version: single-species (human-only) coverage, no upstream regulator analysis, and comparison currently limited to two (occasionally three) concurrent lists. GeneAutomate is available at https://geneautomate.tech/.
bioinformatics2026-07-21v1Extended t-cores for the de novo identification of transposable elements and other inexact repeats from short read RNAseq data
Darmon, S.; Mary, A.; Lacroix, V.Abstract
Transcribed repeats represent a major challenge in the de novo assembly of transcriptomes from short RNA-seq reads. Young transposable elements (TEs) and other inexact repeats create dense and ambiguous regions in the assembly graph, preventing the correct assembly of transcripts. In this paper, we introduce a fully de novo method based on the discovery of dense regions in the compacted De Bruijn graph (DBG) to identify such repeats directly from short reads RNA-seq data, without requiring a reference genome or repeat database. Our approach defines the extended t-cores, subgraphs of the DBG that capture the complex topology induced by highly expressed inexact repeats appearing in RNA-seq reads. Independently of its interest for transcriptome assembly, the proposed method appears to be effective for the de novo identification of repeats in transcriptomes. After classifying cores using sequence-based motifs to distinguish simple repeats from potential TEs, we demonstrate its potential for the de novo discovery of transposable elements. We validate the approach on a Mus musculus dataset using expressed TE consensus sequences, showing that extended t-cores correspond to known expressed TE families. We also illustrate its de novo discovery potential on a non-model species, Canis lupus familiaris, where the method was also able to recover known transposable elements.
bioinformatics2026-07-20v2Leveraging multiplicity in biologically informed neural networks to uncover disease heterogeneity
Gankin, D.; Beltrao, P.Abstract
Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map to named biological entities. Yet BINNs have been hard to train at biobank scale, and the reliability of their interpretations remains largely untested. Here we present a fast BINN implementation trained on UK Biobank genotype and plasma proteomics data from about 500,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph topology, because node degree and layer position influence the scores. Normalization reduces this bias but can weaken enrichment for known disease genes. Second, BINNs show substantial predictive multiplicity, that is, independently trained models with identical architecture and data reach similarly accurate solutions while prioritizing different genes and pathways. Although this multiplicity makes single-model explanations unstable, the range of interpretations can itself reveal disease biology. Across 100 replicate BINNs for type 2 diabetes, we find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways, mirroring known disease heterogeneity. Thus, analyzing the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms.
bioinformatics2026-07-20v1FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations.
Clemente-Larramendi, I.; Hillion, S.; Cornec, D.; Jamin, C.; Foulquier, N.Abstract
Single-cell RNA sequencing (scRNA-seq) enables detailed characterization of cellular heterogeneity, yet understanding the full cellular and regulatory environment of complex tissues remains challenging. In the era of large single-cell atlases, this technology has become increasingly accessible, and datasets have grown in scale and statistical power. As a result, sample representation methods have emerged as a promising strategy to summarize patient-level biological variation. However, most existing approaches rely on unsupervised learning frameworks with ambiguous biological interpretability. Here we present a Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method. FloREN models single-cell data as a heterogeneous network integrating cells and genes together with gene regulatory and cell-cell communication relationships. Through condition-aware embeddings and interpretable attention networks, FloREN enables improved sample stratification and biomarker discovery. In addition, the framework supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
bioinformatics2026-07-20v1Metagenomic profiling of bacterial endosymbionts in wild mutant and permethrin-susceptible head lice shows an expanded microbiota in the resistant strain
Mohammadi, J.; Alipour, H.; Azizi, K.; Kalantari, M.; Moemenbellah-Fard, M. D.Abstract
Background Human head lice (Pediculus humanus capitis de Geer) are known to harbor diverse maternally inherited bacterial symbionts. These endosymbiotic bacteria may contribute to insecticide degradation, potentially helping lice withstand particular environmental pressures. Using next-generation sequencing (NGS), this study investigated the bacterial symbionts present in wild head-louse populations and characterized their phylogenetic relationships in mutant and permethrin-susceptible strains. Methods Head lice specimens were collected from 10 locations across Fars Province, Iran. Following DNA extraction, the samples were analyzed using polymerase chain reaction (PCR), and the resulting amplicons were sequenced to detect mutations in specimens from each location. Lice were then classified according to the presence or absence of mutations and subjected to NGS to characterize the symbiotic bacterial communities in mutant and putatively permethrin-susceptible strains. Results Mutant strains were detected at only three sampling stations. Bioinformatic analysis of the nucleic acid sequences revealed three exons and two introns, with an expected amplicon length of 582 bp. NGS analysis showed that Candidatus Riesia pediculicola (Arsenophonus), belonging to the phylum Proteobacteria, was the predominant bacterial genus. Actinobacteria and Firmicutes were the second- and third-most abundant phyla, respectively. Most of the remaining 25 bacterial taxa were associated with the mutant strain. Additionally, two previously unreported bacterial genera were deposited in GenBank. Conclusions The distinct distribution of Arsenophonus species between susceptible and mutant head-lice strains, along with the greater abundance of Escherichia, Shigella, Lawsonella, and Megamonas in mutant strains, highlights the need for advanced metagenomic analyses to determine how these endosymbionts may help their hosts withstand specific environmental disturbances.
bioinformatics2026-07-20v1GDTR: Layer-wise Settling Depth Reveals Biological Grammar in Genomic Foundation Models
Cho, Y.; Kang, J.; Park, S.; Kim, S.Abstract
Genomic foundation models capture sequence regularities, yet existing interpretability tools rarely ask where in the layer stack a biological grammar becomes stable. We introduce GDTR, the Genomic Deep-Thinking Ratio, a training-free residual-stream lens that assigns each nucleotide token a settling depth : the first layer at which its representation stabilizes against the post-final-norm reference. On Evo 2 7B, splice donor and acceptor sites settle approximately two layers earlier than intronic contexts; enhancer-like cCREs show a smaller but measurable shift; and a chromosome 22 calibration transfers to held-out chromosome 17. Perturbing canonical splice donors shows that the signal is bidirectional: disrupting the central GT motif deepens settling, whereas shuffling the flanking grammar makes the preserved motif settle earlier. Differential GDTR further reveals consequence-associated peak-disruption depths across ClinVar variants, with synonymous substitutions peaking deepest but showing broad class overlap. GDTR therefore provides a layer-wise interpretability axis for genomic foundation models, complementary to existing prediction and variant-scoring tools.
bioinformatics2026-07-20v1EBD-DTI: Episodic Bridge Diffusion for Zero-Shot Cold-Start Drug-Target Interaction Prediction
Liu, J.; Le, J.; Wei, C.; Liu, M.; Yin, Z.Abstract
Predicting drug-target interactions (DTI) for entirely unseen drugs or proteins---the cold-start problem---remains a critical challenge in computational drug discovery. While sequence-based methods naturally support zero-shot generalization, they often ignore relational topology, and existing graph-based approaches either rely on global diffusion that blurs the boundary between inductive and transductive evaluation or require a few known interaction samples at test time (few-shot). We present EBD-DTI, a framework that enables zero-shot inference in graph-based DTI models without requiring any known interactions for unseen entities. The key innovation is episodic cold-start training: at each epoch, a random subset of training entities is masked and treated as pseudo-cold, forcing the model to learn cold-start inference with explicit gradient supervision. A bridge-conditioned local subgraph, together with multi-hop diffusion, provides cold entities with relational context from their nearest observed neighbors. Experiments on three benchmarks (BioSNAP, BindingDB, and DrugBank) demonstrate that EBD-DTI achieves competitive or superior performance compared to state-of-the-art methods under strict zero-shot evaluation, with episodic training improving AUC by up to 12%.
bioinformatics2026-07-20v1Learnable Graph Network Model (LGNM): A Physics Constrained Graph Neural Network with Quantum Hamiltonian Learning
Sharma, B.; Sarkar, C.Abstract
Elastic Network Models (ENMs), particularly the Gaussian Network Model (GNM) and its distance-weighted variant (mENM), predict per-residue protein flexibility from C contact graphs at low computational cost. Their central limitation is the assumption of uniform spring constants, which ignores the chemical identity, burial depth, and evolutionary conservation of individual residue contacts. We introduce the Learnable Graph Network Model (LGNM), a heterogeneous ENM in which per-edge spring constants are parameterised by per-residue flexibility coefficients predicted by a physics-constrained Graph Neural Network (GNN). The GNN is trained on molecular dynamics (MD)-derived root-mean-square fluctuation (RMSF) profiles from 413 proteins in the ATLAS database, using fold-disjoint CATH superfamily splits. The learning objective is an instance of the Quantum Neural PDE (QNPDE) Hamiltonian learning framework, with K = 3 operator types enabling an O(K) quantum gradient. On 91 held-out test proteins, LGNM achieves mean per-protein Pearson correlation r = 0.8549 , versus r = 0.8024 for mENM. The implementation of this methodology is available at https://lgnm.compbiosysnbu.in/ allowing researchers to evaluate flexibility and downstream processes. Keywords: Protein flexibility; Elastic Network Model; Physics-constrained Graph Neural Network; Residue fluctuation.
bioinformatics2026-07-20v1From vehicles to wildlife: transferable deep learning for trajectory generation
Patras, J.; Fablet, R.; Brunel, A.; Roy, A.; Bugoni, L.; Tavares Nunes, G.; Barbraud, C.; Jacoby, J.; Benboudjema, S.; Passuni, G.; Delord, K.; Lanco, S.Abstract
Realistic simulation of animal movement is fundamental to conservation, habitat modeling, and ecological scenario evaluation. Traditional approaches struggle to capture multi-scale trajectory dynamics, while generative deep learning for complete trajectory simulation remains largely unexplored in ecology due to data scarcity. We show that a diffusion model pre-trained on millions of vehicle GPS trajectories can be fine-tuned on hundreds of seabird central-place foraging trips to simultaneously generate ecologically realistic trajectories for five species and six breeding colonies. The fine-tuned model consistently outperforms four state-of-the-art baselines (GAN, VAE, HMM, iSSF) across movement metrics spanning step dynamics, spatial distribution, and behavioral temporality, with comparable or shorter computation times. Domain transfer reaches full performance in 35 minutes versus 7 hours from scratch, and conditioning on species and colony enables generalization to unseen combinations from as few as 10 trajectories. These results establish cross-domain transfer learning as a new paradigm for data-efficient generative animal movement modeling.
bioinformatics2026-07-20v1Interpretable Prediction of Phase Separation and Disease Variant Effects in Intrinsically Disordered Regions
Zhao, M.; Kumar, S.Abstract
Coding mutations within intrinsically disordered regions (IDRs) of proteins are increasingly implicated in human diseases yet remain poorly interpreted by conventional variant-effect predictors that rely on structural stability and conservation-based metrics. Quantifying disruption of IDR-mediated liquid-liquid phase separation (LLPS) offers a biophysically principled approach to interpreting the pathogenic impact of such variants. However, existing LLPS predictors suffer from training biases toward self-separating proteins, show limited performance on partner-dependent phase separation, and often lack interpretability for variant prioritization. We present an interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent. Our two-step classifiers outperform existing methods on independent benchmark datasets, with the largest gains for partner-dependent LLPS proteins. Beyond classification, our framework identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS. Applied to disease-associated variant databases, we found that pathogenic mutations are enriched in predicted phase-separating regions and frequently perturb LLPS propensity scores, implicating mutation-induced LLPS dysregulation as a potential pathogenic mechanism for numerous diseases. Overall, our framework provides an accurate, interpretable approach for identifying phase-separating proteins and linking aberrant phase-separation behavior to disease pathogenesis.
bioinformatics2026-07-20v1DNAS-Bench: Deterministic Nucleic Acid Screener Benchmarking
Wong, H. C.; Kohno, T.; Nivala, J.Abstract
The rapid growth of biotechnology manufacturing for synthetic DNA and proteins has raised concerns that adversaries could exploit commercial synthesis pipelines to create biological weapons. Without effective safeguards, an attacker could seek regulated genetic sequences from synthesis providers; while synthetic DNA is not itself a pathogen or toxin, access to such sequences can lower barriers to downstream misuse, motivating robust order-time screening. To mitigate this risk, Biosecurity Screening Software (BSS) systems have been developed to flag potentially malicious synthesis orders. Here, we propose one of the first deterministic benchmarks for evaluating the robustness of Biosecurity Screening Software. Our framework enables systematic testing of BSS behaviors and potential on specific nucleic-acid sequences and on targeted regions of malicious genomes. Our framework allows for insights into what is being flagged as malicious in BSSs, leading to potential discussions if specific BSS is fit for a specific manufacturing pipeline. We additionally introduce a dataset of manipulated genomes derived from the HHS and USDA Select Agents and Toxins List. When evaluated on this dataset, SeqScreen flags 42% of the sequences as malicious, while Commec flags 10.2%. Across a range of manipulation strategies, we find that simple manipulations, such as padding sequences by adding a repeated nucleotides at 1.5 times the original length, perform nearly as well as more targeted methods, such as embedding malicious sequences within benign genomic context. Padding-based methods trail embedding-based methods by only 0.75 percentage points in average detection rate. Consistent with prior reports from BSS developers and studies, we observe a sharp drop in detection rate when input sequence length falls below a critical threshold, typically between 50 and 100 base pairs (bp). Under our threat model, this implies that an adversary can bypass most existing safeguards by splitting a target genome into fragments shorter than ~50 bp. Fragment-level analysis further reveals that some toxin regions evade detection entirely by SeqScreen, while other malicious genomes remain detectable even when fragmented into 30-50 base-pair segments. We open-source this benchmark to support reproducible evaluation of BSS robustness and to inform the development of next-generation biosecurity screening tools (https://github.com/HenryCWong/DNAS-Bench). For ethical concerns we only open-source the framework while the data is available upon request.
bioinformatics2026-07-20v1multiScaleAC: Cell-Cell interaction with Moran's I as a function of kernel bandwidth
Soupir, A. C.; Hayes, M. T.; Manley, B. J.; Wang, X.; Wrobel, J.; Peres, L. C.; Fridley, B. L.Abstract
Over the last decade, spatial transcriptomic technology has transformed our understanding of tissue architecture including cell-cell interactions within the tumor immune microenvironment. A specific use-case of increasing interest is leveraging the spatial statistical relationship of genes whose protein products are known to be involved in ligand-receptor interactions. One methodological limitation of this approach has been the requirement to choose one radius around a cell as a parameter that can come with selection biases. Rather, interactions between cells vary in strength across a range of spatial scales that single-radius choice may miss. To fill this gap we developed `multiScaleAC` to extended Moran's I, a correlation measure that accounts for locations of values, by employing a Gaussian kernel applied to locations and varying the bandwidth parameter h. The resulting Moranis I\left(h\right) then can be compared between samples using functional data analysis. In the current study, we used simulations to show that our framework has well controlled Type I error due to the use of permutations for assessing significant interactions. We also demonstrate that `multiScaleAC` has high statistical power to identify a significant interaction when a true interaction is simulated (1.00 at bandwidths greater than 5) and increasing power as bandwidth increases when negative interaction is simulated. We found `multiScaleAC` largely captures similar significant ligand-receptor profiles in 8 Visium samples of colon tissue using the same bandwidth as `spatialDM` but without removing low-weight spots from the weight matrix (73.7% - 87.2%). Applying the `multiScaleAC` framework to our previous single-cell spatial transcriptomics data (COL4A1-ITGAV in the stromal compartment of clear cell renal cell carcinoma) followed by functional principal component analysis, we found functional principal component 1 to represent global interaction elevation/depression. Associating functional principal component 1 scores with immunotherapy exposure showed significantly higher scores in stromal tissues exposed to immunotherapy than those naive to immunotherapy, indicating an overall higher interaction of cell expressing COL4A1-ITGAV. These findings recapitulate our previous study while reducing bias in neighbor selections. We believe this is the first study to apply a functional extension of Moran's I in combination with functional data analysis to understand cell-cell interaction over spatial scales.
bioinformatics2026-07-20v1Interpretable Peripheral Blood Cell Classification via Vision-Language Concept Bottleneck and Soft Decision Tree
Chen, K.; Hu, T.Abstract
Motivation: Deep learning classifiers for medical image analysis typically function as black boxes, disclosing neither the image features underlying their predictions nor the reasoning by which individual decisions are reached. Peripheral blood cell classification exemplifies this challenge: experienced laboratory professionals identify cell types through structured morphological criteria---nucleus shape, chromatin texture, nucleus-to-cytoplasm ratio, granularity, and staining properties---yet existing automated systems cannot express their reasoning in these same terms, impeding clinical audit and verification. Results: We present a two-stage interpretable pipeline that addresses both levels of opacity. In the first stage, a frozen domain-adapted vision-language model (ConceptCLIP) projects each cell image onto a 70-dimensional vector of morphological concept scores via zero-shot cosine similarity, eliminating the need for per-image concept annotations. In the second stage, a Soft Decision Tree (SDT) classifies cells solely on these concept scores, producing a deterministic, concept-based decision path for each prediction. On BloodMNIST (eight cell types, 3,421 test images), the full pipeline achieves 94.86% test accuracy---approximately 3 percentage points below the black-box ceiling---while providing fully traceable decision logic. Post-training histological annotation confirms that the learned routing logic aligns with established hematological morphology criteria and reveals an emergent separation of immature granulocyte subtypes (promyelocyte versus metamyelocyte) without subtype supervision, demonstrating that concept-based decision trees can recover clinically meaningful distinctions beyond the granularity of the training labels. Availability and implementation: The source code, trained SDT weights, precomputed concept score data, and inference scripts are publicly available at https://github.com/aquamarineaqua/CLIP-CBM-SoftDecisionTree.
bioinformatics2026-07-20v1A Label-Free Multi-Metric Pipeline for Benchmarking Single-Cell RNA-Sequencing Clustering and Testing the Reproducibility of Cell-Type Heterogeneity
Yousef, Z.; Simone, J.; Klein, D.; Cho, H.; Bhuyan, A.; Bhatt, P.; Wu, J.; Wang, H.; Cai, L.Abstract
A discovered sub-population from single-cell transcriptomic data is only meaningful if it is reproducible, yet clustering is usually done with one method on one embedding and rarely tested. We present a label-free, multi-metric pipeline that reframes clustering as an auditable, methods-blind decision and separates two notions of stability that are commonly conflated: reproducibility under cell resampling (bootstrap) and reproducibility under re-embedding (retraining the representation). The pipeline evaluates seven clustering configurations across cluster counts using five non-redundant quality metrics. As a whole-dataset control on a mouse retinal atlas, it recovers an eight-cell-type annotation at 96.3% accuracy (adjusted Rand index, ARI = 0.91) without labels. We then validate the discovery mode on two cell types with opposite ground truth. On bipolar cells, which have well-established subtypes, the pipeline accepts the sub-structure: across-embedding reproducibility rises with cluster number to a high plateau (mean pairwise ARI ~0.93 near the ~15 known bipolar subtypes), with quality metrics improving in parallel. On rod photoreceptors, treated as homogeneous, it rejects over-clustering: the metric-selected partition passes a bootstrap-stability check but is not reproducible when the embedding is retrained (mean pairwise ARI = 0.69), and the metrics do not improve with cluster number. On synthetic data, the test recovers real structure down to a 5% subpopulation while rejecting null data (high sensitivity and specificity). Bootstrap stability alone is therefore insufficient evidence for sub-population; the across-embedding test discriminates real sub-structure from over-clustering and applies to any cell type as a reproducible alternative to single-method, single-embedding clustering.
bioinformatics2026-07-20v1DPCGS: a computational framework for linking GWAS to single-cell transcriptomics in complex traits and diseases
Liu, C.; Shen, B.; Li, J.; Zhu, R.; Yang, P.; Wu, B.; Xuan, Y.; Yang, S.; Yuan, B.; Yang, N.; Ma, L.; Liu, Q.; Dai, S.; Zhang, Y.Abstract
Complex traits and diseases arise from the interplay between genetic variation and cellular heterogeneity, making it essential to understand how genetic risk manifests at the cellular level. However, connecting genome-wide association studies (GWAS) to specific cell populations remains challenging due to cellular complexity and the prevalence of noncoding variants. Here, we present DPCGS, a computational framework that systematically integrates GWAS summary statistics with single-cell RNA-sequencing (scRNA-seq) data to identify trait-associated cell populations, genes, and regulatory programs. DPCGS is based on the principle that GWAS-prioritized genes should exhibit elevated expression in relevant cells compared with matched controls. Benchmark analyses with simulated datasets showed that DPCGS consistently outperforms existing methods, achieving higher accuracy and sensitivity in detecting trait-relevant cells. Applications to diverse scRNA-seq datasets further validated its robustness, revealing oligodendrocytes and astrocytes as key subpopulations in Alzheimer's disease and macrophages and B cells in asthma. These analyses also highlighted potential molecular regulators, including CD74, FOS, FLI1, and AP-1 transcription factors. Together, these findings establish DPCGS as a versatile framework for dissecting the cellular and molecular basis of complex traits and diseases, with broad implications for biomarker discovery and therapeutic development.
bioinformatics2026-07-20v1Mapping Tumor-Microenvironment dependencies with TMEformer: A spatial foundation framework enabling in silico perturbation
Chen, S.; Zhu, G.; Yang, L.; Wei, X.; Li, S.; Liu, P.; Chen, Q.; Zhang, Z.; Liu, D.; Tang, Y.; Xu, G.; Zhou, M.; Luo, J.; Huang, L.; Chen, B.; Ou, S.; Jiang, J.Abstract
Despite the fundamental role of spatial context in driving tumor progression, most current computational models for virtual perturbation have largely overlooked its importance. Here, we introduce TMEformer, a tumor microenvironment-aware deep learning framework that leverages high-resolution spatial transcriptomics to jointly model intrinsic tumor cell programs and local microenvironmental signals by explicitly incorporating spatial architecture. Validated across diverse tumor spatial transcriptomic cohorts, TMEformer enables virtual perturbations that capture functional dependencies within local cellular ecosystems. Despite being trained on cancer-specific spatial datasets, TMEformer outperforms baseline models pretrained on large-scale corpora in capturing key tumor transitions, including lineage plasticity and the emergence of therapy resistance. Systematic perturbation analyses prioritize tumor-intrinsic transcription factors and TME-derived ligands that drive disease progression, recovering established regulators and revealing novel candidates. Furthermore, TME-derived embeddings improve the spatial stratification of tumor cells and align more closely with pathological architecture. Together, TMEformer establishes a general framework for modeling tumors as spatially coupled, perturbable ecosystems.
bioinformatics2026-07-17v2MAJEC: unified gene, isoform, and locus-level transposable element quantification from RNA-seq
Lim, T.-Y.; Firestone, A. J.Abstract
Background: The study of transposable elements (TEs) has become increasingly central to fields such as cancer biology, immunology, and aging. Accurately quantifying disease- or laboratory-mediated perturbations in these elements is critical to support this expanding research, yet current RNA-seq pipelines struggle with the pervasive overlap between TEs and protein-coding genes. Existing tools either aggregate to the subfamily level with no locus resolution (TEtranscripts), or provide locus-level quantification without modeling gene overlap (Telescope), with the latter attributing over 40% of TE signal to the 1.1% of loci that overlap gene exons. Results: We present MAJEC (Momentum Accelerated Junction Enhanced Counting), a unified Expectation-Maximization (EM) framework that jointly quantifies genes, transcript isoforms, and individual TE loci from BAM alignments in a single pass. Splice junction evidence informs transcript-level priors, enabling MAJEC to probabilistically distinguish genic from TE-derived reads. This approach was independently validated against Salmon and RSEM on isoform quantification benchmarks. The joint feature space reduces exon-overlap contamination of locus-level TE estimates from 43% of total signal (Telescope) to 5% (MAJEC), while preserving subfamily-level accuracy (differential expression r = 0.987 vs TEtranscripts). Using paired biological vignettes, we demonstrate that MAJEC correctly resolves both the false TE reactivation artifacts endemic to TE-only models, and the false gene upregulation artifacts that occur when heuristic rules misassign genuine intragenic TE transcription. Conclusion: MAJEC simultaneously produces the isoform and locus-level resolution that TEtranscripts lacks, with greater accuracy than Telescope, and runs faster than either.
bioinformatics2026-07-17v2Beyond Bisulfite Sequencing: Resolving 5-hmC with Nanopore Sequencing Unmasks the True-5mC Methylation Entropy Landscape
Bertocchi, U.; Katz, E.; Jeffet, J.; Grunwald, A.; Gabay, N.; Deek, J.; Verma, S.; Shwartz, A.; Umschweif-Nevo, G.; Lerer, B.; Roichman, Y.; Ebenstein, Y.Abstract
DNA methylation dynamically regulates cellular function and phenotype. At the tissue level, stochastic variation in methylation patterns, measured as methylation entropy, drives plasticity, development, cancer, and aging. Demethylation is facilitated by erasure of 5-methylcytosine (5mC) via the oxidized intermediate 5-hydroxymethylcytosine (5hmC), but bisulfite sequencing cannot distinguish these modifications, classifying both as 5mC. Using nanopore sequencing with direct detection of 5mC and 5hmC, we quantified how this historical conflation affects genome-wide methylation levels and methylation entropy in kidney cancer and the mouse medial prefrontal cortex. Bisulfite-like analysis introduced systematic, tissue-specific shifts in methylation distributions, influencing biological interpretation. However, these effects were modest in the low-5hmC kidney cancer samples, where pathway-level results remained highly concordant. Our findings demonstrate that True-5mC-based methylation entropy redefines the physical mapping of epigenomes, demonstrating that, in some contexts, what was previously interpreted as stochastic maintenance failure is frequently the structured signature of distinct, mechanistically interpretable cytosine biochemistry.
bioinformatics2026-07-17v2Nextstrain automates real-time phylogenetic analysis of open data for endemic and emerging pathogens
Andrews, K. R.; Chang, J.; Roemer, C.; Hadfield, J.; Lin, V.; Brito, A. F.; Daodu, R.; Joia, I. A.; Kistler, K.; Li, A. W.; Moncla, L. H.; Paredes, M. I.; Kuhnert, D.; Torres, L. M.; Voitl, L.; Aksamentov, I.; Hodcroft, E. B.; Huddleston, J.; McCrone, J. T.; Anderson, J. S.; Sibley, T. R.; Lee, J.; Neher, R. A.; Bedford, T.Abstract
Motivation: Genome sequencing provides an exceptional window into the evolutionary and epidemiological dynamics of endemic and emerging pathogens, and thus allows for better, more targeted, public health interventions. Online genomic surveillance platforms can provide near real-time insight into these dynamics. Results: Nextstrain provides continually updated real-time genomic surveillance for 21 viruses and the bacterial pathogen Mycobacterium tuberculosis, with most analyses relying solely on open sequence data. Each pathogen includes steps to fetch and curate open data, classify sequences using established nomenclature systems, perform phylogenetic analyses, and share the results publicly. These analyses are automated, with most running daily to provide continually updated snapshots of pathogen evolution. Availability and Implementation: All source code is available at https://github.com/nextstrain. Phylogenetic results can be visualized and downloaded at https://nextstrain.org/pathogens, and open sequence data and curated metadata are available at https://nextstrain.org/pathogens/files.
bioinformatics2026-07-17v2SoftHybrid: A Hybrid Imputation Algorithm Optimised for Single-Cell Proteomics Data
Shi, Y.; Davis, S.; Charles, P. D.; Taylor, S.; Dombi, E.; Berridge, G.; Ebner, D.; Fischer, R.Abstract
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive Missing-Not-At-Random (MNAR) sparsity. Existing imputation methods typically target either Missing-At-Random (MAR) or MNAR mechanisms, resulting in a trade-off between replicate consistency and preservation of biological variation, and are largely designed for bulk data. Here, we introduce SoftHybrid, a data-driven imputation framework that jointly models missingness and protein abundance to estimate the probability of MNAR, enabling continuous weighting between MAR- and MNAR-oriented strategies. SoftHybrid requires no external priors (cell type labels, group annotations, predefined missingness assumptions, etc.), enabling fully unsupervised applications. Across ground truth benchmarks and real single-cell proteomics datasets, SoftHybrid outperforms existing methods at low input and matches or exceeds their performance at the mini-bulk level. By preserving proteomic structure and abundance accuracy, it enhances the recovery of biologically meaningful signals. SoftHybrid is implemented as an R package and is freely available on GitHub.
bioinformatics2026-07-17v2Retention, not flux: endpoint confounding caps computational prediction of peptide skin penetration, with a delivery-aware reframing
Komianos, N.; Prakash, P.Abstract
Bioactive peptides are now central to cosmetic and dermatological actives, yet predicting whether a given sequence will reach its site of action in skin remains unsolved. We contend that the dominant framing, predicting a single binary "skin permeability" label from sequence, is ill-posed, and that this, rather than a shortage of modelling power, explains the field's stalled predictive performance. The scope of the claim is narrow: barrier-crossing propensity is a legitimate, learnable function of molecular structure, whereas the vehicle- and endpoint-agnostic binary label that the literature supplies is not. We support this with a first-principles analysis and a study of public-source data. First, the experimental endpoint most commonly reported, transdermal flux into a diffusion-cell receptor compartment (OECD Test Guideline 428), conflates two opposite outcomes (genuine deep delivery and undesired systemic transport) and is, for a cosmetic active, frequently a failure signal rather than a success signal. That receptor flux is an imperfect measure of cutaneous bioavailability is long established in dermatopharmacokinetics; our contribution is to show that the same confound, inherited through scraped labels, is what caps machine learning from sequence. Second, reported "permeability" is a property of the sequence x delivery-vehicle x measurement-compartment triad, two terms of which are usually unrecorded. Third, on public-source data, a physicochemical intrinsic-permeability estimate (Potts-Guy) carries no positive predictive signal for scraped penetration labels (grouped AUC 0.45, 95% CI 0.40-0.51); sequence-only classifiers plateau in the mid-0.70s with diminishing returns as labels accumulate (AUC 0.70-0.77); and the same descriptor pipeline on a clean single-endpoint membrane dataset scores materially higher (AUC 0.83, non-overlapping CI). Our proposed reframing separates barrier-crossing (data-driven, sequence-level) from depth-and-retention (physics-driven, delivery-aware) and treats intrinsic transdermal flux as a regulatory risk axis; we close by proposing a triad-annotated reporting schema and a seed benchmark.
bioinformatics2026-07-17v2