Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation
Moran, J.; Freda, P. J.; Ghosh, A.; Walker, C. T.; Hernandez, M. E.; Moore, J. H.Abstract
Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.
bioinformatics2026-08-27v2Erosion of regenerative regulation: age-associated shifts in the skeletal muscle fiber epigenome and transcriptome
Moo, K. G.; Orchard, P.; Varshney, A.; D'Oliveira Albanus, R.; Manickam, N.; Kinnunen, L.; Lakka, T. A.; Saramies, J.; Laakso, M.; Tuomilehto, J.; Mohlke, K. L.; Boehnke, M.; Scott, L. J.; Koistinen, H. A.; Collins, F. S.; Parker, S. C.Abstract
Skeletal muscle aging is characterized by the deterioration of muscle function, which can lead to negative quality-of-life outcomes including frailty and sarcopenia. While understanding the mechanisms of this process is increasingly important as the global population ages, previous molecular studies of skeletal muscle aging have been limited by statistical power and cell type resolution. In this study, we analyzed single-nucleus gene expression and chromatin accessibility data from 287 human skeletal muscle samples from individuals aged 20-79 years to explore sex- and cell type- specific aging effects. Across 467,126 nuclei from 13 cell types, we identify 384 age-associated genes and 4,061 age-associated chromatin regions. These age-associated molecular features are enriched for functional pathways, including metabolic processes, cell-to-cell communication, and senescence Kyoto Encyclopedia of Genes and Genomes KEGG terms. Age-associated closing chromatin was more common across fiber types and sexes than opening chromatin, and was enriched in active enhancer regions while depleted for active transcription start sites. We observe enrichment for specific transcription factor motifs in closing chromatin, including those of glucocorticoid and androgen receptors, both of which play a key role in the maintenance of healthy skeletal muscle. Together, these findings identify an age-associated regulatory shift, largely invisible in matched transcriptomic data, characterized by closing chromatin which reduces accessibility to hormone receptor binding sites and enhancer regions in the muscle fiber epigenome.
bioinformatics2026-08-27v2HIDE-Deconv: A hierarchical deconvolution framework for multiscale characterization of cellular remodeling
Voelkl, D.; Bolz, S.; Rayford, A.; Sterr, T.; Mensching-Buhr, M.; Seifert, N.; Arp, J.; Tausche, J.; Engel, L.; Schuster, C.; Stevenson, T.; Zacharias, H. U.; Altenbuchinger, M.; Goertler, F.Abstract
Most deconvolution methods estimate cellular composition at a single level of cellular resolution despite biological processes often manifesting within fine-grained cellular subpopulations. We present HIDE-Deconv, a hierarchical deconvolution framework that jointly optimizes cellular compositions across multiple levels of a cell-type hierarchy while maintaining consistency between resolutions. In benchmark experiments, HIDE-Deconv achieved the highest overall predictive performance among evaluated methods. Analyses of lung adenocarcinoma, sepsis, COVID-19 and systemic lupus erythematosus revealed biologically relevant cellular remodeling that remained concealed at broader levels of cellular resolution. HIDE-Deconv is available as an open-source framework at https://github.com/dvoelkl/HIDE-deconv.
bioinformatics2026-08-27v2reactifpTM: an accessible reimplementation of actifpTM
Simpkin, A. J.; Johnson, E.; Rigden, D.Abstract
Motivation: The actual interface pTM score (actifpTM) is a modified version of the ipTM score that limits the calculation to only those residues at the interface. Whilst actifpTM provides an effective interface quality score, a limiting factor is that it makes use of the predicted aligned error (PAE) with probabilities, information that is generated during a ColabFold run, but not output by the package or other model prediction software. The consequent inability to generate actifpTM scores for the results of software such as AlphaFold 2 or AlphaFold 3 has limited its adoption. With reactifpTM we address this problem by providing a standalone tool that can be run on the standard outputs of most model prediction packages. Results: Using the same underlying principles as actifpTM, reactifpTM has been developed to use standard output files from model prediction software (a model and corresponding PAE) to perform an actifpTM-like calculation. ColabFold models were generated for a dataset of 1079 known interfaces in the PDB. A strong correlation was shown between actifpTM and reactifpTM for this dataset. Availability and implementation: reactifpTM is coded in Python. All scripts and associated documentation are available from https://github.com/hlasimpk/reactifptm or https://pypi.org/project/reactifptm.
bioinformatics2026-08-27v1PMPNN-DDG: an accurate machine learning-based {triangleup}{triangleup}G prediction pipeline trained on a novel interpretable feature set extracted from ProteinMPNN
Jani, R.; Ahmed, S.Abstract
An accurate and tractable approximation of the single-point mutation-induced change in protein thermodynamic stability, denoted by DDG, is critical for understanding the genotype-phenotype relationship. Several computational methods have been proposed for this problem; however, limited and error-prone training data and the difficult-to-predict magnitude of structural perturbations make this a challenging task. Consequently, the computational predictors proposed throughout the past decade incrementally improved prediction performance by proposing novel features, combining existing features, task-adapted neural network architectures, loss functions, data augmentation techniques, and pre-training procedures. In this work, we propose PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN. On the S669 independent test set, PMPNN-DDG achieves rF +R = 0.64 and RMSE = 1.45, outperforming all compared baseline methods across the reported evaluation measures. On the Ssym independent test set, it achieves rF +R = 0.81, rF -R = -0.99, and RMSE = 1.10, showing competitive performance relative to the compared baselines. PMPNN-DDG is publicly available at https://github.com/dRanger666/PMPNN-DDG.
bioinformatics2026-08-27v1DeepTMHMM2 enables accurate prediction of transmembrane protein topology and subcellular location
Teufel, F.; Hallgren, J.; Nielsen, H.; Krogh, A.; Tsirigos, K. D.; Winther, O.Abstract
Transmembrane -helical and {beta}-barrel proteins are a ubiquitous component of proteomes. Topology prediction infers how proteins are embedded in lipid bilayers, identifying membrane-spanning segments and their orientation. While recent methods achieve high performance for membrane-spanning segments, they cannot predict re-entrant regions and interfacial helices - membrane-associated segments that partially insert but do not cross the bilayer - nor identify which biological membrane a protein resides in. Here, we present DeepTMHMM2, the first predictor to include re-entrant regions and interfacial helices in its topologies and jointly predict localization across 17 biological membranes. Benchmark results show that DeepTMHMM2 successfully learns to predict the additional elements, while achieving strong performance on canonical -helical and {beta}-barrel topology prediction. Applying DeepTMHMM2 to Swiss-Prot reveals that non-crossing segments are a ubiquitous feature of the transmembrane proteome, with interfacial helices present in nearly a quarter of all -helical transmembrane proteins.
bioinformatics2026-08-27v1Proteome modulation by opposite inotropic drugs in human engineered cardiac tissue revealed by topology-driven cross-modal integration
Staykova, D. K.; Snippert, D.; Wessels, H. J. C. T.; Passier, R.; Conte, F.Abstract
Engineered heart tissues (EHTs) represent an innovative platform enabling physiologically relevant in vitro evaluation of drug-induced cardiac responses. While functional characterization remains central to EHTs, molecular profiling is increasingly used to elucidate mechanisms underlying drug-induced phenotypes. Proteomics provides broad molecular characterization of drug responses at the protein level, yet the complexity, heterogeneity, and high dimensionality of proteomics datasets challenge conventional statistical approaches, which are not designed for cross-modal integration and streamlined multi-omics analysis. In this study, we developed an innovative framework based on topological data analysis (TDA) for the integration of large proteomics profiles and functional readouts to investigate system-level responses to drugs with opposing inotropic effects, epinephrine and doxorubicin. Samples were organized into a topological connectivity network according to multimodal similarity enabling simultaneous exploration of treatments, cardiac function and proteome alterations. Highly correlated features were then used for pathway enrichment analysis, which revealed strong similarities between the enrichment profiles associated with contractile force and epinephrine. These findings are consistent with the positive inotropic effect of epinephrine, whereas doxorubicin exhibited an opposing enrichment profile. Energy homeostasis, mitochondrial translation and proteostasis emerged as the major cellular processes displaying opposite associations with the two inotropic drugs, highlighting a link between cardiac contractility and perturbations in these processes. In conclusion, our TDA-based framework successfully integrated functional and proteomic data to uncover treatment-specific remodeling in EHTs, offering a modular and scalable approach that could be adapted to other in vitro organ models for systems-level mechanistic studies and next-generation drug development.
bioinformatics2026-08-27v1DeMoP: A Language-Model-Guided Mixture-of-Experts Framework for Cancer Prognosis
Tang, C.; Yu, L.; Li, Q.; Xu, L.Abstract
Integrating heterogeneous clinical and molecular data for cancer prognosis remains challenging because their dimensionality, semantics and distributions differ across patients and cohorts. Here we present DeMoP, a language-model-guided mixture-of-experts framework that serializes structured patient profiles as natural-language sequences and learns adaptive prognostic representations from clinical variables, copy-number alterations, and gene descriptions. DeMoP combines a fine-tuned DeBERTa-v3-large encoder, attention-based token pooling, and a residual mixture-of-experts prediction head. In held-out tests from two independent pan-cancer cohorts, GENIE (63,090 patients) and TCGA (4,123 patients), DeMoP outperformed the conventional machine-learning and deep-learning baselines evaluated, achieving AUROCs of 0.939 and 0.805 and class-1 F1 scores of 0.72 in both cohorts. A GENIE-trained model transferred directly to TCGA with an overall class-1 F1 score of 0.62. Gene-level ablations recovered established cancer-associated genes and highlighted less-studied candidates. DeMoP provides a unified approach to heterogeneous biomedical data integration, cross-cohort outcome prediction, and model interpretation.
bioinformatics2026-08-27v1NeuroMesh: A Bottleneck Topology Controller for Missing-Modality Brain Tumor Segmentation - A Mechanistic Pilot Study on BraTS
Kamalakannan, N. K.; Kamalakannan, J.Abstract
Deep segmentation networks can degrade sharply when an expected MRI sequence is unavailable at inference. We present NeuroMesh, a bottleneck controller that combines a gated recurrent unit (GRU) with a graphconvolutional edge-activation mask, designed to adapt a U-Net-style segmentation backbone to missing input. We evaluate NeuroMesh in a pilot study using a 30-patient subset of the BraTS 2020 benchmark (22 training, 4 validation, and 4 held-out test patients) under a prespecified frozentest protocol. On the frozen test set, NeuroMesh has higher tumor-core and enhancing-tumor Dice than a plain U-Net in most evaluated missing-modality conditions, but wholetumor Dice falls from 0.596 to 0.108 when FLAIR is missing, compared with 0.604 to 0.545 for the plain U-Net. Direct analysis of the predicted edge-activation mask shows negligible change across modality-availability conditions. A parameter-light static-gating control reproduces the FLAIR failure mode without recurrence, a failure-signal input, or graph-structured machinery. These results do not support the intended interpretation that the trained controller performs input-conditional topology rewiring at the scale of this pilot. Instead, they expose a discrepancy between architectural intent and realized behavior and identify a specific missing-modality failure mode that warrants further investigation. Given the small validation and test sets, the findings are descriptive and do not establish clinical or population-level generalization.
bioinformatics2026-08-27v1SPC-Clean: A napari Plugin for Reducing Speckle and Isolated Pixel Noise in Fluorescence Microscopy Images
Alirezazadeh, P.; Kirsch, E. M.; Tian, Y.; Bewersdorf, J.; Rittscher, J.; Mergenthaler, P.Abstract
Speckle artifacts and isolated foreground pixels are common in fluorescence microscopy and can interfere with segmentation and subsequent quantitative image analysis. Conventional denoising methods often modify image intensities through filtering or smoothing, potentially altering biologically relevant fluorescence signals. We introduce Sparse Pixel Cluster Cleaning (SPC-Clean), a topology-aware method that removes poorly supported foreground pixels through iterative neighborhood analysis of a thresholded mask. SPC-Clean is deterministic, training-free, preserves original fluorescence intensities for practical microscopy workflows.
bioinformatics2026-08-27v1Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients
Ramesh, P.; Fyta, M.Abstract
Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.
bioinformatics2026-08-27v1Identification of novel HDAC11 inhibitors: In silico & in vitro studies
Paul, M.; Kumar, D. S.; Mishra, S.; Kalle, A. M.Abstract
Histone deacetylases (HDACs) are pivotal epigenetic regulators that modulate diverse cellular pathways by removing acetyl groups from lysine residues on both histone and non-histone proteins. Histone deacetylase 11 (HDAC11), the sole member of class IV HDACs, exhibits both deacetylation and fatty acid deacylation activities. Accumulating evidence implicates HDAC11 as a key epigenetic regulator of fundamental cellular processes, including metabolism, immune responses, and tissue development. Dysregulation of HDAC11 activity has been associated with inflammatory diseases, metabolic disorders, neurodegenerative conditions, and cancer, highlighting its potential as a therapeutic target. Although several HDAC11-specific inhibitors have been identified, none have progressed to clinical development. In this study, we aimed to discover HDAC11-selective inhibitors by integrating in silico and in vitro validation approaches. Homology modelling of the HDAC11 structure was conducted, followed by model validation, structure-based virtual screening, molecular dynamics (MD) simulations, and binding free energy calculations. We identified and validated three lead compounds and their intermediates using biochemical and cell-based assays. Fluorescence-based and HPLC-based enzymatic assays demonstrated potent inhibition of both the deacetylase and deacylase activities of HDAC11, with Inhibitor 6 and Inhibitor 3 exhibiting the strongest effects among the six compounds tested. Further, a decrease in lipid accumulation, reduced stability of the HDAC11 substrate SHMT2, as determined by immunoblot analysis and decreased cell viability, as assessed by MTT assay, confirmed HDAC11 inhibition in cellular models. The study shows that new HDAC11 inhibitors significantly reduce the viability of breast cancer cells and induce apoptosis; inhibitor 6, in particular, showed high potency, similar to the reference compound SIS-17. Flow cytometry showed that treated MDA-MB-231 cells exhibited cell-cycle arrest and increased apoptosis, a finding further confirmed by Annexin V/PI staining. Molecular analysis showed that BAX increased while BCL2 decreased, indicating that apoptotic pathways were activated in novel compound-treated MDA-MB-231 cells. The results suggest that inhibiting HDAC11 is an effective way to induce cancer cell death and provide a basis for further assessment of these compounds as potential treatments for breast cancer. Collectively, this study identifies novel zinc-chelating HDAC11 inhibitors containing a nitro-sp2 group, providing promising candidates for further therapeutic development.
bioinformatics2026-08-27v1AntiSite: Modality Dropout Enables Antibody Paratope Prediction With or Without Structure From a Single Model
Papadopoulos, A. M.; Alvarez, F.; Daras, P.Abstract
Summary: Reliable paratope identification is central to understanding antibody antigen recognition and advancing therapeutic antibody discovery. AntiSite is a unified antibody paratope prediction framework that combines protein language-model sequence embeddings with structure-derived molecular-surface features and, through modality dropout, trains a single checkpoint to predict both with and without a structure. This lets one model support sequence-only inference when no structure is available and structure-aware inference when an antibody structure is provided. Availability and implementation: Source code, trained models and evaluation scripts are freely available at https://github.com/aggelos-michael-papadopoulos/AntiSite. Processed benchmark structures and corrected split metadata are archived on Zenodo at https://doi.org/10.5281/zenodo.21705412.
bioinformatics2026-08-27v1EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-08-26v4CRISPR-HAWK: Haplotype- and Variant-aware Guide Design Toolkit for CRISPR-Cas
Kumbara, A.; Tognon, M.; Carone, G.; Fontanesi, A.; Bombieri, N.; Giugno, R.; Pinello, L.Abstract
Current CRISPR guide RNA design tools rely on reference genomes, overlooking how genetic variation impacts editing outcomes. As genome editing advances toward clinical applications, incorporating population diversity becomes essential for ensuring therapeutic efficacy across diverse populations. We present CRISPR-HAWK, a framework integrating individual- and population-scale variants and haplotypes into gRNA design. Analyzing therapeutic targets across 79,648 genomes reveals that genetic variants substantially alter guide performance. For the clinically approved sickle cell disease therapeutic guide targeting BCL11A, we identify haplotypes that completely abolish predicted cutting activity. Across seven therapeutic loci, 82.5% of guides contain variants modifying on-target activity. Variants also create novel protospacer adjacent motif sites generating individual-specific guides invisible to reference-based design. These findings demonstrate that variant-aware selection is critical for equitable genome editing. CRISPR-HAWK is available at https://github.com/pinellolab/CRISPR-HAWK and https://github.com/InfOmics/CRISPR-HAWK
bioinformatics2026-08-26v3Enrichment-free glycoproteomics harnessing real-time mass defect-driven glycopeptide classification reveals sex differences in murine fucosylation
Zhang, B.; Chau, T. H.; Bienes, K. M.; Arakawa, H.; Hane, M.; Sato, C.; Yokoi, A.; Kaji, H.; Ashwood, C.; Matsui, Y.; Kawahara, R.; Thaysen-Andersen, M.Abstract
Glycopeptide enrichment remains a cornerstone in glycoproteomics, but bias and reproducibility issues continue to hinder biological insight and clinical translation. Employing curated glycoproteomics datasets and machine learning, we trained a glycopeptide classifier to recognize N-glycopeptide precursors through mass defect signatures. Integration of the classifier into a data-dependent acquisition framework facilitated real-time prediction of N-glycopeptides from human serum and revealed sex differences in murine plasma fucosylation opening avenues for enrichment-free glycoproteomics.
bioinformatics2026-08-26v3UMITIC: An unsupervised framework for the joint characterization of cellular phenotypes and spatial neighborhoods in multiplex and hyperplex immunofluorescence imaging data
Sangüesa Recalde, M.; De Andrea, C. E.; Ariz, M.Abstract
Multiplexed imaging technologies enable the simultaneous measurement of dozens of protein markers while preserving context, providing a high-resolution view of tissue organization schemes. However, extracting meaningful insights from these high-dimensional datasets--particularly in hyperplex settings (>20 markers)--remains a major computational challenge, especially in the absence of annotated data. Here, we present UMITIC (Unsupervised Analysis of Multiplex Images via TIssue Characterization), a modular and unsupervised computational framework for the joint characterization of cell phenotypes and tissue neighborhoods from multiplex imaging data. UMITIC integrates three components: (i) CellCut, a strategy that combines nuclear and cytoplasmic predictions to improve the delineation capabilities of the framework; (ii) CellMap, a contrastive learning approach that generates low-dimensional representations of single-cell image crops that are enriched with morphological features; and (iii) TissueNet, a graph neural network that models spatial cell-cell interactions to identify tissue neighborhoods. We evaluated UMITIC across four datasets of increasing complexity to assess its robustness, scalability and biological relevance. With respect to a 7-plex human tonsil dataset, the framework identified canonical immune cell populations and reconstructed well-established anatomical regions. When applied to a 43-plex tonsil image, UMITIC preserved these tissue-level structures while enabling a finer cell subtype stratification process driven by increased marker dimensionality. We further validated our method on a 58-plex colorectal cancer cohort, where UMITIC was able to recover previously reported immune composition differences and spatial organization variations between patient groups with different prognoses. Finally, when an expert-annotated mass cytometry imaging dataset concerning human lung tissue was used, UMITIC achieved higher agreement with the reference tissue annotations than the existing approaches did, demonstrating improved lung microanatomy reconstruction accuracy. Together, these results show that UMITIC enables consistent and interpretable analyses of both cellular phenotypes and tissue architectures across diverse multiplex and hyperplex imaging datasets without the need for manual annotations.
bioinformatics2026-08-26v3Identification of Altered Potassium Channels for Drug Repurposing in Long COVID Patients
George, J. P.; Gaikwad, K. B.; Sharma, J.Abstract
Long COVID (LC) is a complex condition characterized by persistent, chronic multisystem manifestations, with a significant proportion of patients exhibiting neurological symptoms. Human ion channels (HICs), particularly potassium channels, are abundantly expressed in the nervous system and linked to key metabolic processes, making them potential candidates for understanding LC pathophysiology and drug repurposing. Meta-analysis of RNA-Seq datasets from COVID-19-recovered and LC patients was performed to identify altered HICs in LC. Differential gene expression analysis, functional enrichment analysis, and weighted gene co-expression network analysis were performed to uncover key genes, pathways, and co-expression modules consisting of HICs, lipid metabolism-, and immune signaling-related genes. A total of 715 dysregulated genes, including eighteen HICs were identified, among which seven were potassium channels. Three significant modules containing HICs, lipid metabolism-, and immune signaling-related genes were identified and found to be associated with antigen processing and presentation, complement and coagulation cascades, and cytokine-related pathways. Additionally, drug-gene interaction analysis led to identification of approved drugs targeting KCNA6, KCNJ10, KCNN3, and KCNH4 that might provide opportunities for drug repurposing in neurological manifestations in LC. Further experimental validation is required to establish their efficacy and assess their potential for translation into clinical applications for patients with LC.
bioinformatics2026-08-26v2MONTE enables unified pan-cancer tumor purity estimation andmethylation correction from bulk DNA methylation arrays
Kim, M.; Lee, W.-H.; Yao, V.Abstract
Bulk DNA methylation profiling is widely used to study cancer epigenomics in clinical settings, but these measurements aggregate signals from malignant and non-malignant cells, introducing composition-dependent confounding that complicates tumor-intrinsic interpretation and cross-cohort analyses. While existing methods can estimate tumor purity and, in some cases, correct methylation measurements, they typically require cancer-specific reference models, matched normal samples, or predefined probe sets, limiting their applicability to rare cancers, different clinical cohorts, and cross-dataset comparisons. We present MONTE (Methylation-based Observation Normalization and Tumor purity Estimation), a unified, cancer label-free framework for tumor purity inference and CpG-resolved methylation correction from bulk DNA methylation data. MONTE learns probe-wise relationships between methylation and tumor purity using an empirical Bayes-moderated linear model and infers purity in new samples via signal-to-noise weighted aggregation, without requiring matched normals, cancer labels, or predefined probe sets. A single pan-cancer MONTE model outperforms existing cancer-specific methods for purity estimation across 21 cancer types, generalizes across purity references, and runs orders of magnitude faster on full-dataset analyses. MONTE also introduces Bayesian transfer learning, which enables efficient recalibration to alternative purity definitions, validated on three independent external cohorts. Methylation correction with MONTE further amplifies tumor-relevant regulatory signal and improves the reproducibility of differential methylation analyses. By unifying purity estimation and correction in a single flexible, scalable, and interpretable framework, MONTE broadens the accessibility of tumor-intrinsic methylation analysis across cancer types and datasets.
bioinformatics2026-08-26v2KSTITCH links cellular morphology and gene expression in spatial transcriptomics
Kumar, S.; Shi, Y.; Vallius, T.; Day, C.-P.; Absil, P.- A.; Srivastava, A.; Hannenhalli, S.; Gopalan, V.Abstract
In situ spatial (ISS) sequencing can uncover co-variation between cellular morphology and gene expression in vivo. However, a principled and interpretable mathematical representation of morphology has not yet been applied in this context. In particular, current deep learning-based representations of cell images confound a cell's shape with its size. We present an interpretable representation of cellular boundary contours, based on tangent principal component analysis (TPCA) in a Kendall shape manifold, that captures size-independent contour shape features. This approach successfully recovers shape-perturbing genes in an RNAi screen than a previous metric geometry-based approach. We build on TPCA to develop KSTITCH (Kendall Shape-TranscriptomIc Correlation and Harmonization), an approach to reveal covariation between cell morphology with gene expression in ISS datasets. In a Xenium dataset, KSTITCH recovers known morphology-transcriptomic relationships in keratinocytes, macrophages and endothelial cells. Across samples in a melanoma CosMx dataset, KSTITCH reproducibly associates elongated and triangular fibroblasts with proximity to malignant cells and myofibroblast-like transcriptional program. Finally, KSTITCH independently recovers a known link between mesenchymal-like malignant cell states and increased cell area in two melanoma cohorts. KSTITCH can thus yield interpretable morphology-transcriptome relationships across cell types, patients, and spatial transcriptomics platforms. KSTITCH is available at https://github.com/vishakagopalan/kstitch .
bioinformatics2026-08-26v2A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models
Li, L.; Duan, S.; Zha, X.; Ye, F.; Zhang, Y.; Zhang, X.; Cao, Y.; Liu, C.; Fang, B.Abstract
Single-cell drug perturbation models are increasingly used to predict how compounds remodel cellular states, but they are still largely assessed by expression reconstruction. Whether high expression similarity reflects preservation of drug-response signatures remains unclear. Here we present scDrugPerturb-Bench, a mechanism-annotated benchmark that links matched control and drug-treated single-cell RNA-sequencing profiles to literature-curated directional key-gene evidence. The resource covers 181 datasets, 423 annotated response cases, 717 unique key genes and 2.5 million cells. We introduce the Mechanism Fidelity Score (MFS) to evaluate key-gene direction, effect-size recovery, gene-set coherence, mechanism specificity and pathway-level response polarity. Across 12 perturbation-prediction models, 3 baselines and 10 data splits, expression-similarity metrics were weakly aligned with MFS and selected different model configurations. Mechanism-aware selection improved early drug retrieval in a transcriptome-based drug design evaluation, indicating that MFS provides practical information beyond benchmark reporting. Systematic benchmarking revealed limited fidelity to drug-response signatures across cell-line and source-integrated settings. Frozen single-cell foundation model embeddings produced local, metric-dependent gains rather than universal improvements, and source context substantially reshaped model assessment. Hard-negative tests further showed that plausible perturbation responses can arise from non-specific transcriptional shortcuts. These results show that expression reconstruction is an insufficient proxy for preserving drug-response signatures and establish scDrugPerturb-Bench as a benchmark for mechanism-aware evaluation of single-cell drug perturbation models.
bioinformatics2026-08-26v2scDisent: regulatory-aware disentangled representation learning for multi-omic single-cell analysis
Xi, G.Abstract
Single-cell multi-omic technologies measure complementary aspects of cellular identity and regulatory state, yet most integration models compress these signals into one entangled latent space. Such representations are useful for clustering but poorly suited to regulator-centered interpretation or perturbation-oriented analysis. We present scDisent (https://github.com/xiguoren/scDisent), a generative framework that separates expression-associated variables (zexpr) from regulation-associated variables (zreg) and links them through a sparse directed mapping. scDisent combines modality-specific encoding, variational disentanglement, total-correlation and orthogonality regularization, and a Gumbel-gated causal module protected by detach-based gradient isolation. Across benchmark datasets with matched modalities, scDisent achieved the strongest clustering performance among the tested methods while exposing regulatory structure that competing integration models do not represent explicitly. The learned causal atlas remained sparse, perturbation analyses recovered biologically coherent lineage-associated programs, and branch-separation analyses showed that benchmark-label information concentrated in zexpr rather than zreg. These results position scDisent as a multi-omic representation model that improves both integration quality and biological interpretability
bioinformatics2026-08-26v2On the illusion of scRNA-seq batch effect correction
Codice', F.; Fariselli, P.; Raimondi, D.Abstract
Batch correction methods in single-cell RNA sequencing are essential for removing technical variation that can otherwise lead to misleading downstream analyses. The reliability of these methods is typically evaluated using unsupervised metrics. Here, we apply a Machine Learning (ML) technique called probing, formalized as the Batch Probing Score (BPS), to empirically demonstrate across six datasets that the most popular batch correctors fail to fully remove batch signal. In most cases, the batch of origin remains clearly identifiable after correction, even though standard evaluation metrics cannot detect it. We show that existing unsupervised metrics lack the sensitivity and specificity required to capture residual batch signal, whereas ML-based approaches can still detect it. This residual signal can similarly be picked up by downstream analysis tools, potentially leading to biased results. Because BPS is supervised, it directly quantifies batch signal strength by measuring how accurately the batch of origin can be predicted for each sample. It therefore provides an upper bound on the residual ML-actionable batch signal that could otherwise remain unnoticed. Our findings suggest that probing-based metrics should become a standard for assessing batch correction methods in single-cell RNA-seq and other areas of genomics.
bioinformatics2026-08-26v2Evaluating Aggregated Gene Level eQTL Scores
Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.Abstract
Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.
bioinformatics2026-08-26v1Robustness to nuisance perturbations enables unsupervised evaluation of single-cell foundation models
Sallam, A.; Gillis, J.Abstract
Single-cell foundation model evaluations have relied almost exclusively on downstream tasks. While these tasks measure whether an embedding recovers annotated cell types, batches, or trajectories, they cannot determine if that structure is reproducible or merely an artifact of a single noisy draw, a key limitation since incomplete sampling is intrinsic to single-cell measurement. Here, we introduce a fully unsupervised evaluation framework grounded in a fundamental principle: a faithful representation must preserve its neighbourhood structure under nuisance perturbations that mimic technical and sampling variation. Across five scFMs, a PCA baseline, and 39 datasets, we show that models ranked as near-equivalent by standard benchmarks differ nearly twofold in local neighbourhood preservation under a perturbation discarding just 5% of counts. This structural instability is scale-dependent and often masked by visually coherent embeddings. Cluster-level stability under resampling tracks established bio-conservation metrics (Spearman {rho}=0.78), showing that invariance to nuisance perturbations captures representation quality no benchmark measures directly.
bioinformatics2026-08-26v1Automated Detection of Livestock Gastrointestinal Parasite Eggs and Cysts Using YOLOv8-Based Deep Learning
Sarwer, A.Abstract
Parasitic infection is one of the common health problems of livestock in Bangladesh. Due to the country's climate, heavy monsoon rainfall, low biosecurity in farms, and high humidity, along with presence of suitable vector organisms, gastrointestinal parasitism remains widespread in cattle and other livestock. The standard method of diagnosis is microscopic examination of fecal samples, but this depends on manual observation, which is time-consuming and can lead to human error, mainly because many parasite eggs look similar to each other and samples often contain contaminants that can be mistaken for eggs or cysts. In this study, we tried to apply the YOLOv8 deep learning model for automated detection of parasitic eggs and cysts from microscopic images of livestock fecal samples. Images of clinical cases were collected, annotated, and used to train the model in Python, with batch size 16, auto optimizer, learning rate 0.01, momentum 0.937 and weight decay 0.0005. Training was done using Google Colab, and the model was evaluated using precision, recall, F1-score, mAP50, and mAP50-95. The model achieved a precision of 56%, recall of 24%, F1-score of 33.6%, mAP50 of 33%, and mAP50-95 of 22%. The relatively low recall and F1-score indicate that the model still has considerable limitations, largely due to insufficient species-specific training data and presence of image artifacts. Underrepresentation of some parasite species, such as Trichuris spp., in the dataset also caused class imbalance, which affected the model's ability to detect these species reliably. Despite these limitations, the study indicates that YOLOv8 architecture has some potential to be used for detection of parasitic eggs and cysts from microscopic images, and that further work with larger and more balanced datasets may improve performance and applicability in veterinary diagnostics. Keywords: YOLOv8, livestock parasites, deep learning, microscopic image analysis, veterinary diagnostics, Bangladesh
bioinformatics2026-08-26v1Interpretable Forecasting of Kidney Cancer Progression via Generative AI and Symbolic Reasoning
Prol-Castelo, G.; Syrri, E.; Manginas, N.; Manginas, V.; Sanchez-Valle, J.; Katzouris, N.; Paliouras, G.; Valencia, A.; Cirillo, D.Abstract
Predicting cancer stage progression from omics data, and deriving molecular insight into the mechanisms driving it, remains a major challenge, owing in part to the lack of adequate longitudinal data and the interpretability limitations of current forecasting models. Large cancer datasets such as TCGA capture patient profiles cross-sectionally rather than longitudinally, complicating timely treatment decisions as tumors become more invasive. Deep neural networks typically used for forecasting, such as LSTMs, compound this problem by remaining largely opaque and offering clinicians no straightforward way to audit their predictions. Clear cell renal cell carcinoma (ccRCC) illustrates the clinical stakes of both challenges. Five-year survival falls from over 94% at stage I to 28% at stage IV, yet early-stage tumors are often managed under active surveillance, a strategy constrained by sparse molecular evidence of progression risk. Detecting progression in time, meanwhile, demands forecasts clinicians can interpret and trust, not black-box predictions. We address both challenges by combining generative and symbolic AI: a Variational Autoencoder trained on bulk RNA-Seq profiles of 530 TCGA ccRCC patients generates synthetic pseudo-time trajectories that overcome the absence of longitudinal data, while a symbolic rule-induction framework (ASAL) learns finite-state automata from these trajectories, encoding stage transition as human-readable Boolean conditions over gene expression, which a complex event forecasting system (Wayeb) converts into probabilistic forecasts of stage advancement. An independent XGBoost classifier trained on real patients (F1 score = 0.71-0.81) shows a gradual early-to-late probability shift along the synthetic trajectories, absent in non-progressing control trajectories. Pathway enrichment of those trajectories reveals stage-dependent changes in established kidney cancer-related processes, including the TCA cycle and DNA repair. Finally, our symbolic forecaster nearly matches an LSTM baseline (macro F1 = 0.928 vs. 0.964), while additionally offering an inspectable rule set and a probability distribution over transition timing rather than a single opaque score. This work shows that generative and symbolic AI, paired together, can turn cross-sectional cohorts into a transparent, forecast-oriented framework for modeling disease progression, demonstrated here in ccRCC.
bioinformatics2026-08-26v1BatchRefiner: fast, significant improvement in batch integration of single-cell embeddings with ensemble refinement
Schäffer, D. E.; Kang, H.; Aksu, E. D.; Edelman, D.; Berger, B.Abstract
Data from single-cell RNA sequencing (scRNA-seq) and the Assay for Transposase-Accessible Chromatin (scATAC-seq) are high-dimensional, sparse, and undesirably capture technical variability between experiments or batches. Many analysis methods thus seek to produce a low-dimensional cell-by-feature embedding space that groups together biologically similar cells across batches while distancing dissimilar cells. Here, we introduce ensemble refinement for scRNA-seq and scATAC-seq embeddings, inspired by ensemble methods from statistical machine learning, and implement BatchRefiner, a fast post-processing tool to enhance batch integration. We extensively benchmark widely-used scRNA-seq embedding methods on both batch integration and biological conservation over a wide range of datasets, before and after the addition of BatchRefiner. We extend these benchmarking approaches to provide the first comprehensive benchmark of batch integration for scATAC-seq embedding methods, including BatchRefiner. Importantly, we formalize a significance statistic, which we use to demonstrate BatchRefiner's significant improvement in batch integration across a wide range of embedding methods, atlas-scale datasets, and established metrics.
bioinformatics2026-08-26v1Orthology transfer maps only the conserved core of the Varroa destructor proteome and over-calls host absence two times in three
Ryba, S.Abstract
The ectoparasitic mite Varroa destructor is the principal threat to managed honey bees, and a test case for the genome-scale methods applied to non-model organisms, nearly all of which infer from orthology. We reconstructed the first genome-wide protein-interaction network for V. destructor (7,080 proteins, 335,914 interactions), whose modular structure exceeds a degree-preserving null by 368 standard deviations, but whose every edge is interolog-transferred and every node conserved at least to Eukaryota. None of the 791 genes lacking an orthologous group enters it - arithmetic rather than discovery - yet the excluded compartment is large and coherent. It comprises 3,161 genes (30.9% of the proteome), shorter and less annotated than the rest; an annotation-free genome search detects orphans in a tick genome at 4.0% against 70.4% for networked genes. Within the orthology-bearing compartment visibility is non monotonic: the Acari-level bin (74.1%) falls below the Arthropoda-level bin (89.9%). The same logic applied to host comparison yields a benchmarked error: of genes called absent from Apis on group identity alone, 67.4% recover a sequence homologue - against zero for a shuffled null and 1.3% in the presence direction - rising to 78.3% in the least panel-biased stratum. Both figures are properties of the calling rule: under an identity floor the error directions cross near 34% identity; orthology cannot be said to err in either direction without fixing the criterion first. Host divergence resolves into gene absence and residue level substitution, falling in those two compartments respectively. A bee-sparing target map follows as broader impact.
bioinformatics2026-08-26v1UTR-Diffusion: Conditional Diffusion Modeling for Multi-objective and Constrained UTR Design
Dai, C.; Sato, K.Abstract
Motivation: The 5-prime untranslated region (UTR) and the start-codon-proximal region of the coding sequence (CDS) jointly influence translation efficiency and local RNA secondary-structure stability, while synonymous codon choices throughout the CDS shape codon adaptation. Because the encoded protein is often predetermined, practical mRNA design must coordinate these quantitative objectives while preserving specified nucleotide sequences and amino-acid identities. Existing generative approaches typically address continuous-valued targeting, explicit sequence constraints, and codon-usage control separately rather than integrating all three within a single model. Results: We present UTR-Diffusion, a diffusion-based framework for 5-prime UTR and 5-prime UTR-CDS junction design. UTR-Diffusion conditions generation on continuous-valued MRL and MFE targets and supports nucleotide-level constraints, amino-acid-level constraints with synonymous-codon flexibility, and codon-adaptiveness control that modulates the sequence-level codon adaptation index (CAI). Systematic evaluations across dense MRL-MFE target grids showed that generated distributions shifted consistently with both targets, retained substantial diversity, and strictly preserved specified nucleotide sequences and amino-acid identities. Codon-adaptiveness control yielded distinct, monotonically ordered CAI levels that closely followed the specified adaptiveness targets. In comparative benchmarks, UTR-Diffusion outperformed representative existing methods in high-MRL optimization and precise MRL targeting for 5-prime UTR design, and achieved higher MRL, less-negative junction MFE, and higher CAI than peptide-preserving baselines in 50-nt 5-prime UTR-CDS junction design.
bioinformatics2026-08-26v1AI-driven framework modeling perturbation in brain organoids reveals candidate genes for autism
Koh, I. G.; Chang, E.; Choi, Y. S.; Kim, S.-W.; Kim, Y.; Lee, H.; Byeon, G.; Ryu, Y.; Kim, S.; Lee, J.; Park, H.; Sim, H.; Ryu, Y.; Shim, W.; Lee, J.; Salazar, N. B.; de Aquino, M. M.; Engchuan, W.; Zhou, X.; Son, J. H.; Lee, J.; Bong, G.; Kim, I. B.; Han, J. H.; Werling, D. M.; Kim, S. H.; Oh, M.; Kim, M.-S.; Lee, D.; Kim, J.; Lee, Y.-S.; Sun, W.; Kim, E.; Scherer, S. W.; Jeon, M.; Yoo, H. J.; An, J.-Y.Abstract
Autism gene discovery is constrained by the rarity and heterogeneity of damaging variants, requiring large cohorts to identify susceptibility genes. Neural organoids and single-cell foundation models enable perturbation modeling in neurodevelopmental contexts. Here, we show that perturbation-informed foundation modeling of neural organoids can provide functional context for prioritizing candidate genes with genomic and clinical support. We constructed a 3.6-million-cell organoid atlas and trained models to predict genome-wide perturbation responses. Benchmarking 17 models identified a telencephalic neuron-specific model best preserving autism-relevant perturbation structure. Genome-wide profiling revealed two clusters associated with mid-fetal synaptic neuronal processes and early radial glia ubiquitin signaling. These clusters were supported by damaging-variant enrichment and clinical phenotypes across 89,916 family-based samples. Logistic-regression prioritization identified 343 candidates, including 167 in the key clusters, with convergence across TADA signals and recurrent evidence for NBEA and KLHDC10. This framework integrates predicted perturbation effects with genomic evidence to support autism candidate prioritization.
bioinformatics2026-08-26v1Genomic-Based Prediction of Exopolysaccharide Composition and Structure: Insights from Rhizobium and Sinorhizobium Species
Tulumello, J.; Long, J.; Achouak, W.; Garron, M.-L.; Terrapon, N.; Heulin, T.Abstract
Bacterial exopolysaccharides (EPS) are key components in biofilm formation, stress protection, and symbiosis in Rhizobiaceae. While EPS structural diversity is extensive, experimental characterization remains limited. In this study, we experimentally determined and compared four distinct EPS structures produced by ten Rhizobium alamii strains. Using genomic data, we bioinformatically identified supra-operonic clusters (SOCs) responsible for these EPS biosynthesis. We introduced a computational framework to predict, score, and compare EPS SOCs across 84 Rhizobium and Sinorhizobium species, linking gene content to structural and functional EPS diversity. A total of 743 EPS SOCs was selected for network analyses, allowing the identification of 36 major groups of orthologous EPS SOCs, successfully recovering all known EPS biosynthetic loci and two novels SOCs potentially encoding uncharacterized EPS (xEPS-I, xEPS-II). Profiles of EPS SOCs correlated with taxonomical groups, with a single EPS SOC conserved through all 84 genomes and distinct additional EPS SOCs depending on the group, but do not strictly explain symbiotic capacity. Genetic comparisons of transporters (Wzx, Wzy) and glycosyltransferase sequences indicated these proteins as key markers of EPS structure. Overall, this computational framework accurately identified and classified EPS SOCs, providing a scalable, genome-based method for predicting EPS biosynthetic potential in Rhizobiaceae and usable in other microbial genera.
bioinformatics2026-08-26v1CytoGate-Bench: an LLM benchmark for cross-panel cell gating in cytometry
Kim, J.; Lee, B.; Ahn, N.; Ionita, M.; McKeague, M. L.; Lee, M. E.; Jeong, C.-U.; Apostolidis, S. A.; Baxter, A. E.; Shwetank, ; Greenplate, A. R.; Wherry, E. J.; Sohn, K.-A.; Kim, D.Abstract
In cytometry, the workhorse single-cell technology of clinical immunology, every study defines its own antibody panel and cell-type vocabulary, so a classifier trained on one cannot annotate the next. Immunologists instead annotate by manual gating, splitting one parent population at a time on a two-marker plot, down an expert-defined hierarchy. We introduce CytoGate-Bench, a benchmark that reformulates this per-step procedure as a zero-shot, panel-agnostic task for large language models. It comprises 23,646 expert-annotated instances re-curated from 11 public flow- and mass-cytometry cohorts spanning eight marker panels. Across six open- and closed-weight backbones, the strongest formulation draws one rectangular gate per candidate and falls within the range of trained, panel-specialized baselines. It degrades less under distribution shift. Walking the hierarchy stepwise outperforms predicting every cell type at once. Ablations trace the signal to the data distribution shape and curated marker priors. However, adding vision or a self-verification loop systematically tightens gates.
bioinformatics2026-08-26v1Context-dependent regulatory networks connect Alzheimer's disease genetics to microglial inflammatory responses
Fu, T.-T.; Kurkela, M.; Tu, J.; Zhang, J.; Sun, N.; Farrer, L. A.; TCW, J.; Hou, L.Abstract
Inflammation is central to Alzheimer's disease (AD) pathogenesis. Microglia, the resident innate immune cells of the brain, exhibit diverse inflammatory states and are enriched for AD-associated genetic variants within active cis-regulatory elements (CREs). However, the interplay among genetic variants, transcription factor (TF)-CRE-gene programs, and microglial responses across inflammatory and disease contexts remain poorly understood. Here, we develop context-dependent epigenomic networks (cEpiNets), integrating bulk and single-nucleus assay for transposase-accessible chromatin using sequencing (ATAC-seq) to reconstruct regulatory programs across inflammatory, genetic perturbation, and disease contexts. Leveraging TF footprinting and graph embedding, cEpiNets identifies shared and context-specific programs and predicts regulatory circuits in unseen biological contexts. In a SORL1-marked inflammatory microglial state that expands during AD progression, cEpiNets annotates AD risk variants at the SORL1 locus and identifies variants associated with cellular state abundance across donors. Cross-context analysis further identifies ZBTB14, whose inflammation-associated program connects AD risk variant-harboring CREs to target genes and widespread TF remodeling in AD. Donor-level ZBTB14 footprint activity is negatively associated with AD pathology, while combined IFN{gamma}/TNF stimulation represses ZBTB14 and activates a subset of inferred targets. Collectively, cEpiNets bridges genetic variation, regulatory programs, and disease-associated cellular phenotypes to facilitate mechanistic interpretation of complex disease genetics.
bioinformatics2026-08-26v1Tree-aware conditional language modeling recovers mutational patterns of viral evolution
Polunina, P. V.; Maier, W.; Rubin, A. F.Abstract
The evolutionary accessibility of a protein mutation depends on the sequence background in which it arises and its lineage history, yet most protein language models estimate sequence plausibility without explicitly considering the ordered sequence changes through which descendants arise. We developed evoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features. We demonstrated our approach using SARS-CoV-2 spike protein, pairing sequences from early Omicron lineages according to their positions on a mutation-annotated phylogeny, and evaluating model performance on sequence pairs from later lineages. Prompt-masking experiments showed that incorporating phylogenetic context substantially increased reliance on the supplied input information compared with a sequence-only model. Generated descendant sequences accurately reproduced the positional distribution of mutations observed during viral evolution, with strong correlations between predicted and observed mutation-frequency profiles for both the receptor-binding domain (Spearman's {rho} = 0.823) and the full spike protein ({rho} = 0.736). Although prediction accuracy for individual substitutions decreased with increasing evolutionary distance, the model consistently captured aggregate mutational patterns across the spike protein. Model-assigned mutation probabilities were also enriched among substitutions experimentally tolerated in deep mutational scanning assays of Omicron BA.2 receptor-binding domain expression (1.19-fold enrichment) and ACE2 binding (1.04-fold enrichment), despite the model being trained solely on observed ancestor-descendant sequence pairs and associated phylogenetic context features. These results demonstrate that explicitly providing protein language models with phylogenetic context during sequence generation can recover lineage-specific mutational patterns and yields probabilistic predictions consistent with experimentally measured functional constraints. evoPLM-Tree provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.
bioinformatics2026-08-26v1HIDE-Deconv: A hierarchical deconvolution framework for multiscale characterization of cellular remodeling
Goertler, F.; Voelkl, D.; Bolz, S.; Rayford, A.; Stevenson, T.; Sterr, T.; Mensching-Buhr, M.; Seifert, N.; Altenbuchinger, M.; Arp, J.; Schuster, C.; Tausche, J.; Engel, L.; Zacharias, H. U.Abstract
Most deconvolution methods estimate cellular composition at a single level of cellular resolution despite biological processes often manifesting within fine-grained cellular subpopulations. We present HIDE-Deconv, a hierarchical deconvolution framework that jointly optimizes cellular compositions across multiple levels of a cell-type hierarchy while maintaining consistency between resolutions. In benchmark experiments, HIDE-Deconv achieved the highest overall predictive performance among evaluated methods. Analyses of lung adenocarcinoma, sepsis, COVID-19 and systemic lupus erythematosus revealed biologically relevant cellular remodeling that remained concealed at broader levels of cellular resolution. HIDE-Deconv is available as an open-source framework at https://github.com/dvoelkl/HIDE-deconv.
bioinformatics2026-08-26v1Assay concordance sets exact ceilings on what one biological score can predict
Liu, Z.Abstract
Computational models of biology are ranked by averaging one prediction against many experimental realizations of a phenotype that are treated as interchangeable. We show this imposes an exact, model-free ceiling fixed by how much those realizations agree with each other, and that the ceiling depends on the evaluation metric through a single support-function identity. Measuring assay concordance across four public registries, 2,822 MaveDB score sets, 217 ProteinGym assays, two drug screens and 1,150 CRISPR cell lines, we find that two assays of one target agree at 0.56-0.68, and that 541 domains measured twice with different proteases fix assay reliability at 0.897, so 70-90% of every ceiling is irreducible biology rather than noise. Published predictors realize 63% of the achievable on the correlation benchmarks report and 18% on the top-1% selection their users perform. We provide the estimator, the ceilings, and the measurements the field has not made.
bioinformatics2026-08-26v1CCIDeconv: Hierarchical model for deconvolution of subcellular cell-cell interactions in single-cell data
Jayakumar, R.; Panwar, P.; Yang, J. Y. H.; Ghazanfar, S.Abstract
Cell-cell interaction (CCI) underlies several fundamental biological processes, including development, homeostasis and disease progression. Subcellular spatial transcriptomics (sST) provides an opportunity to examine whether CCI-associated signals show compartment-specific patterns within cells. Assessing CCI at subcellular level can help us gain insights into the distinct pathway activation and signalling patterns. We developed a novel approach that deconvolutes CCI into subcellular CCI (sCCI) information from non-spatial single-cell transcriptomics (scRNA- seq) based CCI using a modified CellChat-derived communication score. By estimating communication scores separately for cytoplasmic and nuclear compartments, we identified compartment-associated sCCI. We then deconvolved whole-cell communication scores into subcellular compartments using a hierarchical classification and regression framework, which we call CCIDeconv. To ensure biological fidelity, we integrated protein localization data from the Human Protein Atlas in our deconvolution model. Across nine publicly available human sST datasets, leave-one- dataset-out validation achieved a median composite score of 0.75, with mean R2 values of 0.87 and 0.80 for cytoplasmic- and nuclear-associated scores, respectively. Performance without spatial features approached that of spatial models as the number of training datasets increased, supporting application to non-spatial scRNA-seq data. This highlighted the potential for prediction of sCCI from scRNA-seq, given a sufficiently large number of training datasets. Overall, our method can attribute whole-cell CCI to its subcellular compartments, allowing researchers to dissect sCCI patterns and gain insights into the underlying biology of healthy and disease tissues. Keywords Cell-Cell Communication, Single Cell RNA-seq, Predictive Modeling, Bioinformatics, Transcriptomics, Machine Learning
bioinformatics2026-08-25v3GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations
Andrade, R. L.; Fiuza, T. d. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.Abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearson's $r$ improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTX-GUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
bioinformatics2026-08-25v3Stability-driven multi-omics integration for reproducible latent structure
Guan, H.; Gerwen, M. v.; Kim-Schulze, S.; Colicino, E.; Dolios, G.; Petrick, L.Abstract
High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We propose a Stability-driven framework for multi-omics integration that combines sparse generalized canonical correlation analysis with repeated cross-validation, out-of-sample projection, and systematic evaluation of both component-level and feature-level stability. We apply this framework to untargeted metabolomic and Olink targeted inflammation proteomic profiles in a thyroid cancer case-control cohort (n = 162). Our Stability-driven integration identified reproducible metabolomic and proteomic latent components that showed consistent out-of-sample disease associations and tracked temporally structured changes relative to time to diagnosis. The proposed framework provides a generalizable strategy for identifying reproducible latent structures that improve robustness of biological inference in multi-omics studies.
bioinformatics2026-08-25v3OMIO: A policy-driven Python library for reproducible microscopy image I/O
Musacchio, F.; Antony, H.; Baijal, A.; Crux, S.; Fuhrmann, F.; Gockel, N.; Hoffmann, D. M.; Mercan, D.; Nebeling, F. C.; Wolff, K.; Fuhrmann, M.Abstract
Modern fluorescence and multiphoton microscopy workflows operate within a heterogeneous ecosystem of file formats, partially overlapping metadata standards, and reader-specific conventions. In practice, this frequently leads to silent axis misinterpretations, loss or corruption of physical voxel size information, and laboratory-specific glue code that is fragile, poorly documented, and difficult to reproduce. OMIO, short for Open Microscopy Image I/O, addresses these issues by providing a lightweight, policy-driven image I/O layer for Python that enforces a canonical, OME-compatible data representation at the API boundary. The central contribution of OMIO is the explicit separation of low-level format access from semantic normalization. Existing reader libraries are used as interchangeable backends for extracting pixel data and available metadata, while OMIO enforces axis conventions, metadata interpretation, and fallback decisions in a centralized and auditable policy layer. This design allows heterogeneous microscopy inputs to be converted into a stable representation without propagating backend-specific assumptions into downstream analysis code. The core design principles of OMIO include canonical axis semantics (TZCYX), robust metadata normalization with explicit and auditable fallbacks, memory-aware operation via optional Zarr-based backends, and workflow-level semantics that extend beyond individual files to folder stacks and BIDS-like project structures. This architecture allows OMIO to orchestrate existing reader libraries into a coherent and reproducible I/O pipeline without replacing or duplicating their functionality. OMIO is implemented as an open-source and community-oriented system in which support for additional file formats and metadata conventions can be added incrementally through modular reader backends. By encouraging the contribution of example datasets, backend extensions, and feature requests, OMIO is designed to evolve alongside emerging acquisition systems while preserving strict semantic guarantees at the interface level. The resulting standardized OME-TIFF outputs are immediately suitable for downstream quantitative analysis and interactive inspection in scientific Python workflows, including workflows based on ImageJ and Napari. By standardizing image data at the I/O boundary, OMIO supports FAIR-aligned data sharing and reproducible microscopy analysis while facilitating the development of interoperable downstream tools.
bioinformatics2026-08-25v2From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores
Barbosa Araujo, P. V.; Fiuza, T. d. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.Abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
bioinformatics2026-08-25v2Differential Effects of Incomplete Lineage Sorting and Gene Tree Estimation Error on Gene Tree Distributions and Species Tree Inference
Tahmid, N.; Rhythm, S. I.; Bayzid, M. S.Abstract
Accurate species tree inference from genome-scale data is complicated by gene tree discordance, which can arise both from biological processes such as incomplete lineage sorting (ILS) and from technical factors such as gene tree estimation error (GTEE). While both factors reduce the accuracy of summary methods, their relative impact and characteristic patterns remain poorly understood. Here, we systematically compare the effects of ILS and GTEE by simulating gene tree datasets with comparable overall discordance levels, but with discordance arising exclusively from either ILS or GTEE. Using widely employed summary methods such as ASTRAL and wQFM, we show that GTEE typically has a stronger detrimental effect on species tree accuracy than ILS, even at matched discordance levels. We further characterize the structure of gene tree distributions under these two sources of discordance and show that ILS induces a structured, constrained skew in quartet distributions, whereas GTEE generates more uniform, high-entropy noise that does not diminish with additional genes. Our case study on a widely used avian phylogenomic dataset reveals similar distributional patterns across exons, introns, and ultraconserved elements (UCEs), which differ substantially in their levels of phylogenetic signal. A quartet-based analysis of these gene trees further shows that prioritizing loci with stronger and more consistent quartet support can improve the recovery of established avian clades. Overall, these results provide an empirical framework for a nuanced understanding of how ILS and GTEE shape gene tree distributions and influence species tree inference from limited or noisy gene tree datasets.
bioinformatics2026-08-25v2eSkip2 prioritizes exon-skipping antisense oligonucleotide target regions across exon--intron contexts
Chiba, S.; Kunitake, K.; Shirakaki, S.; Haque, U. S.; Wilton-Clark, H.; Shah, M. N. A.; Leckie, J. N.; Matsui, K.; Uno-Ono, F.; Yokota, T.; Aoki, Y.; Okuno, Y.Abstract
Exon-skipping antisense oligonucleotides (ASOs) can restore productive transcripts, but identifying effective binding regions remains difficult because splicing regulation extends across exons, introns and splice junctions. Here we develop eSkip2, a genome-informed framework that ranks target regions within a unified exon-intron sequence context. eSkip2 combines a genome-pretrained sequence model with ASO-induced exon-skipping data and single-nucleotide-variant splicing perturbations, followed by target-locus adaptation that requires no experimental ASO labels from the locus being designed. Across benchmarks comprising canonical exons and pseudoexons, multiple cell types and chemistries, and exonic, intronic and exon-intron-spanning targets, eSkip2 prioritized active regions and showed a higher median AUROC than applicable exon-restricted models. Prospective application to the combinatorial design of dual-targeting ASOs for DMD exon 46 enriched active candidates near the top of the ranking: the two most active new ASOs ranked within the top three and produced dose-dependent dystrophin restoration in patient-derived cells. These results support eSkip2 as a practical first-pass strategy for reducing experimental search space in exon-skipping ASO discovery.
bioinformatics2026-08-25v2Detecting CYP2C19 deletions from genotyping array signals using neural networks
Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.Abstract
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
bioinformatics2026-08-25v1AFP-R: An Open Resource Dedicated to Antifreeze Proteins
Liu, W.; Zhang, Y.; Xiu, D.; Liu, Y.; Wang, T.; Chai, X.; Qu, H.; Min, Y.; Zhang, Z.Abstract
Antifreeze proteins (AFPs), lower the freezing point via thermal hysteresis activity and/or ice recrystallization inhibition, playing a crucial role in protecting organisms from freezing damage under sub-zero milieu. This property endows them with promising applications in biomedicine and agriculture, ranging from tissue-organ cryopreservation to the development of frost-resistant crops. However, the lack of comprehensive resources dedicated for AFPs hinders further progress in elucidating their functional mechanisms and advancing their applications. Here, we report AFP-R, an online resource comprising AFP-DB and AFP-Predictor. AFP-DB is a comprehensive database with manually curated proteins bearing experimentally validated antifreeze activity derived from published literature, whereas AFP-Predictor is a sequence-based machine-learning model to identify AFPs. AFP-DB stores diverse AFP-related information, including sequences, structures, post-translational modifications, taxonomy and annotations of antifreeze-activity experimental assays. It now holds 186 entries, 607 sub-entries, and 1444 experimental records. AFP-Predictor, an AFP-identification algorithm built on protein language model ESM2 (Evolutionary Scale Modeling2), is trained on data in AFP-DB and outperforms several existing models. This work offers a valuable resource for systematically dissecting the mechanisms underlying AFP antifreeze activity and will facilitate their broader applications.
bioinformatics2026-08-25v1PhageLysData: an evidence-aware and AI-ready dataset of phage lytic enzymes and depolymerases
Medina-Ortiz, D.; Olivera-Nappa, A.; Lienqueo, M. E.; Opazo, R.; Romero, J.Abstract
Bacteriophage lytic enzymes and depolymerases are relevant to phage biology, antimicrobial development, and protein engineering, but their sequence and annotation data remain dispersed across general databases, specialized resources, genome-centred collections, and prediction-oriented datasets. We present PhageLysData, an evidence-aware and AI-ready resource constructed through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The release integrates 807,366 source observations from seven primary resources into 759,105 unique exact-sequence entities, comprising an evidence-supported Core of 11,867 entities, a Prediction Extension of 745,092 prediction-only candidates, and 2,146 Context entities retained for provenance and reference. This architecture preserves broad sequence-space coverage while maintaining a clear distinction between non-predictive and prediction-derived support. Core entities are enriched with harmonized biological annotations, physicochemical properties, independent InterProScan-derived functional annotations, mapped PDB and AlphaFold DB structural assets, and reusable numerical representations. For 11,259 eligible Core sequences, PhageLysData provides embeddings from 11 protein language models together with one-hot encoding under a common representation contract. Release-facing examples demonstrate latent-space exploration, unsupervised clustering, supervised classification, and evidence-aware candidate retrieval without defining a universal predictive benchmark. PhageLysData provides a traceable, versioned, and computationally accessible foundation for protein retrieval, comparative analysis, task-specific dataset construction, and machine-learning applications involving phage lytic enzymes and depolymerases.
bioinformatics2026-08-25v1An inflammation-associated five-gene expression signature stratifies survival and immune states in lung adenocarcinoma: an integrative public-cohort analysis
Zhou, X.; Le, Z.; Song, P.; Xu, Q.; Chen, M.; Liu, X.; Cao, M.; Zhan, S.; Liu, Y.; Zhang, L.Abstract
Background: Inflammation and the tumor immune microenvironment contribute to lung adenocarcinoma (LUAD) progression, but the relationship among inflammation-linked transcriptional heterogeneity, patient survival, and immune-state variation remains incompletely defined. Objective: We aimed to identify inflammation-associated LUAD subtypes, derive a parsimonious survival-stratification signature, and characterize its immune and pathway context across public transcriptomic cohorts. Methods: Expression profiles and clinical data were obtained from TCGA-LUAD, GTEx normal lung, and GEO datasets GSE11969, GSE30219, GSE31210, and GSE40791. A curated set of 596 inflammation-related genes was used for consensus clustering. Differential-expression analysis, functional enrichment, univariate Cox regression, and LASSO-Cox modeling were integrated to construct a gene-expression risk score. The prognostic dataset comprised 730 cases and was randomly divided into training (n=502) and internal-validation (n=228) sets; 85 GSE30219 cases formed an external-validation cohort. Immune-cell enrichment, gene set enrichment analysis (GSEA), gene set variation analysis (GSVA), and pan-cancer analyses were used for biological contextualization. Results: The LUAD-versus-control comparison identified 1,305 differentially expressed genes, including 498 upregulated and 807 downregulated genes. Consensus clustering resolved two inflammation-associated subtypes and 67 subtype-associated genes, of which 64 were higher and 3 were lower in Cluster 1 relative to Cluster 2. Thirty-three genes overlapped between the tumor-control and subtype contrasts. LASSO-Cox regression selected CHRDL1, FDCSP, CXCL13, CYP4B1, and S100P. The 1-, 3-, and 5-year areas under the time-dependent receiver operating characteristic curve were 0.6625, 0.6581, and 0.6658 in the training set; 0.7422, 0.6537, and 0.6761 in internal validation; and 0.6560, 0.6387, and 0.6753 in external validation. Risk groups differed across multiple T-cell, B-cell, natural-killer-cell, myeloid, dendritic-cell, macrophage, and granulocyte signatures. Positive GSEA signals included cell cycle (normalized enrichment score [NES]=2.67; adjusted P=1.42 x 10-), DNA replication (NES=2.52; adjusted P=2.52 x 10-), and mismatch repair (NES=2.20; adjusted P=1.77 x 10-). Conclusions: The five-gene expression score separated LUAD survival groups and captured coordinated proliferative and immune transcriptional states. Its moderate discrimination supports further biological and clinical validation rather than immediate clinical application.
bioinformatics2026-08-25v1Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles
Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.Abstract
Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.
bioinformatics2026-08-25v1UELer: a Jupyter-based framework for interactive exploration of multiplexed imaging datasets
Wu, Y.-L.; Liu, C.-S.; Lenoir, B.; Merz, K.; Dill, M. T.; Hartmann, F. J.Abstract
Summary Multiplexed imaging and spatial proteomics generate complex datasets that require both computational analysis and visual inspection. However, these tasks mostly occur in separate environments because interactive viewers generally require a local display or an additional data server beyond the remote Jupyter sessions itself where large datasets are computationally analyzed. We here present UELer, an interactive viewer that links multi-channel image views with quantitative analysis results directly within Jupyter notebooks, requiring no dedicated infrastructure beyond the notebook session. Cells selected through computational analysis and summary plots can be inspected directly in their tissue context, and selections made in the image can be made available to any downstream analysis. Together, these capabilities support interactive data exploration, iterative cell annotation, and reproducible retrieval of selected regions. Availability and Implementation UELer is a Python package built on ipywidgets and runs in Jupyter environments supporting ipywidgets 8.1 or later, tested in JupyterLab and Visual Studio Code on Linux, macOS, and Windows. It is freely available under GPL-3.0 license and can be installed via pip. Source code and documentation are available at https://github.com/HartmannLab/UELer and https://hartmannlab.github.io/UELer/. An online, no-install version runs remotely via BinderHub (https://mybinder.org/v2/gh/HartmannLab/UELer/main), accessible through the script/run_ueler_binder.ipynb notebook.
bioinformatics2026-08-25v1