Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-08-26v4Enrichment-free glycoproteomics harnessing real-time mass defect-driven glycopeptide classification reveals sex differences in murine fucosylation
Zhang, B.; Chau, T. H.; Bienes, K. M.; Arakawa, H.; Hane, M.; Sato, C.; Yokoi, A.; Kaji, H.; Ashwood, C.; Matsui, Y.; Kawahara, R.; Thaysen-Andersen, M.Abstract
Glycopeptide enrichment remains a cornerstone in glycoproteomics, but bias and reproducibility issues continue to hinder biological insight and clinical translation. Employing curated glycoproteomics datasets and machine learning, we trained a glycopeptide classifier to recognize N-glycopeptide precursors through mass defect signatures. Integration of the classifier into a data-dependent acquisition framework facilitated real-time prediction of N-glycopeptides from human serum and revealed sex differences in murine plasma fucosylation opening avenues for enrichment-free glycoproteomics.
bioinformatics2026-08-26v3UMITIC: An unsupervised framework for the joint characterization of cellular phenotypes and spatial neighborhoods in multiplex and hyperplex immunofluorescence imaging data
Sangüesa Recalde, M.; De Andrea, C. E.; Ariz, M.Abstract
Multiplexed imaging technologies enable the simultaneous measurement of dozens of protein markers while preserving context, providing a high-resolution view of tissue organization schemes. However, extracting meaningful insights from these high-dimensional datasets--particularly in hyperplex settings (>20 markers)--remains a major computational challenge, especially in the absence of annotated data. Here, we present UMITIC (Unsupervised Analysis of Multiplex Images via TIssue Characterization), a modular and unsupervised computational framework for the joint characterization of cell phenotypes and tissue neighborhoods from multiplex imaging data. UMITIC integrates three components: (i) CellCut, a strategy that combines nuclear and cytoplasmic predictions to improve the delineation capabilities of the framework; (ii) CellMap, a contrastive learning approach that generates low-dimensional representations of single-cell image crops that are enriched with morphological features; and (iii) TissueNet, a graph neural network that models spatial cell-cell interactions to identify tissue neighborhoods. We evaluated UMITIC across four datasets of increasing complexity to assess its robustness, scalability and biological relevance. With respect to a 7-plex human tonsil dataset, the framework identified canonical immune cell populations and reconstructed well-established anatomical regions. When applied to a 43-plex tonsil image, UMITIC preserved these tissue-level structures while enabling a finer cell subtype stratification process driven by increased marker dimensionality. We further validated our method on a 58-plex colorectal cancer cohort, where UMITIC was able to recover previously reported immune composition differences and spatial organization variations between patient groups with different prognoses. Finally, when an expert-annotated mass cytometry imaging dataset concerning human lung tissue was used, UMITIC achieved higher agreement with the reference tissue annotations than the existing approaches did, demonstrating improved lung microanatomy reconstruction accuracy. Together, these results show that UMITIC enables consistent and interpretable analyses of both cellular phenotypes and tissue architectures across diverse multiplex and hyperplex imaging datasets without the need for manual annotations.
bioinformatics2026-08-26v3CRISPR-HAWK: Haplotype- and Variant-aware Guide Design Toolkit for CRISPR-Cas
Kumbara, A.; Tognon, M.; Carone, G.; Fontanesi, A.; Bombieri, N.; Giugno, R.; Pinello, L.Abstract
Current CRISPR guide RNA design tools rely on reference genomes, overlooking how genetic variation impacts editing outcomes. As genome editing advances toward clinical applications, incorporating population diversity becomes essential for ensuring therapeutic efficacy across diverse populations. We present CRISPR-HAWK, a framework integrating individual- and population-scale variants and haplotypes into gRNA design. Analyzing therapeutic targets across 79,648 genomes reveals that genetic variants substantially alter guide performance. For the clinically approved sickle cell disease therapeutic guide targeting BCL11A, we identify haplotypes that completely abolish predicted cutting activity. Across seven therapeutic loci, 82.5% of guides contain variants modifying on-target activity. Variants also create novel protospacer adjacent motif sites generating individual-specific guides invisible to reference-based design. These findings demonstrate that variant-aware selection is critical for equitable genome editing. CRISPR-HAWK is available at https://github.com/pinellolab/CRISPR-HAWK and https://github.com/InfOmics/CRISPR-HAWK
bioinformatics2026-08-26v3scDisent: regulatory-aware disentangled representation learning for multi-omic single-cell analysis
Xi, G.Abstract
Single-cell multi-omic technologies measure complementary aspects of cellular identity and regulatory state, yet most integration models compress these signals into one entangled latent space. Such representations are useful for clustering but poorly suited to regulator-centered interpretation or perturbation-oriented analysis. We present scDisent (https://github.com/xiguoren/scDisent), a generative framework that separates expression-associated variables (zexpr) from regulation-associated variables (zreg) and links them through a sparse directed mapping. scDisent combines modality-specific encoding, variational disentanglement, total-correlation and orthogonality regularization, and a Gumbel-gated causal module protected by detach-based gradient isolation. Across benchmark datasets with matched modalities, scDisent achieved the strongest clustering performance among the tested methods while exposing regulatory structure that competing integration models do not represent explicitly. The learned causal atlas remained sparse, perturbation analyses recovered biologically coherent lineage-associated programs, and branch-separation analyses showed that benchmark-label information concentrated in zexpr rather than zreg. These results position scDisent as a multi-omic representation model that improves both integration quality and biological interpretability
bioinformatics2026-08-26v2MONTE enables unified pan-cancer tumor purity estimation andmethylation correction from bulk DNA methylation arrays
Kim, M.; Lee, W.-H.; Yao, V.Abstract
Bulk DNA methylation profiling is widely used to study cancer epigenomics in clinical settings, but these measurements aggregate signals from malignant and non-malignant cells, introducing composition-dependent confounding that complicates tumor-intrinsic interpretation and cross-cohort analyses. While existing methods can estimate tumor purity and, in some cases, correct methylation measurements, they typically require cancer-specific reference models, matched normal samples, or predefined probe sets, limiting their applicability to rare cancers, different clinical cohorts, and cross-dataset comparisons. We present MONTE (Methylation-based Observation Normalization and Tumor purity Estimation), a unified, cancer label-free framework for tumor purity inference and CpG-resolved methylation correction from bulk DNA methylation data. MONTE learns probe-wise relationships between methylation and tumor purity using an empirical Bayes-moderated linear model and infers purity in new samples via signal-to-noise weighted aggregation, without requiring matched normals, cancer labels, or predefined probe sets. A single pan-cancer MONTE model outperforms existing cancer-specific methods for purity estimation across 21 cancer types, generalizes across purity references, and runs orders of magnitude faster on full-dataset analyses. MONTE also introduces Bayesian transfer learning, which enables efficient recalibration to alternative purity definitions, validated on three independent external cohorts. Methylation correction with MONTE further amplifies tumor-relevant regulatory signal and improves the reproducibility of differential methylation analyses. By unifying purity estimation and correction in a single flexible, scalable, and interpretable framework, MONTE broadens the accessibility of tumor-intrinsic methylation analysis across cancer types and datasets.
bioinformatics2026-08-26v2KSTITCH links cellular morphology and gene expression in spatial transcriptomics
Kumar, S.; Shi, Y.; Vallius, T.; Day, C.-P.; Absil, P.- A.; Srivastava, A.; Hannenhalli, S.; Gopalan, V.Abstract
In situ spatial (ISS) sequencing can uncover co-variation between cellular morphology and gene expression in vivo. However, a principled and interpretable mathematical representation of morphology has not yet been applied in this context. In particular, current deep learning-based representations of cell images confound a cell's shape with its size. We present an interpretable representation of cellular boundary contours, based on tangent principal component analysis (TPCA) in a Kendall shape manifold, that captures size-independent contour shape features. This approach successfully recovers shape-perturbing genes in an RNAi screen than a previous metric geometry-based approach. We build on TPCA to develop KSTITCH (Kendall Shape-TranscriptomIc Correlation and Harmonization), an approach to reveal covariation between cell morphology with gene expression in ISS datasets. In a Xenium dataset, KSTITCH recovers known morphology-transcriptomic relationships in keratinocytes, macrophages and endothelial cells. Across samples in a melanoma CosMx dataset, KSTITCH reproducibly associates elongated and triangular fibroblasts with proximity to malignant cells and myofibroblast-like transcriptional program. Finally, KSTITCH independently recovers a known link between mesenchymal-like malignant cell states and increased cell area in two melanoma cohorts. KSTITCH can thus yield interpretable morphology-transcriptome relationships across cell types, patients, and spatial transcriptomics platforms. KSTITCH is available at https://github.com/vishakagopalan/kstitch .
bioinformatics2026-08-26v2On the illusion of scRNA-seq batch effect correction
Codice', F.; Fariselli, P.; Raimondi, D.Abstract
Batch correction methods in single-cell RNA sequencing are essential for removing technical variation that can otherwise lead to misleading downstream analyses. The reliability of these methods is typically evaluated using unsupervised metrics. Here, we apply a Machine Learning (ML) technique called probing, formalized as the Batch Probing Score (BPS), to empirically demonstrate across six datasets that the most popular batch correctors fail to fully remove batch signal. In most cases, the batch of origin remains clearly identifiable after correction, even though standard evaluation metrics cannot detect it. We show that existing unsupervised metrics lack the sensitivity and specificity required to capture residual batch signal, whereas ML-based approaches can still detect it. This residual signal can similarly be picked up by downstream analysis tools, potentially leading to biased results. Because BPS is supervised, it directly quantifies batch signal strength by measuring how accurately the batch of origin can be predicted for each sample. It therefore provides an upper bound on the residual ML-actionable batch signal that could otherwise remain unnoticed. Our findings suggest that probing-based metrics should become a standard for assessing batch correction methods in single-cell RNA-seq and other areas of genomics.
bioinformatics2026-08-26v2Identification of Altered Potassium Channels for Drug Repurposing in Long COVID Patients
George, J. P.; Gaikwad, K. B.; Sharma, J.Abstract
Long COVID (LC) is a complex condition characterized by persistent, chronic multisystem manifestations, with a significant proportion of patients exhibiting neurological symptoms. Human ion channels (HICs), particularly potassium channels, are abundantly expressed in the nervous system and linked to key metabolic processes, making them potential candidates for understanding LC pathophysiology and drug repurposing. Meta-analysis of RNA-Seq datasets from COVID-19-recovered and LC patients was performed to identify altered HICs in LC. Differential gene expression analysis, functional enrichment analysis, and weighted gene co-expression network analysis were performed to uncover key genes, pathways, and co-expression modules consisting of HICs, lipid metabolism-, and immune signaling-related genes. A total of 715 dysregulated genes, including eighteen HICs were identified, among which seven were potassium channels. Three significant modules containing HICs, lipid metabolism-, and immune signaling-related genes were identified and found to be associated with antigen processing and presentation, complement and coagulation cascades, and cytokine-related pathways. Additionally, drug-gene interaction analysis led to identification of approved drugs targeting KCNA6, KCNJ10, KCNN3, and KCNH4 that might provide opportunities for drug repurposing in neurological manifestations in LC. Further experimental validation is required to establish their efficacy and assess their potential for translation into clinical applications for patients with LC.
bioinformatics2026-08-26v2A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models
Li, L.; Duan, S.; Zha, X.; Ye, F.; Zhang, Y.; Zhang, X.; Cao, Y.; Liu, C.; Fang, B.Abstract
Single-cell drug perturbation models are increasingly used to predict how compounds remodel cellular states, but they are still largely assessed by expression reconstruction. Whether high expression similarity reflects preservation of drug-response signatures remains unclear. Here we present scDrugPerturb-Bench, a mechanism-annotated benchmark that links matched control and drug-treated single-cell RNA-sequencing profiles to literature-curated directional key-gene evidence. The resource covers 181 datasets, 423 annotated response cases, 717 unique key genes and 2.5 million cells. We introduce the Mechanism Fidelity Score (MFS) to evaluate key-gene direction, effect-size recovery, gene-set coherence, mechanism specificity and pathway-level response polarity. Across 12 perturbation-prediction models, 3 baselines and 10 data splits, expression-similarity metrics were weakly aligned with MFS and selected different model configurations. Mechanism-aware selection improved early drug retrieval in a transcriptome-based drug design evaluation, indicating that MFS provides practical information beyond benchmark reporting. Systematic benchmarking revealed limited fidelity to drug-response signatures across cell-line and source-integrated settings. Frozen single-cell foundation model embeddings produced local, metric-dependent gains rather than universal improvements, and source context substantially reshaped model assessment. Hard-negative tests further showed that plausible perturbation responses can arise from non-specific transcriptional shortcuts. These results show that expression reconstruction is an insufficient proxy for preserving drug-response signatures and establish scDrugPerturb-Bench as a benchmark for mechanism-aware evaluation of single-cell drug perturbation models.
bioinformatics2026-08-26v2Tree-aware conditional language modeling recovers mutational patterns of viral evolution
Polunina, P. V.; Maier, W.; Rubin, A. F.Abstract
The evolutionary accessibility of a protein mutation depends on the sequence background in which it arises and its lineage history, yet most protein language models estimate sequence plausibility without explicitly considering the ordered sequence changes through which descendants arise. We developed evoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features. We demonstrated our approach using SARS-CoV-2 spike protein, pairing sequences from early Omicron lineages according to their positions on a mutation-annotated phylogeny, and evaluating model performance on sequence pairs from later lineages. Prompt-masking experiments showed that incorporating phylogenetic context substantially increased reliance on the supplied input information compared with a sequence-only model. Generated descendant sequences accurately reproduced the positional distribution of mutations observed during viral evolution, with strong correlations between predicted and observed mutation-frequency profiles for both the receptor-binding domain (Spearman's {rho} = 0.823) and the full spike protein ({rho} = 0.736). Although prediction accuracy for individual substitutions decreased with increasing evolutionary distance, the model consistently captured aggregate mutational patterns across the spike protein. Model-assigned mutation probabilities were also enriched among substitutions experimentally tolerated in deep mutational scanning assays of Omicron BA.2 receptor-binding domain expression (1.19-fold enrichment) and ACE2 binding (1.04-fold enrichment), despite the model being trained solely on observed ancestor-descendant sequence pairs and associated phylogenetic context features. These results demonstrate that explicitly providing protein language models with phylogenetic context during sequence generation can recover lineage-specific mutational patterns and yields probabilistic predictions consistent with experimentally measured functional constraints. evoPLM-Tree provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.
bioinformatics2026-08-26v1HIDE-Deconv: A hierarchical deconvolution framework for multiscale characterization of cellular remodeling
Goertler, F.; Voelkl, D.; Bolz, S.; Rayford, A.; Stevenson, T.; Sterr, T.; Mensching-Buhr, M.; Seifert, N.; Altenbuchinger, M.; Arp, J.; Schuster, C.; Tausche, J.; Engel, L.; Zacharias, H. U.Abstract
Most deconvolution methods estimate cellular composition at a single level of cellular resolution despite biological processes often manifesting within fine-grained cellular subpopulations. We present HIDE-Deconv, a hierarchical deconvolution framework that jointly optimizes cellular compositions across multiple levels of a cell-type hierarchy while maintaining consistency between resolutions. In benchmark experiments, HIDE-Deconv achieved the highest overall predictive performance among evaluated methods. Analyses of lung adenocarcinoma, sepsis, COVID-19 and systemic lupus erythematosus revealed biologically relevant cellular remodeling that remained concealed at broader levels of cellular resolution. HIDE-Deconv is available as an open-source framework at https://github.com/dvoelkl/HIDE-deconv.
bioinformatics2026-08-26v1Assay concordance sets exact ceilings on what one biological score can predict
Liu, Z.Abstract
Computational models of biology are ranked by averaging one prediction against many experimental realizations of a phenotype that are treated as interchangeable. We show this imposes an exact, model-free ceiling fixed by how much those realizations agree with each other, and that the ceiling depends on the evaluation metric through a single support-function identity. Measuring assay concordance across four public registries, 2,822 MaveDB score sets, 217 ProteinGym assays, two drug screens and 1,150 CRISPR cell lines, we find that two assays of one target agree at 0.56-0.68, and that 541 domains measured twice with different proteases fix assay reliability at 0.897, so 70-90% of every ceiling is irreducible biology rather than noise. Published predictors realize 63% of the achievable on the correlation benchmarks report and 18% on the top-1% selection their users perform. We provide the estimator, the ceilings, and the measurements the field has not made.
bioinformatics2026-08-26v1Evaluating Aggregated Gene Level eQTL Scores
Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.Abstract
Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.
bioinformatics2026-08-26v1Genomic-Based Prediction of Exopolysaccharide Composition and Structure: Insights from Rhizobium and Sinorhizobium Species
Tulumello, J.; Long, J.; Achouak, W.; Garron, M.-L.; Terrapon, N.; Heulin, T.Abstract
Bacterial exopolysaccharides (EPS) are key components in biofilm formation, stress protection, and symbiosis in Rhizobiaceae. While EPS structural diversity is extensive, experimental characterization remains limited. In this study, we experimentally determined and compared four distinct EPS structures produced by ten Rhizobium alamii strains. Using genomic data, we bioinformatically identified supra-operonic clusters (SOCs) responsible for these EPS biosynthesis. We introduced a computational framework to predict, score, and compare EPS SOCs across 84 Rhizobium and Sinorhizobium species, linking gene content to structural and functional EPS diversity. A total of 743 EPS SOCs was selected for network analyses, allowing the identification of 36 major groups of orthologous EPS SOCs, successfully recovering all known EPS biosynthetic loci and two novels SOCs potentially encoding uncharacterized EPS (xEPS-I, xEPS-II). Profiles of EPS SOCs correlated with taxonomical groups, with a single EPS SOC conserved through all 84 genomes and distinct additional EPS SOCs depending on the group, but do not strictly explain symbiotic capacity. Genetic comparisons of transporters (Wzx, Wzy) and glycosyltransferase sequences indicated these proteins as key markers of EPS structure. Overall, this computational framework accurately identified and classified EPS SOCs, providing a scalable, genome-based method for predicting EPS biosynthetic potential in Rhizobiaceae and usable in other microbial genera.
bioinformatics2026-08-26v1Robustness to nuisance perturbations enables unsupervised evaluation of single-cell foundation models
Sallam, A.; Gillis, J.Abstract
Single-cell foundation model evaluations have relied almost exclusively on downstream tasks. While these tasks measure whether an embedding recovers annotated cell types, batches, or trajectories, they cannot determine if that structure is reproducible or merely an artifact of a single noisy draw, a key limitation since incomplete sampling is intrinsic to single-cell measurement. Here, we introduce a fully unsupervised evaluation framework grounded in a fundamental principle: a faithful representation must preserve its neighbourhood structure under nuisance perturbations that mimic technical and sampling variation. Across five scFMs, a PCA baseline, and 39 datasets, we show that models ranked as near-equivalent by standard benchmarks differ nearly twofold in local neighbourhood preservation under a perturbation discarding just 5% of counts. This structural instability is scale-dependent and often masked by visually coherent embeddings. Cluster-level stability under resampling tracks established bio-conservation metrics (Spearman {rho}=0.78), showing that invariance to nuisance perturbations captures representation quality no benchmark measures directly.
bioinformatics2026-08-26v1Automated Detection of Livestock Gastrointestinal Parasite Eggs and Cysts Using YOLOv8-Based Deep Learning
Sarwer, A.Abstract
Parasitic infection is one of the common health problems of livestock in Bangladesh. Due to the country's climate, heavy monsoon rainfall, low biosecurity in farms, and high humidity, along with presence of suitable vector organisms, gastrointestinal parasitism remains widespread in cattle and other livestock. The standard method of diagnosis is microscopic examination of fecal samples, but this depends on manual observation, which is time-consuming and can lead to human error, mainly because many parasite eggs look similar to each other and samples often contain contaminants that can be mistaken for eggs or cysts. In this study, we tried to apply the YOLOv8 deep learning model for automated detection of parasitic eggs and cysts from microscopic images of livestock fecal samples. Images of clinical cases were collected, annotated, and used to train the model in Python, with batch size 16, auto optimizer, learning rate 0.01, momentum 0.937 and weight decay 0.0005. Training was done using Google Colab, and the model was evaluated using precision, recall, F1-score, mAP50, and mAP50-95. The model achieved a precision of 56%, recall of 24%, F1-score of 33.6%, mAP50 of 33%, and mAP50-95 of 22%. The relatively low recall and F1-score indicate that the model still has considerable limitations, largely due to insufficient species-specific training data and presence of image artifacts. Underrepresentation of some parasite species, such as Trichuris spp., in the dataset also caused class imbalance, which affected the model's ability to detect these species reliably. Despite these limitations, the study indicates that YOLOv8 architecture has some potential to be used for detection of parasitic eggs and cysts from microscopic images, and that further work with larger and more balanced datasets may improve performance and applicability in veterinary diagnostics. Keywords: YOLOv8, livestock parasites, deep learning, microscopic image analysis, veterinary diagnostics, Bangladesh
bioinformatics2026-08-26v1Interpretable Forecasting of Kidney Cancer Progression via Generative AI and Symbolic Reasoning
Prol-Castelo, G.; Syrri, E.; Manginas, N.; Manginas, V.; Sanchez-Valle, J.; Katzouris, N.; Paliouras, G.; Valencia, A.; Cirillo, D.Abstract
Predicting cancer stage progression from omics data, and deriving molecular insight into the mechanisms driving it, remains a major challenge, owing in part to the lack of adequate longitudinal data and the interpretability limitations of current forecasting models. Large cancer datasets such as TCGA capture patient profiles cross-sectionally rather than longitudinally, complicating timely treatment decisions as tumors become more invasive. Deep neural networks typically used for forecasting, such as LSTMs, compound this problem by remaining largely opaque and offering clinicians no straightforward way to audit their predictions. Clear cell renal cell carcinoma (ccRCC) illustrates the clinical stakes of both challenges. Five-year survival falls from over 94% at stage I to 28% at stage IV, yet early-stage tumors are often managed under active surveillance, a strategy constrained by sparse molecular evidence of progression risk. Detecting progression in time, meanwhile, demands forecasts clinicians can interpret and trust, not black-box predictions. We address both challenges by combining generative and symbolic AI: a Variational Autoencoder trained on bulk RNA-Seq profiles of 530 TCGA ccRCC patients generates synthetic pseudo-time trajectories that overcome the absence of longitudinal data, while a symbolic rule-induction framework (ASAL) learns finite-state automata from these trajectories, encoding stage transition as human-readable Boolean conditions over gene expression, which a complex event forecasting system (Wayeb) converts into probabilistic forecasts of stage advancement. An independent XGBoost classifier trained on real patients (F1 score = 0.71-0.81) shows a gradual early-to-late probability shift along the synthetic trajectories, absent in non-progressing control trajectories. Pathway enrichment of those trajectories reveals stage-dependent changes in established kidney cancer-related processes, including the TCA cycle and DNA repair. Finally, our symbolic forecaster nearly matches an LSTM baseline (macro F1 = 0.928 vs. 0.964), while additionally offering an inspectable rule set and a probability distribution over transition timing rather than a single opaque score. This work shows that generative and symbolic AI, paired together, can turn cross-sectional cohorts into a transparent, forecast-oriented framework for modeling disease progression, demonstrated here in ccRCC.
bioinformatics2026-08-26v1BatchRefiner: fast, significant improvement in batch integration of single-cell embeddings with ensemble refinement
Schäffer, D. E.; Kang, H.; Aksu, E. D.; Edelman, D.; Berger, B.Abstract
Data from single-cell RNA sequencing (scRNA-seq) and the Assay for Transposase-Accessible Chromatin (scATAC-seq) are high-dimensional, sparse, and undesirably capture technical variability between experiments or batches. Many analysis methods thus seek to produce a low-dimensional cell-by-feature embedding space that groups together biologically similar cells across batches while distancing dissimilar cells. Here, we introduce ensemble refinement for scRNA-seq and scATAC-seq embeddings, inspired by ensemble methods from statistical machine learning, and implement BatchRefiner, a fast post-processing tool to enhance batch integration. We extensively benchmark widely-used scRNA-seq embedding methods on both batch integration and biological conservation over a wide range of datasets, before and after the addition of BatchRefiner. We extend these benchmarking approaches to provide the first comprehensive benchmark of batch integration for scATAC-seq embedding methods, including BatchRefiner. Importantly, we formalize a significance statistic, which we use to demonstrate BatchRefiner's significant improvement in batch integration across a wide range of embedding methods, atlas-scale datasets, and established metrics.
bioinformatics2026-08-26v1Orthology transfer maps only the conserved core of the Varroa destructor proteome and over-calls host absence two times in three
Ryba, S.Abstract
The ectoparasitic mite Varroa destructor is the principal threat to managed honey bees, and a test case for the genome-scale methods applied to non-model organisms, nearly all of which infer from orthology. We reconstructed the first genome-wide protein-interaction network for V. destructor (7,080 proteins, 335,914 interactions), whose modular structure exceeds a degree-preserving null by 368 standard deviations, but whose every edge is interolog-transferred and every node conserved at least to Eukaryota. None of the 791 genes lacking an orthologous group enters it - arithmetic rather than discovery - yet the excluded compartment is large and coherent. It comprises 3,161 genes (30.9% of the proteome), shorter and less annotated than the rest; an annotation-free genome search detects orphans in a tick genome at 4.0% against 70.4% for networked genes. Within the orthology-bearing compartment visibility is non monotonic: the Acari-level bin (74.1%) falls below the Arthropoda-level bin (89.9%). The same logic applied to host comparison yields a benchmarked error: of genes called absent from Apis on group identity alone, 67.4% recover a sequence homologue - against zero for a shuffled null and 1.3% in the presence direction - rising to 78.3% in the least panel-biased stratum. Both figures are properties of the calling rule: under an identity floor the error directions cross near 34% identity; orthology cannot be said to err in either direction without fixing the criterion first. Host divergence resolves into gene absence and residue level substitution, falling in those two compartments respectively. A bee-sparing target map follows as broader impact.
bioinformatics2026-08-26v1UTR-Diffusion: Conditional Diffusion Modeling for Multi-objective and Constrained UTR Design
Dai, C.; Sato, K.Abstract
Motivation: The 5-prime untranslated region (UTR) and the start-codon-proximal region of the coding sequence (CDS) jointly influence translation efficiency and local RNA secondary-structure stability, while synonymous codon choices throughout the CDS shape codon adaptation. Because the encoded protein is often predetermined, practical mRNA design must coordinate these quantitative objectives while preserving specified nucleotide sequences and amino-acid identities. Existing generative approaches typically address continuous-valued targeting, explicit sequence constraints, and codon-usage control separately rather than integrating all three within a single model. Results: We present UTR-Diffusion, a diffusion-based framework for 5-prime UTR and 5-prime UTR-CDS junction design. UTR-Diffusion conditions generation on continuous-valued MRL and MFE targets and supports nucleotide-level constraints, amino-acid-level constraints with synonymous-codon flexibility, and codon-adaptiveness control that modulates the sequence-level codon adaptation index (CAI). Systematic evaluations across dense MRL-MFE target grids showed that generated distributions shifted consistently with both targets, retained substantial diversity, and strictly preserved specified nucleotide sequences and amino-acid identities. Codon-adaptiveness control yielded distinct, monotonically ordered CAI levels that closely followed the specified adaptiveness targets. In comparative benchmarks, UTR-Diffusion outperformed representative existing methods in high-MRL optimization and precise MRL targeting for 5-prime UTR design, and achieved higher MRL, less-negative junction MFE, and higher CAI than peptide-preserving baselines in 50-nt 5-prime UTR-CDS junction design.
bioinformatics2026-08-26v1AI-driven framework modeling perturbation in brain organoids reveals candidate genes for autism
Koh, I. G.; Chang, E.; Choi, Y. S.; Kim, S.-W.; Kim, Y.; Lee, H.; Byeon, G.; Ryu, Y.; Kim, S.; Lee, J.; Park, H.; Sim, H.; Ryu, Y.; Shim, W.; Lee, J.; Salazar, N. B.; de Aquino, M. M.; Engchuan, W.; Zhou, X.; Son, J. H.; Lee, J.; Bong, G.; Kim, I. B.; Han, J. H.; Werling, D. M.; Kim, S. H.; Oh, M.; Kim, M.-S.; Lee, D.; Kim, J.; Lee, Y.-S.; Sun, W.; Kim, E.; Scherer, S. W.; Jeon, M.; Yoo, H. J.; An, J.-Y.Abstract
Autism gene discovery is constrained by the rarity and heterogeneity of damaging variants, requiring large cohorts to identify susceptibility genes. Neural organoids and single-cell foundation models enable perturbation modeling in neurodevelopmental contexts. Here, we show that perturbation-informed foundation modeling of neural organoids can provide functional context for prioritizing candidate genes with genomic and clinical support. We constructed a 3.6-million-cell organoid atlas and trained models to predict genome-wide perturbation responses. Benchmarking 17 models identified a telencephalic neuron-specific model best preserving autism-relevant perturbation structure. Genome-wide profiling revealed two clusters associated with mid-fetal synaptic neuronal processes and early radial glia ubiquitin signaling. These clusters were supported by damaging-variant enrichment and clinical phenotypes across 89,916 family-based samples. Logistic-regression prioritization identified 343 candidates, including 167 in the key clusters, with convergence across TADA signals and recurrent evidence for NBEA and KLHDC10. This framework integrates predicted perturbation effects with genomic evidence to support autism candidate prioritization.
bioinformatics2026-08-26v1CytoGate-Bench: an LLM benchmark for cross-panel cell gating in cytometry
Kim, J.; Lee, B.; Ahn, N.; Ionita, M.; McKeague, M. L.; Lee, M. E.; Jeong, C.-U.; Apostolidis, S. A.; Baxter, A. E.; Shwetank, ; Greenplate, A. R.; Wherry, E. J.; Sohn, K.-A.; Kim, D.Abstract
In cytometry, the workhorse single-cell technology of clinical immunology, every study defines its own antibody panel and cell-type vocabulary, so a classifier trained on one cannot annotate the next. Immunologists instead annotate by manual gating, splitting one parent population at a time on a two-marker plot, down an expert-defined hierarchy. We introduce CytoGate-Bench, a benchmark that reformulates this per-step procedure as a zero-shot, panel-agnostic task for large language models. It comprises 23,646 expert-annotated instances re-curated from 11 public flow- and mass-cytometry cohorts spanning eight marker panels. Across six open- and closed-weight backbones, the strongest formulation draws one rectangular gate per candidate and falls within the range of trained, panel-specialized baselines. It degrades less under distribution shift. Walking the hierarchy stepwise outperforms predicting every cell type at once. Ablations trace the signal to the data distribution shape and curated marker priors. However, adding vision or a self-verification loop systematically tightens gates.
bioinformatics2026-08-26v1Context-dependent regulatory networks connect Alzheimer's disease genetics to microglial inflammatory responses
Fu, T.-T.; Kurkela, M.; Tu, J.; Zhang, J.; Sun, N.; Farrer, L. A.; TCW, J.; Hou, L.Abstract
Inflammation is central to Alzheimer's disease (AD) pathogenesis. Microglia, the resident innate immune cells of the brain, exhibit diverse inflammatory states and are enriched for AD-associated genetic variants within active cis-regulatory elements (CREs). However, the interplay among genetic variants, transcription factor (TF)-CRE-gene programs, and microglial responses across inflammatory and disease contexts remain poorly understood. Here, we develop context-dependent epigenomic networks (cEpiNets), integrating bulk and single-nucleus assay for transposase-accessible chromatin using sequencing (ATAC-seq) to reconstruct regulatory programs across inflammatory, genetic perturbation, and disease contexts. Leveraging TF footprinting and graph embedding, cEpiNets identifies shared and context-specific programs and predicts regulatory circuits in unseen biological contexts. In a SORL1-marked inflammatory microglial state that expands during AD progression, cEpiNets annotates AD risk variants at the SORL1 locus and identifies variants associated with cellular state abundance across donors. Cross-context analysis further identifies ZBTB14, whose inflammation-associated program connects AD risk variant-harboring CREs to target genes and widespread TF remodeling in AD. Donor-level ZBTB14 footprint activity is negatively associated with AD pathology, while combined IFN{gamma}/TNF stimulation represses ZBTB14 and activates a subset of inferred targets. Collectively, cEpiNets bridges genetic variation, regulatory programs, and disease-associated cellular phenotypes to facilitate mechanistic interpretation of complex disease genetics.
bioinformatics2026-08-26v1CCIDeconv: Hierarchical model for deconvolution of subcellular cell-cell interactions in single-cell data
Jayakumar, R.; Panwar, P.; Yang, J. Y. H.; Ghazanfar, S.Abstract
Cell-cell interaction (CCI) underlies several fundamental biological processes, including development, homeostasis and disease progression. Subcellular spatial transcriptomics (sST) provides an opportunity to examine whether CCI-associated signals show compartment-specific patterns within cells. Assessing CCI at subcellular level can help us gain insights into the distinct pathway activation and signalling patterns. We developed a novel approach that deconvolutes CCI into subcellular CCI (sCCI) information from non-spatial single-cell transcriptomics (scRNA- seq) based CCI using a modified CellChat-derived communication score. By estimating communication scores separately for cytoplasmic and nuclear compartments, we identified compartment-associated sCCI. We then deconvolved whole-cell communication scores into subcellular compartments using a hierarchical classification and regression framework, which we call CCIDeconv. To ensure biological fidelity, we integrated protein localization data from the Human Protein Atlas in our deconvolution model. Across nine publicly available human sST datasets, leave-one- dataset-out validation achieved a median composite score of 0.75, with mean R2 values of 0.87 and 0.80 for cytoplasmic- and nuclear-associated scores, respectively. Performance without spatial features approached that of spatial models as the number of training datasets increased, supporting application to non-spatial scRNA-seq data. This highlighted the potential for prediction of sCCI from scRNA-seq, given a sufficiently large number of training datasets. Overall, our method can attribute whole-cell CCI to its subcellular compartments, allowing researchers to dissect sCCI patterns and gain insights into the underlying biology of healthy and disease tissues. Keywords Cell-Cell Communication, Single Cell RNA-seq, Predictive Modeling, Bioinformatics, Transcriptomics, Machine Learning
bioinformatics2026-08-25v3GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations
Andrade, R. L.; Fiuza, T. d. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.Abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearson's $r$ improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTX-GUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
bioinformatics2026-08-25v3Stability-driven multi-omics integration for reproducible latent structure
Guan, H.; Gerwen, M. v.; Kim-Schulze, S.; Colicino, E.; Dolios, G.; Petrick, L.Abstract
High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We propose a Stability-driven framework for multi-omics integration that combines sparse generalized canonical correlation analysis with repeated cross-validation, out-of-sample projection, and systematic evaluation of both component-level and feature-level stability. We apply this framework to untargeted metabolomic and Olink targeted inflammation proteomic profiles in a thyroid cancer case-control cohort (n = 162). Our Stability-driven integration identified reproducible metabolomic and proteomic latent components that showed consistent out-of-sample disease associations and tracked temporally structured changes relative to time to diagnosis. The proposed framework provides a generalizable strategy for identifying reproducible latent structures that improve robustness of biological inference in multi-omics studies.
bioinformatics2026-08-25v3OMIO: A policy-driven Python library for reproducible microscopy image I/O
Musacchio, F.; Antony, H.; Baijal, A.; Crux, S.; Fuhrmann, F.; Gockel, N.; Hoffmann, D. M.; Mercan, D.; Nebeling, F. C.; Wolff, K.; Fuhrmann, M.Abstract
Modern fluorescence and multiphoton microscopy workflows operate within a heterogeneous ecosystem of file formats, partially overlapping metadata standards, and reader-specific conventions. In practice, this frequently leads to silent axis misinterpretations, loss or corruption of physical voxel size information, and laboratory-specific glue code that is fragile, poorly documented, and difficult to reproduce. OMIO, short for Open Microscopy Image I/O, addresses these issues by providing a lightweight, policy-driven image I/O layer for Python that enforces a canonical, OME-compatible data representation at the API boundary. The central contribution of OMIO is the explicit separation of low-level format access from semantic normalization. Existing reader libraries are used as interchangeable backends for extracting pixel data and available metadata, while OMIO enforces axis conventions, metadata interpretation, and fallback decisions in a centralized and auditable policy layer. This design allows heterogeneous microscopy inputs to be converted into a stable representation without propagating backend-specific assumptions into downstream analysis code. The core design principles of OMIO include canonical axis semantics (TZCYX), robust metadata normalization with explicit and auditable fallbacks, memory-aware operation via optional Zarr-based backends, and workflow-level semantics that extend beyond individual files to folder stacks and BIDS-like project structures. This architecture allows OMIO to orchestrate existing reader libraries into a coherent and reproducible I/O pipeline without replacing or duplicating their functionality. OMIO is implemented as an open-source and community-oriented system in which support for additional file formats and metadata conventions can be added incrementally through modular reader backends. By encouraging the contribution of example datasets, backend extensions, and feature requests, OMIO is designed to evolve alongside emerging acquisition systems while preserving strict semantic guarantees at the interface level. The resulting standardized OME-TIFF outputs are immediately suitable for downstream quantitative analysis and interactive inspection in scientific Python workflows, including workflows based on ImageJ and Napari. By standardizing image data at the I/O boundary, OMIO supports FAIR-aligned data sharing and reproducible microscopy analysis while facilitating the development of interoperable downstream tools.
bioinformatics2026-08-25v2From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores
Barbosa Araujo, P. V.; Fiuza, T. d. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.Abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
bioinformatics2026-08-25v2Differential Effects of Incomplete Lineage Sorting and Gene Tree Estimation Error on Gene Tree Distributions and Species Tree Inference
Tahmid, N.; Rhythm, S. I.; Bayzid, M. S.Abstract
Accurate species tree inference from genome-scale data is complicated by gene tree discordance, which can arise both from biological processes such as incomplete lineage sorting (ILS) and from technical factors such as gene tree estimation error (GTEE). While both factors reduce the accuracy of summary methods, their relative impact and characteristic patterns remain poorly understood. Here, we systematically compare the effects of ILS and GTEE by simulating gene tree datasets with comparable overall discordance levels, but with discordance arising exclusively from either ILS or GTEE. Using widely employed summary methods such as ASTRAL and wQFM, we show that GTEE typically has a stronger detrimental effect on species tree accuracy than ILS, even at matched discordance levels. We further characterize the structure of gene tree distributions under these two sources of discordance and show that ILS induces a structured, constrained skew in quartet distributions, whereas GTEE generates more uniform, high-entropy noise that does not diminish with additional genes. Our case study on a widely used avian phylogenomic dataset reveals similar distributional patterns across exons, introns, and ultraconserved elements (UCEs), which differ substantially in their levels of phylogenetic signal. A quartet-based analysis of these gene trees further shows that prioritizing loci with stronger and more consistent quartet support can improve the recovery of established avian clades. Overall, these results provide an empirical framework for a nuanced understanding of how ILS and GTEE shape gene tree distributions and influence species tree inference from limited or noisy gene tree datasets.
bioinformatics2026-08-25v2eSkip2 prioritizes exon-skipping antisense oligonucleotide target regions across exon--intron contexts
Chiba, S.; Kunitake, K.; Shirakaki, S.; Haque, U. S.; Wilton-Clark, H.; Shah, M. N. A.; Leckie, J. N.; Matsui, K.; Uno-Ono, F.; Yokota, T.; Aoki, Y.; Okuno, Y.Abstract
Exon-skipping antisense oligonucleotides (ASOs) can restore productive transcripts, but identifying effective binding regions remains difficult because splicing regulation extends across exons, introns and splice junctions. Here we develop eSkip2, a genome-informed framework that ranks target regions within a unified exon-intron sequence context. eSkip2 combines a genome-pretrained sequence model with ASO-induced exon-skipping data and single-nucleotide-variant splicing perturbations, followed by target-locus adaptation that requires no experimental ASO labels from the locus being designed. Across benchmarks comprising canonical exons and pseudoexons, multiple cell types and chemistries, and exonic, intronic and exon-intron-spanning targets, eSkip2 prioritized active regions and showed a higher median AUROC than applicable exon-restricted models. Prospective application to the combinatorial design of dual-targeting ASOs for DMD exon 46 enriched active candidates near the top of the ranking: the two most active new ASOs ranked within the top three and produced dose-dependent dystrophin restoration in patient-derived cells. These results support eSkip2 as a practical first-pass strategy for reducing experimental search space in exon-skipping ASO discovery.
bioinformatics2026-08-25v2Detecting CYP2C19 deletions from genotyping array signals using neural networks
Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.Abstract
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
bioinformatics2026-08-25v1AFP-R: An Open Resource Dedicated to Antifreeze Proteins
Liu, W.; Zhang, Y.; Xiu, D.; Liu, Y.; Wang, T.; Chai, X.; Qu, H.; Min, Y.; Zhang, Z.Abstract
Antifreeze proteins (AFPs), lower the freezing point via thermal hysteresis activity and/or ice recrystallization inhibition, playing a crucial role in protecting organisms from freezing damage under sub-zero milieu. This property endows them with promising applications in biomedicine and agriculture, ranging from tissue-organ cryopreservation to the development of frost-resistant crops. However, the lack of comprehensive resources dedicated for AFPs hinders further progress in elucidating their functional mechanisms and advancing their applications. Here, we report AFP-R, an online resource comprising AFP-DB and AFP-Predictor. AFP-DB is a comprehensive database with manually curated proteins bearing experimentally validated antifreeze activity derived from published literature, whereas AFP-Predictor is a sequence-based machine-learning model to identify AFPs. AFP-DB stores diverse AFP-related information, including sequences, structures, post-translational modifications, taxonomy and annotations of antifreeze-activity experimental assays. It now holds 186 entries, 607 sub-entries, and 1444 experimental records. AFP-Predictor, an AFP-identification algorithm built on protein language model ESM2 (Evolutionary Scale Modeling2), is trained on data in AFP-DB and outperforms several existing models. This work offers a valuable resource for systematically dissecting the mechanisms underlying AFP antifreeze activity and will facilitate their broader applications.
bioinformatics2026-08-25v1PhageLysData: an evidence-aware and AI-ready dataset of phage lytic enzymes and depolymerases
Medina-Ortiz, D.; Olivera-Nappa, A.; Lienqueo, M. E.; Opazo, R.; Romero, J.Abstract
Bacteriophage lytic enzymes and depolymerases are relevant to phage biology, antimicrobial development, and protein engineering, but their sequence and annotation data remain dispersed across general databases, specialized resources, genome-centred collections, and prediction-oriented datasets. We present PhageLysData, an evidence-aware and AI-ready resource constructed through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The release integrates 807,366 source observations from seven primary resources into 759,105 unique exact-sequence entities, comprising an evidence-supported Core of 11,867 entities, a Prediction Extension of 745,092 prediction-only candidates, and 2,146 Context entities retained for provenance and reference. This architecture preserves broad sequence-space coverage while maintaining a clear distinction between non-predictive and prediction-derived support. Core entities are enriched with harmonized biological annotations, physicochemical properties, independent InterProScan-derived functional annotations, mapped PDB and AlphaFold DB structural assets, and reusable numerical representations. For 11,259 eligible Core sequences, PhageLysData provides embeddings from 11 protein language models together with one-hot encoding under a common representation contract. Release-facing examples demonstrate latent-space exploration, unsupervised clustering, supervised classification, and evidence-aware candidate retrieval without defining a universal predictive benchmark. PhageLysData provides a traceable, versioned, and computationally accessible foundation for protein retrieval, comparative analysis, task-specific dataset construction, and machine-learning applications involving phage lytic enzymes and depolymerases.
bioinformatics2026-08-25v1An inflammation-associated five-gene expression signature stratifies survival and immune states in lung adenocarcinoma: an integrative public-cohort analysis
Zhou, X.; Le, Z.; Song, P.; Xu, Q.; Chen, M.; Liu, X.; Cao, M.; Zhan, S.; Liu, Y.; Zhang, L.Abstract
Background: Inflammation and the tumor immune microenvironment contribute to lung adenocarcinoma (LUAD) progression, but the relationship among inflammation-linked transcriptional heterogeneity, patient survival, and immune-state variation remains incompletely defined. Objective: We aimed to identify inflammation-associated LUAD subtypes, derive a parsimonious survival-stratification signature, and characterize its immune and pathway context across public transcriptomic cohorts. Methods: Expression profiles and clinical data were obtained from TCGA-LUAD, GTEx normal lung, and GEO datasets GSE11969, GSE30219, GSE31210, and GSE40791. A curated set of 596 inflammation-related genes was used for consensus clustering. Differential-expression analysis, functional enrichment, univariate Cox regression, and LASSO-Cox modeling were integrated to construct a gene-expression risk score. The prognostic dataset comprised 730 cases and was randomly divided into training (n=502) and internal-validation (n=228) sets; 85 GSE30219 cases formed an external-validation cohort. Immune-cell enrichment, gene set enrichment analysis (GSEA), gene set variation analysis (GSVA), and pan-cancer analyses were used for biological contextualization. Results: The LUAD-versus-control comparison identified 1,305 differentially expressed genes, including 498 upregulated and 807 downregulated genes. Consensus clustering resolved two inflammation-associated subtypes and 67 subtype-associated genes, of which 64 were higher and 3 were lower in Cluster 1 relative to Cluster 2. Thirty-three genes overlapped between the tumor-control and subtype contrasts. LASSO-Cox regression selected CHRDL1, FDCSP, CXCL13, CYP4B1, and S100P. The 1-, 3-, and 5-year areas under the time-dependent receiver operating characteristic curve were 0.6625, 0.6581, and 0.6658 in the training set; 0.7422, 0.6537, and 0.6761 in internal validation; and 0.6560, 0.6387, and 0.6753 in external validation. Risk groups differed across multiple T-cell, B-cell, natural-killer-cell, myeloid, dendritic-cell, macrophage, and granulocyte signatures. Positive GSEA signals included cell cycle (normalized enrichment score [NES]=2.67; adjusted P=1.42 x 10-), DNA replication (NES=2.52; adjusted P=2.52 x 10-), and mismatch repair (NES=2.20; adjusted P=1.77 x 10-). Conclusions: The five-gene expression score separated LUAD survival groups and captured coordinated proliferative and immune transcriptional states. Its moderate discrimination supports further biological and clinical validation rather than immediate clinical application.
bioinformatics2026-08-25v1Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles
Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.Abstract
Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.
bioinformatics2026-08-25v1UELer: a Jupyter-based framework for interactive exploration of multiplexed imaging datasets
Wu, Y.-L.; Liu, C.-S.; Lenoir, B.; Merz, K.; Dill, M. T.; Hartmann, F. J.Abstract
Summary Multiplexed imaging and spatial proteomics generate complex datasets that require both computational analysis and visual inspection. However, these tasks mostly occur in separate environments because interactive viewers generally require a local display or an additional data server beyond the remote Jupyter sessions itself where large datasets are computationally analyzed. We here present UELer, an interactive viewer that links multi-channel image views with quantitative analysis results directly within Jupyter notebooks, requiring no dedicated infrastructure beyond the notebook session. Cells selected through computational analysis and summary plots can be inspected directly in their tissue context, and selections made in the image can be made available to any downstream analysis. Together, these capabilities support interactive data exploration, iterative cell annotation, and reproducible retrieval of selected regions. Availability and Implementation UELer is a Python package built on ipywidgets and runs in Jupyter environments supporting ipywidgets 8.1 or later, tested in JupyterLab and Visual Studio Code on Linux, macOS, and Windows. It is freely available under GPL-3.0 license and can be installed via pip. Source code and documentation are available at https://github.com/HartmannLab/UELer and https://hartmannlab.github.io/UELer/. An online, no-install version runs remotely via BinderHub (https://mybinder.org/v2/gh/HartmannLab/UELer/main), accessible through the script/run_ueler_binder.ipynb notebook.
bioinformatics2026-08-25v1MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data
Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.Abstract
Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.
bioinformatics2026-08-25v1MultiFlow: coupled flow matching for predicting single-cell multiomic perturbation responses in unseen cellular contexts
Wang, H.; Zhang, C.; Zhang, M.; Nie, X.; Liu, Q.Abstract
Predicting cellular responses to perturbation requires resolving coordinated changes across molecular layers, yet most single-cell perturbation models focus on transcriptional responses alone. Here we present MultiFlow, a coupled flow-matching framework that unifies generation and perturbation prediction of paired gene expression and chromatin accessibility. By learning coupled RNA-ATAC flows conditioned on perturbation and control-derived cellular-state representation, MultiFlow enables prediction of coordinated multiomic responses in unseen cellular contexts. Across multiomic generation benchmarks, MultiFlow accurately reproduced paired RNA-ATAC states and their population distributions. In multiomic perturbation benchmarks, MultiFlow achieved the strongest overall performance in predicting both gene-expression and chromatin-accessibility responses, outperforming competing modality-specific perturbation-prediction methods. Joint multiomic modeling further preserved perturbation-induced RNA-ATAC coordination, including concordant peak-gene effects and cross-modal cellular neighborhood structure. These results establish coupled flow matching as a unified generative framework for modeling paired multiomic states and predicting coordinated perturbation responses across cellular contexts. Code and tutorial for MultiFlow are available at https://github.com/liuq-lab/MultiFlow.
bioinformatics2026-08-25v1ClustoCell reveals cell states and their markers from single-cell transcriptomes
Salavaty, A.; Foroutan, M.; Pretel, N. P.; Egelberg, J.; Parish, I. A.; Huntington, N. D.; Beltran, H.; Sandhu, S.; Molania, R.Abstract
Accurate identification of cell types and states is essential for reliable single-cell RNA-sequencing analyses, yet current methods remain sensitive to continuous biological states, data preprocessing choices, and reference selection. Here we present ClustoCell, a reference-free method that resolves cell identity using within-cell transcriptional architecture. By stratifying gene expression of each cell into high and medium tiers, ClustoCell constructs cell-cell similarity graphs that prioritize intrinsic expression structure over global variance. Across 450 datasets spanning over 24 million cells, ClustoCell recovered expert annotations with high concordance (92%). Benchmarked against state-of-the-art methods, ClustoCell identifies more stable and coherent cell types and states, avoids excessive partitioning of closely related cells, and improves the identification of cell type-specific markers. From transcriptional structure alone, ClustoCell resolves rare and transitional cell states, distinguishes malignant from non-malignant cells, and refines expert cell annotations. Applied to immunotherapy datasets, ClustoCell uncovered coordinated pre-treatment immune circuits linking T cell states to PD-1 responsiveness in a tumour-type-specific manner. ClustoCell provides an interpretable and scalable foundation for single-cell analysis and translational profiling.
bioinformatics2026-08-25v1ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation
Moran, J.; Freda, P. J.; Ghosh, A.; Hernandez, M. E.; Moore, J. H.Abstract
Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.
bioinformatics2026-08-25v1HuMMANet: A Harmonized Cross-Study Resource for Integrative Analysis of Human Gut Microbiome Metabolome Associations
Verma, S.; Arora, N.; Ajay, C. P.; Singh, P.; Mallick, H.; Ghosh, T. S.Abstract
Deciphering gut microbiome to host metabolome interaction is critical for understanding how microbial communities generate bioactive signals that shape host physiology and disease. Progress, however, has been hindered by inconsistent metabolite annotations, poor interoperability across studies, and the absence of integrated resources placing microbiome-derived metabolites within their functional, microbial, physiological, and clinical context. Here we present HuMMANet (Human Microbiome Metabolome Annotation Network), a harmonized resource integrating 46 paired gut microbiome metabolome studies (59 study-units; 14,405 samples; 13 disease categories plus a healthy/control reference category) with a scalable metabolite-harmonization framework. HuMMANet resolves heterogeneous annotations through a multi-stage workflow spanning RefMet, HMDB, PubChem, Metabolomics Workbench, SMPDB, MiMeDB 2.0, GNPS/microbeMASST, DrugBank, and DrugCentral, yielding a reference atlas of 54,914 unique metabolites, annotated with standardized chemical identifiers, biochemical pathways, microbial producer associations, physiological distributions, disease links, and structural relationships to approved therapeutics, a unified reference framework for microbiome metabolome research. Applying HuMMANet to a multi-cohort integration of adult serum and fecal metabolomes, we identified 519 serum and 322 fecal metabolites reproducibly associated with gut microbial community composition (PERMANOVA, P < 0.05 in at least 50% of studies in which detected), enriched for specific biomolecular classes and pathways. Cross-referencing these against Health Associated Core Keystone (HACK) taxa revealed 58 serum and 25 fecal metabolites (HACK positive) whose taxon-level associations tracked positively with the taxon specific HACK indices. These reproducible metabolomic signatures of microbiome health included indole3propionic acid, a gut barrier-protective microbial tryptophan metabolite, and 3phenylpropionate. Drug similarity annotation within HuMMANet linked 16 of this serum and 13 fecal HACK positive metabolites to therapeutics used in neurological, inflammatory, and vascular disease. Conversely, 38 serum and 65 fecal metabolites, including imidazole propionate and long-chain acylcarnitines such as ACar 18:0, showed HACK negative signatures previously associated with dysbiosis-linked disease. GNPS/microbeMASST and MiMeDB 2.0 annotations further traced subsets of these metabolites to putative bacterial producers. HuMMANet thus provides a standardized framework for reproducible microbiome metabolome integration, enabling cross study discovery and translational prioritization of conserved microbiome derived metabolic signatures across human populations and disease states.
bioinformatics2026-08-25v1Addressing technical variations in ATAC-seq data and improving motif accessibility analyses
Wang, J.; Sonder, E.; Domcke, S.; Robinson, M. D.; Germain, P.-L.Abstract
Tagmentation-based methods such as ATAC-seq and Cut&Tag have provided easy ways to profile the epigenome in low-input samples and even single cells. In this contribution, we discuss forms of bias (i.e. technical variations) in tagmentation-based data, in particular ATAC-seq, and introduce three R/bioconductor packages to facilitate bulk and single-cell epigenomic data analysis, with a special focus on motif accessibility analysis. The weightedMotifAccess package uses weight models to enable motif accessibility analysis, including transcription factor footprint information. The betterChromVAR package provides a novel, analytical re-implementation of the popular chromVAR method that offers substantial speed improvements, eliminates stochasticity, and offers additional features. Based on this, we also propose a method, CVnorm, that outperforms alternatives in normalizing technical bias in peak count data. The computational efficiency of these tools further enables a new framework for systematically investigating synergistic and antagonistic interactions between transcription factor motifs. Finally, the epiwraps package streamlines the visualization, normalization, and summarization of epigenomic data.
bioinformatics2026-08-25v1NetSyn: prokaryotic genomic context exploration of protein families
Stam, M.; Langlois, j.; Chevalier, C.; Mainguy, J.; Reboul, G.; Bastard, K.; Medigue, C.; Vallenet, D.Abstract
Background: The growing availability of large prokaryotic genomic datasets presents an opportunity to discover new metabolic pathways and enzymatic reactions useful for industrial or synthetic biological applications. Efforts to identify new enzyme functions in this vast number of sequences cannot be achieved without bioinformatics tools and the development of new strategies. Standard methods for assigning a biological function to a gene are based on sequence similarity. However, complementary approaches rely on mine databases to identify conserved gene clusters (i.e. syntenies). In prokaryotic genomes, genes involved in the same pathway are frequently encoded in a single locus with an operonic organisation. This genomic context conservation is considered as a reliable indicator of functional relationships, and is therefore a promising approach for improving the gene function prediction. Methods. Here we present NetSyn (Network Synteny), a tool to group protein sequences based on the conservation of their genomic context rather than solely on sequence similarity. From a list of protein sequence identifiers, NetSyn searches corresponding genome entries to retrieve neighboring genes. Corresponding protein sequences are grouped into families to define homology relationships and compute a synteny conservation score between the different extracted genomic contexts. A network is then created in which the nodes represent the input proteins and the edges indicate that two proteins share a conserved synteny. Finally, the network is partitioned into clusters grouping proteins with similar genomic contexts, using a community detection algorithm. Results. As a proof of concept, we used NetSyn on two different datasets. The first one is the BKACE protein family (formerly named DUF849) which has previously been divided into isofunctional sub-families. NetSyn was able to go a step further by providing additional sub-families beyond those already described. The second dataset corresponds to a set of non-homologous proteins belonging to three different glycoside hydrolase (GH) families. These GHs are known to work cooperatively in a Polysaccharide-Utilization Loci (PUL) and are therefore grouped together in the same genomic contexts. NetSyn was able to identify a locus grouping 3 GHs, involved in the degradation of xyloglucan, in 162 prokaryotic genomes. Discussion. By highlighting conserved synteny in distantly related prokaryotic species, NetSyn enables functional links between proteins to be established beyond sequence similarity alone. We showed that NetSyn is efficient for exploring large prokaryotic protein families, enabling the definition of isofunctional groups and the identification of functional interactions between non-homologous enzymes. These features enable the prediction of new genomic structures that have not yet been experimentally characterized. Finally, NetSyn is also useful for pinpointing annotation errors that have been propagated across databases, and for suggesting annotations on proteins lacking functional prediction. NetSyn is freely available at https://github.com/labgem/netsyn.
bioinformatics2026-08-24v5Scaling genome annotation across the eukaryotic tree of life with OrionGeno
Liu, L.; Cai, X.; Wang, S.; Deng, Y.; Wu, Y.; Pan, Y.; Wang, J.; Zhang, C.; Xia, H.; Tan, N.; Su, K.; Liu, Y.; Zhou, X.; Liu, L.; Wei, T.; Zhang, Y.; Li, Q.; Li, Y.; Yin, P.; Xu, X.Abstract
The rapid expansion of eukaryotic genome sequencing has created an urgent demand for accurate and scalable genome annotation. Existing ab initio methods often struggle to reconstruct complex gene architectures and generalize across distant lineages, limiting their use for large-scale annotation. Here we present OrionGeno, a phylogeny-aware deep learning model for end-to-end eukaryotic genome annotation. OrionGeno integrates phylogenetic context, long-range sequence modeling and joint prediction of gene structures and repetitive elements to annotate exons, introns, untranslated regions and repeats directly from genomic sequences. Applied to chromosome-level eukaryotic genomes from NCBI that lack annotations, OrionGeno generates annotations for more than 5,300 genomes, substantially expanding public annotation resources. Across diverse eukaryotic lineages, OrionGeno outperforms state-of-the-art methods at the exon, gene, protein-sequence, and protein-structural levels. It also identifies candidate protein-coding loci absent from reference protein-coding annotations in well-curated genomes. Together with a web platform and integrated annotation database, OrionGeno provides a scalable and accessible framework for translating genome assemblies into functional biological resources and supporting large-scale biodiversity initiatives such as the Earth BioGenome Project.
bioinformatics2026-08-24v2Interpolating and Extrapolating Node Counts in Colored Compacted de Bruijn Graphs for Pangenome Diversity
Parmigiani, L.; Peterlongo, P.Abstract
A pangenome is a collection of taxonomically related genomes, often from the same species, serving as a representation of their genomic diversity. The study of pangenomes, or pangenomics, aims to quantify and compare this diversity, which has significant relevance in fields such as medicine and biology. Originally conceptualized as sets of genes, pangenomes are now commonly represented as pangenome graphs. These graphs consist of nodes representing genomic sequences and edges connecting consecutive sequences within a genome. Among possible pangenome graphs, a common option is the compacted de Bruijn graph. In our work, we focus on the colored compacted de Bruijn graph, where each node is associated with a set of colors that indicate the genomes traversing it. In response to the evolution of pangenome representation, we introduce a novel method for comparing pangenomes by their node counts, addressing two main challenges: the variability in node counts arising from graphs constructed with different numbers of genomes, and the large influence of rare genomic sequences. We propose an approach for interpolating and extrapolating node counts in colored compacted de Bruijn graphs, adjusting for the number of genomes. To tackle the influence of rare genomic sequences, we apply Hill numbers, a well-established diversity index previously utilized in ecology and metagenomics for similar purposes, to proportionally weight both rare and common nodes according to the frequency of genomes traversing them.
bioinformatics2026-08-24v2Integration of proteomic data from cell lines and tumors
Ta, C. Q.; Auth, J. M.; Schilling, M.; Klingmueller, U.; Raue, A.Abstract
Cancer cell lines are widely used in preclinical research, yet the clinical translation of findings from cell lines remains limited. Identifying cell lines that best resemble patient tumors requires integration of molecular profiles across biologically distinct sample types. Advances in transcriptomic integration have demonstrated the potential of deep learning for aligning data across different sample types. However, comparable approaches for proteomic data integration remain lacking, potentially because of the prevalence of missing values in proteomic datasets. Here, we introduce ProtInt, a deep learning-based framework that integrates proteomic data by combining principles from proteomic imputation and transcriptomic integration methods. We applied ProtInt to integrate label-free proteomic profiles from 771 cancer cell lines and 550 treatment-naive tumors, and showed that ProtInt outperformed batch correction and transcriptomic integration methods. Comparison of the cell line proteomes before and after integration revealed recurrent increase of proteins associated with immune reaction and reduction of proteins involved in mitochondrial gene expression as proteomes of cell lines were adapted to resemble tumors. These results establish ProtInt as a framework for joint analysis of proteomic datasets across distinct sample types and may facilitate the identification of cell lines best suited for clinically relevant studies.
bioinformatics2026-08-24v2Virtual-cell models compress unseen intervention geometry through a target-specific generalization bottleneck
Huang, Y.; Wang, H.; Wilson, P. C.Abstract
Predictive models of cellular perturbation are often judged by how closely they reconstruct molecular states after unseen interventions. We show that high state-level similarity can coexist with loss of the relationships that distinguish perturbations, a failure we term Intervention Geometry Compression (IGC). Across established models and perturbation settings, unseen interventions show weakened global and local geometry, reduced between-intervention variance and spectral collapse. The failure is not primarily explained by response-space capacity. Instead, diagnostic projections localize much of the missing geometry to a small number of residual response directions learned from seen interventions; these directions outperform complexity-matched random subspaces and replicate in an independent Jiang perturbation resource. Polarity captures part, but not all, of this continuous orientation signal. Time-resolved analyses further show that correct trajectory entry markedly improves downstream propagation, while a held target's own early empirical response rapidly reveals endpoint orientation. Finally, same-target empirical anchoring transfers intervention identity across contexts far more effectively than increasing exposure to other interventions. These results identify intervention-coordinate assignment as an information bottleneck in virtual-cell generalization and support a design principle: empirically anchor intervention identity, then use models to generalize anchored effects across cellular contexts.
bioinformatics2026-08-24v1Thal-Kak: unifying biomolecular structure predictors reveals a sampling-selection gap
Bae, J.; Jo, S.; Kim, Y.; Kim, D.; Kim, K.; Park, S.; Park, S.; Myung, S.; Shin, H.; Kim, M. H.; Kang, M.; Baek, M.Abstract
Complementary all-atom structure predictors sample different solutions, but how to allocate a fixed sampling budget across them and select the best output remains unclear. Thal-Kak unifies five released predictors under shared upstream inputs and a common schema. Across FoldBench and CASP16, model mixing improves oracle sampling over single-model runs, but selection remains a bottleneck because confidence scores do not transfer across models and existing quality-assessment methods cannot resolve this gap.
bioinformatics2026-08-24v1Point-in-time evidence and cross-area clinical precedent anticipate clinical entry across 100 focal areas: retrospective validation of the Intangia triage layer
Elliott, T. O.; Molnar, S.; Peeters, G.; Collart, O.Abstract
Early-opportunity teams face a combinatorial problem: once a focal target, mechanism or indication is fixed, the space of plausible partners runs to thousands of candidates per area. Intangia's triage layer ranks that space from point-in-time evidence (how much literature, patent and clinical activity a candidate pairing has accumulated, and whether the partner already has clinical precedent in other contexts) so that review starts where clinical activity is most likely to begin next. This preprint validates that capability retrospectively across 100 focal areas spanning drug targets, mechanisms and disease indications, replaying 24.1 million historically scored combination-years with every area scored by a model trained on the other 99 and never on itself. The headline is operational. At a twenty-partner review shortlist per focal area, the median area's four-year first-alert precision is 0.234, against a matched random-ranker median of 0.008: roughly one in four shortlisted partners subsequently entered the focal clinical context within four years, about 38 times each area's own background rate (95% CI 31 to 45). A panel-level permutation puts the result at p = 0.0005. Discrimination generalises: the full 13-feature specification reaches a median leave-one-focal-out ROC-AUC of 0.922 (95% CI 0.911 to 0.929), with no area below chance and all 100 areas beating their strongest count-based baseline. Shortlisted entrants are anticipated with a median observed lead of two years within the evaluation window, and three years (interquartile range one to five) once the window cap is removed and every realised entrant is counted. The core ranking is carried by two interpretable signal families: cumulative co-occurrence counts and leave-one-area-out clinical precedent. Burst detection serves a complementary role: it supplies the time-stamped, source-specific momentum evidence attached to every recommendation (what is accelerating, and why now) rather than additional ranking power. A conditional view of the same landscape ranks candidates with no cross-area precedent against one another, enriched relative to matched random ranking, supporting a lower-yield emerging-opportunities capability. Two worked examples, PD-1 combination immunotherapy and CTLA-4, are point-in-time historical replays of the same architecture in familiar territory, showing what an alert looked like with the dated evidence behind it. The endpoint throughout is first clinical entry, not clinical success; prospective validation is the next stage.
bioinformatics2026-08-24v1