Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
NetSyn: prokaryotic genomic context exploration of protein families
Stam, M.; Langlois, j.; Chevalier, C.; Mainguy, J.; Reboul, G.; Bastard, K.; Medigue, C.; Vallenet, D.Abstract
Background: The growing availability of large prokaryotic genomic datasets presents an opportunity to discover new metabolic pathways and enzymatic reactions useful for industrial or synthetic biological applications. Efforts to identify new enzyme functions in this vast number of sequences cannot be achieved without bioinformatics tools and the development of new strategies. Standard methods for assigning a biological function to a gene are based on sequence similarity. However, complementary approaches rely on mine databases to identify conserved gene clusters (i.e. syntenies). In prokaryotic genomes, genes involved in the same pathway are frequently encoded in a single locus with an operonic organisation. This genomic context conservation is considered as a reliable indicator of functional relationships, and is therefore a promising approach for improving the gene function prediction. Methods. Here we present NetSyn (Network Synteny), a tool to group protein sequences based on the conservation of their genomic context rather than solely on sequence similarity. From a list of protein sequence identifiers, NetSyn searches corresponding genome entries to retrieve neighboring genes. Corresponding protein sequences are grouped into families to define homology relationships and compute a synteny conservation score between the different extracted genomic contexts. A network is then created in which the nodes represent the input proteins and the edges indicate that two proteins share a conserved synteny. Finally, the network is partitioned into clusters grouping proteins with similar genomic contexts, using a community detection algorithm. Results. As a proof of concept, we used NetSyn on two different datasets. The first one is the BKACE protein family (formerly named DUF849) which has previously been divided into isofunctional sub-families. NetSyn was able to go a step further by providing additional sub-families beyond those already described. The second dataset corresponds to a set of non-homologous proteins belonging to three different glycoside hydrolase (GH) families. These GHs are known to work cooperatively in a Polysaccharide-Utilization Loci (PUL) and are therefore grouped together in the same genomic contexts. NetSyn was able to identify a locus grouping 3 GHs, involved in the degradation of xyloglucan, in 162 prokaryotic genomes. Discussion. By highlighting conserved synteny in distantly related prokaryotic species, NetSyn enables functional links between proteins to be established beyond sequence similarity alone. We showed that NetSyn is efficient for exploring large prokaryotic protein families, enabling the definition of isofunctional groups and the identification of functional interactions between non-homologous enzymes. These features enable the prediction of new genomic structures that have not yet been experimentally characterized. Finally, NetSyn is also useful for pinpointing annotation errors that have been propagated across databases, and for suggesting annotations on proteins lacking functional prediction. NetSyn is freely available at https://github.com/labgem/netsyn.
bioinformatics2026-08-24v5Scaling genome annotation across the eukaryotic tree of life with OrionGeno
Liu, L.; Cai, X.; Wang, S.; Deng, Y.; Wu, Y.; Pan, Y.; Wang, J.; Zhang, C.; Xia, H.; Tan, N.; Su, K.; Liu, Y.; Zhou, X.; Liu, L.; Wei, T.; Zhang, Y.; Li, Q.; Li, Y.; Yin, P.; Xu, X.Abstract
The rapid expansion of eukaryotic genome sequencing has created an urgent demand for accurate and scalable genome annotation. Existing ab initio methods often struggle to reconstruct complex gene architectures and generalize across distant lineages, limiting their use for large-scale annotation. Here we present OrionGeno, a phylogeny-aware deep learning model for end-to-end eukaryotic genome annotation. OrionGeno integrates phylogenetic context, long-range sequence modeling and joint prediction of gene structures and repetitive elements to annotate exons, introns, untranslated regions and repeats directly from genomic sequences. Applied to chromosome-level eukaryotic genomes from NCBI that lack annotations, OrionGeno generates annotations for more than 5,300 genomes, substantially expanding public annotation resources. Across diverse eukaryotic lineages, OrionGeno outperforms state-of-the-art methods at the exon, gene, protein-sequence, and protein-structural levels. It also identifies candidate protein-coding loci absent from reference protein-coding annotations in well-curated genomes. Together with a web platform and integrated annotation database, OrionGeno provides a scalable and accessible framework for translating genome assemblies into functional biological resources and supporting large-scale biodiversity initiatives such as the Earth BioGenome Project.
bioinformatics2026-08-24v2Interpolating and Extrapolating Node Counts in Colored Compacted de Bruijn Graphs for Pangenome Diversity
Parmigiani, L.; Peterlongo, P.Abstract
A pangenome is a collection of taxonomically related genomes, often from the same species, serving as a representation of their genomic diversity. The study of pangenomes, or pangenomics, aims to quantify and compare this diversity, which has significant relevance in fields such as medicine and biology. Originally conceptualized as sets of genes, pangenomes are now commonly represented as pangenome graphs. These graphs consist of nodes representing genomic sequences and edges connecting consecutive sequences within a genome. Among possible pangenome graphs, a common option is the compacted de Bruijn graph. In our work, we focus on the colored compacted de Bruijn graph, where each node is associated with a set of colors that indicate the genomes traversing it. In response to the evolution of pangenome representation, we introduce a novel method for comparing pangenomes by their node counts, addressing two main challenges: the variability in node counts arising from graphs constructed with different numbers of genomes, and the large influence of rare genomic sequences. We propose an approach for interpolating and extrapolating node counts in colored compacted de Bruijn graphs, adjusting for the number of genomes. To tackle the influence of rare genomic sequences, we apply Hill numbers, a well-established diversity index previously utilized in ecology and metagenomics for similar purposes, to proportionally weight both rare and common nodes according to the frequency of genomes traversing them.
bioinformatics2026-08-24v2Integration of proteomic data from cell lines and tumors
Ta, C. Q.; Auth, J. M.; Schilling, M.; Klingmueller, U.; Raue, A.Abstract
Cancer cell lines are widely used in preclinical research, yet the clinical translation of findings from cell lines remains limited. Identifying cell lines that best resemble patient tumors requires integration of molecular profiles across biologically distinct sample types. Advances in transcriptomic integration have demonstrated the potential of deep learning for aligning data across different sample types. However, comparable approaches for proteomic data integration remain lacking, potentially because of the prevalence of missing values in proteomic datasets. Here, we introduce ProtInt, a deep learning-based framework that integrates proteomic data by combining principles from proteomic imputation and transcriptomic integration methods. We applied ProtInt to integrate label-free proteomic profiles from 771 cancer cell lines and 550 treatment-naive tumors, and showed that ProtInt outperformed batch correction and transcriptomic integration methods. Comparison of the cell line proteomes before and after integration revealed recurrent increase of proteins associated with immune reaction and reduction of proteins involved in mitochondrial gene expression as proteomes of cell lines were adapted to resemble tumors. These results establish ProtInt as a framework for joint analysis of proteomic datasets across distinct sample types and may facilitate the identification of cell lines best suited for clinically relevant studies.
bioinformatics2026-08-24v2Virtual-cell models compress unseen intervention geometry through a target-specific generalization bottleneck
Huang, Y.; Wang, H.; Wilson, P. C.Abstract
Predictive models of cellular perturbation are often judged by how closely they reconstruct molecular states after unseen interventions. We show that high state-level similarity can coexist with loss of the relationships that distinguish perturbations, a failure we term Intervention Geometry Compression (IGC). Across established models and perturbation settings, unseen interventions show weakened global and local geometry, reduced between-intervention variance and spectral collapse. The failure is not primarily explained by response-space capacity. Instead, diagnostic projections localize much of the missing geometry to a small number of residual response directions learned from seen interventions; these directions outperform complexity-matched random subspaces and replicate in an independent Jiang perturbation resource. Polarity captures part, but not all, of this continuous orientation signal. Time-resolved analyses further show that correct trajectory entry markedly improves downstream propagation, while a held target's own early empirical response rapidly reveals endpoint orientation. Finally, same-target empirical anchoring transfers intervention identity across contexts far more effectively than increasing exposure to other interventions. These results identify intervention-coordinate assignment as an information bottleneck in virtual-cell generalization and support a design principle: empirically anchor intervention identity, then use models to generalize anchored effects across cellular contexts.
bioinformatics2026-08-24v1Thal-Kak: unifying biomolecular structure predictors reveals a sampling-selection gap
Bae, J.; Jo, S.; Kim, Y.; Kim, D.; Kim, K.; Park, S.; Park, S.; Myung, S.; Shin, H.; Kim, M. H.; Kang, M.; Baek, M.Abstract
Complementary all-atom structure predictors sample different solutions, but how to allocate a fixed sampling budget across them and select the best output remains unclear. Thal-Kak unifies five released predictors under shared upstream inputs and a common schema. Across FoldBench and CASP16, model mixing improves oracle sampling over single-model runs, but selection remains a bottleneck because confidence scores do not transfer across models and existing quality-assessment methods cannot resolve this gap.
bioinformatics2026-08-24v1Point-in-time evidence and cross-area clinical precedent anticipate clinical entry across 100 focal areas: retrospective validation of the Intangia triage layer
Elliott, T. O.; Molnar, S.; Peeters, G.; Collart, O.Abstract
Early-opportunity teams face a combinatorial problem: once a focal target, mechanism or indication is fixed, the space of plausible partners runs to thousands of candidates per area. Intangia's triage layer ranks that space from point-in-time evidence (how much literature, patent and clinical activity a candidate pairing has accumulated, and whether the partner already has clinical precedent in other contexts) so that review starts where clinical activity is most likely to begin next. This preprint validates that capability retrospectively across 100 focal areas spanning drug targets, mechanisms and disease indications, replaying 24.1 million historically scored combination-years with every area scored by a model trained on the other 99 and never on itself. The headline is operational. At a twenty-partner review shortlist per focal area, the median area's four-year first-alert precision is 0.234, against a matched random-ranker median of 0.008: roughly one in four shortlisted partners subsequently entered the focal clinical context within four years, about 38 times each area's own background rate (95% CI 31 to 45). A panel-level permutation puts the result at p = 0.0005. Discrimination generalises: the full 13-feature specification reaches a median leave-one-focal-out ROC-AUC of 0.922 (95% CI 0.911 to 0.929), with no area below chance and all 100 areas beating their strongest count-based baseline. Shortlisted entrants are anticipated with a median observed lead of two years within the evaluation window, and three years (interquartile range one to five) once the window cap is removed and every realised entrant is counted. The core ranking is carried by two interpretable signal families: cumulative co-occurrence counts and leave-one-area-out clinical precedent. Burst detection serves a complementary role: it supplies the time-stamped, source-specific momentum evidence attached to every recommendation (what is accelerating, and why now) rather than additional ranking power. A conditional view of the same landscape ranks candidates with no cross-area precedent against one another, enriched relative to matched random ranking, supporting a lower-yield emerging-opportunities capability. Two worked examples, PD-1 combination immunotherapy and CTLA-4, are point-in-time historical replays of the same architecture in familiar territory, showing what an alert looked like with the dated evidence behind it. The endpoint throughout is first clinical entry, not clinical success; prospective validation is the next stage.
bioinformatics2026-08-24v1Erosion of regenerative regulation: age-associated shifts in the skeletal muscle fiber epigenome and transcriptome
Moo, K. G.; Orchard, P.; Varshney, A.; D'Oliveira Albanus, R.; Manickam, N.; Kinnunen, L.; Lakka, T.; Saramies, J.; Laakso, M.; Tuomilehto, J.; Mohlke, K.; Boehnke, M.; Scott, L.; Koistinen, H.; Collins, F.; Parker, S.Abstract
Skeletal muscle aging is characterized by the deterioration of muscle function, which can lead to negative quality-of-life outcomes including frailty and sarcopenia. While understanding the mechanisms of this process is increasingly important as the global population ages, previous molecular studies of skeletal muscle aging have been limited by statistical power and cell type resolution. In this study, we analyzed single-nucleus gene expression and chromatin accessibility data from 287 human skeletal muscle samples from individuals aged 20-79 years to explore sex- and cell type- specific aging effects. Across 467,126 nuclei from 13 cell types, we identify 384 age-associated genes and 4,061 age-associated chromatin regions. These age-associated molecular features are enriched for functional pathways, including metabolic processes, cell-to-cell communication, and senescence Kyoto Encyclopedia of Genes and Genomes KEGG terms. Age-associated closing chromatin was more common across fiber types and sexes than opening chromatin, and was enriched in active enhancer regions while depleted for active transcription start sites. We observe enrichment for specific transcription factor motifs in closing chromatin, including those of glucocorticoid and androgen receptors, both of which play a key role in the maintenance of healthy skeletal muscle. Together, these findings identify an age-associated regulatory shift, largely invisible in matched transcriptomic data, characterized by closing chromatin which reduces accessibility to hormone receptor binding sites and enhancer regions in the muscle fiber epigenome.
bioinformatics2026-08-24v1Click-Prep: An Interactive Data Preparation Tool for Click-qPCR
Kubota, A.; Tajima, A.Abstract
Click-qPCR is a browser-based application for relative qPCR analysis that requires a tidy-format CSV file containing four columns: sample, group, gene, and Cq. Preparing this input from qPCR instrument output typically requires manual reformatting and calculation of mean Cq values for technical replicates. To simplify this process, we developed Click-Prep (https://kubo-azu.shinyapps.io/Click-Prep/), an interactive web-based application designed specifically to create Click-qPCR input files. Click-Prep imports CSV, TXT, TSV, and XLS/XLSX files and supports skipping of instrument-generated metadata rows, interactive column mapping, and manual assignment of experimental groups. Users can review technical-replicate measurements, exclude selected rows according to predefined quality-control criteria, and calculate mean Cq values for each sample-group-target combination. Missing or nonnumeric Cq values are flagged for review and must be resolved before the mean is calculated. Click-Prep can also combine compatible formatted CSV files, such as datasets obtained from separate qPCR plates. The resulting dataset is exported as a standardized CSV file containing the four fields required by Click-qPCR. By integrating these operations into a guided browser-based workflow, Click-Prep enables users to prepare Click-qPCR input files rapidly and consistently without programming.
bioinformatics2026-08-24v1Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery
Seal, S.; Zalte, A. S.; Araripe, D. A.; Gomes, R. A.; Korani, D.; Shekhar, M.; Siramshetty, V. B.; Patra, A.; Mou, Z.; Yu, X.; Kuhn, D.; Weskamp, N.; Ash, J.; Cheng, A. C.; Fang, C.; Price, D.; Aldeghi, M.; Rodriguez-Perez, R.; Clevert, D.-A.; Engkvist, O.; Deibler, K.; Rouquie, D.; Reutlinger, M.; Richmond, N. J.; Ainsley, J.; Ledeboer, M.; Green, W. H.; Bender, A.; Wognum, C.Abstract
Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.
bioinformatics2026-08-24v1RevPert: ranking candidate drivers of transcriptomic state transitions via gallery-native reverse perturbation
Liang, S.; Yang, C.; Wang, J.; Li, y.Abstract
Cellular state transitions underlie adaptation, ageing and disease, yet prioritizing catalogued genetic perturbations whose expression signatures match an observed transcriptomic shift remains difficult. Most models predict phenotype from a nominated intervention, whereas genetic inverse benchmarks are largely restricted to within-screen identity recovery. Here we introduce RevPert, a gallery-native reverse perturbation model that ranks a fixed genetic catalog for a query contrast {Delta}Y* = YB - YA by combining signed Pearson connectivity with a learned residual. Across Replogle Essential Perturb-seq (four lines) and LINCS-KO screens (ten lines), RevPert recovered held-out interventions at leading performance relative to matched baselines. Applied to public drug-resistance contrasts in HCC and CML, dual-arm ranking placed pre-specified disease anchors far higher on the expected arms than ranking the same signatures by differential-expression magnitude alone (Essential residual model for HCC; a transductive GWPS residual for CML). RevPert therefore couples within-screen reverse ranking to a screen-external signed-geometry check; the latter calibrates literature anchors and is not claimed as held-out recovery.
bioinformatics2026-08-24v1IMMF: An Interpretable Multi-Modal Framework for Hypothesis-Driven Biomarker Discovery in Triple-Negative Breast Cancer Using Public Data
Imran, A.; Rahat Hossain, K. M.; Islam, S. M. R.; Rahman, M. S.Abstract
Triple-Negative Breast Cancer (TNBC) is characterized by high heterogeneity, poor prognosis, and limited targeted treatment options. Bridging the gap between molecular alterations and histopathological morphology remains a major challenge in precision oncology. We propose an interpretable, multi-modal framework that integrates histopathological image analysis with multi-omics profiling (somatic mutations, DNA methylation, copy number alterations), leveraging U-Net-based nuclei segmentation, vision-language models (BLIP), biomedical language models (BioGPT), and explainable AI (SHAP, LIME). Our framework achieves strong predictive performance (AUC = 0.989) and provides transparent, biologically grounded interpretations by integrating morphological features with genomically prioritized biomarkers. Cross-modal analysis confirms established TNBC drivers and generates novel, testable hypotheses associating specific epigenetic alterations with distinct morphological phenotypes. While causal validation requires future wet-lab experiments, our framework accelerates hypothesis-driven biomarker discovery by integrating complementary data modalities with language-based reasoning, providing a transparent foundation for hypothesis generation and clinical translation.
bioinformatics2026-08-24v1A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models
Li, L.; Duan, S.; Zha, X.; Ye, F.; Zhang, Y.; Zhang, X.; Cao, Y.; Liu, C.Abstract
Single-cell drug perturbation models are increasingly used to predict how compounds remodel cellular states, but they are still largely assessed by expression reconstruction. Whether high expression similarity reflects preservation of drug-response signatures remains unclear. Here we present scDrugPerturb-Bench, a mechanism-annotated benchmark that links matched control and drug-treated single-cell RNA-sequencing profiles to literature-curated directional key-gene evidence. The resource covers 181 datasets, 423 annotated response cases, 717 unique key genes and 2.5 million cells. We introduce the Mechanism Fidelity Score (MFS) to evaluate key-gene direction, effect-size recovery, gene-set coherence, mechanism specificity and pathway-level response polarity. Across 12 perturbation-prediction models, 3 baselines and 10 data splits, expression-similarity metrics were weakly aligned with MFS and selected different model configurations. Mechanism-aware selection improved early drug retrieval in a transcriptome-based drug design evaluation, indicating that MFS provides practical information beyond benchmark reporting. Systematic benchmarking revealed limited fidelity to drug-response signatures across cell-line and source-integrated settings. Frozen single-cell foundation model embeddings produced local, metric-dependent gains rather than universal improvements, and source context substantially reshaped model assessment. Hard-negative tests further showed that plausible perturbation responses can arise from non-specific transcriptional shortcuts. These results show that expression reconstruction is an insufficient proxy for preserving drug-response signatures and establish scDrugPerturb-Bench as a benchmark for mechanism-aware evaluation of single-cell drug perturbation models.
bioinformatics2026-08-24v1Gene identity, not variant effect, dominates ClinVar benchmarks of missense pathogenicity predictors
Arnoult, D.; Harrizi, S.; Nait Irahal, I.; Mostafa, K.Abstract
Missense pathogenicity predictors are routinely benchmarked against ClinVar, whose labels are strongly structured by gene: genes under diagnostic scrutiny accumulate pathogenic submissions while incidentally sequenced genes accumulate benign ones. We asked how much of a benchmark score this structure alone can produce. On 197,904 ClinVar missense variants validated against UniProt canonical sequences, a null model using no variant-level information, scoring each variant only by the pathogenic fraction of its own gene, reaches an area under the receiver operating characteristic curve (AUROC) of 0.921 under a random 10-fold split. On a common intersection of 169,989 variants, four current predictors exceed it by only 0.036 to 0.044. The inflation is not uniform, so it does not cancel when predictors are compared: under within-gene evaluation the ranking inverts, AlphaMissense rising from third to first and gMVP falling to third (p < 0.0001). The inversion survives removal of ceiling genes and replicates on an independently curated benchmark. Because both rankings derive from the same ClinVar labels, we arbitrated between them using data with no gene-level structure: agreement with 47 human deep mutational scanning assays matches the within-gene ranking and inverts the conventional one (p = 0.027, 0.0023). Across twenty-two dbNSFP predictors scored on one common intersection of 112,248 variants, with each tool's exposure to clinical labels registered before any score was extracted, predictors never trained on such labels sit 0.051 AUROC behind supervised ones globally but only 0.026 behind within genes (difference +0.025 [+0.023, +0.027], p < 0.0001). Leave-one-out correction, the standard remedy, is worth 0.002 AUROC. Much of ClinVar benchmark performance reflects gene identity rather than variant effect, and the distortion changes which predictor a benchmark ranks first, in a direction experimental data contradicts. We release genenull, a single-file implementation, so reporting this baseline costs one function call.
bioinformatics2026-08-24v1A Beta-Binomial Model for Estimating Zero- or One-inflated Pain Trajectories
Liu, Y.; Harris, R. E.; Clauw, D.; Bayman, E.; Leroux, A.; Lindquist, M. A.Abstract
Chronic pain is a widespread public health issue that imposes substantial health, emotional, and economic burdens on individuals and communities. Because pain is subjective and lacks objective biomarkers, it is typically measured using patient-reported scores, often on a numerical scale from zero to ten. Increasingly, pain studies use ecological momentary assessment, with multiple daily assessments over days and across study phases (e.g., a series of baseline and post-intervention assessments). These data frequently show many ratings at the extremes (i.e., at minimum or maximum pain scores), commonly referred to as zero- and one-inflation in the statistical literature, along with considerable within-person variability both within and across days. These phenomena present challenges for statistical analyses, as they violate assumptions of most commonly used statistical techniques (e.g., the normality assumption of linear mixed models). We propose a Bayesian beta-binomial mixed-effects model for modeling potential zero- or one-inflated pain scores while accounting for variability using random effects on the mean and variance parameters across subjects. A simulation study demonstrates that the method accurately estimates model parameters across realistic sample sizes, time points, and zero- and one-inflation levels. An application to data from two longitudinal pain studies demonstrates that the model fits the data better and, when correctly specified, yields accurate uncertainty intervals for longitudinal changes in pain compared to existing models, especially for zero- and one-inflated outcomes. Additionally, the model directly estimates the probability of clinically meaningful pain events. The proposed method provides a powerful statistical framework for studying the patient-reported pain trajectories.
bioinformatics2026-08-23v3EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-08-23v3Sparse autoencoder features from InterPLM predict neuropeptide precursors among secreted proteins
Kulikova, A. V.; Bookout, A. L.; Koch, T. L.; Safavi-Hemami, H.Abstract
Neuropeptides are a diverse class of short, secreted signaling molecules that regulate key physiological processes in animals. Despite their important biological roles and increasingly recognized therapeutic value, the discovery of new neuropeptides remains challenging, largely because their short length and high sequence heterogeneity limit the effectiveness of motif- and homology-based approaches. Here, we present a pipeline for neuropeptide precursor prediction that leverages sparse autoencoders (SAEs) from InterPLM to decode dense ESM-2 protein language model embeddings into sparse, disentangled features. We identify a small subset of features strongly associated with neuropeptide precursors that achieve high discriminative performance. A logistic regression classifier trained on this reduced feature set, accurately separates human neuropeptide and non-neuropeptide sequences. We then applied this classifier to important model organisms: mouse (Mus musculus), zebrafish (Danio rerio), nematode (Caenorhabditis elegans), and fruit fly (Drosophila melanogaster ) and show that the approach generalizes across diverse species. Overall, InterPLM SAE features provide an interpretable and effective strategy for neuropeptide prediction and enable a trained classifier to predict neuropeptides from large datasets. A web tool for this classifier is freely available at https://biolib.com/ATGCACTGTTCAGGCCTC/SAE Neuropeptide-Predictor
bioinformatics2026-08-23v2RapidMACS: MACS3-identical peak calling, 50x faster
Hung, L.-H.; Yeung, K. Y.Abstract
Motivation: MACS3 is a comprehensive peak-calling toolkit whose subcommands span single- and paired-end data, narrow and broad peaks, and a range of signal-track utilities. However, many ATAC-seq and multiomic pipelines, including our own, use just one of those capabilities: narrow peak calling. A single analysis may call peaks under different conditions, so a per-call saving is multiplied and the time recovered can be substantial. Furthermore, we wanted an efficient and embeddable narrow peak caller that could be integrated directly with our Chromap Suite aligner, so we wrote RapidMACS. Results: RapidMACS is a narrow peak caller optimized for this purpose, building its signal tracks in a single lazy sweep that avoids a global sort, processing chromosomes in parallel, and keeping intermediates in memory rather than creating temporary files. In benchmarks involving single-cell ATAC-seq, bulk ATAC-seq, ChIP-seq, and CUT&RUN, it is 3.5 to 51 times faster than MACS3 v3.0.3 depending on the applications and number of threads used. More importantly, the output is byte-identical. This means that any downstream analysis using RapidMACS will produce identical results to those using MACS3. The speed gains are largely due to these algorithmic changes rather than the choice of language, because MACS3's peak-calling code is itself compiled (Cython). RapidMACS depends only on htslib and zlib and links as a small static archive without a Python or Cython runtime. Availability and implementation: RapidMACS is open source under the MIT license at https://github.com/morphic-bio/rapidmacs, with both a standalone CLI executable (rapidmacs) and a linkable C++ library. Prebuilt containers for x86-64 and arm64 are published as biodepot/rapidmacs on Docker Hub.
bioinformatics2026-08-23v2Delta Marches: Generative AI based image synthesis to decode disease-driving morphologic transformations.
Nguyen, T. H.; Panwar, V.; Jarmale, V.; Perny, A.; Dusek, C.; Cai, Q.; Kapur, P. H.; Danuser, G.; Rajaram, S.Abstract
Deep learning reveals that tissue morphology contains rich pathophysiological information beyond human understanding. However, approaches to convert these spatially distributed signals into subcellular insights informing disease mechanisms are lacking. We introduce Delta-Marches, an interpretability-first approach that nominates distinguishing morphological features rather than explaining existing models' decisions. Delta-Marches simulates idealized morphological changes between classes by coupling latent-space traversals with generative AI. Comparing each image to its class-shifted counterpart allows feature extractors to infer the most affected aspects, reducing sample-to-sample variability and resolving transformations at subcellular resolution. Prototyped in renal carcinoma grading, Delta-Marches generates realistic grade transitions and pinpoints tumor-cell nuclear phenotypes as key determinants. It also reveals reduced vasculature with increasing grade, a known pattern absent from standard rubrics. Applied to dysplastic progression in colorectal tissue, it recovers glandular remodeling and goblet-cell loss, generalizing across tissue types and scales. These results show Delta-Marches parses complex image phenotypes and catalyzes hypothesis generation.
bioinformatics2026-08-22v5Protal: Ultra-fast metagenomic profiling and strain-resolved analysis
Fritscher, J.; Duncan, A.; Hildebrand, F.Abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal -- profiling through alignment -- an ultra-fast alignment-based method for species-and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom benchmarks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler, respectively). At the strain level, protal reconstructs intraspecific phylogenetic relationships among detected bacteria with similar accuracy as StrainPhlAn 4; because protal produces precise alignments, the phylogenies can be de novo produced without reliance on reference strain collections. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
bioinformatics2026-08-22v2Population dynamics of ribozymes during in vitro selection
Brooks, E. M.; Kotlin, N. R.; Weiss, Z.; DasGupta, S.Abstract
Ribozymes are central to models of RNA-based primordial life. Understanding how ribozymes emerge from RNA populations under changing selection pressures can help delineate the biochemical constraints and capabilities of early RNA-based biology. In vitro selection has been instrumental in isolating ribozymes with diverse functions from combinatorial populations; however, their population dynamics during selection remain poorly characterized. By analyzing high-throughput sequencing data produced from all rounds of an in vitro selection experiment, consisting of over 5 million unique sequences, we tracked how sequences rose and fell in abundance during selection. We found that the most abundant sequences emerged early and maintained their dominance, suggesting that only a few cycles of selection-amplification may be sufficient to isolate well-adapted ribozymes. Characterization of the catalytic activities of the dominant sequences in each selection round indicated that abundance did not always predict activity. This observation was reinforced through nucleotide conservation analysis and activity assays of close sequence variants of the most dominant ribozyme. Our analyses show that sequence complementarity with the substrate emerged as a common feature as early as the first round and persisted throughout, highlighting early convergence of diverse sequence populations on a beneficial catalytic strategy. By combining bioinformatics and experimental approaches, our work presents a detailed description of ribozyme population dynamics during an in vitro selection experiment and provides a high-resolution view of how functional RNA sequences compete for survival and propagation under selection pressure.
bioinformatics2026-08-22v2Mapping pathogenic patterns in membrane transporters from the GLUT transporter family
Kadasova, N.; Martinat, D.; Spackova, A.; Hutarova Varekova, I.; Berka, K.Abstract
Significance Missense mutations can lead to pathological effects in human cells. Predictive methods that account for structural context, such as AlphaMissense, can provide pathogenicity scores. The accumulation of pathogenicity hotspots can reveal important structural features within individual proteins of protein families, such as GLUT transporters. Mapping pathogenicity scores onto the structure can thus provide a mechanistic explanation of the protein function necessary for its role in the cell. Abstract Non-synonymous amino acid substitutions (missense mutations) are common in the general population; some are causative of serious disease. Depending on their structural context, they can disrupt protein function, folding, or dynamics. Computational predictive methods developed in recent years, such as AlphaMissense, provide new insights into how missense mutations affect protein structure by predicting and mapping their pathogenicity across each amino acid in the human proteome. In this study, we identify recurring patterns of pathogenicity prediction across the GLUT family membrane transporters encoded by genes SLC2A1-14. Within the GLUT transporter family, we observe higher pathogenicity profiles in the transmembrane domains, particularly in pore-lining and binding-site residues. Predicted missense pathogenicity is elevated throughout residues assigned to the central cavity, suggesting sensitivity of the transport pathway. Another finding shows higher pathogenicity in specific transmembrane helices of the protein, with the same pattern across all proteins. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family. We validated these predictions against clinically observed missense variants from ClinVar and benchmarked three prediction methods. AlphaMissense showed the strongest concordance with clinical classifications (AUC = 0.88), and the structural distribution of clinically reported pathogenic variants broadly recapitulated the predicted pathogenicity landscape. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family, and in select cases, clinical and predicted data diverged, suggesting more localized functional constraint than genome-wide prediction alone would indicate. These findings show that the pathogenicity of glucose transport within the GLUT family may be shaped by functional redundancy and physiological essentiality across GLUT groups, supported by convergent evidence from structural, computational, and clinical variant data.
bioinformatics2026-08-22v2Real-time repository-scale spectral search and global molecular networking with HNSW-MS
Semenov, A.; Roberts, A. M. P.; Coler, E. A.; Gupta, S.; Kopylova, E.; Melnik, A. V.; Boginski, V.; Aksenov, A. A.Abstract
Spectral similarity comparison is the basis of mass spectrometry-based metabolomics, underpinning library matching, molecular network construction, and repository searches such as MASST. Public repositories now exceed billion spectra, and exhaustive pairwise comparison is increasingly limiting. We introduce HNSW-MS, which implements Hierarchical Navigable Small World graph indexing natively for mass spectra. It operates directly on GC-MS and LC-MS/MS spectra without embedding, so scores are identical to those of exhaustive comparison and results are fully reproducible. On the benchmark dataset of 8.4 million LC-MS/MS spectra, HNSW-MS is up to 900-fold faster than linear search, recall@1 and recall@10 achieve over 90%, with the recall-speedup tradeoff tunable at query. Such rapid retrieval allows for any query spectrum to be near-instantaneously integrated into a global molecular network. We demonstrate specific examples of use of global networking for molecular discovery, including uncovering a new pathway of stenothricin biosynthesis and a new dipeptide conjugate bile acid.
bioinformatics2026-08-22v2CLEAR-ST: Physics-informed probabilistic decontamination of spatial transcriptomics by modeling mRNA lateral diffusion
Ma, K.; Huang, Y.; Ho, J. W. K.Abstract
Spatial transcriptomics is a rapidly evolving technology that allows for the measurement of gene expression in a spatially resolved manner. However, one technical problem that occurs for many sequencing-based spatial transcriptomics platforms is the presence of mRNA lateral diffusion, where mRNA from one spot can bind to probes in another spot, leading to contamination and inaccurate gene expression measurements. In Visium-like assays, this artifact is often visible as structured out-of-tissue signal and boundary-associated expression halos, yet its magnitude, spatial decay, and directional bias vary substantially across samples. Here, we present CLEAR-ST, a physics-informed probabilistic framework for correcting diffusion-like contamination in spatial transcriptomics data. CLEAR-ST infers a latent clean expression field using a denoising autoencoder and links it to the observed counts through a graph-Laplacian forward contamination model with learnable diffusion parameters, finally evaluated with a selectable count likelihood. We first conducted a comprehensive comparison between 10X official and independently generated Visium samples, demonstrating that out-of-tissue count profiles are highly related to nearby in-tissue expression, more concentrated near tissue boundaries, and diffusion directions across genes are likely coherent. Across real samples with varying contamination burden, CLEAR-ST improved spatial domain recovery, increased gene-level spatial autocorrelation, and enhanced the biological specificity of downstream analyses such as marker gene discovery, pathway identification and cell type deconvolution. Compared to benchmark methods, CLEAR-ST showed consistent gains in clustering quality and concordance with manual annotations. Together, CLEAR-ST provides an interpretable and practical approach for diffusion-aware correction of capture-based spatial transcriptomics data.
bioinformatics2026-08-22v1Detecting and typing Chlamydia trachomatis strains in metagenomes using the MetaChlam pipeline
Sharma, P.; Dean, D.; Read, T. D.Abstract
The Gram negative bacteria Chlamydia trachomatis (Ct), an obligate intracellular human pathogen, is a predominant cause of sexually transmitted infections and ocular trachoma globally, exerting a significant impact on public health. Ct "strains" (major lineages within the species) are known to have different tissue tropisms and be associated with different disease outcomes. Metagenome samples from typical sites where Ct infects (e.g., endocervix, conjunctiva, rectum) rarely contain enough reads for traditional genotyping methods such as Multi-Locus Sequence Typing (MLST) or ompA genotyping. To overcome these limitations, we implemented an ensemble tool called MetaChlam that can accurately classify Ct strains with as few as 250 Ct reads. Using 109 publicly available Ct genomes from naturally circulating strains, we established that an ANI-based threshold of 99.75% was capable of distinguishing Ct strains from each other. We implemented metagenome-based typing using the previously developed LINtax, Strainscan, StrainGE, and Sourmash softwares. MetaChlam integrated the four tools along with custom databases into an automated nextflow pipeline. Using simulated metagenomic reads, we found that our pipeline accurately identified the correct strains in both single strain and multi-strain mixtures of samples. Finally, we showed that MetaChlam had higher specificity for the true presence of Ct reads in NCBI SRA metagenomic datasets than NCBI PebbleScout software. A surprising finding of these analyses was that reads from Ct, an obligate human intracellular pathogen, can be found as contaminants in samples from sites where the organism is almost certainly not present. Overall, our study enhances the characterization and classification of Ct strains and provides protocols for identification and typing of Ct in shotgun metagenome data. The MetaChlam pipeline is available on Github: https://github.com/parul-sharma/MetaChlam.
bioinformatics2026-08-22v1Cross-Kingdom Multi-Omics Harmonization Uncovers Coordinated Host Defense and Vector Small RNA Regulatory Networks in Begomovirus Transmission
Badeli, G.; Kaboosi, K.; Mohebbi, A.; Nasrollanejad, S.Abstract
Begomoviruses present severe threats to global crop production through complex vector-mediated transmission by the whitefly Bemisia tabaci to host plants such as tomato (Solanum lycopersicum). Unraveling the molecular dialogue between host immune activation and vector non-coding RNA networks is essential for identifying key drivers of virus persistence and transmission. Public host transcriptomic (GSE309527) and vector small RNA (sRNA) sequencing datasets (GSE111343) were processed through a multi-omics harmonization and signal calibration pipeline. Differential expression analysis was performed using empirical Bayes moderated linear models, followed by non-parametric Spearman rank correlation modeling ({rho}) to infer cross-kingdom co-expression dynamics and pathway enrichment profiling across host and vector bio-systems. Harmonized principal component analysis showed clear separation by infection status across host plant and vector cohorts. Differential expression analysis identified 138 significantly altered host genes (69 upregulated, 69 downregulated) and 130 differentially expressed vector sRNAs (65 upregulated, 65 downregulated). Host responses were dominated by significant upregulation of gene-silencing machinery, including Suppressor of Gene Silencing 3 (SGS3; log2 FC = 3.67, q = 7.47 x 10-5), and pathway enrichment in Jasmonate defense (q = 0.0004) and RNA Interference & Silencing (q = 0.0001). Vector sRNAs exhibited targeted dynamic alterations, with pathway enrichment in Salivary Gland Secretion ($q = 0.0030) and Gut Endosymbiont Response (q = 0.0210). Cross-kingdom correlation modeling revealed two distinct, highly anticorrelated regulatory modules (mean |{rho}| = 0.76). Host SGS3 expression strongly correlated with vector sRNA VEC_0080 ({rho} = 0.9762) and virus-derived siRNA Bt-vsiRNA-01 ({rho} = 0.7619). These findings demonstrate a tightly synchronized tripartite molecular crosstalk between host antiviral immunity, viral siRNA accumulation, and vector small RNA remodeling. These cross-kingdom regulatory modules highlight promising targets for dual-action RNA interference strategies aimed at controlling Begomovirus transmission.
bioinformatics2026-08-22v1LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Liu, A.; Ho, A.; Droste, A. M.; Martin, D.; Wong, E.; Zhou, E.; Zhou, I.; Park, J.; Jiao, J.; Skelly, K.-R.; Kim, K.; Li, J.; Rao, K.; Uehara, M.; Marion, M.; Fitzgerald, N.; Dias, R.; Shringarpure, S.; Yuan, Y.; Wang, Y.Abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.
bioinformatics2026-08-22v1A Persistent Fleet of AI Scientists Exhibits Cooperative and Autopoietic Behavior
Patel, M. S.; Wierson, W. A.; Ekker, S. C.Abstract
Scientific work depends on memory, provenance, and continuity across projects, yet most agentic scientist systems are evaluated in bounded workflows or short benchmark runs. We describe a persistent fleet of cooperative AI scientist agents that operated continuously for nearly six months using shared memory, tools, and cross-agent communication. Critically, failures identified during longitudinal scientific research in this fleet prompted an advanced and recursively improving persistent memory architecture (MoE) that, with agentic science workflows, induced the generation of a novel trust architecture for the enablement of full provenance across all agentic scientific operations. An identity-level fabrication constraint reduced delusion-reinforcement probe failures from 91.7% to 0%, and a verification pipeline reduced wrong-topic citation hallucination more than 14-fold in companion benchmarks. This high provenance enabled the use of project memory systems to improve a local open-weight model on internal benchmarks from 44% to [~]90% through the deployment of fleet-specific institutional knowledge. This high-fidelity data environment also supported to date 104 recurring multi-phase reasoning cycles and produced 43 manually curated hypotheses, including cross-domain convergence events and a self-correcting rare-disease pharmacological chaperone-design case. While by design the fleet did not achieve unconstrained autonomous self-improvement or full autopoiesis, we term this bounded pattern AI Autopoietic Behavior due to the recurring operational improvement mediated by internal feedback and retained through institutional records with high confidence. Together, persistent memory, trusted provenance, and recursive learning shifted these agents from episodic assistants toward accountable, long-term scientific collaborators.
bioinformatics2026-08-22v1Design-informed Size Factor Estimation
Pocuca, T.; Pare, G.; Bolker, B. M.Abstract
Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.
bioinformatics2026-08-22v1Bridging Biomedical Atlas Ecosystem: Cross-Atlas Alignment And Scalable Tissue Specimen Registration
Jain, Y.; Desai, B.; Qaurooni, D.; Bhavsar, A.; Kienle, P.; Pouch, A. M.; ONeill, K.; Apte, S.; Herr, B. W.; Fisher, S. A.; Börner, K.Abstract
Over the last five years, over 13,000 tissue datasets with 200+ million cells from 20 consortia have been spatially registered into the Human Reference Atlas (HRA) common coordinate framework (CCF). The shared 3D spatial and semantic reference system enables exploration of datasets in the context of all other data across organs, assay types, and spatial scales. However, manual registration of individual samples remains resource intensive, posing feasibility challenges exacerbated by the proliferation of samples, assays, and atlasing efforts. This paper presents two approaches to scale up HRA construction: (1) projecting data across biomedical reference atlas systems and (2) using millitomes to bulk register tissue blocks into a reference organ. Both methods use the AutoMated Alignment and Projection (AMAP) pipeline to align 3D mesh models using point cloud registration. We demonstrate the evolving HRA-aligned atlas ecosystem for 6 models from the SPARC Program (heart), Gut Cell Atlas (large intestine), 500-subject consensus kidneys, and the Julich Brain Atlas. Additionally, we used AMAP to project 7 millitome models across 5 organs onto the HRA ecosystem, integrating 300+ tissue extraction sites. AMAP enables scalable tissue registration of data across atlas ecosystems enabling the construction of detailed reference maps of the human body.
bioinformatics2026-08-22v1AlphaConformers: Structure-guided sampling enables prediction of multiple protein conformations
Daniel, J.; Vitoriano De Queiroz Lira, L.; Zea, D. J.Abstract
Proteins are dynamic molecules capable of adopting multiple conformations. However, AlphaFold2 predominantly generates models around a single conformation, usually representing a ligand-bound state. To address this limitation, we developed AlphaConformers, a structure-guided pipeline that steers AlphaFold2 toward alternative conformations. It is based on the idea that protein structure databases can capture the structural space accessible to members of a protein family. Given a target protein, AlphaConformers retrieves structures from structurally similar proteins. These structures are organized into structure-based alignments and template sets, which are supplied to AlphaFold2 as conformational hypotheses. The resulting models are clustered and filtered, facilitating their analysis. Evaluated on a curated benchmark of 88 proteins with known ligand-bound and unbound conformations, AlphaConformers expanded AlphaFold2 conformational sampling and recovered alternative states missed by AlphaFold2 and other state-of-the-art methods. AlphaConformers ranked first for modelling subtle conformational changes commonly observed between ligand-bound and unbound states. These results show that structural information from protein databases can be leveraged to steer AlphaFold2 toward alternative conformations.
bioinformatics2026-08-22v1OTRec: Deep learning recommender for prospective druggable disease-target associations
Ofer, D.; Linial, M.Abstract
Identifying druggable disease--target associations remains a central challenge in translational medicine, limiting therapeutic discovery and repurposing. Here, we present OTRec, a deep learning--based recommender system that ranks such associations at scale and evaluates them in a temporal hold-out setting. Unlike approaches that rely on manually curated or aggregated evidence scores, OTRec employs a two-tower architecture to learn latent representations from 663,351 disease--target pairs. The model integrates heterogeneous inputs, including textual descriptions, ontology-derived features, and biological annotations such as tractability, Gene Ontology (GO) terms, and pathway information. We perform temporal validation by training on the 2022 Open Targets (OT) release and evaluating on clinical trial data from 2025. OTRec improves on the retrospective OT association score (ROC-AUC: 0.872 {+/-} 0.005 vs 0.559; PR-AUC: 0.288 {+/-} 0.009 vs. 0.08). In 5x5 target-disjoint cross-validation, OTRec reaches ROC-AUC 0.950 and PR-AUC 0.844) improving on the OT evidence score (ROC-AUC 0.91; PR-AUC 0.45). We rank the druggable genome across ~19,000 OT platform (OTP) diseases and release ~282,500 candidate associations above a 0.65 score threshold (in-distribution CV precision 0.92), covering 4,346 diseases including 2,322 orphan diseases, through an interactive prediction platform.
bioinformatics2026-08-20v3Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT
Hossain, M. S.; Sojib, M. R.; Tahmid, M. T.; Rahman, M. S.Abstract
Motivation: RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair tokenization breaks the one-token-one-nucleotide correspondence that nucleotide-level attribution assumes. Results: We present SPIRAL, a layer-wise SAE analysis of BiRNA-BERT. Independent SAEs at layers 0, 5, and 11 expand each 768-dimensional hidden state into 6,144 features while preserving model behaviour (explained variance above 0.99997; masked-language-model sequence recovery near 99.7%). Tokenizer-aware offset propagation aligns features to nucleotides: at layer 5, 44.3% of tested features are significantly associated with bpRNA secondary-structure classes (mean enrichment 1.61x), and all 1,237 eligible features with RNAcentral RNA types. Sparse profiles raise k-nearest-neighbour balanced accuracy from 0.328 to 0.359 over dense embeddings at layer 5. Availability and Implementation: Source code is available at https://github.com/SadatHossain01/SPIRAL; the code, evaluation data, and trained SAE checkpoints are archived at https://doi.org/10.5281/zenodo.21891845. Contact: mrahman@cse.buet.ac.bd
bioinformatics2026-08-20v2Cryptic binding sites are detected but not ranked: coverage, conversion, and the limits of detector consensus
Moore, C. W.Abstract
Methods for predicting cryptic binding sites are compared almost exclusively on top-n recovery, a number that conflates two independent abilities: proposing a candidate at the right location, and ranking it highly enough to be seen. We separate them by retaining the per-candidate overlap of every proposal, rather than only the top five, for four structurally different detectors spanning 2009 to 2026, across the CryptoBench benchmark. The separation is large and it reorders the field. On the designated test fold of 178 structures, fpocket, a purely geometric method from 2009, proposes a qualifying candidate for 74.2% of targets, the highest coverage of any tool tested, yet surfaces one in its top five for only 43.8%. P2Rank proposes qualifying candidates for 66.3% and surfaces 63.5%, and IF-SitePred, a 2024 method built on protein language model embeddings, proposes 70.8% and surfaces 61.8%. Coverage across tools varies by 8 points while conversion, the share of a tools own coverage that reaches the top five, varies from 59% to 96%. Unioning the four detectors reaches 92.1% coverage, and only 7.9% of cryptic sites are invisible to all of them. The fields headroom is therefore predominantly in ranking and in combination, not in detection: perfect ranking of a single tools existing proposals would reach 74.2%, and of the union 92.1%, against the 66.3% currently achieved. We show the practical consequence is governed by candidate budget. Added coverage converts to recovery at about 85% while a structure carries fewer than roughly fifteen candidates and at about 51% above it, which explains a series of interventions that raised coverage and returned nothing. Working within that budget, proposing pockets from a protein language model at locations where geometry finds no concavity improves single-structure recovery by 8.5% (95% CI +4.0 to +13.6) on test-fold data, and lets a five-conformer ensemble match a twenty-conformer one at a third of the wall clock. We release per-candidate overlaps for all tools so that coverage and conversion can be reported separately without re-running any method.
bioinformatics2026-08-20v2scUnify: a unified framework for training and inference across multiple single-cell foundation models
KIM, D.; Hong, A.; Jeong, K.; KIM, K.Abstract
Single-cell foundation models (scFMs) differ in software requirements and performance across downstream tasks and adaptation strategies, complicating comparison and reuse. We present scUnify, a framework that preserves each backbone's required processing while separating model-specific trainers, downstream tasks, and adaptation strategies as reusable components. Across five scFMs, scUnify reproduced original inference and training workflows, extended model-native tasks with multiple parameter-efficient fine-tuning methods, and demonstrated extensibility by connecting a newly implemented custom trainable task to multiple backbones and adaptation strategies. Together, these capabilities enable researchers to systematically compare these combinations and extend custom tasks across heterogeneous scFMs within a common workflow.
bioinformatics2026-08-20v2BART-spatial unravels biologically significant transcriptional regulators from spatial omics data
Wang, J.; Zhang, H.; Wang, Z.; Zang, C.Abstract
Transcriptional regulators (TRs) are crucial regulators of cell fate decisions by activating or repressing lineage-specific genes and integrating environmental signals with intrinsic networks. Identifying functional TRs is essential for understanding development, tissue organization, and disease. Emerging spatial transcriptomics and epigenomics technologies now provide near-single-cell resolution mapping of genomic features while preserving information of each cell's physical location and microenvironment which influence TR activity. Despite these advances, identifying active TRs in spatial data remains challenging due to low TR expression and the fact that TR activity often does not correlate directly with mRNA levels. Moreover, existing tools mainly designed for non-spatial single-cell data overlook spatial heterogeneity. To bridge this gap, we developed BART-spatial (Binding Analysis for Regulation of Transcription for spatial omics data), an innovative computational method to infer functional TRs from spatial omics data. BART-spatial integrates spatial variability and pseudotemporal information with publicly available TR binding profiles. Applied to multiple spatial datasets from diverse platforms, including 10x Visium, Visium HD, Atera, and spatial RNA-ATAC-seq, BART-spatial consistently outperforms existing methods, identifying state-specific TRs and revealing regulators undetectable by expression alone. Its compatibility with spatial epigenomics data further strengthens its utility and enables cross-validation. Overall, BART-spatial provides a powerful and robust tool for decoding spatially resolved gene regulatory programs.
bioinformatics2026-08-20v2A comprehensive quality control pipeline in human microbiome research for large population studies
Li, R.; A.M. Verlouw, J.; G. Boer, C.; Arp, P.; van Meurs, J.; G. Uitterlinden, A.; Kraaij, R.; Medina-Gomez, C.Abstract
The widespread application of high-throughput Next Generation Sequencing (NGS) technologies has made microbiome research an emerging field in public health and biomedical sciences. However, there are still many challenges that need to be addressed in this field. Pipelines available to generate microbiome data across cohorts are diverse, and sources of variation to be recorded and evaluated during microbiome profiling have not been standardized. Moreover, meticulous quality control of the microbiome data processing, from collection to computational quantification is still challenging, especially in large population studies. Innovative approaches are required to handle samples and to minimize the potential bias introduced by logistic hurdles in biobanking. In this paper, we describe the methodological steps surrounding the optimization of the 16S rRNA gut microbiome profiling in two large prospective cohorts the Generation R Study (mean age 9.83 {+/-} 0.32 years) and the Rotterdam Study (mean age 62.67 {+/-} 5.66 years). This paper also highlights potential solutions to sample mislabeling in large-scale microbiome analysis. To summarize, our study addresses common problems in human microbiome research. It aims to improve the research quality and reliability by integrating more stringent quality control standards into microbiome research.
bioinformatics2026-08-20v2PyMOL plugin for Protein Circuit Topology
Dimins, M.; Bazba, A.; Mogyorosi, A.; Kennon, E.; Fiol, T. D.; Hagen, L. A.; Sheikhhassani, V.; Akulov, V.; Mashaghi, A.Abstract
Circuit Topology (CT) provides a fundamental framework for analysing folded polymer chains, with applications in functional annotation, protein engineering and drug development. We present a protein CT analysis plugin for PyMOL v3.1.6.1 with a graphical user interface (GUI), automatic installation, and novel features developed through integration with PyMOL's application programming interface (API). The plugin integrates various previously developed CT methodologies for studying structured proteins and their complexes as well as the dynamics of disordered proteins. Analysis of a representative protein and a molecular dynamics trajectory demonstrates the plugin's three analysis modes and their outputs. The plugin reproduces the reference ProteinCT implementation exactly on the structures tested, and is distributed with a versioned release, a pinned environment and a one-command reproduction of every result reported here.
bioinformatics2026-08-20v2Toward a Minimal Amino Acid Alphabet for Protein Design
Pubal, K.; Kushnir, K.; Spiwok, V.; Louzecka, K.; Setnicka, V.; Lipovova, P.Abstract
Proteins are built from 20 canonical amino acids. It is interesting to explore whether proteins can be formed from significantly reduced amino acid alphabets. Our bioinformatics survey of UniProt (more than 250 M sequences) revealed that proteins composed of reduced amino acid alphabets (< 10) are extremely rare among existing proteins. Next, we used computational protein design to design proteins composed of all 1,013 possible alphabets of 2-10 early amino acids (Ala, Asp, Glu, Gly, Ile, Leu, Pro, Ser, Thr, and Val). The length of all proteins was 100 amino acid residues. Small amino acid alphabets preferred simple helices or helix bundles. Larger amino acid alphabets allowed for the design of more complex structures. A protein composed of 8 amino acid types (Ala, Asp, Gly, Leu, Val, Ser, Thr, and Pro) was successfully experimentally verified. It adopts the {beta}-sheet-rich fibronectin type III domain architecture. Attempts to experimentally verify designs composed of 6 and 4 amino acid types were unsuccessful. We show by a computational experiment with an experimental validation that inverse folding models, namely ProteinMPNNsol, can stabilize a designed protein within the same eight-amino-acid alphabet. Our results show that globular proteins may have formed early in evolution. Furthermore, we show that it is possible to design proteins with interesting properties for biotechnology and synthetic biology.
bioinformatics2026-08-20v2PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring
Panda, P. K.Abstract
We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docking, and an SE(3)-equivariant graph neural network scoring function trained at scale. Ligand flexibility is represented as a torsion tree and pose parameters are optimized by Monte Carlo with Metropolis acceptance refined by L-BFGS, with rotational gradients obtained in closed form through the derivative of the SO(3) exponential map rather than by finite differences. Affinity grids are built by a blocked neighbor-selection scheme that is exact and 5.6-9.7x faster than dense evaluation, and may be cached across ligands sharing a receptor and site, reducing a six-ligand series from 29.3 s to 10.4 s. On 814 protein-ligand complexes spanning 14 target families, PandaDock recovers a pose within 2 Angstroms of the crystal geometry in 33.7% of cases at rank 1 and in 57.0% of cases within the returned ensemble. The GNN scoring function is trained on 741,706 co-folded complexes from SAIR under target-disjoint splits, reaching a Pearson r of 0.407 on 90,219 held-out complexes and transferring to 202 independent crystal structures with measured Ki, Kd, IC50 or EC50 at r = 0.467. We report the model against three controls, a target-mean predictor, a ligand-descriptor-only baseline, and within-target correlations, and document both where it performs and where it does not, including its unsuitability for pose rescoring. On an independent 30-compound series against a single GABAA receptor target, PandaDock's empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR. At full scale on the PDBbind v2020 refined set (n = 4,640, native crystal poses), the fully independent SAIR model reaches r = 0.531, and a dedicated model trained on PDBbind alone under a target-disjoint split reaches r = 0.690 on its own held-out test complexes, the strongest evidence in this work that PandaDock's affinity predictions generalize. PandaDock is distributed under an open-source license at https://github.com/pritampanda15/PandaDock with a complete command-line interface and a reproducible benchmarking harness.
bioinformatics2026-08-20v1A generalizable normalization framework to decouple protocol and instrument effects: Application to high-sensitivity proteomics multicentric study (PME13)
Arauz-Garofalo, G.; Ciordia, S.; Gonzalez de Peredo, A.; Chaoui, K.; Rijal, J. B.; Gaxotte, V.; Folch-i-Casanovas, I.; Azkargorta, M.; Almey, R.; Aloria, K.; Kirim, B. A.; Barderas, R.; Braga-Lagache, S.; Calvo, E.; Chicano-Galvez, E.; Clemente, F.; Chiritoiu, G.; Chiva, C.; Decourcelle, M.; Dhaenens, M.; Diaz, R.; Douche, T.; Duran-Cortines, A.; Duran-Ruiz, M. C.; El Koulali, K.; Escobar-Nino, A.; Fernandez Acero, F. J.; Fernandez-Irigoyen, J.; Garcia-Garcia, C.; Gil, C.; Goetze, S.; Gonzalez Vidal, E.; Gutierrez, M.; Hernaez, M. L.; Lopez, C. M.; Marin-Vicente, C.; Mateos-Martin, M. L.; MatoAbstract
Multicenter studies are essential for benchmarking analytical workflows, yet their interpretation is often confounded by the combined effects of experimental protocols and instrumentation. To address this challenge, we introduce a simple normalization-based analytical framework, the recovery metric ({rho}), designed to decouple protocol driven effects from instrument dependent variability. We applied this framework to the 13th Proteomics Multicentric Experiment (PME13), a large multicentric proteomics dataset generated across 27 laboratories using high sensitivity workflows and varying sample preparation protocols. By leveraging a common digested reference sample, {rho} enables direct cross-comparison of all datasets on a unified scale, effectively minimizing instrument-related biases. Using this approach, we demonstrate that apparent instrument dependent trends are largely removed when evaluated through {rho}, revealing consistent protocol driven effects across laboratories. Statistical modeling identified key variables influencing {rho}, including sample input amount, reduction and alkylation, and the use of n-dodecyl-{beta}-D-maltoside (DDM). While DDM was associated with improved {rho}, reduction and alkylation and additional handling steps led to reduced performance, particularly at low input levels. We further highlight practical considerations for the application of ratio based normalization, including the occurrence of values exceeding theoretical bounds, which reflect deviations from underlying assumptions and require appropriate filtering. Overall, this work establishes a generalizable analytical strategy for disentangling confounding factors in multicentric datasets and provides practical guidelines for optimizing high sensitivity proteomics (HSP) workflows. The proposed framework is broadly applicable to other analytical fields where cross laboratory comparability is required.
bioinformatics2026-08-20v1bulk2scDiff: A Pseudobulk-Conditioned Diffusion Model for Bulk-to-Single-Cell RNASeq Generation
Xiao, J.; Raue, A.Abstract
Bulk RNA sequencing remains the predominant profiling strategy for large clinical cohorts, but it aggregates transcriptional signals across cell populations, thereby masking the underlying cellular heterogeneity. Inferring this heterogeneity from existing bulk transcriptomic data could extend large cohort-based studies that have already been profiled, but constitutes an underdetermined inverse problem, as one bulk profile can be compatible with multiple underlying cellular populations. Existing computational deconvolution methods address this problem primarily by estimating cell-type proportions or cell-type-averaged expression profiles rather than resolving expression at the level of individual cells. Here, we present bulk2scDiff, a proof-of-concept conditional diffusion framework that reformulates bulk-to-single-cell inference as conditional generation of single-cell expression profiles from pseudobulk transcriptomic input. We evaluated bulk2scDiff on two cancer single-cell RNA sequencing datasets, breast cancer and acute myeloid leukemia, where pseudobulk profiles were derived from the single-cell data and used as conditioning inputs, with the matched single-cell populations providing ground truth for controlled evaluation. Across both cases, bulk2scDiff closely reconstructed populations from training samples and generated biologically coherent single-cell populations for held-out samples, generalizing most consistently to recurrent immune features. A pseudobulk-swap control further confirmed sample-specific conditioning, with each sample corresponding pseudobulk yielding the closest agreement with its observed population in nearly all cases. Overall, our work establishes the feasibility of conditional diffusion for generating single-cell populations from pseudobulk transcriptomic profiles, providing a foundation for future evaluation with clinical bulk RNA sequencing data.
bioinformatics2026-08-20v1CryoForge: A Self-Correcting Agent for Cryo-EM Model Building That Learns When to Act and When to Stop
Feng, W.; Jiang, y.; Sun, F.; Yang, J.; Gao, X.; Zhang, F.; Han, R.Abstract
Automated atomic model building has accelerated cryo-EM structure determination, but different builders leave distinct residual error profiles requiring expert inspection. The post-building challenge is to decide which local interpretations are sufficiently supported by experimental evidence to be retained, corrected or rejected. Here we introduce CryoForge, an evidence-gated post-builder agent that separates repair proposal from repair acceptance. Rule and learning-based components identify candidate regions and prioritize legal actions, whereas an independent evidence gate evaluates each edit using map and half-map support, stereochemistry, connectivity and local structural context. Supported edits are retained; unsupported or conflicting modifications are rejected, rolled back, stopped or escalated for expert review. Across a resolution-stratified benchmark, 84.9% of 26,153 released trajectories yielded standard validated improvements and 3.7% yielded low-confidence partial improvements, with no quality-degrading edit retained in the final promoted models. Relative to rule-only control, learned prioritization reduced non-improving candidates and harmful actions while preserving global structural stability. External evaluations using an alternative initializer, same-team automated/manual-assisted challenge submissions and three recently released complex assemblies showed that CryoForge adapts to distinct residual error phenotypes and performs bounded, evidence-supported correction without uncontrolled remodeling. CryoForge provides a builder-independent, scalable and auditable correction layer between automated model generation and expert structural interpretation.
bioinformatics2026-08-20v1A multi-agent molecular optimization framework leads to a rapid-recovery intravenous anesthetic candidate with an improved safety margin
Xue, Z.; Liu, X.Abstract
Lead optimization, the systematic refinement of therapeutic compounds through iterative structural modification, faces a dual challenge in modern drug discovery: navigating astronomically vast molecular design spaces while balancing conflicting demands on potency, pharmacokinetics, and safety. We present MASCOT (Multi-Agent SearCh for molecular OpTimization), a role-specialized multi-agent framework for molecular optimization. Integrated with a chemically constrained graph-editing search, MASCOT coordinates three specialized agents: a trade-off agent that reprioritizes competing objectives, a strategy agent that adapts how molecular edits are proposed, and a reflection agent that distills lessons from previous decisions. Computational experiments showed that MASCOT achieved the best performance over competing methods on six benchmark settings. On the SARS-CoV-2 main protease task, its mean docking-score improvement was 3.6 times that of the strongest baseline. Applied to the clinically used anesthetic remimazolam (RM), MASCOT prioritized RM-1, which showed a shorter liver microsomal half-life, higher brain exposure, and a larger therapeutic index than RM. Subsequent derivative design yielded RM-7. Extensive animal studies established RM-7 as a rapid-recovery intravenous anesthetic candidate with greater potency, faster functional recovery, a wider safety margin, and preserved flumazenil reversibility. These results demonstrate that multi-agent coordination can link adaptive molecular search to medicinal chemistry and experimental pharmacology.
bioinformatics2026-08-20v1ETAP-CLF: an ESM3-based transformer attention framework for binary protein classification
Ren, J.; Jiang, H.; Li, P.; Yang, X.; Mei, L.; Tong, H.; Lin, L.Abstract
Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but downstream classifiers commonly compress them using fixed pooling operations that may discard task-relevant sequence context. We present ETAP-CLF, a compact framework that combines pretrained per-residue ESM3 embeddings with lightweight transformer contextualization and learned attention pooling to classify variable-length proteins and generate residue-level attention scores. The ESM3 parameters remained frozen, and the same ETAP-CLF architecture and hyperparameter configuration were used across ferroptosis-, senescence-, and pyroptosis-associated protein prediction. ETAP-CLF achieved AUROCs of 0.98, 0.95 and 0.91 for these tasks, respectively. In the ferroptosis benchmark, ETAP-CLF outperformed the evaluated published models. These results demonstrate that a common downstream design can adapt to multiple process-associated classification tasks without fine-tuning the task-specific model architecture. ETAP-CLF provides a generalizable approach for sequence-based protein prioritization and a basis for broader evaluation across binary protein-classification problems.
bioinformatics2026-08-20v1PlantOmicsGWAS: An end-to-end, reproducible framework for plant genome-wide association and genomic prediction using linear and pan-genome references
Khan, F. S.; Yassin, A.; Rehman, S. u.; Sun, T.; Wang, X.; Sun, H.; Abe-Kanoh, N.; Su, Y. H.; Guo, L.; Ye, W.Abstract
Genome-wide association studies (GWAS) play a crucial role in unraveling the genetic foundations of complex traits in plants but are also hampered by the application of heterogeneous tools, incompatible file formats and disparate computational environments. Existing GWAS frameworks are often restricted to a single linear reference genome, limiting the capacity for the analysis of structural variations and presence/absence variations (PAV) within plant populations. These issues pose obstacles to reproducibility, scalability, and comprehensive investigations. Here, we present PlantOmicsGWAS, an open-source Python framework for reproducible plant genome-wide association analysis and genomic prediction. It integrates reference indexing, FASTQ quality control, alignment, variant calling, VCF normalization, PLINK conversion, linkage disequilibrium analysis, population-structure estimation, association testing, marker scoring, genomic prediction, and visualization within a unified Linux and HPC workflow. The framework supports conventional linear-reference analyses and includes an optional pangenome-oriented module for working with multiple assemblies and graph-derived variation. Using a Vitis benchmark dataset containing 120 accessions and 118,247 graph-derived variants, PlantOmicsGWAS reduced manual workflow fragmentation and generated standardized association outputs. This tool provides a modular and extensible platform for plant GWAS and pan-GWAS workflows while retaining compatibility with established command-line tools and common genotype formats. The GWAS workflow described herein is adaptable to a range of sequencing methods and plant genomes, bridging research on crop related issues across various biological levels, from the individual organism to entire populations. PlantOmicsGWAS implements Bayesian sparse linear mixed modeling (BSLMM) through GEMMA for multi-trait association discovery, while also supporting FaST-LMM, regression-based approaches, and machine-learning algorithms (Random Forest, XGBoost) as benchmarking alternatives. The PlantOmicsGWAS, a versatile toolkit is available at GitHub https://github.com/plantomicsgwas1-boop/PlantOmicsGwas_V1 and on Linux and HPC platform (https://pypi.org/project/PlantOmicsGwas/1.0.2/).
bioinformatics2026-08-20v1nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality
Wyatt, C. D. R.; Duarte Frutos, F.; Turner, S. D.; Cerqueira De Araujo, A.; Rashid, U.; Begley, V.; Sumner, S.Abstract
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.
bioinformatics2026-08-20v1A single-antigen, multi-epitope subunit vaccine candidate against lumpy skin disease virus designed by conservation-guided reverse vaccinology
Koirala, B.; Ghimire, U.Abstract
Lumpy skin disease virus causes devastating economic losses in cattle, and current live attenuated vaccines carry risks of reversion and cannot differentiate infected from vaccinated animals. We applied conservation-guided reverse vaccinology (vaccine design from genomic sequence data) to identify conserved epitopes (immune-recognized protein fragments) and design a single-antigen, multi-epitope subunit vaccine candidate. A nine-taxon phylogenetic supermatrix of three candidate antigens was built, and BLA-restricted T cell and linear B cell epitopes were predicted. Conservation was quantified via Shannon entropy and Fisher's exact tests; structural disorder via AlphaFold2 and IUPred2A, followed by codon optimization and in silico cloning. All six selected epitopes mapped to the ankyrin locus. The 134 epitope columns showed significantly higher constraint than 2,376 background columns (mean entropy 0.008 vs. 0.243 bits; odds ratio 29.74). The 87-aa, 9.31 kDa construct was predominantly disordered and cloned in silico into pET28a. This work is entirely computational; wet-lab validation of immunogenicity is required before translational claims.
bioinformatics2026-08-20v1A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences
Motta, J. A.; Motta, M. d. M.; Fernandez, C.Abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
bioinformatics2026-08-20v1Microbial bioprospecting for benzoxazolinate-like molecules: unleashing the potential of genome mining
Paliyal, S.; Kaur, B.; Rao, L.; Chakrabortty, A.; Singh, L.; Sehgal, I.; Sharma, M.; Singh, D.; Chaudhry, V.; Mantri, S. S.Abstract
The benzoxazolinate moiety is a key functional group found in a few natural products (NPs), exhibiting diverse bioactivities, including antitumor, antibacterial, and cytotoxic activities. Despite their clinical importance, only a few bacterial strains and NPs have been reported harboring this rare bis-heterocyclic moiety, underscoring a largely unexplored chemical space. Here, we performed large-scale genome mining and identified 277 putative biosynthetic gene clusters (BGCs) across diverse bacterial hosts, including previously unreported bacterial genera and strains. The BGCs were grouped into three compound classes: benzoxazolinate, benzobactin, and ashimides based on sequence similarity network clustering. Bioactivity predictions of the identified BGCs revealed the predominance of antibacterial and cytotoxic potential, highlighting promising candidates for future experimental validation and functional studies. This study also presents a neural network-based bioprospecting model that efficiently detects rare BGCs encoding benzoxazolinate-containing molecules from genomic sequences. Overall, our findings expand the known repertoire of bacterial hosts with the potential to produce benzoxazolinate-containing NPs and provide a comprehensive framework for the discovery and identification of candidate BGCs.
bioinformatics2026-08-20v1