Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Modeling Patient-Reported Pain Trajectories with Frequent Minimum and Maximum Scores
Liu, Y.; Harris, R. E.; Clauw, D.; Bayman, E.; Leroux, A.; Lindquist, M. A.Abstract
Chronic pain is a widespread public health issue that imposes substantial health, emotional, and economic burdens on individuals and communities. Because pain is subjective and lacks objective biomarkers, it is typically measured using patient-reported scores, often on a numerical scale from zero to ten. Increasingly, pain studies use ecological momentary assessment, with multiple daily assessments over days and across study phases (e.g., a series of baseline and post-intervention assessments). These data frequently show many ratings at the extremes (i.e., at minimum or maximum pain scores), commonly referred to as zero- and one-inflation in the statistical literature, along with considerable within-person variability both within and across days. These phenomena present challenges for statistical analyses, as they violate assumptions of most commonly used statistical techniques (e.g., the normality assumption of linear mixed models). We propose a Bayesian beta-binomial mixed-effects model for modeling potential zero- or one-inflated pain scores while accounting for variability using random effects on the mean and variance parameters across subjects. A simulation study demonstrates that the method accurately estimates model parameters across realistic sample sizes, time points, and zero- and one-inflation levels. An application to data from two longitudinal pain studies demonstrates that the model fits the data better and, when correctly specified, yields accurate uncertainty intervals for longitudinal changes in pain compared to existing models, especially for zero- and one-inflated outcomes. Additionally, the model directly estimates the probability of clinically meaningful pain events. The proposed method provides a powerful statistical framework for studying the patient-reported pain trajectories.
bioinformatics2026-09-03v4CoTRA: an integrated R/Shiny framework for transparent bulk and single-cell RNA-seq analysis
Seemab, U.; Vainionpaa, K.; Tanoli, Z.; Leinonen, H. O.Abstract
Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.
bioinformatics2026-09-03v2Architectonic Spandrels in the Origin of Enzymatic Function
Poley-Gil, M.; Fernandez-Martin, M.; Banka, A.; Heinzinger, M.; Rost, B.; Valencia, A.; Parra, R. G.Abstract
How new molecular functions emerge during protein evolution remains a fundamental question in molecular biology. The energy landscape theory states that proteins are minimally frustrated, i.e. they have minimized their internal conflicts, to allow robust folding. Yet, not all energetic conflicts are eliminated, with functional regions such as catalytic residues and ligand-binding sites being often enriched in frustrated interactions, trading localized stability for biological activity. However, it is still uncertain whether this functional frustration is an evolutionary adaptation, positively selected despite its energetic cost or an inevitable physical byproduct of the fold architecture. Here, we combine reverse folding, structure prediction, and sequence analysis with local frustration profiling to address this long-standing question. Unexpectedly, we found that reverse folding algorithms are unable to energetically minimize evolutionary conserved frustration at specific residues, even when detrimental to overall structural stability. We propose that these frustration hotspots act as architectural spandrels, inherent physical constraints of the fold that evolution subsequently co-opts for function. Our findings connect biophysical constraints and evolutionary selection, providing a new framework to understand how functional specificity emerges in protein landscapes.
bioinformatics2026-09-03v2MetaClaw: an auditable AI agent for end-to-end, multi-directional metagenomic and multi-omics analysis
Zhang, H.; Li, Z.; Lagniton, P. N. P.; Wang, Z.; Zhao, L.; Li, W.; Duan, p.; Jiang, X.; Ning, K.Abstract
End-to-end omics analysis requires more than selecting tools: a usable agent must bind data correctly, execute long workflows without blocking, preserve provenance and recover the biological conclusions that motivate an analysis. Existing LLM-driven bioinformatics agents automate parts of this process, but their operational dependencies and conclusion-level validity are often unclear. Here we present MetaClaw, an auditable agent that maps a user request to a registered workflow, executes standardized upstream processing on FlowHub, and runs study-specific downstream analyses in network-isolated OpenClaw containers. A YAML registry and an explicit plan-submit-poll-finalise lifecycle record file bindings, parameters, scripts, environments and outputs in per-job bundles. Across the full cohorts of four published studies (769 metagenomic profiles), MetaClaw recovered 4/4 sorghum marker groups, 3/3 RRMS features, 4/5 canonical CRC markers among the top 20 classifier features and 5/5 permafrost marker groups. In 45 model-by-prompt runs, upstream completion was consistent whereas downstream validity depended on the backend and instruction detail; three decoy-tested endpoints showed no significant differences. In 48 ablation sessions, removing the registry, planning loop or manifest caused distinct losses, with registry removal increasing time, tool calls and token cost. MetaClaw therefore connects standardized upstream execution, local analytical flexibility and conclusion-level validation in a rerunnable framework for metagenomic and microbiome multi-omics analysis.
bioinformatics2026-09-03v2DAQplugin: Interactive Deep Learning-Based Validation of Cryo-EM Protein Models in ChimeraX
Terashi, G.; Zhu, H.; Kihara, D.Abstract
Although an increasing number of protein structures are determined by cryogenic electron microscopy (cryo-EM), structure modeling frequently suffers from residue misassignments and sequence register shifts, particularly in regions with ambiguous density. Here, we present DAQplugin, a ChimeraX plugin for real-time evaluation of protein models against cryo-EM density maps using the deep-learning-based residue-wise model quality (DAQ) score. Unlike existing validation tools that are typically applied after model construction, DAQplugin enables interactive validation during model building and refinement. DAQ has been shown to accurately identify residue assignment errors, including sequence register shifts, as well as local conformational modeling errors. DAQplugin also provides guidance for correcting sequence register shifts by suggesting alternative residue placements along the backbone. The plugin is computationally efficient and runs on standard CPUs without requiring GPU hardware, enabling deep-learning-based validation on ordinary laptops during interactive model building, model-map fitting, and refinement. DAQplugin facilitates more accurate interpretation of cryo-EM density maps and improve the reliability assessment of protein structure models.
bioinformatics2026-09-03v2kamino: fast proteome-wide variant calling for amino acid phylogenomics
Derelle, R.; Lees, J. A.; Chindelevitch, L.Abstract
Amino acid-based phylogenetics usually relies on first clustering and aligning orthologous proteins. This approach is powerful but computationally demanding. Here, we present kamino, a reference- and alignment-free method that rapidly builds amino acid phylogenomic alignments directly from proteomes. As with similar algorithms, homologous regions are identified through shared sequences flanking variable regions. The method uses local changes in recoded k-mer occupancy to efficiently identify variable positions within these homologous regions and extract the corresponding pseudo-aligned sequences. It generates phylogenetically informative alignments across diverse prokaryotic and eukaryotic datasets. Phylogenetic analyses show that it accurately recovers Mycobacterium tuberculosis lineages, most curated GTDB taxa, and relationships consistent with published Drosophila and mammalian phylogenies, while producing signals broadly similar to BUSCO-based approaches. Runtimes are comparable to genome-based alignment-free methods and several orders of magnitude faster than classical marker-based pipelines, with moderate memory requirements. The method performs well across a broad range of divergence levels, from within-species comparisons to family-level prokaryotic and phylum-level eukaryotic datasets. kamino therefore provides a fast and simple route from proteomes to phylogenomic alignments across a broad range of evolutionary scales. The program is implemented in Rust and freely available at https://github.com/rderelle/kamino.
bioinformatics2026-09-03v2AI semantics for biomedical data integration
McLaughlin, J.; Puig-Barbe, A.; Ibrahim, A.; Pava, D.; Pendlington, Z. M.; Matentzoglu, N.; Sollis, E.; Foreman, A.; Wilson, R.; Lopez Gomez, F.; Harris, L.; Adeleye, Y.; Kaur, S.; Meldal, B.; Smedley, D.; Parkinson, H.Abstract
Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge. Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries.
bioinformatics2026-09-03v2Non-Covalent Poly(ADP-ribose) Signaling Organizes a Circadian E3 Ligase Network in the Brain
Harikrishnan, S.; Kang, S.-U.Abstract
Although PAR biology has traditionally been studied through covalent PARylation, non-covalent PAR-binding proteins provide an additional mechanism for interpreting transient PAR signals and converting them into downstream regulatory programs. Among these effectors, E3 ubiquitin ligases are uniquely positioned to couple PAR sensing to selective ubiquitination, thereby integrating stress signaling with proteostatic control. Because circadian systems depend heavily on temporally coordinated protein turnover, we hypothesized that PAR-binding E3 ligases may form a circadian-structured regulatory layer within the brain. To test this, we integrated GTEx v10 brain transcriptomics, GWAS Catalog gene-mapped associations, CIRCA circadian phase annotations, and Human Protein Atlas single-cell transcriptomic resources to characterize the organization of PAR-binding E3 ubiquitin ligases across neural tissues. Across the brain, the E3 ligase repertoire was broadly deployed yet regionally structured, with cerebellar and cortical enrichment patterns preserved within the PAR-binding subset. Representative ligases spanning circadian regulation, DNA repair, and neurodegeneration-relevant pathways displayed distinct abundance and regional-variability archetypes across GTEx brain regions. Human genetic analyses demonstrated that E3 ligases associated with cognition-, neurodegeneration-, and sleep/circadian-related phenotypes were disproportionately PAR-binding, supporting convergence between PAR-responsive ubiquitin regulation and disease-relevant biology. Circadian phase analyses further revealed that PAR-binding ligases occupy structured, non-random circadian windows within the broader E3 background, including distinct co-phasing relationships with BMAL1 and CRY1. Finally, cell-type enrichment analyses identified microglia as the dominant compartment for circadian-linked and PAR-binding circadian E3 weighting within the brain E3 program. Together, these findings support a systems-level framework in which non-covalent PAR-binding E3 ubiquitin ligases constitute a brain-deployed, circadian-organized regulatory layer that couples PAR signaling to time-dependent ubiquitin control in neural systems.
bioinformatics2026-09-03v1AnnFlux: object-conditioned neural stochastic differential equations for single-cell perturbation dynamics
Choi, H.; Byeon, G.; Park, H.; Park, J.; Lim, S.; An, J.-Y.Abstract
Single-cell perturbation profiling measures responses to genetic and chemical interventions, yet most models learn a static map, ignoring how populations move over time and how perturbations combine. AnnFlux, an object-conditioned stochastic differential equation, learns a drift field in latent cell-state space. Conditioning on the perturbing object makes the field queryable one object at a time, yielding per-object drifts comparable across genes and drugs. By learning a drift field tailored to each perturbation context, it interpolates a held-out timepoint in an epithelial-mesenchymal transition time course and predicts unseen perturbations. Beyond point estimates, AnnFlux improves distributional fidelity and predicts responses to held-out perturbation combinations. An IFN-response signature predicted by AnnFlux was associated with TLS proximity in an independent pan-cancer spatial atlas. This framework maps perturbation-driven cell-state evolution as continuous trajectories and represents unseen perturbations using prior-knowledge embeddings.
bioinformatics2026-09-03v1Towards Sparse Causal Features for Zero-shot Mutation Effect Prediction in a Protein Language Model
Mohanty, S.; Phutela, M.; Green, A. G.Abstract
Protein language models (pLMs) such as ESM-2 achieve strong zero-shot mutation-effect prediction, yet the internal computations supporting these predictions remain poorly understood. We introduce a sparse feature circuit framework that combines sparse autoencoders, integrated-gradients attribution, and activation patching to identify the latent features that causally mediate zero-shot mutation effect prediction in ESM-2 650M. We evaluate this framework over 67 mutations ranging from strongly deleterious to weakly deleterious in the DNAJA1 J-domain, where ESM-2 predictions agree strongly with deep mutational scanning measurements. We find that circuits selected by indirect effect recover the model's predictions more efficiently and provide more informative biological explanations than those selected by raw activation changes, showing that activation magnitude does not necessarily reflect causal importance. We find that related substitutions reuse substantial portions of their recovered circuits, ranging from 40% to 75%, and that the shared features often represent residues in three-dimensional contact with the mutation site. To our knowledge, our work provides the first causal, feature-level account of zero-shot mutation effect prediction in a pLM.
bioinformatics2026-09-03v1BioIMA: a one-click desktop tool for standardized extraction of phenotypic traits from biological images
Qumu, X.; Dan, X.; Feng, J.; Cui, Y.; Gong, Y.; Hou, Y.; Lai, Q.; Wang, Z.; Zhang, Y.; Zhu, Y.; Yu, Y.; Zhang, F.; Todesco, M.; Wang, J.Abstract
Standardized extraction of quantitative phenotypes from images is increasingly important across plant biology, from ecological and evolutionary studies to genetics, breeding, and functional genomics. However, as large image datasets are increasingly used for trait analysis, many biologically relevant traits, including size, shape, color, and spatial patterning, are still measured manually or using fragmented semi-automated workflows. These limitations reduce throughput, reproducibility, and accessibility, especially for researchers without computational expertise. Here, we present BioIMA, an open-source desktop tool for rapid and standardized phenotyping from biological images. BioIMA integrates foundation model-based segmentation with automated trait computation, allowing users to extract quantitative measurements from images through an intuitive graphical interface and without model training. To validate its performance, we quantified a set of knot morphological traits in two Populus species, as these measurements are typically time-consuming to perform manually. Automatic measurements showed strong agreement with manual ImageJ-based measurements (R2 > 0.95), while reducing per-image processing time by approximately 75% (from ~15 s to ~4 s). BioIMA was further applied to diverse plant datasets, including Helianthus and Rhododendron images with varying morphologies and background conditions. Although developed for plant phenotyping, BioIMA may also be extended to other biological samples where region-based size, shape, or color traits are of interest. By combining accessibility and standardization in a lightweight local application, BioIMA provides a practical community resource for image-based phenotyping in ecological and evolutionary studies.
bioinformatics2026-09-03v1Cell-type separability predicts annotation accuracy and outweighs algorithm choice: a factorial benchmark across seven scRNA paradigms
Wardhana, O.; Zeng, Z.; Lu, X.Abstract
Automated cell-type annotation is a prerequisite for most single-cell RNA-sequencing (scRNA-seq) analyses, but the rapid proliferation of methods spanning marker-based, correlation-based, classical machine-learning, deep-learning, semi-supervised, large-language-model (LLM), and transformer foundation-model paradigms has outpaced head-to-head evaluation. Existing benchmarks rely on convenience samples of real datasets in which cell count, class imbalance, cell-type number, and differential-expression strength co-vary uncontrollably, precluding causal attribution of performance to any dataset property. To resolve this, we benchmarked 63 tools across seven paradigms using a Taguchi L9(34) orthogonal array that varies four dataset properties independently, progressively reconfiguring experimental control across five phases: fully controlled simulation, within-platform and cross-platform real-data validation, database-connected and LLM-based annotation under ontology-aware scoring, and fine-tuned foundation models. Using standardized oracle inputs and Cohen's {kappa}, we found that, within the ranges tested, the major paradigms achieved comparable accuracy. Accuracy was predicted near-linearly by the separability of cell types in a shared expression embedding, measured as k-nearest-neighbor (kNN) purity, a relationship that held across sequencing platforms and in fine-tuned foundation models. We attributed the vast majority of {kappa} variance to dataset structure and only a small share to tool identity. Computational cost traded against workflow accessibility rather than accuracy: accessible correlation-based and LLM-based approaches performed competitively, while foundation models matched them only after fine-tuning. Because our oracle design isolates algorithmic capability from upstream noise, these results reframe how methods should be selected: the field's near-term gains lie in strengthening infrastructure--prioritizing tool accessibility, standardized evaluation, and robustness to pipeline variation.
bioinformatics2026-09-03v1Bravais Lattice Sampling: Geometry-Guided Sparse Probing for Connected-Component Detection in 3D Discretized Spaces
Carrascoza, F.Abstract
We introduce Bravais Lattice Sampling (BLS), a two-phase method for detecting connected high-density regions in three-dimensional space. BLS places probe sites on a Bravais lattice scaled to the expected nearest-neighbour distance dNN of the target structures, then recovers cluster boundaries by depth-first expansion seeded only from occupied probes, replacing the exhaustive raster scan that conventional connected-component labelling uses to discover seeds. The spacing between probe sites is set from the covering radius of the lattice, which is what allows the method to state in advance the size below which a cluster may escape detection. The second phase, an expansion refinement activated only on probes that return an occupied voxel, verifies every edge, so the components returned are true connected components. BLS versatility allows for selection of different Bravais lattice unit cells to match the target structure; for amorphous, non-crystalline shapes, BLS can default to a simple face-centred cubic unit cell, where the expected minimum cluster size is the only parameter that needs to be set. The current BLS implementation has been developed as a post-processing tool for molecular dynamics trajectories, and was tested for searching water ice clusters of different morphologies. BLS returns component counts and maximum cluster sizes identical to exhaustive-labeller algorithms, with 100% recall; it runs at about 0.94 times the cost of depth-first search, and at 0.84 to 0.90 times the cost of the fastest other labeller in our benchmark set. This algorithm, although implemented by us for molecular dynamics applications, could be of interest in other domain areas where searching for high-density elements in 3D space is relevant.
bioinformatics2026-09-03v1From neuropeptide and receptor annotation to ligand-receptor pairing: a sequence- and structure-based framework for mapping the neuropeptide-receptor interactome in Gryllus bimaculatus
Margaretha, F.; Sakamoto, M.; Seike, H.; Nagata, S.; Nakamura, Y.; Mochizuki, T.Abstract
Neuropeptides and their G protein-coupled receptors (GPCRs) control much of insect physiology and behaviour, but in Gryllus bimaculatus, an emerging model and edible insect, receptor sequence similarity hinders the mapping of which peptide each GPCR activates. We re-annotated a chromosome-scale genome (BUSCO 95.3%, from 86.7% on insecta_odb12) with comprehensive curation of 48 neuropeptide precursor families (51 loci, including seven not previously identified) and 134 candidate GPCRs (66 rhodopsin-class, 68 secretin-class), providing a near complete neuropeptide-receptor interactome catalogue. We modelled all 15,946 peptide-receptor pairs with AlphaFold3 and Boltz-2 and scored each interface with pLDDT and ipSAE. Ranking these scores, and cross-checking the top candidate for each family against a receptor phylogeny of known ligand specificity, gave a confident, phylogenetically related receptor for 27 of 35 curated receptor groups. These matches confirm the structural scorings with existing deorphanization data and propose receptors for peptides with no prior functional evidence. The annotation, curated peptide and receptor sets, and ranked complexes are available through CricketBase (https://cricket.annotation.jp), a genome browser with a structure viewer of peptide-receptor complexes, providing a resource for G. bimaculatus endocrinology and a workflow to deorphanize GPCRs in other non-model insects.
bioinformatics2026-09-03v1Comprehensive characterization of genomic, transcriptomic and epigenomic artifacts introduced in formalin-fixed, paraffin-embedded tissues.
Zmuda, E. J.; Covington, K. R.; Kandoth, C.; Ng, C. K. Y.; Weisenberger, D. J.; Bowlby, R.; Chu, A.; Corbett, R.; Brooks, D.; Cherniack, A. D.; Murray, B.; Ling, S.; Zhao, W.; Shinbrot, E.; Lim, R. S.; Bootwalla, M. S.; Hinoue, T.; Hu, J.; Kakkar, N.; Lawrence, M.; McLellan, M.; Miller, C.; Morton, D.; Mungall, A. J.; Muzny, D.; Sougnez, C.; Stojanov, P.; Xi, L.; Schultz, N.; Akbani, R.; Curtis, C.; Fulton, B.; Gibbs, R. A.; Getz, G.; Spellman, P.; Marra, M.; Ladanyi, M.; Tarnuzzer, R.; Sofia, H.; The Cancer Genome Atlas Research Network, ; Hutter, C.; Wheeler, D. A.; Gastier-Foster, J.; LairAbstract
Genomic, transcriptomic and epigenomic characterization has accelerated the discovery of clinically-relevant alterations in cancer, predominantly using fresh frozen (FF) specimens. However, clinical molecular pathology laboratories prefer formalin-fixed paraffin-embedded (FFPE) methods, known to introduce artifacts at the nucleic acid level, over fresh frozen methods. Extending the multi-platform analysis to FFPE specimens for comprehensive clinical molecular diagnosis requires a thorough understanding of the consequence of formalin-fixation. We present a detailed multi-platform characterization of FFPE preservation using paired FF specimens as the 'gold standard'. DNA and RNA were obtained from 38 patients across 6 cancer types using a FFPE optimized co-isolation. The impact of FFPE on exome sequencing was dependent on filtering, where a minimum coverage or supporting read filter can mitigate FFPE-specific false positives. Copy number alterations, MSI assessment, mutational signatures, and DNA methylation were comparable between FFPE and FF. FFPE biases in RNA expression can be overcome when using biology-relevant genes and we describe a novel consequence of FFPE on miRNA species diversity. Collectively, this data provides a broad view of FFPE artifact and offers best practices for overcome these biases.
bioinformatics2026-09-03v1Atlantis: An integrative database for human proteome structural and functional sites
De Oliveira Rosa, N.; Ferronato, P.; Varisco, M.; Matic, M.; Ruscio, M.; Miglionico, P.; Raimondi, F.Abstract
Understanding protein mechanisms in health and disease requires characterizing the functional roles of individual amino acid residues. To explore the role of residues and their mutations, we have developed Atlantis, a database that integrates structural and functional information at the human proteome residue level. A graph database enables complex queries and the retrieval of integrated information for multiple functional analysis of protein systems. A Model Context Protocol (MCP) connector allows the interrogation of the resource through Large Language Models (LLMs) or agentic frameworks for biomedical research. Atlantis annotates over 11M residues across 20k human proteins, identifying hundreds thousands intra- and inter-protein contacts in PDB as well as AlphaFoldDB structures. We also provide the possibility to analyze and integrate predicted 3D complexes inputted by the user, and we showcased these features on hundreds of AlphaFold-multimer complexes of GPCRs and LRRK2 interaction networks. The tool is freely accessible at https://atlantis.bioinfolab.sns.it/.
bioinformatics2026-09-03v1Joint ancestry inference reveals the landscape of archaic introgression in admixed populations
Medina Tretmanis, J.; Anorve-Garibay, V.; Peede, D.; Banuelos, M. M.; Avila Arcos, M. C.; Jay, F.; Huerta-Sanchez, E.Abstract
Studying the evolutionary history of archaic segments in recently admixed individuals requires inferring both continental and archaic ancestry in admixed genomes. Here, we present TRACTINATOR, the first deep-learning method for simultaneous inference of continental and archaic ancestry in admixed human genomes. The model combines SNP sequences, population allele-frequency information, and S* statistics to improve both inference tasks. By learning relationships between haplotypes and population allele frequencies, TRACTINATOR can generalize across genomic regions and even across different genomic datasets. We train our model using both real and synthetic data, and show that augmenting with synthetic data improves accuracy for both continental and archaic ancestry inference. Finally, we apply TRACTINATOR to admixed Latin American populations from the 1,000 Genomes Project, revealing how archaic ancestry is distributed within chromosomal segments of African, European and Indigenous American ancestry in Latin American individuals. For candidates of adaptive introgression, we also infer whether the archaic haplotype was introduced via European or Indigenous American ancestors.
bioinformatics2026-09-03v1FibrilNet maps conserved and tissue-specific molecular environments across systemic amyloidoses
Guzzi, P. H.; Ugo, L. H.; Carbonari, V.; Lio', P.; Veltri, P.Abstract
Systemic amyloidoses are initiated by distinct amyloidogenic precursor proteins but frequently contain recurrent extracellular, complement, lipid-transport and matrix-remodelling components. Whether these recurrent proteins form a conserved systems-level environment across amyloid diseases, and how strongly that environment depends on precursor and tissue context, remains unresolved. We developed FibrilNet, a network framework that integrates experimentally defined amyloid proteomes with a human protein protein interaction graph and Gene Ontology derived semantic information. FibrilNet compares topology-only random walk with restart (RWR) with ontology aware semantic RWR in frozen leave-one-out module reconstruction and precursor-seeded prioritization tasks. The human graph contains 17,997 proteins and 925,977 physical interactions, with a 9-dimensional semantic representation of interaction context. In expanded cardiac transthyretin amyloidosis (ATTR), semantic-RWR increased mean reciprocal rank (MRR) from 0.00167 to 0.05015 and Recall@100 from 0.0199 to 0.3377, improving 132 of 151 held-out targets. Significant semantic gains were also observed in renal serum amyloid A amyloidosis (AA) and leukocyte chemotactic factor 2 amyloidosis (ALECT2). Across compact ATTR, light-chain amyloidosis (AL), AA and ALECT2 modules, APCS, VTN and TIMP3 formed a direct four-disease recurrent core, while APOE occurred in three of four modules. A tissue-aware ATTR analysis showed limited overlap between cardiac and neurologic modules (19 shared proteins; Jaccard 0.0569). In the hTTR-A97S peripheral-nerve model, semantic-RWR significantly improved reconstruction of the 202-protein mapped neurologic module, with the strongest evidence concentrated in the downregulated proteomic program. TTR-seeded propagation improved with semantic information but remained weak in absolute terms, separating precursor identity from the distributed downstream molecular environment. These results support a multilayer model in which a restricted conserved amyloid environment coexists with precursor-, tissue- and disease-specific organization
bioinformatics2026-09-03v1An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.
Villalba, M. P. C.; Bustamante, F. E.Abstract
Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.
bioinformatics2026-09-03v1Beyond benchmark accuracy: machine-learning turnover-number predictors require system-level validation
Rimon Martinez, M. J.; Lottermoser, J.; Bouillon, A. T. C.; Vranken, W. F.; Zehetner, L.; Zanghellini, J.Abstract
Enzyme turnover numbers (kcat) are essential for kinetic models and enzyme-constrained genome-scale metabolic models (ecGEMs), but measured values are sparse and therefore increasingly estimated using machine learning (ML). Although these predictors are commonly evaluated by global regression metrics, their practical utility depends on how errors propagate through downstream models. We benchmarked six current kcat predictors on a curated BRENDA-derived dataset and five of them on EnzyExtract. To assess the influence of training-set proximity, we compared each benchmark dataset with the available training data for each predictor. We then used the predicted kcat values to parameterize ecGEMs of Saccharomyces cerevisiae and evaluated growth predictions across 19 conditions. We find that benchmark accuracy is moderate even on the BRENDA-derived dataset and drops sharply on EnzyExtract, where all predictors achieve R2 values of 0.20 or lower. This decline is accompanied by substantially lower overlap between the benchmark and training datasets, with exact sequence matches ranging from 24% to 78% for BRENDA, compared with 9% to 26% for EnzyExtract. However, that overlap alone does not explain differences in generalization across predictors. Moreover, downstream performance is also not explained by benchmark ranking. Across 19 conditions, none of the tool-specific ecGEMs consistently reproduces the experimentally observed variation in growth. In glucose minimal medium, the weakest benchmark performer yields the most accurate growth prediction in the downstream ecGEMs, whereas higher-ranked predictors produce larger deviations in growth. We trace this mismatch to localized errors at high-leverage positions in yeast's metabolic network, where underpredicted mitochondrial ADP/ATP carrier turnover numbers restrict adenine nucleotide exchange and impose an apparent limitation on cytosolic ATP supply. Relaxing this constraint shifts predicted growth toward the experimental reference. Thus, ML-derived kcat values can affect not only quantitative growth predictions but also the phenotype a mechanistic model appears to identify. These results argue for application-driven validation of biological parameter predictors in the downstream systems they are intended to support.
bioinformatics2026-09-03v1spatialMET: an open and scalable framework for spatial metabolomics analysis
Mekonnen, Y. A.; Ospina, O. E.; Rubio, V.; Welsh, E.; Uddin, R.; Ackerman, H. D.; Soupir, A.; Cox, J. E.; Fridley, B. L.; Flores, E. R.; Koomen, J.; Stewart, P. A.Abstract
Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.
bioinformatics2026-09-03v1Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach
Kalkhoven, E.; Baak, R. E.; Hooiveld, G. J. E. J.; Schrauwen, P.; Hoeks, J.; Raymakers, R.; van der Stolpe, A.; Kersten, S.Abstract
Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.
bioinformatics2026-09-03v1DAG-HEART: Directed Acyclic Graph-Guided Health Equity-Aware Representation Transfer Learning Framework for Breast Cancer
Baek, M.; Wang, J.; Wan, S.Abstract
Breast cancer outcome prediction remains challenging for underrepresented populations because genomic datasets are demographically imbalanced and conventional multi-omics integration largely relies on undirected molecular similarity. We developed DAG-HEART, a directed acyclic graph-guided multi-omics transfer-learning framework that extends our previous transfer learning strategy with data augmentation. Using TCGA-BRCA mRNA, miRNA, and DNA-methylation data, DAG-HEART was evaluated for progression-free interval prediction in a data-minority group. DAG-guided nonlinear integration consistently improved predictive performance relative to direction-agnostic and correlation-based representations, while biologically motivated directional constraints generally outperformed reversed or unconstrained structures. Recurrently selected features converged on extracellular-matrix and regulatory pathways and supported clinically meaningful risk stratification. DAG-HEART provides an interpretable strategy for combining directed multi-omics structure with transfer learning under data imbalance across racial groups.
bioinformatics2026-09-03v1Macrophage signature-based prediction of cancer treatment response using MIL-attention
Madgwick, M.; Witham, S.; Occhetta, M.; Haneklaus, M.; Camanzi, B.; Smyrnakis, M.; Gardiner, L.-J.Abstract
Predicting immunotherapy response from single-cell data remains difficult due to patient-level labels, extreme class imbalance, and highly heterogeneous macrophage states. We present a Multiple Instance Learning (MIL) framework that treats each patient as a bag of macrophage embeddings derived from a single-cell RNA foundation model. The architecture incorporates an attention-based pooling mechanism with reduced model complexity, dropout-enhanced regularization and explicit attention penalties to improve stability in small-sample regimes. To address imbalanced clinical datasets, MIL outputs are optimized with a combined focal loss and supervised contrastive objective that simultaneously sharpens class boundaries and improves representation clustering. Across three cancer datasets, this approach outperforms pseudobulk aggregation, embedding baselines and standard MIL variants. Attention-weighted attribution and transcriptional regulatory analysis reveal distinct macrophage programs, interferon and antigen-presentation networks in responders versus hypoxia-linked regulatory modules in non-responders. This shows the potential of MIL to uncover predictive and mechanistically interpretable immune states.
bioinformatics2026-09-03v1An alignment-last approach enables rapid transcriptomic biomarker discovery in large cohorts
Narmanli, E.; Lanau, A.; Neacsu, M.; Koshkina, M. K.; Fumeron, P.; Martin, P.; Servant, N.; Perrin-Gilbert, N.; Waterfall, J. J.Abstract
Canonical transcriptomic analysis requires committing from the outset to a reference genome or transcriptome, which imposes a predefined feature set, usually annotated genes or isoforms. Alignment and annotation dilute the signal through feature-level aggregation, discard any sequence absent from the reference, and require reprocessing the entire dataset for each new question (mutations, fusions, transposable elements). Here, we introduce the alignment-last paradigm, in which the read becomes the unit of comparison across samples, and alignment is deferred to annotate only the relevant sequences. Querying the merome, a reference-free cohort k-mer index, with just a handful of reads (about 0.01% of a sample's) reveals the cohort's transcriptomic structure in bulk and single-cell data. At single-cell resolution, these reads outperform genes for cell classification and rediscover, without supervision, a transposable-element signature (VL30) of exhausted T cells. Finally, unsupervised read-level differential analysis recovers established lncRNA biomarkers; uncovers new prognostic transposable-element reads in adrenocortical carcinoma and sarcomas; and extracts signals even from reads that fail to align.
bioinformatics2026-09-03v1Morphologic intratumoral heterogeneity from routine whole-slide histopathology is prognostic for survival in primary central nervous system lymphoma: development in the LOC Network and international external validation
Rincon de la Rosa, L.; Barillot, N.; Hernandez-Verdin, I.; Velasco, R.; Mathon, B.; Eimer, S.; Vignes, J. R.; Rousseau, A.; Paillassa, J.; Ahle, G.; Lerintiu, F.; Uro-Coste, E.; Oberic, L.; Tabouret, E.; Appay, R.; Gauchotte, G.; Taillandier, L.; Marolleau, J.-P.; Adam, C.; Ursu, R.; Cuzzubbo, S.; Schmitt, A.; Nichelli, L.; Pons-Escoda, A.; Charlotte, F.; Davi, F.; Le Garff-Tavernier, M.; Choquet, S.; Soussain, C.; Vidal, N.; Gonzalez-Barca, E.; Climent, F.; Lopez, P.; Drieux, F.; Veresezan, E.-L.; Heming, M.; Meyer zu Hörste, G.; Grauer, O.; Jardin, F.; Mokhtari, K.; Houillier, C.; Hoang-XuaAbstract
Background: Clinical scores incompletely capture outcomes in primary central nervous system lymphoma (PCNSL). We quantified morphologic heterogeneity in pretreatment hematoxylin and eosin (H\&E) whole slides. Patients and methods: Three independent cohorts of immunocompetent, HIV- and EBV-negative patients treated recently were analyzed: LOC 2023 (122 slides), phase III BLOCAGE-01 (245 slides; NCT02313389), and external Barcelona (BCN; 41 slides). UNI embeddings, prototype learning, spatial metrics, and elastic-net Cox regression defined ITH-C. Results: Models achieved bootstrap-corrected concordance of 0.797--0.834. Age-, sex-, and KPS-adjusted ITH-C HRs were 1.29 (95\% CI 1.01--1.64), 1.27 (1.07--1.51), and 2.13 (1.35--3.37), respectively. Adding ITH-C increased MSKCC C-index from 0.671 to 0.717, 0.560 to 0.593, and 0.588 to 0.706. Spatial transcriptomics linked ITH-C to immune programs. Conclusions: Routine H\&E encodes prognostic spatial heterogeneity in PCNSL. ITH-C complements clinical scores, supporting prospective risk stratification.
bioinformatics2026-09-03v1Accessible and reproducible deployment reveals the practical boundaries of single-cell foundation models
Hou, S.; Yang, P.; Ma, W.; Xiang, J.; Wang, J. X.; Wan, H.; Ma, Y.; Zhou, X.Abstract
Single-cell foundation models (scFMs) have been widely promoted as a unifying paradigm for transcriptomic analysis, yet whether large-scale pretraining translates into reproducible biological advantages remains unclear. Their adoption is further hindered by heterogeneous implementations, preprocessing requirements, and computational environments. Here we develop a unified, automated, and reproducible framework for standardized deployment and controlled evaluation of scFMs across datasets, computational environments, training regimes, and downstream analyses, substantially lowering the technical barriers to their use. Leveraging this framework, we systematically investigate thirteen scFMs alongside established methods across nearly one hundred datasets spanning diverse biological contexts. Our analyses reveal clear practical boundaries to scFM utility. First, increased model scale, architectural complexity, pretraining corpus size, or input encoding does not consistently translate into superior downstream performance. Instead, measurable properties of embedding geometry provide a model-agnostic, representation-level explanation for differences in zero-shot performance across diverse model families. Second, the benefits of pretrained representations depend strongly on the biological and supervision regime: scFMs provide their clearest advantages under extremely limited supervision, particularly for rare-cell annotation and open-set detection of source-absent cell states, whereas established methods remain competitive or preferable in most other settings. Task-matched analyses further show that scFM representations transfer inconsistently to spatial-domain recovery, while their gene embeddings capture broad functional relatedness without reliably recovering context-specific regulatory relationships. Together, these results establish that scFM utility is neither universal nor determined simply by model scale alone, but varies with learned representation geometry, biological context, and supervision. By combining reproducible deployment with large-scale empirical and mechanistic investigation, our framework provides a principled foundation for determining when foundation-model pretraining offers genuine practical value and when simpler approaches remain sufficient.
bioinformatics2026-09-02v2CROWN: Curated Repository Of Well-resolved Noncovalent interactions
Poelmans, R.; Van Eynde, W.; Bruncsics, B.; Bruncsics, B.; Arany, A.; Moreau, Y.; Voet, A. R.Abstract
The development of machine-learning models for protein-ligand interactions is constrained by the quality and diversity of the available structural data. Existing resources force researchers into a trade-off: carefully curated collections such as PDBBind and HiQBind offer high structural reliability but cover only a narrow slice of the Protein Data Bank (PDB), whereas large-scale resources such as PLINDER provide broad coverage with minimal quality control. We present CROWN (Curated Repository Of Well-resolved Non-covalent interactions), a machine-learning-ready dataset that reconciles scale and rigor through a fully automated preprocessing pipeline. Starting from the PDB database, CROWN applies a series of interleaved quality filters and processing stages that address crystallographic resolution, ligand identity, pocket completeness, structural repair, interaction quality, and protonation at physiological pH. The pipeline finishes with a constrained energy-minimization step built on custom flat-bottomed restraints - a step absent from all existing protein-ligand datasets - that balances crystallographic evidence against the relaxation of intramolecular strain. By reconciling the heterogeneous refinement practices of different depositions without distorting the experimentally observed binding geometry, this step yields a structurally uniform collection of 178,263 complexes, representing a roughly four-fold increase in protein diversity over PDBBind and HiQBind. Rather than organizing the data around sparsely available, bias-prone binding affinities, CROWN adopts a geometry-centric design philosophy that treats the three-dimensional arrangement of atoms at the binding interface as a self-consistent source of information. To demonstrate its value as a training resource, we trained two knowledge-based scoring functions on CROWN and benchmarked them on CASF-2016: relative to HiQBind-trained counterparts, CROWN-trained models showed markedly improved ranking power (mean Spearman correlation rising from 0.509 to 0.637) and docking power (top-1 near-native pose recovery of 0.785 versus 0.724). Because CROWN imposes no requirement for affinity labels, it can in principle support any model that learns from or is evaluated against protein-ligand complex structures. We anticipate that it will serve as a broadly useful resource for tasks such as the training of binder generation, protein design or protein folding models conditioned on bound ligands, the development of scoring functions or benchmarking of interaction-prediction methods.
bioinformatics2026-09-02v2DESPOT: Direction-Enhanced Scoring POTentials
Poelmans, R.; Bruncsics, B.; Arany, A.; Van Eynde, W.; Shemy, A.; Moreau, Y.; Voet, A. R.Abstract
Knowledge-based potentials (KBPs) remain among the most reliable and interpretable scoring functions for protein-ligand interactions, yet most share two structural limitations. They assume that the space around each protein atom is isotropic, and their interaction-conditioned reference state cannot represent regions of space that are preferentially left empty. We introduce DESPOT (Direction-Enhanced Scoring POTentials), an all-atom anisotropic KBP that overcomes both. DESPOT classifies atoms into isotropic, axially symmetric, and fully anisotropic symmetry classes from their hybridization and bonding environment, and discretizes the surrounding interaction space using the according symmetry. By adopting a positionally averaged reference state and using a void ligand atom type, it learns, for every point around a protein atom, the probability that the point is occupied by a given ligand atom type or preferentially left empty - a ligand-independent description that naturally encodes steric exclusion. This occupancy-conditioned potential captures the precise, atom-level placement of ligand atoms; we pair it with a complementary geometry-conditioned, residue-level formulation (DESPOT-screen, in the spirit of KORP-PL) and combine the two inverse-Boltzmann scores into a consensus score, DESPOT-combo. Derived from 110,943 curated, energy-minimized complexes drawn from the CROWN database and evaluated on the CASF-2016 benchmark, DESPOT achieves competitive scoring power (Pearson r = 0.61), while DESPOT-combo attains best-in-class docking power (89.5% top-1 success); all anisotropic DESPOT variants significantly outperform isotropic KBPs and established empirical scoring functions in virtual screening. Anisotropy is decisive for rejecting geometrically implausible poses, and uniting the atom-level precision of DESPOT with the implicit flexibility tolerance of the residue-level score yields the most consistent performance across tasks. Because the same occupancy-conditioned potentials can be evaluated over an empty grid, DESPOT generates molecular interaction fields as well, unifying pose scoring with direction-aware binding-site characterization within a single interpretable model.
bioinformatics2026-09-02v2High-Resolution Subtyping of Pediatric Low-Grade Glioma Using an Integrated Meta-Clustering Framework
Tuerhanbayi, B.; Wang, J.; Wan, S.Abstract
Pediatric low-grade glioma (pLGG) is the most common type of brain tumor in children, accounting for approximately 30% of all central nervous system tumors in children. pLGG has multiple molecular subtypes that differ in disease progression, recurrence patterns, and treatment responses. Conventional wet lab approaches including molecular profiling and histopathological studies for pLGG characterization are time consuming, costly, and laborious. Recently, methods based on artificial intelligence (AI) or machine learning (ML) have been widely used for pLGG molecular categorization, but most of them can only identify two or three pLGG subtypes. To more comprehensively characterize the molecular subtypes of pLGG and their potential biological and therapeutic significance, we develop an integrated meta-clustering approach, namely Meta-pLGG, that can explore high resolution molecular subtypes and their transcriptional heterogeneity for pLGG. Specifically, we first performed multiple rounds of random projection (RP) to generate dimension-reduced feature vectors from pLGG transcriptomics data, each of which was subsequently clustered by different clustering algorithms including hierarchical clustering, K-means, Self-Organizing Maps (SOM), Non-negative Matrix Factorization (NMF), Gaussian Mixture Model (GMM), and Spectral Clustering, as base clustering methods. Then, to yield robust clustering performance, we integrated the clustering results of these RP based individual clustering algorithms by adopting a weighted meta-clustering (wMetaC) approach. Results based on 532 pLGG patients suggested that our proposed approach demonstrated superior stability and discriminative powers for higher resolution pLGG subtyping compared to conventional approaches. Based on consensus matrix analysis, we identified two major pLGG mega-subtypes, with one further subdivided into three subgroups and the other into two. Then, we performed cluster specific differential gene expression analysis, molecular pathway analysis, and gene-drug-disease association analysis. The results showed that the identified five subgroups exhibited significant subtype-specific transcriptomic heterogeneity. In summary, our meta-clustering approach demonstrated much higher performance and robustness in identifying higher resolution molecular subtypes of pLGG, revealing the molecular heterogeneity within pLGG and potentially providing new insights for more precise molecular subtyping and precision therapy.
bioinformatics2026-09-02v1Interactive downstream proteomics analysis with MiraProt using Mueller cell proteomes from equine recurrent uveitis
Schmalen, A.; Fleischer, A. B.; Riedel, B. M.; Deeg, C. A.Abstract
Mass spectrometry-based proteomics requires downstream analysis of processed protein abundance data, including data inspection, filtering, statistical testing, functional enrichment, protein set comparison, network analysis, and visualization. MiraProt was developed as a modular, metadata-aware R Shiny platform that integrates these steps in a single interactive workflow for processed protein-level proteomics data. Its metadata-aware design enables identifiers, sample information, experimental conditions, transformations, and derived data columns to be defined during data preparation and reused consistently across downstream analyses. To demonstrate its use, we reanalyzed a previously published label-free proteomic dataset of primary retinal Mueller cells from healthy horses and horses with equine recurrent uveitis (ERU). ERU is a naturally occurring autoimmune eye disease of horses characterized by recurrent intraocular inflammation triggered by autoreactive T-cells. Mueller cells are specialized retinal macroglia with various functions such as maintaining retinal ion homeostasis and supporting retinal neuron metabolism. Of 193 proteins with an adjusted p-value [≤] 0.05, 187 also showed at least a twofold abundance difference between ERU-derived and control Mueller cells. Functional enrichment highlighted nuclear RNA processing, chromatin-associated structures, DNA and RNA binding, interferon responses, and cell-cycle-associated programs. Gene set enrichment analysis identified positive enrichment of Interferon Alpha Response, Interferon Gamma Response, and MYC-, E2F-, and G2M-associated gene sets. Network analysis of shared proteins further linked this signature to DNA replication, mitotic checkpoint control, and RNA processing. ERU-derived Mueller cells also showed increased abundance of MHC class II-associated proteins. Together, these findings identified an interferon-responsive, cell-cycle-associated, and MHC class II-associated Mueller cell protein signature in ERU and generated experimentally testable hypotheses for further mechanistic studies. MiraProt provides an accessible, metadata-aware framework for reproducible downstream exploration of processed proteomic datasets and prioritization of candidate proteins and pathways for experimental follow-up.
bioinformatics2026-09-02v1A critical evaluation of Gene Ontology priors in biologically-informed neural networks
Verlaan, T.; Lieftinck, M. A.; Mwine, W.; Reinders, M. J. T.Abstract
Biologically-informed neural networks (BINNs) embed prior knowledge such as the Gene Ontology (GO) into their architecture to produce structurally interpretable representations, yet whether and how this prior improves performance or interpretation remains unclear. Here, we introduce GONNECT, a BINN incorporating GO into an autoencoder. We evaluate GO constraints in the encoder, decoder, or both on RNA-seq tumour samples from The Cancer Genome Atlas (TCGA), comparing against published BINNs (OntoVAE and VEGA), randomized-prior controls, and an unconstrained baseline. Across metrics, GO structure adds little to reconstruction or latent-space organization, frequently matched by randomized or unconstrained models. Its value lies in node activations, particularly in the encoder, where they correlate with a gene set enrichment analysis (GSEA)-derived reference. GONNECT-SL introduces regularized connections outside GO, but these soft links are unstable across seeds and concentrate where the ontology is sparse, appearing to compensate for the priors constraints rather than reveal new biology. They recover near-unconstrained reconstruction, keeping encoder activations interpretable. We identify the soft-link encoder as most promising. Our results clarify what biological priors contribute: their value lies not in the identity of the imposed connections or in improved performance, but in organizing activations into biologically meaningful units that can be interrogated directly.
bioinformatics2026-09-01v3A Visually Interpretable Histopathology-Based Immune Model Predicts T-effector Biology and Response to Immune checkpoint inhibition in Clear Cell Renal Cell Carcinoma Clinical Trial and Contemporary Real-World Datasets
Perny, A.; Jarmale, V.; Jasti, J.; Zhong, H.; Christie, A. L.; Miyata, J.; Nielsen, A. W.; Kontoyiannis, P.; Rakheja, D.; Modrusan, Z.; Huseni, M.; Kadel, W.; Brugarolas, J.; Kapur, P.; Rajaram, S.Abstract
Immune checkpoint inhibitors (ICI) are central to the treatment of metastatic clear cell renal cell carcinoma (ccRCC), yet only a subset of patients derive durable benefit, and clinically deployable predictive biomarkers remain an unmet need. RNA-based T-effector signatures capture cytotoxic immune biology and have been associated with ICI response in clinical trial cohorts; however, their clinical implementation is limited by the marked spatial heterogeneity of ccRCC, as well as cost, long turnaround time, sample quality requirements, and limited accessibility. Here, we developed a visually interpretable deep learning (DL) model that predicts a T-cell-enriched immune score directly from hematoxylin and eosin (H&E)-stained whole-slide images. To overcome the inability of H&E morphology alone to distinguish lymphocyte subsets, we trained the model using multimodal spatial supervision from CD8, PAX8, and ERG IHC, which respectively identified cytotoxic T-cell-rich regions, tumor cells, and endothelial cells, thereby constraining immune predictions to relevant tumor microenvironmental niches. The resulting H&E DL Immune score was validated by pathologist review, comparison with held-out CD8 IHC annotations, and independent datasets. The H&E DL Immune score correlated with T-effector RNA scores across independent institutional and IMmotion150 clinical trial cohorts (spearman correlations of 0.726; p=5.90x10-15 and 0.706; p=4.04x10-19). As a proof of principle, the score was used to characterize associations with key biological features across large cohorts, including sarcomatoid differentiation, BAP1 and PBRM1 mutation status, and additional transcriptomic signatures. In IMmotion150 clinical trial cohort, a median-dichotomized H&E DL Immune score, similar to RNA-based T-effector score, was significantly associated with clinical benefit from atezulumab therapy. In contemporary institutional cohorts of patients treated with frontline ipilimumab plus nivolumab or in initial 3 lines of nivolumab monotherapy, patients in the top quartile of H&E DL Immune score had significantly longer progression-free survival. Collectively, these findings support a scalable and interpretable H&E-based biomarker that captures T-effector biology and can help identify patients with ccRCC more likely to benefit from ICIs.
bioinformatics2026-09-01v2An Integrated Transcriptomic Landscape of Lung Cancer Identifies Tumor Clusters of Biological Significance
Arora, S.; Suresh, L.; Thirmanne, H. N.; Jensen, M.; Glatzer, G.; Fatherree, J.; Konnick, E.; Levine, K.; Brooks, A. N.; Houghton, A. M.; Pritchard, C.; MacPherson, D.; Berger, A.; Holland, E. C.Abstract
Lung cancer encompasses multiple histological entities with substantial molecular heterogeneity that remain incompletely resolved at population scale. Here, we constructed a unified reference landscape of lung cancer by analyzing raw RNA sequencing data from 1,824 tumors spanning adenocarcinoma (n=966), squamous cell carcinoma (n=628), small cell lung cancer (n=150), and unclassified non small cell lung cancer (n=80). Following batch correction, samples were analyzed using consensus clustering and visualized with PaCMAP to generate a molecular atlas annotated with clinical and biological metadata. Rather than segregating by pathological diagnosis, tumors organized along conserved transcriptional axes defined by tumor-intrinsic biology including proliferative or metabolic programs and immune-infiltrated states. Consensus clustering resolved nine robust molecular clusters, including an adenocarcinoma-associated subgroup, a neuroendocrine-like adenocarcinoma marked by ASCL1 activation, immune-associated regions, and bifurcation of both small cell and squamous carcinomas into biologically distinct states. Spatially restricted expression of selected clinically relevant transcripts nominated state-specific therapeutic hypotheses requiring future functional and clinical validation. Projection of patient tumors and patient-derived xenografts onto the atlas demonstrated preservation of transcriptional identity and enabled quantitative assessment of model fidelity. This integrated framework organizes lung cancer as a structured continuum of transcriptional states and provides a reference resource for biological interpretation and future translational studies.
bioinformatics2026-09-01v2Spatial Transcriptomics As Rasterized Image Tensors (STARIT) characterizes cell states with subcellular molecular heterogeneity
Velazquez, D.; Hallinan, C.; An, R.; Clifton, K.; Fan, J.Abstract
Abstract Imaging-based spatially resolved transcriptomics (imSRT) technologies provide high-throughput molecular-resolution spatial characterization of genes within cells. Conventional analysis methods to identify cell-types and states in imSRT data rely on gene count matrices derived from tallying the number of mRNA molecules detected for each gene per segmented cell, thereby overlooking subcellular heterogeneity that can be useful in defining cell states. To take advantage of the molecular-resolution information in imSRT data and potentially identify cell-states based on subcellular heterogeneity, we developed STARIT (Spatial Transcriptomics As Rasterized Image Tensors). STARIT converts transcripts within segmented cells in imSRT data into an image-based tensor representation that can be combined with deep learning computer vision models for downstream analysis. Using simulated and real imSRT data, we demonstrate that STARIT distinguishes transcriptionally distinct cell-types and further separates cell states based on subcellular transcript localization, which conventional gene count analysis fails to capture. By providing a standardized framework to encode subcellular molecular information in imSRT data, STARIT will enable deeper insights into subcellular heterogeneity and enhance the identification and characterization of cell-types and states that are overlooked by gene count representations.
bioinformatics2026-09-01v2Large Language Models for Accessible Reporting of Bioinformatics Analyses in Interdisciplinary Contexts
Yu, L.; Kim, D.; Cao, Y.; Shu, M. W. S.; Shen, M.; Liang, X.; Gu, J.; Jayakumar, R.; Ding, W.; Yang, F.; Zhang, X.; Kim, J.; Yang, P.; Yang, J. Y. H.Abstract
Health and life scientists frequently rely on quantitative experts to perform complex data analyses, yet interpretation of these results often becomes a bottleneck due to communication barriers across disciplines. Large Language Models (LLMs) have the potential to act as intermediaries that translate analytical outputs into accessible scientific narratives, but their ability to reliably interpret outputs from real-world biomedical data analyses across disciplines remains unclear. In this study, we investigate how state-of-the-art LLMs behave when integrated into real bioinformatics analysis workflows as report-generation assistants for interdisciplinary teams. We evaluated LLM-generated summaries using both automated and human evaluation frameworks to ensure holistic evaluation. Automated assessment employed multiple choice questions designed using Bloom's taxonomy to assess multiple levels of understanding, while human evaluation tasked scientists to score summaries for factual consistency, lack of harmfulness, comprehensiveness, and coherence. All generally produced readable and largely safe summaries, confirming their value for first-pass translation of technical analyses, however frequently misinterpreted visualisations, produced verbose summaries and rarely offered novel insights beyond what was already contained in the analytics. Our findings suggest that LLMs are best suited for easing interdisciplinary communication rather than replacing domain expertise and human oversight remains essential to guarantee accuracy, interpretative depth, and the generation of genuinely novel scientific insights.
bioinformatics2026-09-01v2PathFold: Predicting the Entire Protein Folding Pathway from Protein Sequence Alone
Zhang, Z.; Ibtehaz, N.; Kagaya, Y.; Xu, Z.; Punuru, P.; Kihara, D.Abstract
Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured {Phi}-values.
bioinformatics2026-09-01v1CyChat: a conversational Cytoscape app for no-code, reproducible network analysis
Liebold, J.; Stahl, M.; Schulze, J.-O.; Razavi, M. M.; Bader, G. B.; Kurtz, S.; Baumbach, J.Abstract
Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.
bioinformatics2026-09-01v1The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI
Nelen, J.; Khan, O.; Adams, E.; Aschenbrenner, J. C.; Thompson, W.; Ebrahim, A.; Capkin, E.; Vallee, C.; OpenBind, ; Shotton, E. J.; Griffen, E. J.; Chodera, J. D.; Deane, C. M.; von Delft, F.; AlQuraishi, M.; Imrie, F.Abstract
High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.
bioinformatics2026-09-01v1GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler
Uzum, A. S.; Haliloglu, T.Abstract
Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.
bioinformatics2026-09-01v1Sphingolipid metabolism-related genes as key regulatory hubs in white smoke inhalation induced lung injury
Meng, F.; Xin, H.; Li, R. R.Abstract
Objective White smoke inhalation injury (WSI) causes severe acute lung damage with no specific therapy currently available. Sphingolipid metabolism is implicated in pulmonary inflammation, but its transcriptional regulatory landscape in WSI remains unexplored. This study aimed to identify key sphingolipid metabolism related genes and evaluate their regulatory roles and therapeutic potential in WSI. Methods We established a rat model of WSI and performed integrated bulk RNA sequencing, weighted gene coexpression network analysis (WGCNA), and single-cell RNA sequencing (scRNAseq) to screen for differentially expressed sphingolipid metabolism-related genes (DESRGs). Protein-protein interaction (PPI) network with four centrality algorithms was used to prioritize hub genes. In silico gene knockout and molecular docking were conducted to assess regulatory functions and identify potential drug candidates. Results We identified 22 DESRGs that were predominantly enriched in DNA replication and cell cycle pathways rather than canonical sphingolipid metabolic processes. PPI consensus prioritized three hub genes--Top2a, Ttk, and Ccna2--with Top2a exhibiting the highest expression in epithelial cells and significant downregulation after smoke exposure. ScRNAseq revealed immune cell infiltration and epithelial differentiation trajectories. Virtual knockout showed that Top2a depletion affected the largest transcriptomic fraction (~0.4%) and was enriched in lysosome biogenesis, innate immunity, phagocytosis, and lipid catabolism. Molecular docking identified thalidomide as a high affinity ligand for Top2a (Vina score: -8.5 kcal/mol). Conclusion Our multiomics integrative framework identifies Top2a as a central regulatory hub linking sphingolipid associated inflammation to epithelial responses in WSI, and nominates thalidomide as a potential drug repurposing candidate. These findings provide prioritized targets for future translational investigation.
bioinformatics2026-09-01v1RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution
Yang, X.; Hao, N.; Zhao, R.; Angel, S.; Tan, Y.; Lian, C. G.; Zhou, L.; Olson, D.; Yu, K.-H.; Ruiz de Luzuriaga, A.; Wan, G.Abstract
Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.
bioinformatics2026-09-01v1The DYNAM-O Toolbox: Characterizing Individualized Neural Signatures in Sleep EEG
He, M.; Saremsky, S. R.; Noamany, H.; Chen, S.; Prerau, M. J.Abstract
Conventional sleep electroencephalography (EEG) measures often rely on predefined bands, thresholds, and averages that incompletely capture transient oscillatory dynamics across an entire night. Here, we introduce the Dynamic Oscillation (DYNAM-O) Toolbox, an open-source, cross-platform (MATLAB, Python, and Rust) software package for data-driven characterization of individualized neural dynamics in sleep EEG. DYNAM-O identifies transient oscillations as time-frequency peaks on multitaper spectrograms using a novel multi-resolution procedure, computes intrinsic and sleep-state-dependent extrinsic features for each event, and represents the overnight distributions of tens of thousands of TF-peaks as feature histograms spanning oscillation frequency, slow oscillation power, and slow oscillation phase. This distributional representation preserves continuous brain-state variation that could be obscured by averaging within conventional sleep stages. The toolbox further provides Gaussian and spline basis-based dimensionality reduction, visualization, and whole-histogram statistical testing tools to support both exploratory and hypothesis-driven analyses. To demonstrate its use for group-level inference, we analyzed overnight C3-channel EEG from 133 adults (71 females, 72 males; ages 20-35 years) in the Cleveland Family Study. Whole-histogram and parameterized-mode analyses reproduced the established higher center frequency of fast-spindle activity in females and additionally revealed greater low-alpha transient oscillatory activity in females, a pattern outside the conventional sleep spindle range. By completing the analysis cycle from TF-peak extraction to statistical inference, DYNAM-O provides an accessible and interpretable framework for studying individualized sleep physiology and identifying subtle, reproducible electrophysiological patterns.
bioinformatics2026-09-01v1Calibration-free compression brings Evo 2 to its full million-token context on a single GPU
Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.Abstract
Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.
bioinformatics2026-09-01v1Intelligent differential ion mobility spectrometry (iDMS): A deep neural network that predicts optimal space-resolved ion mobility parameters for isomeric monoglycosphingolipids
Nguyen-Tran, T.; Shi, X. X.; Hashimoto-Roth, E.; Organ, M. G.; Lavallee-Adam, M.; Perkins, T. J.; Bennett, S. A. L.Abstract
Simultaneous quantification of monoglycosphingolipid stereoisomers is required to monitor changes in defective enzymatic pathways linked to diseases such as Gaucher Disease, Parkinson's Disease, and Krabbe Disease. Resolution of beta-glucosyl and beta-galactosyl epimers cannot be achieved by standard liquid chromatography, electrospray ionization, tandem mass spectrometry (LC-ESI-MS/MS). Separation becomes possible when field asymmetric ion mobility spectrometry (FAIMS), also known as differential mobility mass spectrometry (DMS), is added as an orthogonal separation technique to LC. FAIMS/DMS separates epimeric ion clusters in a high versus low electric field (separation voltage, SV) then redirects the target epimeric ions to the mass spectrometer through the application of a direct current (compensation voltage, CoV). Resolving SVs and CoVs must be manually determined for each lipid. Manual derivation is a labour-intensive process that requires pure synthetic standards, limiting the number of stereoisomers a user can include in an assay. To address this problem, we introduce here intelligent DMS (iDMS). iDMS is an in silico supervised neural network model that learns the ion mobility relationships between SV and CoV and the monoglycosphingolipid structural features of sugar headgroup, N-acyl chain length, and N-acyl degree of unsaturation. iDMS predicts the SV and CoV combinations capable of resolving any stereoisomer pair from a training dataset of composed of measured signal intensities across a range of SVs and CoVs of 12 lipids. This machine learning alternative to manual DMS optimization promises to accelerate the deployment of multiple-reaction-monitoring mode (MRM) RPLC-ESI-DMS-MS/MS assays for the routine and rapid quantification of biologically relevant monoglycosphingolipid stereoisomers.
bioinformatics2026-09-01v1From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology
qin, y.; Pang, J.; Zhang, X.Abstract
Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.
bioinformatics2026-09-01v1Mural-VISTA: A tool for mural cell-vessel interaction assessment and multiscale single-cell topo-morphological analysis
Zeng, H.; Hu, M.; Phng, L.-K.; Matsunaga, Y. T.Abstract
Three-dimensional (3D) mural cell morphology is heterogeneous and coupled to vessel geometry, however, measurements from two-dimensional (2D) maximum intensity projections (MIP) obscure overlapping processes and cell-vessel contacts. Accordingly, we developed Mural-VISTA, a semi-automated Python workflow for mural cell-vessel interaction and single-cell topo-morphology analysis of reconstructed surface meshes. This workflow integrates mesh pretreatment, interactive centerline extraction, hierarchical segmentation of cell soma, main axis and secondary processes (branches), and extraction of 36 multiscale (cell process segment level, process level, and whole cell level) topo-morphological and vessel-referenced metrics. Mural-VISTA identified morphological changes in pericytes and vascular smooth muscle cells (vSMCs) with altered RhoA activity. Constitutive active RhoA (RhoA CA) over-expression reduced branch complexity and increased process alignment in both cell types, while increased whole-cell and branch solidity only in vSMCs. Dominant negative RhoA (RhoA DN) over-expression increased branch abundance and reduced branch solidity in pericytes but not vSMCs, suggesting cell-type specific effect of reduced RhoA activity. In conclusion, Mural-VISTA enables quantitative 3D profiling of mural cell architecture and its spatial relationship with the vessel.
bioinformatics2026-09-01v1Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R
Zeng, Z.; Wang, Y.Abstract
Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.
bioinformatics2026-09-01v1Automatic bioinformatic software named entity recognition from literature
Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.Abstract
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.
bioinformatics2026-09-01v1Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach
Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.Abstract
Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.
bioinformatics2026-09-01v1