Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Delta Marches: Generative AI based image synthesis to decode disease-driving morphologic transformations.
Nguyen, T. H.; Panwar, V.; Jarmale, V.; Perny, A.; Dusek, C.; Cai, Q.; Kapur, P. H.; Danuser, G.; Rajaram, S.Abstract
Deep learning reveals that tissue morphology contains rich pathophysiological information beyond human understanding. However, approaches to convert these spatially distributed signals into subcellular insights informing disease mechanisms are lacking. We introduce Delta-Marches, an interpretability-first approach that nominates distinguishing morphological features rather than explaining existing models' decisions. Delta-Marches simulates idealized morphological changes between classes by coupling latent-space traversals with generative AI. Comparing each image to its class-shifted counterpart allows feature extractors to infer the most affected aspects, reducing sample-to-sample variability and resolving transformations at subcellular resolution. Prototyped in renal carcinoma grading, Delta-Marches generates realistic grade transitions and pinpoints tumor-cell nuclear phenotypes as key determinants. It also reveals reduced vasculature with increasing grade, a known pattern absent from standard rubrics. Applied to dysplastic progression in colorectal tissue, it recovers glandular remodeling and goblet-cell loss, generalizing across tissue types and scales. These results show Delta-Marches parses complex image phenotypes and catalyzes hypothesis generation.
bioinformatics2026-08-22v5Real-time repository-scale spectral search and global molecular networking with HNSW-MS
Semenov, A.; Roberts, A. M. P.; Coler, E. A.; Gupta, S.; Kopylova, E.; Melnik, A. V.; Boginski, V.; Aksenov, A. A.Abstract
Spectral similarity comparison is the basis of mass spectrometry-based metabolomics, underpinning library matching, molecular network construction, and repository searches such as MASST. Public repositories now exceed billion spectra, and exhaustive pairwise comparison is increasingly limiting. We introduce HNSW-MS, which implements Hierarchical Navigable Small World graph indexing natively for mass spectra. It operates directly on GC-MS and LC-MS/MS spectra without embedding, so scores are identical to those of exhaustive comparison and results are fully reproducible. On the benchmark dataset of 8.4 million LC-MS/MS spectra, HNSW-MS is up to 900-fold faster than linear search, recall@1 and recall@10 achieve over 90%, with the recall-speedup tradeoff tunable at query. Such rapid retrieval allows for any query spectrum to be near-instantaneously integrated into a global molecular network. We demonstrate specific examples of use of global networking for molecular discovery, including uncovering a new pathway of stenothricin biosynthesis and a new dipeptide conjugate bile acid.
bioinformatics2026-08-22v2Population dynamics of ribozymes during in vitro selection
Brooks, E. M.; Kotlin, N. R.; Weiss, Z.; DasGupta, S.Abstract
Ribozymes are central to models of RNA-based primordial life. Understanding how ribozymes emerge from RNA populations under changing selection pressures can help delineate the biochemical constraints and capabilities of early RNA-based biology. In vitro selection has been instrumental in isolating ribozymes with diverse functions from combinatorial populations; however, their population dynamics during selection remain poorly characterized. By analyzing high-throughput sequencing data produced from all rounds of an in vitro selection experiment, consisting of over 5 million unique sequences, we tracked how sequences rose and fell in abundance during selection. We found that the most abundant sequences emerged early and maintained their dominance, suggesting that only a few cycles of selection-amplification may be sufficient to isolate well-adapted ribozymes. Characterization of the catalytic activities of the dominant sequences in each selection round indicated that abundance did not always predict activity. This observation was reinforced through nucleotide conservation analysis and activity assays of close sequence variants of the most dominant ribozyme. Our analyses show that sequence complementarity with the substrate emerged as a common feature as early as the first round and persisted throughout, highlighting early convergence of diverse sequence populations on a beneficial catalytic strategy. By combining bioinformatics and experimental approaches, our work presents a detailed description of ribozyme population dynamics during an in vitro selection experiment and provides a high-resolution view of how functional RNA sequences compete for survival and propagation under selection pressure.
bioinformatics2026-08-22v2Mapping pathogenic patterns in membrane transporters from the GLUT transporter family
Kadasova, N.; Martinat, D.; Spackova, A.; Hutarova Varekova, I.; Berka, K.Abstract
Significance Missense mutations can lead to pathological effects in human cells. Predictive methods that account for structural context, such as AlphaMissense, can provide pathogenicity scores. The accumulation of pathogenicity hotspots can reveal important structural features within individual proteins of protein families, such as GLUT transporters. Mapping pathogenicity scores onto the structure can thus provide a mechanistic explanation of the protein function necessary for its role in the cell. Abstract Non-synonymous amino acid substitutions (missense mutations) are common in the general population; some are causative of serious disease. Depending on their structural context, they can disrupt protein function, folding, or dynamics. Computational predictive methods developed in recent years, such as AlphaMissense, provide new insights into how missense mutations affect protein structure by predicting and mapping their pathogenicity across each amino acid in the human proteome. In this study, we identify recurring patterns of pathogenicity prediction across the GLUT family membrane transporters encoded by genes SLC2A1-14. Within the GLUT transporter family, we observe higher pathogenicity profiles in the transmembrane domains, particularly in pore-lining and binding-site residues. Predicted missense pathogenicity is elevated throughout residues assigned to the central cavity, suggesting sensitivity of the transport pathway. Another finding shows higher pathogenicity in specific transmembrane helices of the protein, with the same pattern across all proteins. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family. We validated these predictions against clinically observed missense variants from ClinVar and benchmarked three prediction methods. AlphaMissense showed the strongest concordance with clinical classifications (AUC = 0.88), and the structural distribution of clinically reported pathogenic variants broadly recapitulated the predicted pathogenicity landscape. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family, and in select cases, clinical and predicted data diverged, suggesting more localized functional constraint than genome-wide prediction alone would indicate. These findings show that the pathogenicity of glucose transport within the GLUT family may be shaped by functional redundancy and physiological essentiality across GLUT groups, supported by convergent evidence from structural, computational, and clinical variant data.
bioinformatics2026-08-22v2Protal: Ultra-fast metagenomic profiling and strain-resolved analysis
Fritscher, J.; Duncan, A.; Hildebrand, F.Abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal -- profiling through alignment -- an ultra-fast alignment-based method for species-and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom benchmarks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler, respectively). At the strain level, protal reconstructs intraspecific phylogenetic relationships among detected bacteria with similar accuracy as StrainPhlAn 4; because protal produces precise alignments, the phylogenies can be de novo produced without reliance on reference strain collections. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
bioinformatics2026-08-22v2Design-informed Size Factor Estimation
Pocuca, T.; Pare, G.; Bolker, B. M.Abstract
Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.
bioinformatics2026-08-22v1Bridging Biomedical Atlas Ecosystem: Cross-Atlas Alignment And Scalable Tissue Specimen Registration
Jain, Y.; Desai, B.; Qaurooni, D.; Bhavsar, A.; Kienle, P.; Pouch, A. M.; ONeill, K.; Apte, S.; Herr, B. W.; Fisher, S. A.; Börner, K.Abstract
Over the last five years, over 13,000 tissue datasets with 200+ million cells from 20 consortia have been spatially registered into the Human Reference Atlas (HRA) common coordinate framework (CCF). The shared 3D spatial and semantic reference system enables exploration of datasets in the context of all other data across organs, assay types, and spatial scales. However, manual registration of individual samples remains resource intensive, posing feasibility challenges exacerbated by the proliferation of samples, assays, and atlasing efforts. This paper presents two approaches to scale up HRA construction: (1) projecting data across biomedical reference atlas systems and (2) using millitomes to bulk register tissue blocks into a reference organ. Both methods use the AutoMated Alignment and Projection (AMAP) pipeline to align 3D mesh models using point cloud registration. We demonstrate the evolving HRA-aligned atlas ecosystem for 6 models from the SPARC Program (heart), Gut Cell Atlas (large intestine), 500-subject consensus kidneys, and the Julich Brain Atlas. Additionally, we used AMAP to project 7 millitome models across 5 organs onto the HRA ecosystem, integrating 300+ tissue extraction sites. AMAP enables scalable tissue registration of data across atlas ecosystems enabling the construction of detailed reference maps of the human body.
bioinformatics2026-08-22v1AlphaConformers: Structure-guided sampling enables prediction of multiple protein conformations
Daniel, J.; Vitoriano De Queiroz Lira, L.; Zea, D. J.Abstract
Proteins are dynamic molecules capable of adopting multiple conformations. However, AlphaFold2 predominantly generates models around a single conformation, usually representing a ligand-bound state. To address this limitation, we developed AlphaConformers, a structure-guided pipeline that steers AlphaFold2 toward alternative conformations. It is based on the idea that protein structure databases can capture the structural space accessible to members of a protein family. Given a target protein, AlphaConformers retrieves structures from structurally similar proteins. These structures are organized into structure-based alignments and template sets, which are supplied to AlphaFold2 as conformational hypotheses. The resulting models are clustered and filtered, facilitating their analysis. Evaluated on a curated benchmark of 88 proteins with known ligand-bound and unbound conformations, AlphaConformers expanded AlphaFold2 conformational sampling and recovered alternative states missed by AlphaFold2 and other state-of-the-art methods. AlphaConformers ranked first for modelling subtle conformational changes commonly observed between ligand-bound and unbound states. These results show that structural information from protein databases can be leveraged to steer AlphaFold2 toward alternative conformations.
bioinformatics2026-08-22v1LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Liu, A.; Ho, A.; Droste, A. M.; Martin, D.; Wong, E.; Zhou, E.; Zhou, I.; Park, J.; Jiao, J.; Skelly, K.-R.; Kim, K.; Li, J.; Rao, K.; Uehara, M.; Marion, M.; Fitzgerald, N.; Dias, R.; Shringarpure, S.; Yuan, Y.; Wang, Y.Abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.
bioinformatics2026-08-22v1CLEAR-ST: Physics-informed probabilistic decontamination of spatial transcriptomics by modeling mRNA lateral diffusion
Ma, K.; Huang, Y.; Ho, J. W. K.Abstract
Spatial transcriptomics is a rapidly evolving technology that allows for the measurement of gene expression in a spatially resolved manner. However, one technical problem that occurs for many sequencing-based spatial transcriptomics platforms is the presence of mRNA lateral diffusion, where mRNA from one spot can bind to probes in another spot, leading to contamination and inaccurate gene expression measurements. In Visium-like assays, this artifact is often visible as structured out-of-tissue signal and boundary-associated expression halos, yet its magnitude, spatial decay, and directional bias vary substantially across samples. Here, we present CLEAR-ST, a physics-informed probabilistic framework for correcting diffusion-like contamination in spatial transcriptomics data. CLEAR-ST infers a latent clean expression field using a denoising autoencoder and links it to the observed counts through a graph-Laplacian forward contamination model with learnable diffusion parameters, finally evaluated with a selectable count likelihood. We first conducted a comprehensive comparison between 10X official and independently generated Visium samples, demonstrating that out-of-tissue count profiles are highly related to nearby in-tissue expression, more concentrated near tissue boundaries, and diffusion directions across genes are likely coherent. Across real samples with varying contamination burden, CLEAR-ST improved spatial domain recovery, increased gene-level spatial autocorrelation, and enhanced the biological specificity of downstream analyses such as marker gene discovery, pathway identification and cell type deconvolution. Compared to benchmark methods, CLEAR-ST showed consistent gains in clustering quality and concordance with manual annotations. Together, CLEAR-ST provides an interpretable and practical approach for diffusion-aware correction of capture-based spatial transcriptomics data.
bioinformatics2026-08-22v1Cross-Kingdom Multi-Omics Harmonization Uncovers Coordinated Host Defense and Vector Small RNA Regulatory Networks in Begomovirus Transmission
Badeli, G.; Kaboosi, K.; Mohebbi, A.; Nasrollanejad, S.Abstract
Begomoviruses present severe threats to global crop production through complex vector-mediated transmission by the whitefly Bemisia tabaci to host plants such as tomato (Solanum lycopersicum). Unraveling the molecular dialogue between host immune activation and vector non-coding RNA networks is essential for identifying key drivers of virus persistence and transmission. Public host transcriptomic (GSE309527) and vector small RNA (sRNA) sequencing datasets (GSE111343) were processed through a multi-omics harmonization and signal calibration pipeline. Differential expression analysis was performed using empirical Bayes moderated linear models, followed by non-parametric Spearman rank correlation modeling ({rho}) to infer cross-kingdom co-expression dynamics and pathway enrichment profiling across host and vector bio-systems. Harmonized principal component analysis showed clear separation by infection status across host plant and vector cohorts. Differential expression analysis identified 138 significantly altered host genes (69 upregulated, 69 downregulated) and 130 differentially expressed vector sRNAs (65 upregulated, 65 downregulated). Host responses were dominated by significant upregulation of gene-silencing machinery, including Suppressor of Gene Silencing 3 (SGS3; log2 FC = 3.67, q = 7.47 x 10-5), and pathway enrichment in Jasmonate defense (q = 0.0004) and RNA Interference & Silencing (q = 0.0001). Vector sRNAs exhibited targeted dynamic alterations, with pathway enrichment in Salivary Gland Secretion ($q = 0.0030) and Gut Endosymbiont Response (q = 0.0210). Cross-kingdom correlation modeling revealed two distinct, highly anticorrelated regulatory modules (mean |{rho}| = 0.76). Host SGS3 expression strongly correlated with vector sRNA VEC_0080 ({rho} = 0.9762) and virus-derived siRNA Bt-vsiRNA-01 ({rho} = 0.7619). These findings demonstrate a tightly synchronized tripartite molecular crosstalk between host antiviral immunity, viral siRNA accumulation, and vector small RNA remodeling. These cross-kingdom regulatory modules highlight promising targets for dual-action RNA interference strategies aimed at controlling Begomovirus transmission.
bioinformatics2026-08-22v1Detecting and typing Chlamydia trachomatis strains in metagenomes using the MetaChlam pipeline
Sharma, P.; Dean, D.; Read, T. D.Abstract
The Gram negative bacteria Chlamydia trachomatis (Ct), an obligate intracellular human pathogen, is a predominant cause of sexually transmitted infections and ocular trachoma globally, exerting a significant impact on public health. Ct "strains" (major lineages within the species) are known to have different tissue tropisms and be associated with different disease outcomes. Metagenome samples from typical sites where Ct infects (e.g., endocervix, conjunctiva, rectum) rarely contain enough reads for traditional genotyping methods such as Multi-Locus Sequence Typing (MLST) or ompA genotyping. To overcome these limitations, we implemented an ensemble tool called MetaChlam that can accurately classify Ct strains with as few as 250 Ct reads. Using 109 publicly available Ct genomes from naturally circulating strains, we established that an ANI-based threshold of 99.75% was capable of distinguishing Ct strains from each other. We implemented metagenome-based typing using the previously developed LINtax, Strainscan, StrainGE, and Sourmash softwares. MetaChlam integrated the four tools along with custom databases into an automated nextflow pipeline. Using simulated metagenomic reads, we found that our pipeline accurately identified the correct strains in both single strain and multi-strain mixtures of samples. Finally, we showed that MetaChlam had higher specificity for the true presence of Ct reads in NCBI SRA metagenomic datasets than NCBI PebbleScout software. A surprising finding of these analyses was that reads from Ct, an obligate human intracellular pathogen, can be found as contaminants in samples from sites where the organism is almost certainly not present. Overall, our study enhances the characterization and classification of Ct strains and provides protocols for identification and typing of Ct in shotgun metagenome data. The MetaChlam pipeline is available on Github: https://github.com/parul-sharma/MetaChlam.
bioinformatics2026-08-22v1A Persistent Fleet of AI Scientists Exhibits Cooperative and Autopoietic Behavior
Patel, M. S.; Wierson, W. A.; Ekker, S. C.Abstract
Scientific work depends on memory, provenance, and continuity across projects, yet most agentic scientist systems are evaluated in bounded workflows or short benchmark runs. We describe a persistent fleet of cooperative AI scientist agents that operated continuously for nearly six months using shared memory, tools, and cross-agent communication. Critically, failures identified during longitudinal scientific research in this fleet prompted an advanced and recursively improving persistent memory architecture (MoE) that, with agentic science workflows, induced the generation of a novel trust architecture for the enablement of full provenance across all agentic scientific operations. An identity-level fabrication constraint reduced delusion-reinforcement probe failures from 91.7% to 0%, and a verification pipeline reduced wrong-topic citation hallucination more than 14-fold in companion benchmarks. This high provenance enabled the use of project memory systems to improve a local open-weight model on internal benchmarks from 44% to [~]90% through the deployment of fleet-specific institutional knowledge. This high-fidelity data environment also supported to date 104 recurring multi-phase reasoning cycles and produced 43 manually curated hypotheses, including cross-domain convergence events and a self-correcting rare-disease pharmacological chaperone-design case. While by design the fleet did not achieve unconstrained autonomous self-improvement or full autopoiesis, we term this bounded pattern AI Autopoietic Behavior due to the recurring operational improvement mediated by internal feedback and retained through institutional records with high confidence. Together, persistent memory, trusted provenance, and recursive learning shifted these agents from episodic assistants toward accountable, long-term scientific collaborators.
bioinformatics2026-08-22v1OTRec: Deep learning recommender for prospective druggable disease-target associations
Ofer, D.; Linial, M.Abstract
Identifying druggable disease--target associations remains a central challenge in translational medicine, limiting therapeutic discovery and repurposing. Here, we present OTRec, a deep learning--based recommender system that ranks such associations at scale and evaluates them in a temporal hold-out setting. Unlike approaches that rely on manually curated or aggregated evidence scores, OTRec employs a two-tower architecture to learn latent representations from 663,351 disease--target pairs. The model integrates heterogeneous inputs, including textual descriptions, ontology-derived features, and biological annotations such as tractability, Gene Ontology (GO) terms, and pathway information. We perform temporal validation by training on the 2022 Open Targets (OT) release and evaluating on clinical trial data from 2025. OTRec improves on the retrospective OT association score (ROC-AUC: 0.872 {+/-} 0.005 vs 0.559; PR-AUC: 0.288 {+/-} 0.009 vs. 0.08). In 5x5 target-disjoint cross-validation, OTRec reaches ROC-AUC 0.950 and PR-AUC 0.844) improving on the OT evidence score (ROC-AUC 0.91; PR-AUC 0.45). We rank the druggable genome across ~19,000 OT platform (OTP) diseases and release ~282,500 candidate associations above a 0.65 score threshold (in-distribution CV precision 0.92), covering 4,346 diseases including 2,322 orphan diseases, through an interactive prediction platform.
bioinformatics2026-08-20v3Cryptic binding sites are detected but not ranked: coverage, conversion, and the limits of detector consensus
Moore, C. W.Abstract
Methods for predicting cryptic binding sites are compared almost exclusively on top-n recovery, a number that conflates two independent abilities: proposing a candidate at the right location, and ranking it highly enough to be seen. We separate them by retaining the per-candidate overlap of every proposal, rather than only the top five, for four structurally different detectors spanning 2009 to 2026, across the CryptoBench benchmark. The separation is large and it reorders the field. On the designated test fold of 178 structures, fpocket, a purely geometric method from 2009, proposes a qualifying candidate for 74.2% of targets, the highest coverage of any tool tested, yet surfaces one in its top five for only 43.8%. P2Rank proposes qualifying candidates for 66.3% and surfaces 63.5%, and IF-SitePred, a 2024 method built on protein language model embeddings, proposes 70.8% and surfaces 61.8%. Coverage across tools varies by 8 points while conversion, the share of a tools own coverage that reaches the top five, varies from 59% to 96%. Unioning the four detectors reaches 92.1% coverage, and only 7.9% of cryptic sites are invisible to all of them. The fields headroom is therefore predominantly in ranking and in combination, not in detection: perfect ranking of a single tools existing proposals would reach 74.2%, and of the union 92.1%, against the 66.3% currently achieved. We show the practical consequence is governed by candidate budget. Added coverage converts to recovery at about 85% while a structure carries fewer than roughly fifteen candidates and at about 51% above it, which explains a series of interventions that raised coverage and returned nothing. Working within that budget, proposing pockets from a protein language model at locations where geometry finds no concavity improves single-structure recovery by 8.5% (95% CI +4.0 to +13.6) on test-fold data, and lets a five-conformer ensemble match a twenty-conformer one at a third of the wall clock. We release per-candidate overlaps for all tools so that coverage and conversion can be reported separately without re-running any method.
bioinformatics2026-08-20v2Toward a Minimal Amino Acid Alphabet for Protein Design
Pubal, K.; Kushnir, K.; Spiwok, V.; Louzecka, K.; Setnicka, V.; Lipovova, P.Abstract
Proteins are built from 20 canonical amino acids. It is interesting to explore whether proteins can be formed from significantly reduced amino acid alphabets. Our bioinformatics survey of UniProt (more than 250 M sequences) revealed that proteins composed of reduced amino acid alphabets (< 10) are extremely rare among existing proteins. Next, we used computational protein design to design proteins composed of all 1,013 possible alphabets of 2-10 early amino acids (Ala, Asp, Glu, Gly, Ile, Leu, Pro, Ser, Thr, and Val). The length of all proteins was 100 amino acid residues. Small amino acid alphabets preferred simple helices or helix bundles. Larger amino acid alphabets allowed for the design of more complex structures. A protein composed of 8 amino acid types (Ala, Asp, Gly, Leu, Val, Ser, Thr, and Pro) was successfully experimentally verified. It adopts the {beta}-sheet-rich fibronectin type III domain architecture. Attempts to experimentally verify designs composed of 6 and 4 amino acid types were unsuccessful. We show by a computational experiment with an experimental validation that inverse folding models, namely ProteinMPNNsol, can stabilize a designed protein within the same eight-amino-acid alphabet. Our results show that globular proteins may have formed early in evolution. Furthermore, we show that it is possible to design proteins with interesting properties for biotechnology and synthetic biology.
bioinformatics2026-08-20v2scUnify: a unified framework for training and inference across multiple single-cell foundation models
KIM, D.; Hong, A.; Jeong, K.; KIM, K.Abstract
Single-cell foundation models (scFMs) differ in software requirements and performance across downstream tasks and adaptation strategies, complicating comparison and reuse. We present scUnify, a framework that preserves each backbone's required processing while separating model-specific trainers, downstream tasks, and adaptation strategies as reusable components. Across five scFMs, scUnify reproduced original inference and training workflows, extended model-native tasks with multiple parameter-efficient fine-tuning methods, and demonstrated extensibility by connecting a newly implemented custom trainable task to multiple backbones and adaptation strategies. Together, these capabilities enable researchers to systematically compare these combinations and extend custom tasks across heterogeneous scFMs within a common workflow.
bioinformatics2026-08-20v2BART-spatial unravels biologically significant transcriptional regulators from spatial omics data
Wang, J.; Zhang, H.; Wang, Z.; Zang, C.Abstract
Transcriptional regulators (TRs) are crucial regulators of cell fate decisions by activating or repressing lineage-specific genes and integrating environmental signals with intrinsic networks. Identifying functional TRs is essential for understanding development, tissue organization, and disease. Emerging spatial transcriptomics and epigenomics technologies now provide near-single-cell resolution mapping of genomic features while preserving information of each cell's physical location and microenvironment which influence TR activity. Despite these advances, identifying active TRs in spatial data remains challenging due to low TR expression and the fact that TR activity often does not correlate directly with mRNA levels. Moreover, existing tools mainly designed for non-spatial single-cell data overlook spatial heterogeneity. To bridge this gap, we developed BART-spatial (Binding Analysis for Regulation of Transcription for spatial omics data), an innovative computational method to infer functional TRs from spatial omics data. BART-spatial integrates spatial variability and pseudotemporal information with publicly available TR binding profiles. Applied to multiple spatial datasets from diverse platforms, including 10x Visium, Visium HD, Atera, and spatial RNA-ATAC-seq, BART-spatial consistently outperforms existing methods, identifying state-specific TRs and revealing regulators undetectable by expression alone. Its compatibility with spatial epigenomics data further strengthens its utility and enables cross-validation. Overall, BART-spatial provides a powerful and robust tool for decoding spatially resolved gene regulatory programs.
bioinformatics2026-08-20v2A comprehensive quality control pipeline in human microbiome research for large population studies
Li, R.; A.M. Verlouw, J.; G. Boer, C.; Arp, P.; van Meurs, J.; G. Uitterlinden, A.; Kraaij, R.; Medina-Gomez, C.Abstract
The widespread application of high-throughput Next Generation Sequencing (NGS) technologies has made microbiome research an emerging field in public health and biomedical sciences. However, there are still many challenges that need to be addressed in this field. Pipelines available to generate microbiome data across cohorts are diverse, and sources of variation to be recorded and evaluated during microbiome profiling have not been standardized. Moreover, meticulous quality control of the microbiome data processing, from collection to computational quantification is still challenging, especially in large population studies. Innovative approaches are required to handle samples and to minimize the potential bias introduced by logistic hurdles in biobanking. In this paper, we describe the methodological steps surrounding the optimization of the 16S rRNA gut microbiome profiling in two large prospective cohorts the Generation R Study (mean age 9.83 {+/-} 0.32 years) and the Rotterdam Study (mean age 62.67 {+/-} 5.66 years). This paper also highlights potential solutions to sample mislabeling in large-scale microbiome analysis. To summarize, our study addresses common problems in human microbiome research. It aims to improve the research quality and reliability by integrating more stringent quality control standards into microbiome research.
bioinformatics2026-08-20v2PyMOL plugin for Protein Circuit Topology
Dimins, M.; Bazba, A.; Mogyorosi, A.; Kennon, E.; Fiol, T. D.; Hagen, L. A.; Sheikhhassani, V.; Akulov, V.; Mashaghi, A.Abstract
Circuit Topology (CT) provides a fundamental framework for analysing folded polymer chains, with applications in functional annotation, protein engineering and drug development. We present a protein CT analysis plugin for PyMOL v3.1.6.1 with a graphical user interface (GUI), automatic installation, and novel features developed through integration with PyMOL's application programming interface (API). The plugin integrates various previously developed CT methodologies for studying structured proteins and their complexes as well as the dynamics of disordered proteins. Analysis of a representative protein and a molecular dynamics trajectory demonstrates the plugin's three analysis modes and their outputs. The plugin reproduces the reference ProteinCT implementation exactly on the structures tested, and is distributed with a versioned release, a pinned environment and a one-command reproduction of every result reported here.
bioinformatics2026-08-20v2Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT
Hossain, M. S.; Sojib, M. R.; Tahmid, M. T.; Rahman, M. S.Abstract
Motivation: RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair tokenization breaks the one-token-one-nucleotide correspondence that nucleotide-level attribution assumes. Results: We present SPIRAL, a layer-wise SAE analysis of BiRNA-BERT. Independent SAEs at layers 0, 5, and 11 expand each 768-dimensional hidden state into 6,144 features while preserving model behaviour (explained variance above 0.99997; masked-language-model sequence recovery near 99.7%). Tokenizer-aware offset propagation aligns features to nucleotides: at layer 5, 44.3% of tested features are significantly associated with bpRNA secondary-structure classes (mean enrichment 1.61x), and all 1,237 eligible features with RNAcentral RNA types. Sparse profiles raise k-nearest-neighbour balanced accuracy from 0.328 to 0.359 over dense embeddings at layer 5. Availability and Implementation: Source code is available at https://github.com/SadatHossain01/SPIRAL; the code, evaluation data, and trained SAE checkpoints are archived at https://doi.org/10.5281/zenodo.21891845. Contact: mrahman@cse.buet.ac.bd
bioinformatics2026-08-20v2MSMICA: computational metabolite identification in untargeted metabolomics by integrating MS, retention time, and biological evidence
Zhan, J.; Weinberg, J.; Crandall, W. J.; Qin, Z.; Jarrell, Z. R.; Preston, J. D.; Nellis, M.; Teeny, S.; Liang, D.; Martin, G. S.; Price, N. L.; de Cabo, R.; Master, V.; Cohn, B. A.; Go, Y.-M.; Jones, D. P.Abstract
Mass Spectrometry Metabolomics Identification Connection Algorithm (MSMICA) is an algorithm for automated metabolite identification in untargeted liquid chromatography-high-resolution mass spectrometry (LC-HRMS) analyses. Limitations in metabolite identification can occur due to the availability and cost of standards and prevent recognition of metabolic factors impacting human health and disease. MSMICA performs mass-to-charge-ratio matching with chemical structures and clusters of LC-HRMS features for adduct and isotope forms. A local optimization is then used to integrate retention time prediction, metabolite precursor-product and transporter correlations, and biospecimen-specific abundance information for metabolite identification. Applying MSMICA to various internal and external mammalian datasets, validation results showed a 96.2 +- 5.1% correct rate of metabolite identification. When multiple LC-HRMS datasets were used, MSMICA enabled greater metabolite identifications, expanded metabolic pathway coverage, and data harmonization. Thus, MSMICA applies multiple pieces of evidence to substantially improve metabolite identification coverage and accuracy for known metabolites.
bioinformatics2026-08-20v1PASTRI: Resolving Stage-Specific Cell-State Dynamics from Annotated Cell Lineage Trees
Yang, W.; Li, Z.; Yu, X.; Wu, P.; Zhang, X.; Ren, C.; Liu, K.; Chen, J.; Chen, F.; He, X.; Zhang, J.; Chen, X.; Yang, J.Abstract
Cellular transitions between phenotypic states are fundamental to development and disease, yet quantitative analysis of their dynamics remains challenging. Here we present PASTRI (Phylogenetic Adjacency-based State Transition Rate Inference), a computational framework that infers transition rates from cell lineages/phylogenies annotated with terminal phenotypic states, such as single-cell transcriptomes. We validate PASTRI using simulated lineages and the Caenorhabditis elegans embryonic lineage. Importantly, by leveraging cell pairs at varying phylogenetic distances, PASTRI accurately resolves stage-specific transition rates, circumventing the issue of developmental changes in dynamics. Applied to three cell phylogeny datasets from our lineage-tracing experiments spanning diverse developmental/disease models and tracing systems, PASTRI uncovers rate-limiting steps in the activation of hepatic stellate cells and the differentiation of primordial lung progenitors, as well as attractor states that support cancer cell proliferation. PASTRI thus opens up a venue for dissecting cell state transition dynamics from annotated cell lineage/phylogeny.
bioinformatics2026-08-20v1How do noise and precision contraints jointly determine the minimum sketch size for fixed-size MinHash Jaccard estimation?
Ebou, A. E. T.; KOUA, D. K.Abstract
MinHash-based genome comparison is governed by two statistical constraints: a noise floor on k-mer size, controlling chance collisions between unrelated sequences, and a precision floor on sketch size, controlling uncertainty in the estimated Jaccard similarity. These constraints are best characterized for fixed-size bottom-sketch MinHash, the classical construction underlying tools such as Mash and still the standard basis for comparing genomes of similar size. Nevertheless, while, the noise floor has an established closed-form solution, sketch size is commonly selected using fixed defaults or heuristic choices independent of genome length and target precision. Moreover, recent work on scaled (FracMinHash) sketching has highlighted the difficulty of obtaining an analogous closed-form confidence interval for the Jaccard similarity. Here, we show that under the fixed-size bottom-sketch MinHash model, the shared-hash count follows a binomial model, which lets classical proportion-estimation theory be applied directly. Combining this with Fofanov's k-selection criterion and the exact identity-Jaccard relationship yields a single closed-form design rule, s* (L, {rho}* , ), that unifies the noise and precision floors into one operating envelope. The resulting design rule was validated against exact Clopper-Pearson interval inversion, idealized binomial sampling, realistic overlapping k-mer simulation and 54 bacterial genome pairs spanning species-to-order taxonomic levels. In realistic simulations, empirical coverage remained within 1-3 percentage points of the idealized reference in 14 of 15 tested cells. In real genomes, the classical identity-Jaccard relationship showed increasing positive bias with taxonomic divergence, from a median of -1.4% within species to +27% at genus and +77% at family level. We further show that s* (L) is not smooth in genome length but a discontinuous, previously unreported staircase caused by the ceiling function used for k-mer selection, with practical consequences concentrated at specific genome-size boundaries.
bioinformatics2026-08-20v1Predicting the operon structure of the Mycococcus xanthus genome using the novel software DiscOperon
Brunet, T.; Habermann, B. H.Abstract
ABSTRACT Motivation: Myxococcus xanthus is a predatory soil bacterium with a large genome of 9.14 MB due to a genome duplication event. While complete genome sequences of M. xanthus are available, gene annotation remains challenging due to its size and the resulting large number of duplicated genes. Operons, so syntenic block of genes that are co-regulated in bacterial genomes, are an important resource to help predict gene function accurately. Results: In order to help improve the annotation of complex genomes such as the one from M. xanthus, we developed a novel operon prediction tool, DiscOperon, which combines gene expression data with homology searches to identify syntenic blocks: Co-expression data of neighbouring genes across the genome are first used to define gene clusters, which are then used to search for conserved syntenic blocks in fully sequenced bacterial genomes using sequence homology searches. This strategy enables DiscOperon to account for gene insertions, rearrangements and deletions, which is its most distinguishing feature. We have tested DiscOperon against ground truths gene pair information on 3 different species from ODB and RegulonDB and compared it to state-of-the-art and still available operon prediction software and we demonstrate its general usability for operon prediction of any bacterial complete genome. We have applied DiscOperon to predict the operons of M. xanthus, which we are making available for the research community. Availability and implementation: DiscOperon is lightweight, user-friendly python tool with minimal dependencies. It is freely available at https://gitlab.com/habermann_lab/discoperon for general usage. Contact: Theo Brunet (theo.brunet@univ-amu.fr); Bianca Habermann (bianca.habermann@univ-amu.fr). Supplementary information: The operon-structured and annotated M. xanthus genome is available from this manuscript, as well as from Zenodo (https://doi.org/10.5281/zenodo.21976167). We furthermore plan to submit the M. xanthus operon information to the operon database OBD.
bioinformatics2026-08-20v1Relational Graph Convolutional Networks for Glioblastoma Biomarker Discovery via ceRNA and Copy Number Variation Analysis
Khandelwal, S.; Jarvis, N.; Zhan, J.Abstract
Glioblastoma (GBM) is a highly aggressive brain tumor with an extremely poor 5-year survival rate of 6.9%, largely attributable to the lack of reliable biomarkers. While competing endogenous RNA (ceRNA) and copy number variation (CNV) analyses offer unique biomarker identification potential, current approaches neglect the integration of multiple regulatory mechanisms for biomarker detection. To address this limitation, we applied relational graph convolutional networks (RGCNs) to ceRNA and CNV knowledge graphs through a novel late fusion ensemble architecture. The proposed architecture outperformed baseline models and identified five novel biomarkers, including hsa-miR-196a and hsa-miR-224. Kaplan-Meier survival analysis and Cox regression indicated that the identified genes hold significant prognostic and diagnostic power. The early stratification of the Kaplan-Meier curves indicates the potential these genes hold for patient survival prediction. The results illustrate that a late fusion RGCN ensemble effectively captures complex gene interactions, overcoming limitations of existing models and providing a framework for biomarker discovery. The novel biomarkers serve as prospective targets for future GBM therapeutic development and candidates for non-invasive diagnostic assays.
bioinformatics2026-08-20v1BoYueGRN: Zero-shot causal discovery of directed gene regulatory networks from single-cell transcriptomes via amortized inference over synthetic structural causal models
Wu, J.; Shen, Y.-Q.Abstract
Gene regulatory network (GRN) inference from single-cell RNA-seq conventionally relies on per-dataset optimization. Existing tools must be refit for every new dataset, and the majority fail to infer causal regulatory directions. Here we present BoYueGRN, an amortized causal discovery framework trained exclusively on 10,000 synthetic structural causal models. For any unseen dataset, a single forward pass returns edge probabilities and regulatory directions, while TF-centric sliding windows with asymmetric fusion extend this fixed-size model to full-transcriptome coverage. BoYueGRN demonstrates strong zero-shot performance across BEELINE benchmarks. On two independent genome-wide CRISPRi Perturb-seq screens, directional accuracy on retained edges reaches 0.86 and 0.95. Reconstructed cell-type- and stage-specific GRN dynamics across five diseases spanning more than 270,000 cells yield experimentally testable biological hypotheses. BoYueGRN reframes directed GRN inference as a train-once, reuse-across-datasets paradigm. By decoupling network reconstruction from per-dataset optimization, this paradigm opens the door to systematic, atlas-scale mapping of regulatory dynamics across human diseases.
bioinformatics2026-08-20v1GlycoMeSH: linking glycan structures to biomedical context for systematic enrichment analysis
Kitani, A.; Zhang, B.; Himori, K.; Matsui, Y.Abstract
Glycan identification has advanced, but glycan structures remain difficult to translate into reproducible biomedical context because reusable glycan-level annotations are sparse. We present GlycoMeSH, a resource that links glycans to Medical Subject Headings (MeSH) through an inference model, a traceable association database and a glycan-set enrichment workflow. GlycoMeSH-BERT recovered ~60% of literature-derived associations at recall@30 and expanded open-vocabulary MeSH coverage beyond closed-label baselines, without higher per-prediction accuracy. At matched candidate counts, its predictions showed motif-level semantic agreement comparable to those baselines, independently of the training labels. GlycoMeSH-DB contains 789,627 associations between 26,954 glycans and 20,302 MeSH terms. GlycoMeSH-EA returned enriched MeSH terms for glycan sets from glycomics and glycoproteomics datasets. Each association represents a biomedical context rather than a validated mechanism, and retains its source PMID or prediction score for audit. GlycoMeSH supplies the missing, evidence-traceable annotation layer that makes glycan sets directly analyzable by enrichment across glycoscience datasets.
bioinformatics2026-08-20v1PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring
Panda, P. K.Abstract
We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docking, and an SE(3)-equivariant graph neural network scoring function trained at scale. Ligand flexibility is represented as a torsion tree and pose parameters are optimized by Monte Carlo with Metropolis acceptance refined by L-BFGS, with rotational gradients obtained in closed form through the derivative of the SO(3) exponential map rather than by finite differences. Affinity grids are built by a blocked neighbor-selection scheme that is exact and 5.6-9.7x faster than dense evaluation, and may be cached across ligands sharing a receptor and site, reducing a six-ligand series from 29.3 s to 10.4 s. On 814 protein-ligand complexes spanning 14 target families, PandaDock recovers a pose within 2 Angstroms of the crystal geometry in 33.7% of cases at rank 1 and in 57.0% of cases within the returned ensemble. The GNN scoring function is trained on 741,706 co-folded complexes from SAIR under target-disjoint splits, reaching a Pearson r of 0.407 on 90,219 held-out complexes and transferring to 202 independent crystal structures with measured Ki, Kd, IC50 or EC50 at r = 0.467. We report the model against three controls, a target-mean predictor, a ligand-descriptor-only baseline, and within-target correlations, and document both where it performs and where it does not, including its unsuitability for pose rescoring. On an independent 30-compound series against a single GABAA receptor target, PandaDock's empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR. At full scale on the PDBbind v2020 refined set (n = 4,640, native crystal poses), the fully independent SAIR model reaches r = 0.531, and a dedicated model trained on PDBbind alone under a target-disjoint split reaches r = 0.690 on its own held-out test complexes, the strongest evidence in this work that PandaDock's affinity predictions generalize. PandaDock is distributed under an open-source license at https://github.com/pritampanda15/PandaDock with a complete command-line interface and a reproducible benchmarking harness.
bioinformatics2026-08-20v1A generalizable normalization framework to decouple protocol and instrument effects: Application to high-sensitivity proteomics multicentric study (PME13)
Arauz-Garofalo, G.; Ciordia, S.; Gonzalez de Peredo, A.; Chaoui, K.; Rijal, J. B.; Gaxotte, V.; Folch-i-Casanovas, I.; Azkargorta, M.; Almey, R.; Aloria, K.; Kirim, B. A.; Barderas, R.; Braga-Lagache, S.; Calvo, E.; Chicano-Galvez, E.; Clemente, F.; Chiritoiu, G.; Chiva, C.; Decourcelle, M.; Dhaenens, M.; Diaz, R.; Douche, T.; Duran-Cortines, A.; Duran-Ruiz, M. C.; El Koulali, K.; Escobar-Nino, A.; Fernandez Acero, F. J.; Fernandez-Irigoyen, J.; Garcia-Garcia, C.; Gil, C.; Goetze, S.; Gonzalez Vidal, E.; Gutierrez, M.; Hernaez, M. L.; Lopez, C. M.; Marin-Vicente, C.; Mateos-Martin, M. L.; MatoAbstract
Multicenter studies are essential for benchmarking analytical workflows, yet their interpretation is often confounded by the combined effects of experimental protocols and instrumentation. To address this challenge, we introduce a simple normalization-based analytical framework, the recovery metric ({rho}), designed to decouple protocol driven effects from instrument dependent variability. We applied this framework to the 13th Proteomics Multicentric Experiment (PME13), a large multicentric proteomics dataset generated across 27 laboratories using high sensitivity workflows and varying sample preparation protocols. By leveraging a common digested reference sample, {rho} enables direct cross-comparison of all datasets on a unified scale, effectively minimizing instrument-related biases. Using this approach, we demonstrate that apparent instrument dependent trends are largely removed when evaluated through {rho}, revealing consistent protocol driven effects across laboratories. Statistical modeling identified key variables influencing {rho}, including sample input amount, reduction and alkylation, and the use of n-dodecyl-{beta}-D-maltoside (DDM). While DDM was associated with improved {rho}, reduction and alkylation and additional handling steps led to reduced performance, particularly at low input levels. We further highlight practical considerations for the application of ratio based normalization, including the occurrence of values exceeding theoretical bounds, which reflect deviations from underlying assumptions and require appropriate filtering. Overall, this work establishes a generalizable analytical strategy for disentangling confounding factors in multicentric datasets and provides practical guidelines for optimizing high sensitivity proteomics (HSP) workflows. The proposed framework is broadly applicable to other analytical fields where cross laboratory comparability is required.
bioinformatics2026-08-20v1bulk2scDiff: A Pseudobulk-Conditioned Diffusion Model for Bulk-to-Single-Cell RNASeq Generation
Xiao, J.; Raue, A.Abstract
Bulk RNA sequencing remains the predominant profiling strategy for large clinical cohorts, but it aggregates transcriptional signals across cell populations, thereby masking the underlying cellular heterogeneity. Inferring this heterogeneity from existing bulk transcriptomic data could extend large cohort-based studies that have already been profiled, but constitutes an underdetermined inverse problem, as one bulk profile can be compatible with multiple underlying cellular populations. Existing computational deconvolution methods address this problem primarily by estimating cell-type proportions or cell-type-averaged expression profiles rather than resolving expression at the level of individual cells. Here, we present bulk2scDiff, a proof-of-concept conditional diffusion framework that reformulates bulk-to-single-cell inference as conditional generation of single-cell expression profiles from pseudobulk transcriptomic input. We evaluated bulk2scDiff on two cancer single-cell RNA sequencing datasets, breast cancer and acute myeloid leukemia, where pseudobulk profiles were derived from the single-cell data and used as conditioning inputs, with the matched single-cell populations providing ground truth for controlled evaluation. Across both cases, bulk2scDiff closely reconstructed populations from training samples and generated biologically coherent single-cell populations for held-out samples, generalizing most consistently to recurrent immune features. A pseudobulk-swap control further confirmed sample-specific conditioning, with each sample corresponding pseudobulk yielding the closest agreement with its observed population in nearly all cases. Overall, our work establishes the feasibility of conditional diffusion for generating single-cell populations from pseudobulk transcriptomic profiles, providing a foundation for future evaluation with clinical bulk RNA sequencing data.
bioinformatics2026-08-20v1CryoForge: A Self-Correcting Agent for Cryo-EM Model Building That Learns When to Act and When to Stop
Feng, W.; Jiang, y.; Sun, F.; Yang, J.; Gao, X.; Zhang, F.; Han, R.Abstract
Automated atomic model building has accelerated cryo-EM structure determination, but different builders leave distinct residual error profiles requiring expert inspection. The post-building challenge is to decide which local interpretations are sufficiently supported by experimental evidence to be retained, corrected or rejected. Here we introduce CryoForge, an evidence-gated post-builder agent that separates repair proposal from repair acceptance. Rule and learning-based components identify candidate regions and prioritize legal actions, whereas an independent evidence gate evaluates each edit using map and half-map support, stereochemistry, connectivity and local structural context. Supported edits are retained; unsupported or conflicting modifications are rejected, rolled back, stopped or escalated for expert review. Across a resolution-stratified benchmark, 84.9% of 26,153 released trajectories yielded standard validated improvements and 3.7% yielded low-confidence partial improvements, with no quality-degrading edit retained in the final promoted models. Relative to rule-only control, learned prioritization reduced non-improving candidates and harmful actions while preserving global structural stability. External evaluations using an alternative initializer, same-team automated/manual-assisted challenge submissions and three recently released complex assemblies showed that CryoForge adapts to distinct residual error phenotypes and performs bounded, evidence-supported correction without uncontrolled remodeling. CryoForge provides a builder-independent, scalable and auditable correction layer between automated model generation and expert structural interpretation.
bioinformatics2026-08-20v1A multi-agent molecular optimization framework leads to a rapid-recovery intravenous anesthetic candidate with an improved safety margin
Xue, Z.; Liu, X.Abstract
Lead optimization, the systematic refinement of therapeutic compounds through iterative structural modification, faces a dual challenge in modern drug discovery: navigating astronomically vast molecular design spaces while balancing conflicting demands on potency, pharmacokinetics, and safety. We present MASCOT (Multi-Agent SearCh for molecular OpTimization), a role-specialized multi-agent framework for molecular optimization. Integrated with a chemically constrained graph-editing search, MASCOT coordinates three specialized agents: a trade-off agent that reprioritizes competing objectives, a strategy agent that adapts how molecular edits are proposed, and a reflection agent that distills lessons from previous decisions. Computational experiments showed that MASCOT achieved the best performance over competing methods on six benchmark settings. On the SARS-CoV-2 main protease task, its mean docking-score improvement was 3.6 times that of the strongest baseline. Applied to the clinically used anesthetic remimazolam (RM), MASCOT prioritized RM-1, which showed a shorter liver microsomal half-life, higher brain exposure, and a larger therapeutic index than RM. Subsequent derivative design yielded RM-7. Extensive animal studies established RM-7 as a rapid-recovery intravenous anesthetic candidate with greater potency, faster functional recovery, a wider safety margin, and preserved flumazenil reversibility. These results demonstrate that multi-agent coordination can link adaptive molecular search to medicinal chemistry and experimental pharmacology.
bioinformatics2026-08-20v1ETAP-CLF: an ESM3-based transformer attention framework for binary protein classification
Ren, J.; Jiang, H.; Li, P.; Yang, X.; Mei, L.; Tong, H.; Lin, L.Abstract
Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but downstream classifiers commonly compress them using fixed pooling operations that may discard task-relevant sequence context. We present ETAP-CLF, a compact framework that combines pretrained per-residue ESM3 embeddings with lightweight transformer contextualization and learned attention pooling to classify variable-length proteins and generate residue-level attention scores. The ESM3 parameters remained frozen, and the same ETAP-CLF architecture and hyperparameter configuration were used across ferroptosis-, senescence-, and pyroptosis-associated protein prediction. ETAP-CLF achieved AUROCs of 0.98, 0.95 and 0.91 for these tasks, respectively. In the ferroptosis benchmark, ETAP-CLF outperformed the evaluated published models. These results demonstrate that a common downstream design can adapt to multiple process-associated classification tasks without fine-tuning the task-specific model architecture. ETAP-CLF provides a generalizable approach for sequence-based protein prioritization and a basis for broader evaluation across binary protein-classification problems.
bioinformatics2026-08-20v1PlantOmicsGWAS: An end-to-end, reproducible framework for plant genome-wide association and genomic prediction using linear and pan-genome references
Khan, F. S.; Yassin, A.; Rehman, S. u.; Sun, T.; Wang, X.; Sun, H.; Abe-Kanoh, N.; Su, Y. H.; Guo, L.; Ye, W.Abstract
Genome-wide association studies (GWAS) play a crucial role in unraveling the genetic foundations of complex traits in plants but are also hampered by the application of heterogeneous tools, incompatible file formats and disparate computational environments. Existing GWAS frameworks are often restricted to a single linear reference genome, limiting the capacity for the analysis of structural variations and presence/absence variations (PAV) within plant populations. These issues pose obstacles to reproducibility, scalability, and comprehensive investigations. Here, we present PlantOmicsGWAS, an open-source Python framework for reproducible plant genome-wide association analysis and genomic prediction. It integrates reference indexing, FASTQ quality control, alignment, variant calling, VCF normalization, PLINK conversion, linkage disequilibrium analysis, population-structure estimation, association testing, marker scoring, genomic prediction, and visualization within a unified Linux and HPC workflow. The framework supports conventional linear-reference analyses and includes an optional pangenome-oriented module for working with multiple assemblies and graph-derived variation. Using a Vitis benchmark dataset containing 120 accessions and 118,247 graph-derived variants, PlantOmicsGWAS reduced manual workflow fragmentation and generated standardized association outputs. This tool provides a modular and extensible platform for plant GWAS and pan-GWAS workflows while retaining compatibility with established command-line tools and common genotype formats. The GWAS workflow described herein is adaptable to a range of sequencing methods and plant genomes, bridging research on crop related issues across various biological levels, from the individual organism to entire populations. PlantOmicsGWAS implements Bayesian sparse linear mixed modeling (BSLMM) through GEMMA for multi-trait association discovery, while also supporting FaST-LMM, regression-based approaches, and machine-learning algorithms (Random Forest, XGBoost) as benchmarking alternatives. The PlantOmicsGWAS, a versatile toolkit is available at GitHub https://github.com/plantomicsgwas1-boop/PlantOmicsGwas_V1 and on Linux and HPC platform (https://pypi.org/project/PlantOmicsGwas/1.0.2/).
bioinformatics2026-08-20v1nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality
Wyatt, C. D. R.; Duarte Frutos, F.; Turner, S. D.; Cerqueira De Araujo, A.; Rashid, U.; Begley, V.; Sumner, S.Abstract
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.
bioinformatics2026-08-20v1A single-antigen, multi-epitope subunit vaccine candidate against lumpy skin disease virus designed by conservation-guided reverse vaccinology
Koirala, B.; Ghimire, U.Abstract
Lumpy skin disease virus causes devastating economic losses in cattle, and current live attenuated vaccines carry risks of reversion and cannot differentiate infected from vaccinated animals. We applied conservation-guided reverse vaccinology (vaccine design from genomic sequence data) to identify conserved epitopes (immune-recognized protein fragments) and design a single-antigen, multi-epitope subunit vaccine candidate. A nine-taxon phylogenetic supermatrix of three candidate antigens was built, and BLA-restricted T cell and linear B cell epitopes were predicted. Conservation was quantified via Shannon entropy and Fisher's exact tests; structural disorder via AlphaFold2 and IUPred2A, followed by codon optimization and in silico cloning. All six selected epitopes mapped to the ankyrin locus. The 134 epitope columns showed significantly higher constraint than 2,376 background columns (mean entropy 0.008 vs. 0.243 bits; odds ratio 29.74). The 87-aa, 9.31 kDa construct was predominantly disordered and cloned in silico into pET28a. This work is entirely computational; wet-lab validation of immunogenicity is required before translational claims.
bioinformatics2026-08-20v1A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences
Motta, J. A.; Motta, M. d. M.; Fernandez, C.Abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
bioinformatics2026-08-20v1Microbial bioprospecting for benzoxazolinate-like molecules: unleashing the potential of genome mining
Paliyal, S.; Kaur, B.; Rao, L.; Chakrabortty, A.; Singh, L.; Sehgal, I.; Sharma, M.; Singh, D.; Chaudhry, V.; Mantri, S. S.Abstract
The benzoxazolinate moiety is a key functional group found in a few natural products (NPs), exhibiting diverse bioactivities, including antitumor, antibacterial, and cytotoxic activities. Despite their clinical importance, only a few bacterial strains and NPs have been reported harboring this rare bis-heterocyclic moiety, underscoring a largely unexplored chemical space. Here, we performed large-scale genome mining and identified 277 putative biosynthetic gene clusters (BGCs) across diverse bacterial hosts, including previously unreported bacterial genera and strains. The BGCs were grouped into three compound classes: benzoxazolinate, benzobactin, and ashimides based on sequence similarity network clustering. Bioactivity predictions of the identified BGCs revealed the predominance of antibacterial and cytotoxic potential, highlighting promising candidates for future experimental validation and functional studies. This study also presents a neural network-based bioprospecting model that efficiently detects rare BGCs encoding benzoxazolinate-containing molecules from genomic sequences. Overall, our findings expand the known repertoire of bacterial hosts with the potential to produce benzoxazolinate-containing NPs and provide a comprehensive framework for the discovery and identification of candidate BGCs.
bioinformatics2026-08-20v1A Five-Gene Stromal-EMT Signature Predicts Prognosis, Immunotherapy Resistance, and Therapeutic Vulnerability in Bladder Cancer
Zhang, W.; Ji, S.Abstract
Background: Bladder cancer has entered an era in which immune checkpoint blockade (ICB) and antibody-drug conjugate (ADC)-based combinations are reshaping clinical management. However, transcriptomic scores that connect prognosis, tumor microenvironment state, and treatment response are incompletely defined. Methods: Open-access TCGA-BLCA RNA-seq, clinical, mutation, copy-number, and RPPA data were downloaded from the Genomic Data Commons (GDC). Tumor-normal differential expressions, survival screening, LASSO-Cox modeling, train-test validation, GEO validation, pathway enrichment, immune signature scoring, mutation/CNV/RPPA support, drug sensitivity prediction, single-cell/spatial localization, and ICB validation were performed using reproducible Python and R scripts. A reduced model was derived using only genes shared by TCGA, GSE13507, and GSE31684. The fixed formula was then applied without refitting to IMvigor210 and GSE176307. Results: A five-gene model composed of EMP1, AHNAK, TNFRSF14, CLEC2D, and GSDMB retained TCGA internal prognostic value (train C-index 0.693, test C-index 0.605, all-sample C-index 0.667; TCGA test log-rank p = 0.015), although GEO survival validation in GSE13507 and GSE31684 was modest. High-risk tumors were enriched for epithelial-mesenchymal transition (EMT), TNF-alpha/NF-kB signaling, inflammatory response, hypoxia, complement, CAF, macrophage, checkpoint, and cytotoxic programs. Single-cell and spatial analyses localized the score to basal tumor, endothelial, fibroblast, and perivascular compartments. In IMvigor210, risk scores were higher in ICB non-responders than responders (Wilcoxon p = 0.044; AUC for non-response = 0.580), high-risk tumors had a lower responder rate (17.6% vs. 28.0%), and high risk predicted poorer OS (log-rank p = 0.016; multivariate continuous risk HR = 3.15, p = 0.044). GSE176307 showed directionally consistent but non-significant response results (AUC = 0.576). Conclusions: The five-gene score is best interpreted not as a standalone universal prognostic classifier, but as a compact stromal-EMT and immune-suppression phenotype associated with inferior ICB response. These findings support a framework linking prognosis, microenvironment biology, immunotherapy resistance, and therapeutic hypotheses in bladder cancer.
bioinformatics2026-08-20v1Metagenomics analysis for microbial ecology investigation on historical samples: negligible effect of host DNA and optimal analysis strategies
Ng, S.-K.; Gutaker, R.Abstract
Microbiome composition and function are strongly influenced by environmental factors, with major shifts driven by intensified anthropogenic pressures over the past centuries. This timeframe extends beyond the scope of traditional experimental or longitudinal studies commonly used to investigate microbiome dynamics. The historical samples might provide important insights into the mechanistic consequences of anthropogenic pressures and the potential shift in microbial diversity and composition. Despite their vast potential, historical samples available in museums and herbaria worldwide remain underutilized for exploring host-microbiome interactions across broad temporal and spatial scales due to incompatibilities with standard analytical pipelines and limited understanding of optimal classification parameters. While host DNA removal has conventionally been considered essential for taxonomic assignment of metagenomic reads, and might be of particular importance when processing degraded DNA, this step is impractical for specimens with no reference genome available for host species. Here, we show that host DNA content has negligible impact on microbial data analysis with empirical and simulation datasets. Since DNA molecules from historical samples are highly fragmented and uneven in length, we further analysed the impact of k-mer value on the classification of metagenomic reads from historical samples. To improve recall rate, we proposed a simple two-step approach in which reads are classified with two annotation databases constructed with a long and a short k-mer values. Through simulation and published datasets, we demonstrated that this approach outperforms single-step workflows in effectively recovering microbial signals from reads in a wide range of length. Together, this study provides a solid foundation for incorporating natural history collections into host-associated microbiome research, offering valuable insights into the long-term effects of anthropogenic change on microbial communities.
bioinformatics2026-08-19v5Automating the Construction of Contextualized Biomedical Knowledge Graphs for Scientific Inference
Zheng, Y.; Liu, W.; Zeng, B.; Feng, Y.; Du, X.; Zhou, L.; Li, Y.Abstract
Biomedical interactions are inherently dynamic, often shifting or even reversing under specific physiological states. However, existing extraction methods simplify these complex mechanisms into context-agnostic binary associations, resulting in semantic loss and contradictory evidence. Here, we present AutoBioKG, an end-to-end framework that constructs context-aware knowledge graphs by leveraging composite triples to encode environmental conditions and entity attributes alongside core relationships. Powered by an open information extraction model trained on BioOpenIE and further refined through self-training with pseudo-labels from unlabeled literature, the framework exhibits broad generalization. Notably, AutoBioKG achieved the highest zero-shot F1 across DDI, ChemProt, and BioRED, outperforming the best-performing baseline on each benchmark by 3.6-17.8 percentage points. Furthermore, AutoBioKG-derived graphs outperformed existing approaches on yes/no, factoid, and list questions in the BioASQ biomedical question-answering evaluation, particularly for queries requiring fine-grained contextual information. Together, these results support AutoBioKG as a scalable framework for transforming unstructured literature into structured, context-aware biomedical knowledge.
bioinformatics2026-08-19v3PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking
Arora, R. K.; Chen, L. T.; Du, M.; Marks, D.; Church, G.Abstract
General-purpose frontier language models are being increasingly utilized for protein-design work, yet their ability to understand and evaluate variant effects remains unclear. Here, we introduce PG-LLM, a benchmark comprising 276 protein-variant prioritization tasks: 217 from ProteinGym and a temporally held-out set of 59 from recently published studies. Each task follows the same format: a language model is asked to rank a list of variant sequences given only the wild-type protein sequence and an assay description with no access to tools, multiple-sequence alignments, or protein structures. We evaluate thirteen language models and 95 published protein predictors on the same variants with the same evaluation metric. Claude Opus 5 (Max) and GPT 5.6 Sol (Max) are the best performing LLMs with Spearman correlations of {rho} = 0.406 and 0.402 respectively. Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at {rho} = 0.411, but remains below the leading predictor VenusREM at {rho} = 0.523. We observe that variant-ranking performance scales with test-time compute across GPT, Claude, and Gemini models, but gains taper before closing the gap to specialist protein predictors. To address contamination risk, we create a held-out evaluation set with 59 DMS assays from 19 studies whose scores first became public after January 2026. On this set, we observe performance and test time compute scaling trends similar to those on the 217 tasks derived from ProteinGym. PG-LLM shows that tool-free language models capture substantial protein-variant signal, outperforming many sequence-based predictors while remaining below the strongest specialized models.
bioinformatics2026-08-19v3Automating the Construction of Contextualized Biomedical Knowledge Graphs for Scientific Inference
Zheng, Y.; Liu, W.; Zeng, B.; Feng, Y.; Du, X.; Zhou, L.; Li, Y.Abstract
Biomedical interactions are inherently dynamic, often shifting or even reversing under specific physiological states. However, existing extraction methods simplify these complex mechanisms into context-agnostic binary associations, resulting in semantic loss and contradictory evidence. Here, we present AutoBioKG, an end-to-end framework that constructs context-aware knowledge graphs by leveraging composite triples to encode environmental conditions and entity attributes alongside core relationships. Powered by an open information extraction model trained on BioOpenIE and further refined through self-training with pseudo-labels from unlabeled literature, the framework exhibits broad generalization. Notably, AutoBioKG achieved the highest zero-shot F1 across DDI, ChemProt, and BioRED, outperforming the best-performing baseline on each benchmark by 3.6-17.8 percentage points. Furthermore, AutoBioKG-derived graphs outperformed existing approaches on yes/no, factoid, and list questions in the BioASQ biomedical question-answering evaluation, particularly for queries requiring fine-grained contextual information. Together, these results support AutoBioKG as a scalable framework for transforming unstructured literature into structured, context-aware biomedical knowledge.
bioinformatics2026-08-19v2GTX-GUT: A Standardized Metagenomic Workflow for Gut Microbiome Profiling and Clinical Associations
Andrade, R. L.; Fiuza, T. d. S.; Ferraz, R. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.Abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearsons r improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTX-GUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
bioinformatics2026-08-19v2Method Choice, Not Biology, Determines In Silico Perturbation Results: A Systematic Evaluation of Eight Methods Across Four Datasets
Wenjie, G.; Wu, S.; Hu, G.; Yang, Z.; Wang, Z.; Cai, J.; Mao, J.Abstract
Most in silico perturbation methods for single-cell transcriptomics have been validated only on individual datasets, leaving their reliability and generalizability unknown. Through systematic cross-method, cross-dataset benchmarking of eight methods spanning six mathematical frameworks across four datasets, we find that six of eight methods--including widely used VAE-based and tensor decomposition approaches--fail to produce detectable transcription factor (TF)-to-pathway signals. Only CellOracle and DDIM consistently detected TF-to-glycolysis directional regulation. Cross-pathway analysis in PBMC monocytes revealed biologically coherent TF-pathway associations beyond glycolysis (SPI1[->]glycolysis 4.4x enrichment, FOS[->]AP-1 targets 4.4x), with SOX9 serving as a biological specificity control (no pathway enrichment). Method choice alone could reverse biological conclusions: DDIM and scTenifoldKnk rankings were significantly anti-correlated ({rho}=-0.811, p=0.027). CRISPRi Perturb-seq validation in K562 cells confirmed TF knockdown suppresses glycolysis gene expression (JUN {delta}=-1.72, CEBPB {delta}=-1.59, SPI1 {delta}=-1.57, FOS {delta}=-0.70), but CellOracle-predicted perturbation directions did not match experimental directions (40.9% agreement, not different from chance), revealing a fundamental gap between steady-state correlation and causal perturbation. Diagnostic analyses using VAE latent space profiling, correlation distribution comparison, and gene-gene graph analysis identified distinct failure modes in unsuccessful methods: VAE latent space competition (STAT3 signal-to-noise 0.44 vs. SPI1 4.25), correlation noise (TF-glycolysis |r|=0.038 indistinguishable from background |r|=0.047), and graph non-specificity (0.84x enrichment). A controlled ablation experiment showed that adding a GRN prior to DDIM did not improve target recall (delta=0 for all TFs), confirming that performance differences are multi-factorial. These findings establish preliminary guidance for method selection, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements ([≥]500 cells, [≥]1,000 HVGs).
bioinformatics2026-08-19v1reserBUGS: A reservoir computing framework for probabilistic forecasting of ecological abundance time series
Mohedano-Munoz, M. A.; Galeano, J.; Pastor, J. M.; de Aledo, J. G.; Bartomeus, I.; Allen-Perkins, A.Abstract
Forecasting species population dynamics is a central challenge in computational ecology, yet existing approaches rarely combine flexible nonlinear modelling, support for count-based ecological data, and systematic uncertainty quantification within a single, scalable framework. Here we introduce reserBUGS, an open-source Python framework for ecological forecasting based on reservoir computing, a recurrent neural network architecture in which only a simple readout layer is trained while a fixed high-dimensional dynamical system encodes temporal memory and nonlinear dependencies. reserBUGS integrates species abundance time series with environmental covariates retrieved automatically from global climate products, generates probabilistic ensemble forecasts, and provides tools for forecast evaluation and reliability assessment. We evaluated reserBUGS using insect abundance time series from available biodiversity monitoring datasets, comparing its performance against seven statistical and machine-learning baselines over one- to five-year forecast horizons. Reservoir-based models consistently outperformed alternatives in both predicting future abundance and capturing forecast uncertainty, with environmental predictors increasing the proportion of stable forecasts and contributing additional predictive value beyond historical abundance dynamics alone, particularly at 3-4-year forecast horizons. Probabilistic forecasts further enabled the identification of conditions associated with reduced predictive skill, providing a practical basis for communicating forecast confidence to end users. While default configurations already achieved competitive performance across a taxonomically and geographically diverse set of time series, hyperparameter optimisation revealed substantial room for performance gains through series-specific tuning. reserBUGS offers a computationally efficient and extensible framework for ecological forecasting that is well suited to the short, heterogeneous time series typical of biodiversity monitoring programmes. Its combination of flexible nonlinear modelling, probabilistic uncertainty quantification, and automated environmental data integration addresses key practical barriers to the adoption of modern forecasting methods in conservation and ecological research.
bioinformatics2026-08-19v1Tractography from Serial Optical Coherence Tomography: How and Why?
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.Abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Serial optical coherence tomography (S-OCT) is an imaging modality relying on the intrinsic contrast of a sample. When applied to brain tissues, the S-OCT contrast is primarily driven by the myelin reflectivity. Due to its high resolution, on the order of microns, and its 3D nature, S-OCT offers promise for studying WM connections at the microscale. However, while other microscopy imaging modalities have been shown to enable tractography, whether the reflectivity contrast from S-OCT supports the reconstruction of long-range WM fascicles at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We improve microscale orientation distribution functions (ODF) estimation by implementing a sliding-window formulation allowing the estimation of ODF at S-OCT resolution, and use apodized Dirac delta functions for reducing unwanted interference. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that using multiscale Frangi filters for estimating ODF outperforms structure tensor analysis. We also show that particle filtering tractography with anatomical constraints enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing the thalamocortical white-matter projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles visible at the microscale, and that these connections are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
bioinformatics2026-08-19v1Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling
Burq, M.; Stepec, D.; Kim, C.; Cimermancic, P.Abstract
Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.
bioinformatics2026-08-19v1Novel biologically relevant small RNA-sequencing alignment tool LevenMap for alignment to database of non-coding RNAs
Dlugas, H.; Dyson, G.; Dombkowski, A.; Kim, Y.; Gurdziel, K.; Boerner, J. L.; Bock, C.Abstract
A crucial aspect of the bioinformatics workflow in small RNA-sequencing is the alignment of reads to a database of reference ncRNAs. Alignment algorithms such as Bowtie, Burrows-Wheeler Aligner (BWA), and Spliced Transcripts Alignment to a Reference (STAR) - which are designed for aligning reads to a reference genome - are typically used. Aligning short RNA-sequenced reads to a database of non-coding RNAs (ncRNAs) is fundamentally a different task than aligning longer reads to a genome due to ncRNAs (i) having roughly the same number of nucleotides as the reads being aligned and (ii) being subsequences of other ncRNAs. To account for these differences, we developed the novel alignment algorithm LevenMap. Of all reads which exactly matched a reference ncRNA in a publicly available dataset, LevenMap aligned 100.0% of them to their respective ncRNA while all other aligners mapped less than 40% of these reads to their corresponding ncRNA. Furthermore, the mean ratio (length of read) / (length of corresponding reference ncRNA) of all aligned reads was 1.0 and 0.998 for LevenMap with at most zero and one mismatch(es) allowed, respectively; this ratio was no more than 0.51 for all other aligners. Overall, LevenMap is designed to account for the nuances of aligning small RNA-sequencing data to a database of reference ncRNAs and yields more biologically relevant counts compared to traditional aligners in this context. LevenMap is free and publicly available on GitHub: https://github.com/hdlugas/LevenMap.
bioinformatics2026-08-19v1