Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
From Abandoned Scripts to FAIR Community Pipelines: Rescuing Orphan Bioinformatics Workflows with nf-core - Lessons from Light-Sheet Fluorescence Microscopy
Schwitalla, C.; Kuhn Cuellar, L.; Hoertenhuber, M.; Grote, N.; Woller, T.; Lamberti, I.; Pavie, B.; Kuestner, T.; Kyere, F. A.; Curtin, I.; Stein, J. L.; Nahnsen, S.Abstract
Background Research software is essential for modern data analysis but is often developed and maintained by a small number of researchers. When developers leave, software may become orphaned, limiting reuse and risking the loss of valuable domain knowledge and computational methods. While the FAIR Principles for Research Software (FAIR4RS) provide an essential foundation for improving the reuse of research software, compliance with these principles alone does not guarantee practical reusability. Here, we investigate whether orphaned scientific software can be systematically rescued and transformed into sustainable, reusable workflows using established software engineering practices and community standards. Findings We re-engineered the abandoned MATLAB-based NuMorph toolkit for large-scale light-sheet microscopy image analysis into nf-core/lsmquant, a Nextflow-based workflow developed according to nf-core community guidelines. The re-engineered workflow preserved the original scientific methods at comparable computational cost while improving the software's FAIRness, portability, and reproducibility. Integration into the nf-core ecosystem provides a community-driven framework that supports software sustainability through distributed maintenance and shared development practices, while the modular workflow architecture simplified adaptation of nf-core/lsmquant to additional light-sheet microscopy datasets beyond the original application Conclusion Our work demonstrates that orphaned scientific software can be successfully rescued through systematic re-engineering guided by FAIR and software sustainability principles. By transforming a legacy codebase into a community-maintained workflow, we preserve valuable domain-specific methods while improving usability, maintainability, and reproducibility. This approach provides a practical strategy for recovering orphan research software and integrating it into modern, reusable research ecosystems. Keywords Light-sheet fluorescence microscopy, Nextflow, nf-core, re-engineering, sustainable software
bioinformatics2026-08-11v4RingNet: An Interactive Platform for Multi-Modal Data Visualization in Networks
Zhang, L.; Lai, X.Abstract
The exponential growth of data in biomedicine has created an urgent need for intuitive visualization tools. These tools must be able to effectively represent complex biological networks and remain accessible to domain experts without extensive computational training. Current network visualization approaches often require specialized programming skills and/or cannot handle the scale and complexity of modern biomedical datasets, which creates significant barriers to biological discovery. We develop RingNet, a web-based interactive visualization tool that integrates computational efficiency with flexible, user-driven exploration. This tool addresses the community's need to visualize multi-modal datasets within a single, compact network representation, as well as identify patterns of interest in complex data. RingNet uses an R backend for network computation and coordinate optimization. This generates JSON data structures that feed into a JavaScript and HTML frontend, which provides real-time, interactive visualization functions. It offers dynamic layout adjustments, node and edge filtering, and customizable color schemes for representing data. It can export reproducible, publication-ready figures in SVG and PNG formats. In our case studies, we use RingNet to visualize breast cancer patients' omics profiles in a gene regulatory network and a cell-to-cell communication network in atopic dermatitis. This demonstrates RingNet's ability to reveal biological relationships across multiple data modalities. RingNet lowers the barrier to exploring, analyzing, and communicating data-driven findings, thereby accelerating research.
bioinformatics2026-08-11v4A bio-informatics approach to identify new drug targets in multidrug-resistant bacteria
Bramhill, I.; Chiam, A. J.; de Jong-Hoogland, D.; Ulmschneider, M. B.Abstract
Antibiotic resistance poses a global health crisis. In order to develop new antibiotic agents, it is crucial to identify drug targets in multidrug-resistant bacteria. Criteria for such a target are an -helical, essential membrane protein, that is non-homologues with the human membrane proteome, and present across multiple bacterial species. Using a stepwise subtractive genomics approach, the membrane protein F0F1 ATP synthase subunit C was identified as a non-human analogues drug target that is present in 11 bacterial species.
bioinformatics2026-08-11v3A disease dynamics atlas forecasts patient states and maps molecular programs in ALS
Li, Z.; Gao, C.; Kong, J.; Fu, Y.; Wen, S.; Li, G.; Cao, Y.; Fu, Y.; Zhang, H.; Jia, S.; Liu, X.; Yang, J.; Cai, L.; Yan, F.; Liu, X.; Tian, L.Abstract
ALS progression is multidimensional, yet fragmented records and scalar outcomes obscure how patients move through disease states and how those states relate to molecular variation. MEDSTREM converts patient-held medical-record images into standardised longitudinal data, enabling bottom-up cohort construction. Using MEDSTREM-structured records from more than 8,000 AskHelpU participants together with PRO-ACT and Answer ALS, we developed DynaALS, the ALS Disease Dynamics Atlas. DynaALS represents ALS as a dynamic patient-state manifold that captures distinct directions of deterioration and patient movement between them over time. DynaALS retrieved population-referenced future states and decoded them into multidimensional clinical profiles without requiring patient-specific longitudinal histories. Motor-neuron RNA and chromatin profiles linked DynaALS states to developmental and regulatory programs, while neuromuscular-organoid single-cell multi-omics converged on a neural-developmental Netrin-DCC signalling axis across interacting cell types. By coupling MEDSTREM-enabled data construction to dynamic state modelling, DynaALS establishes a transferable patient-state engine for predictive and biologically interpretable disease models.
bioinformatics2026-08-11v2Spurious correlation inflates performance in single-cell perturbation prediction
Nicol, P. B.; Shivakumar, S.; Irizarry, R.Abstract
The increasing number of computational methods designed to predict the effects of genetic perturbations on cellular gene expression profiles has led to a need for rigorous evaluation metrics. Recent benchmarking studies rely on correlation or cosine similarity of differential expression relative to a shared population of control cells. We show that these metrics are systematically inflated by statistical bias induced by reusing the same control population to define both quantities being compared. As a result, even non-informative methods can appear to perform well, particularly in datasets with limited numbers of control cells. Reanalysis of published datasets using a simple control-splitting procedure that removes this bias leads to a substantial reduction in performance previously attributed to biological signal.
bioinformatics2026-08-11v2Spliformer-V2 enables multi-tissue prediction and interpretation of splice-altering genetic variants
Tang, X.; Shao, M.; Lei, H.; Ma, X.; Guo, J.; Shen, Y.; Wu, Q.; Dong, Y.; Zeng, Y.; Gitler, A.; Chen, Y.; Abrahao, A.; Zinman, L.; Rogaeva, E.; Chen, Y.; Ichida, J.; Zhang, M.Abstract
Precise regulation of pre-mRNA splicing underlies transcriptomic diversity and is disrupted in aging and disease, yet tissue-specific splice-altering genetic variants remain poorly resolved. Here, we present Spliformer-V2, a SegmentNT-based deep learning model for predicting and interpreting variant effects on RNA splicing across human tissues. We generated a diploid sequence resolved RNA splice map from paired whole-genome-sequencing and RNA-seq data across 12 central nervous system (CNS) and 6 peripheral tissues for model development. Spliformer-V2 outperformed SpliceTransformer, Pangolin and AlphaGenome in predicting splice-site usage, identified tissue-specific splicing regulatory motifs, and revealed tissue vulnerability to pathogenic splice-altering variants. Analyses of loci associated with 8 neurological diseases prioritized CNS-specific mis-splice-vulnerable genes. In 1,405 amyotrophic lateral sclerosis (ALS) genomes, Spliformer-V2 nominated rare splice-altering variants enriched in PTPRN2, which showed reduced expression in TDP-43-depleted neurons. PTPRN2 overexpression rescued C9ORF72-patient derived motor neuron degeneration and modulated TDP-43 mislocalization, indicating it as a potential therapeutic modifier in ALS.
bioinformatics2026-08-11v2Structure-aware Graph Learning Predicts RNA Editability Across Tissues and Species
Rosenwsser, Z.; Levitt, M.; Levanon, E. Y.; Oren, G.Abstract
Programmable A-to-I RNA editing using endogenous ADAR enzymes is emerging as a therapeutic strategy, but editability remains difficult to predict because ADAR recognition depends on double-stranded RNA geometry and stability rather than sequence alone. We present AdarEdit, a structure-explicit graph-attention framework that represents each dsRNA substrate as a nucleotide graph with backbone and base-pair edges. The framework includes a baseline model and a bio-aware model, with the latter augmenting this representation with typed interactions and a motif-sensitive sequence branch. We trained and evaluated both models on high-confidence inverted Alu duplexes (n = 884) with secondary structures predicted by RNAfold and editing levels measured across 8,603 GTEx RNA-seq samples spanning 47 tissues. Across five tissue contexts, the baseline and bio-aware models achieved strong held-out performance (test F1 = 0.814-0.869, AUROC = 0.869-0.933) and outperformed a matched structure-string baseline on the Liver split. The same graph representation retained predictive ability in evolutionarily distant non-Alu species (sea urchin, acorn worm, and octopus), suggesting conserved principles of ADAR substrate recognition. Finally, attention profiles and in silico mutagenesis recapitulated known biochemical constraints, including suppression by an upstream guanosine, and revealed longer-range asymmetric structural influences on editing. Because Alu duplexes are edited predominantly by ADAR1, AdarEdit is geared primarily to the ADAR1 regime. ADAR1 is particularly relevant to therapeutic editing given its broad tissue expression. The sources of this work are available at our repository: https://github.com/Scientific-Computing-Lab/AdarEdit
bioinformatics2026-08-11v2Fast retrieval of structurally similar antibodies from large sequence databases with AbSLang
Wang, E. J. D.; Spoendlin, F. C.; Greenshields-Watson, A.; Taylor, C. R.; Deane, C. M.Abstract
The first steps in antibody therapeutic discovery involve identification of sequences with desirable binding properties. A way of finding these lead molecules is through the search of large sequence databases. Current methods, due to the size of databases, rely on germline or complementarity-determining-region (CDR) sequence identities, overlooking structurally similar antibodies with divergent sequences which can have identical binding properties . To address this, we introduce AbSLang, a model trained for pairwise CDR RMSD prediction using a contrastive learning approach. We demonstrate that AbSLang has comparable accuracy to exact RMSD calculation after explicit structure prediction with state-of-the-art models. Building on this model, we implemented AbSLang-search, a pipeline for retrieval of structurally similar antibodies from large sequence databases. AbSLang-search is highly compute efficient and allows to search datasets with 10 million sequences in less than 2 seconds.
bioinformatics2026-08-11v1Learning Shared Residue Backgrounds and Modification-Specific Offsets for PTM Site Prediction
Pokharel, S.; Bhusal, B.Abstract
Post-translational modifications (PTMs) are chemical changes added to proteins after translation. These changes affect protein function and regulation, and their disruption is linked to disease-associated mechanisms. Because experimentally validating all possible modification sites is impractical, many computational predictors have been developed for PTM site prediction. In this work, we study whether a shared model can represent common residue-background patterns while learning modification-specific background-to-positive offsets. This framing is especially relevant for residues such as lysine (K), which can be acetylated, ubiquitinated, methylated, or sumoylated depending on the surrounding protein context. We propose an anchor-guided rectified flow matching framework for multi-type PTM site prediction from protein language model embeddings. For each PTM--residue pair, the model builds residue-background anchors from PTM-compatible unannotated residues and positive anchors from experimentally annotated modified residues. Given a candidate residue and target modification type, the model compares the residue embedding with these anchor sets and uses a rectified flow module to estimate a modification-conditioned background-to-positive offset. This offset is combined with anchor-based features and used for site scoring. We evaluate the framework on a dbPTM-derived benchmark covering six commonly studied PTMs: phosphorylation, acetylation, ubiquitination, methylation, sumoylation, and N-linked glycosylation. In the shared-model setting, our approach achieves a macro AUPRC of 0.4195, improving over the gated multi-anchor baseline of 0.4154, while independently trained per-modification models achieve 0.4353. These results suggest that multi-type PTM prediction can be modeled within a single shared framework by combining residue-background anchors with modification-conditioned offset features.
bioinformatics2026-08-11v1Systematic assessment of the biological impact of cellular deconvolution on downstream analyses of disease transcriptomes
Mitra, S.; Ibrahim, M.; Narayanan, M.Abstract
Background Cellular deconvolution methods estimate cell type proportions from bulk RNA seq data, typically using single cell RNA seq derived signatures, enabling separation of disease associated transcriptional changes into composition driven and cell intrinsic effects. However, these approaches depend on model assumptions and the stability of cell type signatures, and it remains unclear how deconvolution related uncertainties influence downstream analyses and biological conclusions. Results We systematically evaluated the effect of cell type correction on disease relevant transcriptomic insights, using Alzheimer's disease (AD) as a model and the Mount Sinai Brain Bank cohort as a primary dataset. Applying dtangle, selected after comparison with another deconvolution approach, we estimated cell type proportions across four brain regions and assessed how correction reshaped differential gene expression and pathway enrichment. Cell type correction (CTC) markedly altered differentially expressed gene (DEG) profiles in a region dependent manner: the superior temporal gyrus lost all significant signals, while the frontal pole gained DEGs with improved cross region concordance. At the pathway level, correction shifted enrichment from synaptic loss and immune activation toward suppression of stress response and immune regulatory programs, suggesting that composition changes partly obscure cell intrinsic regulatory signals. Overlap with AD genome wide association study loci and replication in an independent cohort indicated that cell intrinsic changes are more consistently validated than composition driven changes. Notably, KCNN2 and RIMS1, not currently recognized as canonical AD biomarkers, emerged as robust transcriptional signatures, potentially reflecting both compositiondriven and cell intrinsic dysregulation and warranting further investigation. Conclusions Parallel evaluation of uncorrected and CTC analyses distinguishes composition driven from cell intrinsic transcriptional effects and highlights robust disease signatures in heterogeneous tissues such as the brain.
bioinformatics2026-08-11v1Data-Centric Evaluation of Protein Function Prediction Pipelines
Soto-Garcia, N.; Murillo-Acevedo, N.; Garcia Vinuesa, J.; Islas-Avila, A. L.; D. Davari, M.; Murgas, L.; Hassanin, A.; Orostica, K.; Gonzalez-Puelma, J.; Navarrete, M.; Rebollar-Martinez, A.; Uribe-Paredes, R.; Cadet, F.; Medina-Ortiz, D.Abstract
Performance estimates in protein function prediction depend not only on model choice but also on upstream decisions that define the learning problem. Using antioxidant protein classification as a controlled case study, we evaluated how dataset harmonisation, protein representation, redundancy control, and partitioning strategy affect protein machine learning pipelines. We integrated 18,804 records from 12 publicly available dataset entries into a curated consensus dataset of 4,193 protein sequences. One-hot encoding and six pretrained protein language model representations were evaluated as model inputs and as similarity spaces for redundancy reduction and distance-aware splitting. Representation choice substantially altered dataset geometry, retained dataset size, class balance, and downstream evaluation. At representation-specific p90 thresholds, one-hot encoding retained the complete dataset, whereas pretrained embeddings retained between 5% and 25% of sequences. Distance-aware partitioning reduced apparent performance relative to random splitting by up to 0.15 MCC before redundancy control, while this difference narrowed after similarity filtering. Selected configurations nevertheless maintained high performance under stricter evaluation, reaching an MCC of 0.84. These findings show that performance estimates should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.
bioinformatics2026-08-11v1ASPIRE: the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems
McLaughlin, R. J.; Chen, S.; Nag, A.; Noonan, A. J. C.; Bartolomeu, C.; Borden, S. A.; Lam, S.; Myers, R.; Hallam, S. J.Abstract
Microbial communities inhabiting the respiratory tract contribute to health status through interactions with host physiology, immune function, and local environmental conditions. Advances in small subunit ribosomal RNA (SSU or 16S rRNA) gene amplicon sequencing enable culture-independent profiling of microbial communities as amplicon sequence variants (ASVs), revealing links between microbial dysbiosis and respiratory diseases, and the use of mass spectrometry to measure volatile organic compounds (VOCs) in exhaled breath shows emerging promise for biomarker discovery. Here we present ASPIRE, the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems, an accessible Nextflow workflow for processing, analyzing, and interpreting linked ASV-VOC data from respiratory microbiome studies. ASPIRE is designed to support scalable comparative analysis across respiratory sample types while preserving intermediate file outputs for inspection and reuse within a standardized file structure.
bioinformatics2026-08-11v1AnchorR: A QuPath and R interface for collaborative exploration of spatial transcriptomics and histology
Morris, C. A.; Bastian, W. C.; Cui, Y.; Kurago, Z.; Douglass, E. F.Abstract
Single-cell spatial transcriptomics can connect molecular cell states with tissue morphology, but this promise depends on accurate registration to histopathology. In serial sections, however, tissue borders often differ because of sectioning artifacts, staining variability, and field-of-view acquisition, limiting conventional area-based registration. We developed AnchorR, an expert-guided workflow for coarse-grained alignment of hematoxylin and eosin (H&E) images with CosMx Spatial Molecular Imaging data. Bioinformaticians first define and color-code cell types in Seurat, and pathologists then identify corresponding internal landmarks using QuPath overlays. AnchorR combines these paired landmarks to estimate affine transformations, quantify residual error, and support visual quality control and anchor refinement. Using six oral pre-cancerous tissue sections, we identified 60 cross-modal landmarks. Fitting each section independently reduced mean landmark error from 121.5 m with a single whole-slide transformation to 14.6 m. Cross-validation further showed that increasing the number of anchors improved robustness, with nine-anchor fits achieving approximately 20 m error, or about one cell diameter. AnchorR is designed to complement automated computer-vision methods by providing reliable tissue-level alignment when border mismatch makes global registration difficult. By creating a shared workspace for pathologists and bioinformaticians, it operationalizes an expert-in-the-loop approach and makes feature-based multimodal registration accessible without specialized computer-vision expertise or high-performance computing.
bioinformatics2026-08-11v1Convergent biology, divergent drivers: a cross-species comparison of human and canine invasive urothelial carcinoma
Cho, H.; Mochel, J. P.; Corbett, M. P.; Olivieira, L. J.; Allenspach, K.; Zdyrski, C.; Pawlak, A.; Johnson, B. A.; Douglass, E. F.Abstract
Traditional animal models are often inbred and genetically uniform. This makes them powerful for controlled experiments, but it limits how well they represent the patient-to-patient variation seen in real-world disease. Comparative oncology seeks to address this gap by studying naturally occurring cancers in outbred companion animals, especially dogs. Canine medicine offers two important advantages: first, prospective trials can often be completed faster than in humans and second, dogs are already part of the translational pipeline through pharmacokinetic and toxicology studies. Here, we assessed the transcriptional fidelity of human and canine invasive urothelial carcinoma in primary tumors and patient-derived organoids. We then used single-cell and spatial data to resolve the underlying cellular organization. Despite strong species and platform differences, human and canine tumors preserved the same major luminal-basal structure and a similar tumor microenvironment. The two species reached this shared biology through different recurrent mutations. These included FGFR3 alterations in humans and BRAF alterations in dogs, which converged on overlapping pathways and a luminal phenotype. Human and canine organoids also underwent a similar shift in culture. Both became more proliferative and metabolic while losing inflammatory programs. Thus, organoids preserved important tumor biology while introducing predictable platform effects. Single-cell and spatial analyses showed that the luminal-basal axis reflects a gradient of cell states organized around the tumor-stroma boundary, rather than two discrete tumor types. This helps explain why bulk RNA-sequencing subtypes are reproducible but coarse. Together, these findings define where canine and human bladder cancer agree, where they differ, and how dogs can support parallel therapeutic and diagnostic development.
bioinformatics2026-08-11v1AbPACER: parent-aware, affinity-label-blind prioritization of affinity-matured scFv clones from phage-display NGS
Chung, A. J.; Park, B. Y.; Park, E.-B.; Han, J.-H.Abstract
Background: Affinity-maturation phage-display next-generation sequencing (NGS) yields more paired single-chain variable fragment clones than can be characterized experimentally, creating a fixed-budget prioritization problem. Read counts provide empirical support rather than direct affinity labels. We developed AbPACER (Antibody Parent-Aware Contextual Evidence Ranker), an affinity-label-blind neural ranker combining parent-relative mutation descriptors, frozen antibody-language-model context, and NGS evidence from related clones. AbPACER is campaign-adaptive rather than zero-shot: for each campaign, it is fitted to paired sequences and round-resolved R1-R3 counts before returning a 384-candidate assay list. We evaluated it in two retrospective phage-display campaigns and separately assessed its supervised mean-squared-error adaptation on AlphaSeq, denoted AbPACER-MSE. Results: From frozen top-5% candidate sets containing 16,323 Fas-associated factor 1 (FAF1) and 7,487 vascular endothelial growth factor receptor (VEGFR) clones, each method ranked the complete target-specific set and selected 384 candidates. In FAF1, AbPACER recovered 2.00 +/- 0.00 of seven retrospective panel clones, recovering two in every seed, compared with 1/7 by total count, 1.00 +/- 0.00 by Ens-Grad CNN, 1.67 +/- 1.15 by A2Binder-HL, and 1.33 +/- 0.58 by AbAffinity. In VEGFR, AbPACER recovered 2.33 +/- 0.58 of three panel clones, the highest observed learned-method mean, whereas total count recovered 3/3. No learned method was uniformly best at broader hypothetical budgets. On the public AlphaSeq common split of 11,670 fixed-test variants, AbPACER-MSE recovered 187.0 +/- 2.6 of the true top-384, closely matching AbAffinity (188.0 +/- 2.6) and exceeding A2Binder (175.7 +/- 6.4) and Ens-Grad CNN (154.0 +/- 6.1). AbPACER-MSE updated 1.378 million task-specific parameters, compared with 651.04 million for AbAffinity, and achieved Pearson 0.687 +/- 0.003 and Spearman 0.652 +/- 0.002. Conclusions: AbPACER provides a campaign-specific, parent-aware framework for fixed-budget prioritization from affinity-label-blind phage-display NGS data. At the 384-candidate endpoint, it showed the highest mean recovery among learned methods in both retrospective campaigns. AbPACER-MSE closely matched AbAffinity in true top-384 recovery while updating substantially fewer task-specific parameters. These results motivate prospective evaluation of sequence-conditioned reranking as a complement to count-based prioritization.
bioinformatics2026-08-11v1MSGPCA: Multi-Slice Graph PCA for replicate-aware Spatial Omics analysis
Chakraborty, A.; Neelon, B.; Lawson, A.; Angel, P.; Chung, D.; Seal, S.Abstract
As spatial transcriptomics (ST) and spatial proteomics (SP) technologies mature, experimental designs are increasingly moving beyond single-slice analyses toward multi-slice studies involving one or more donors and experimental conditions. Although these designs enable the identification of reproducible spatial signals, they also introduce substantial biological heterogeneity, particularly when integrating non-serial slices or anatomically distinct regions. If not modeled carefully, such variation can blur slice-specific tissue structure, mask conserved molecular patterns, and limit the discovery of biologically relevant latent structure. Although dimension reduction is essential for representing high-dimensional molecular data in a lower-dimensional space, existing multi-slice methods typically enforce a globally shared representation that inadequately accommodates slice-level heterogeneity. To address this limitation, we propose Multi-Slice Graph Principal Component Analysis (MSGPCA), which decomposes molecular variation into shared spatial factors conserved across slices and slice-specific factors that capture local tissue microarchitecture. In downstream analyses, MSGPCA-derived representations recover spatial tissue structure, denoise molecular profiles, and reveal biologically interpretable metafeatures associated with shared and slice-specific biology. In a mass spectrometry imaging dataset comprising nonserial slices of ductal carcinoma in situ (DCIS) and invasive breast cancer (IBC), the shared factors captured broad biological differences across tissue regions, whereas the slice-specific factors revealed intratumoral spatial variation within the IBC microenvironment. In human dorsolateral prefrontal cortex ST data, MSGPCA recovered laminar cortical architecture across adjacent slices, closely aligning with expert pathologist annotations. Together, these findings demonstrate that MSGPCA resolves shared tissue architecture while preserving local microenvironmental variation in complex multi-slice spatial omics datasets.
bioinformatics2026-08-11v1A Practice on Antibody Hydrophobic Interaction Chromatography Retention Time Prediction using Pre-Trained Large Language Model Fine-Tuning
Wang, B.; Cai, B.; Chen, H.; Xia, H.; Wang, B.; Liu, J.; Han, L.; Wang, R.Abstract
Hydrophobicity is a critical property associated with the risk of non-specific binding, and it is commonly assessed using hydrophobic interaction chromatography retention time. Several computational approaches have been developed to predict antibody developability based on pre-trained language models. Such models can be fine-tuned with limited labeled antibody sequences and, in principle, do not require structural information, which is often challenging to obtain. Nevertheless, few studies have achieved strong performance in hydrophobicity prediction without incorporating structural features. Here, we present a case study of fine-tuning the pre-trained model IgBert to predict antibody hydrophobicity. Using Herceptin as a reference, we performed hydrophobic interaction chromatography retention time experiments and generated Herceptin-adjusted datasets. The fine-tuned model achieved a best R2 of 0.916, underscoring the critical role of rigorous data quality control. We also synthesized and validated 20 commercially available antibody sequences, and the results showed that the predicted hydrophobic properties were correctly reflected. Our findings provide practical guidance and highlight considerations for future applications of fine-tuned pre-trained language models in antibody hydrophobicity prediction.
bioinformatics2026-08-11v1orthoSynAssign: refine orthogroups using synteny information
Tsai, C.-H.; Pina Paez, C. G.; Stajich, J. E.Abstract
Accurately identifying orthogroups is crucial for precise phylogenetic reconstruction, but clustering-based methods often generate complex, many-to-many orthogroups that include confounding paralogs. Incorporating synteny offers a robust strategy to refine these clusters into high-granularity, single-copy orthologs. We introduce orthoSynAssign, a user-friendly, high-performance rewrite of the orthogroup refinement tool OrthoRefine, combining an intuitive Python interface with a core computing engine written in Rust. This hybrid architecture ensures straightforward installation, seamless data parsing, and exceptional computational efficiency. Evaluated against the Yeast Gene Order Browser (YGOB) dataset, orthoSynAssign demonstrated outstanding performance, substantially elevating the Area Under the Precision-Recall Curve. Furthermore, multi-threading benchmarks across 193 Eurotiomycetes genomes confirmed strong scalability, drastically reducing execution runtime while maintaining a strictly bounded, thread-independent memory footprint. Ultimately, orthoSynAssign provides a reliable and scalable framework for high-throughput phylogenomic workflows.
bioinformatics2026-08-11v1Virtual-cell verification enables self-auditing AI discovery for immune rejuvenation
You, Y.; Fan, X.; Li, G.; Deng, W.; Fu, Y.; Hu, H.; Ren, W.; Lu, S.; Han, G.; Shao, J.; Zheng, S.; Zhou, K.; Kong, J.; Chen, J.; Liu, X.; Tian, L.Abstract
Artificial-intelligence agents propose drug-discovery hypotheses faster than experiments can test them, yet their conclusions are rarely verified, against the underlying biology, the predicted perturbation, or the agent's own scoring logic. We close this verification gap with an agentic framework built on three verifiers. First, PACE, a phenotype verifier, resolves immune aging into ten directionally scored, cell-type-resolved gene-set modules, selected for cross-cohort stability across four PBMC cohorts, and outperforms five established aging clocks in an independent in-house aging cohort of 434 elderly donors. Second, CellQ, a virtual-cell verifier built with multi-modal LLM, compresses each single-cell transcriptome into eight discrete tokens aligned to a language model's vocabulary through residual vector quantization; it attains state-of-the-art perturbation prediction and uniquely resolves the weak, module-level shifts that differential-expression recovery misses. Third, an Analyzer-Planner-Auditor agent verifies its own scoring logic: screening 110 compounds in primary human PBMCs, it found aged-down modules more reversible than aged-up modules and revised its objective from an equal-weight mean to a balance-constrained minimum, a self-correction that generalized to an independent 13-compound T-cell assay. By verifying its predictions and its own objective against experiment, the framework points beyond hypothesis-generating AI toward self-correcting AI scientists whose objectives could continuously evolve.
bioinformatics2026-08-11v1KiMA: Kinematic Motion Analysis for Spinal Cord Injury Research
Kumaran, M.; N R, S. S.; Venkatesh, I.Abstract
Accurate quantification of locomotor recovery is essential for evaluating therapeutic outcomes in spinal cord injury (SCI) models. Manual scoring systems remain observer-dependent, and commercial gait-analysis platforms are costly and proprietary. Markerless pose-estimation tools such as DeepLabCut generate accurate body-part coordinates, but converting these coordinates into biologically meaningful locomotor parameters typically requires custom programming and multiple external tools. We developed KiMA (Kinematic Motion Analysis), an open-source, browser-based suite for integrated analysis of rodent gait and hindlimb kinematics. KiMA accepts DeepLabCut coordinate files and performs automated coordinate parsing, stick-figure reconstruction, frame-by-frame movement inspection, and single- and multi-sample analysis, with dedicated workflows for ladder and rung analysis, footfall detection, and CatWalk gait analysis. The platform quantifies joint angles (metatarsophalangeal, ankle, knee, hip, and pelvic), stride length, stride width, cadence, stance and swing durations, paw-contact events, swing clearance, and locomotor symmetry, and supports cohort-level comparisons, correlation analysis, principal component analysis, and export of processed datasets and publication-quality figures. Because KiMA runs entirely within a standard web browser, it requires no software installation or local programming environment, supporting cross-platform accessibility and data privacy. By unifying gait quantification, visualization, and multivariate analysis in a single interface, KiMA lowers the computational barrier to markerless locomotor analysis and helps researchers detect subtle functional recovery after SCI.
bioinformatics2026-08-11v1Enhanced Detection of Age-related Macular Degeneration in Low-quality Retinal Images via Noise-Augmented YOLO and Adaptive Attention Mechanisms
Bai, X.; Kishimoto, K.; Sugiyama, O.; TAMURA, H.Abstract
This study aims to improve the detection performance of age-related macular degeneration (AMD) in low-quality retinal images. Background: AMD is a leading cause of vision loss among older adults globally, and accurate detection is crucial for clinical management. However, low-quality optical coherence tomography (OCT) images significantly com-promise diagnostic accuracy. Objective: To enhance AMD detection in low-quality images using noise-augmented data augmentation and an improved YOLO deep learning model. Methods: Public datasets from UCSD and Duke University were utilized; the training dataset comprised 24,980 OCT images (high-quality and noise-augmented low-quality), while the testing dataset included 1,000 images (584 AMD, 416 normal). The model is based on the YOLOv8n framework, integrated with Squeeze-and-Excitation blocks (SEblock) and Adaptive Sparse Self-Attention (ASSA), with an addition-al 160*160 detection layer for detecting small lesions. Evaluation metrics included accuracy, sensitivity, specificity, and F2-score. Results: The proposed model achieved an accuracy of 99.02%, sensitivity of 98.17%, specificity of 100%, and an F2-score of 98.50% on the Duke dataset. Detection rates were significantly improved compared to traditional methods, particularly in low-quality images, with a detection rate of 89.60%, markedly superior to original YOLOv8n (55.10%) and classical models like ResNet50. Conclusion: The enhanced model, employing noise-augmented training data and improved attention mechanisms, demonstrates excellent AMD detection capabilities in low-quality OCT images, showing broad potential for clinical applications.
bioinformatics2026-08-11v1SeqDesk: a sequencing-facility management system for standards-compliant and FAIR (meta)data submission
Muench, P. C.; Robertson, G.; McHardy, A. C.Abstract
Achieving FAIR compliance requires both standardized metadata and infrastructure for data deposition, yet in practice a large fraction of sequencing studies is still published without the persistent, standards-compliant metadata that reuse depends on. Collecting MIxS-compliant metadata is complex: environment-specific checklists can contain hundreds of fields, and the effort is magnified when metadata is assembled retrospectively at publication time rather than captured throughout the project. We developed SeqDesk, an open-source data management system for sequencing facilities that is designed so that FAIR-compliant public data is produced as the natural output of routine operations. Its current scope is microbial sequencing data, covering metagenomes as well as isolate genomes, for which it supports the corresponding MIxS checklists. SeqDesk gives a sequencing facility a configurable order-and-tracking system for sequencing projects, captures and validates MIxS-compliant metadata aligned with ENA checklists at project initiation, runs bioinformatics analyses through Nextflow pipelines, and brokers submission to the European Nucleotide Archive, all within the institution's own infrastructure. By embedding standards-compliant metadata capture into the sequencing-facility workflow rather than bolting it on at submission, SeqDesk shortens the path from sample to reusable public data. The underlying checklist model is generic, so support can be extended to further data types and metadata standards beyond the microbial domain. SeqDesk is free and open source under the Apache 2.0 licence and available at https://seqdesk.org, with a live demonstration at https://seqdesk.org/#demo.
bioinformatics2026-08-11v1Moirai: single-cell trajectory inference grounded in gene-level expression dynamics
Fijn, A. H. B.; S. Jeuken, G.Abstract
Underlying the development of multicellular organisms is the process of cell differentiation, which is governed by the concerted and sequential change in gene expression. Various methods have been developed that employ scRNA-seq data to infer the position of a cell along a pseudo-temporal axis and identify relevant genes involved in the process. These trajectory inference methods typically rely on global transcriptomic changes and mathematical methods. However, overemphasis on large-scale transcriptomic changes may impair sensitivity to identify branching points and convergent trajectories, which are rather governed by small-scale transcriptional events. Motivated by this, we developed Moirai, a graph-based trajectory inference method that identifies gene expression patterns that change dynamically over a developmental continuum and leverages these to define a common pseudotime axis between all cells. In doing so, Moirai shifts the focus to individual gene dynamics, which enhances its ability to detect putative branching points that are masked by global transcriptomic similarities. We apply Moirai to four developmental datasets, where we demonstrate its ability to recover gene expression patterns of genes with a known involvement in the respective developmental process, motivating their use for defining a cell's pseudotime. We furthermore show that Moirai can robustly infer gene expression patterns across different embedding approaches, highlighting the value of moving the focus of the inference process to the small-scale transcriptional dynamics.
bioinformatics2026-08-11v1Supervised Deep Learning for Efficient Cryo-EM Image Alignment in Drug Discovery with cryoPARES
Sanchez-Garcia, R.; Berndt, A.; Apelbaum, A.; Reeks, J.; Williams, P. A.; Poelking, C.; Deane, C.; Saur, M.Abstract
Cryo-Electron Microscopy (cryo-EM) is a pivotal tool for determining 3D structures of biological macromolecules. Current workflows are computationally demanding and require manual intervention, creating bottlenecks for high-throughput applications like structure-based drug discovery. In such contexts, where all protein samples can be assumed to be equivalent at resolutions relevant for image alignment, information about particle poses from previous refinements could be reused. Existing methods, however, ignore this prior knowledge, aligning each dataset from scratch. We present cryoPARES, a deep learning pose estimation method trained on pre-aligned datasets. Our method not only provides accurate angular predictions significantly faster than traditional approaches but also introduces automated particle pruning capabilities that eliminate manual intervention. Together with its single-pass operation, these features enable near real-time reconstructions that provide feedback during data acquisition. We demonstrate cryoPARES's effectiveness through rapid structural determination of seven ligand-bound complexes across four distinct protein targets. We also release three fragment-bound cryo-EM datasets.
bioinformatics2026-08-10v5Impact of the N-glycosylation on full-length IgG2 and IgG4 antibodies: a comparative study using molecular dynamics simulations.
LEON FOUN LIN, R.; Bellaiche, A.; Diharce, J.; Etchebest, C.Abstract
Like other proteins, monoclonal antibodies - important biodrugs- are subject to post translational modifications, especially the N-glycosylations. However, the effect of the N-glycosylations remains poorly studied and atomistic details about their influence are rarely available. . Moreover, the few existing studies focus on the prevalent immunoglobulin G1. To go further in the understanding of the impact of glycosylations, we have carried out a comparative exploration of the effect of N-glycosylations on two different classes of antibodies, namely Mab231, an IgG2 and the pembrolizumab, an IgG4 . The two antibodies differ by their sequences, their length, their 3D structure but also by the location and composition of the glycans. In the present work, detailed and important information were gained through molecular dynamics simulations where both monoclonal antibodies were studied without and with the presence of their glycans. The results of 1.5 microseconds of sampling for each system show that glycosylation does not drastically alter the overall conformational landscape of either antibody, whatever the metrics considered. However, it measurably modulates local flexibility, inter-domain correlated motions, and the relative orientation of the Fab arms with respect to the Fc domain, with statistically significant shifts in key geometric descriptors. Importantly, contact analysis reveals that glycan interactions extend beyond the Fc region to reach Fab residues. The allosteric network calculations demonstrate that the influence of Fc-bound glycans propagates even until the Fab framework regions in both mAbs, which could impact the antigen binding. The nature and magnitude of these effects are subclass-dependent, reflecting differences in glycan composition, hinge architecture, and three-dimensional organization Our findings challenge the prevailing view that Fc glycosylation uniformly promotes CH2 domain opening. More importantly, it underscores the necessity of considering full-length structures and IgG subclass diversity in glyco-engineering strategies.
bioinformatics2026-08-10v5Baseline regulatory programs in larval and adult neural progenitors converge towards an injury-induced state after spinal cord injury
Cosacak, M. I.; Severinov, D.; Westphal, M.; Heilemann, K.; Baerhold, D.; Bretschneider, A.; Reinhardt, S.; Roscito, J. G.; Becker, T.; Becker, C. G.; Poetsch, A. R.Abstract
Regeneration after spinal cord injury requires progenitor cells to convert injury-associated signals into coordinated remodeling of gene regulatory programs. Mammalian spinal progenitors show limited neurogenic output after injury, whereas zebrafish regenerate spinal neurons and recover motor function. To investigate the regulatory changes that allow ependymo-radial glia (ERG) cells, the progenitor cells of the zebrafish spinal cord, to generate new neurons, we combined single-nucleus gene expression and chromatin accessibility profiling across embryonic, larval, and adult stages with topic-based gene regulatory network (GRN) inference. We found that larval and adult ERGs enter the injury response from distinct regulatory baselines: larval progenitors are characterized by a gliogenic program, whereas adult progenitors maintain a comparatively quiescent state. Following injury, both populations gradually change their baseline programs and shift towards a lesion-associated module marked by stress-responsive and chromatin-associated regulators, including jun, hmga1a, hmga2, ybx1, and foxj1a. The shift away from homeostatic states is supported by decreased expression of the Notch-associated regulators nuclear factor I A (nfia) and hey1 in larvae, while in adults, downregulation of the same nuclear factor and other TFs such as bhlhe41 is associated with quiescence exit. Pathway analysis showed stage-specific alterations after injury, characterized predominantly by extracellular signaling and cytoskeletal reorganization in larvae and by metabolic and translational remodeling in adults. Despite divergence from the homeostatic states, injury-induced larval and adult GRNs remain distinct from embryonic hERG regulatory programs. Thus, larval and adult progenitors follow different trajectories from their baselines towards a related lesion-reactive state, in which shared regeneration-associated features are acquired within respective contexts.
bioinformatics2026-08-10v2Accelerating String Comparison in RLZ Compressed Sequences via LCE Jumps
Varki, R.; Boucher, C.Abstract
Relative Lempel-Ziv (RLZ) is an effective compression method for large, repetitive collections; however, the fundamental primitives required to elevate it from a passive archival format to a tractable representation for compressed construction have yet to be fully established. In this paper, we introduce an algorithmic framework for structurally comparing and lexicographically sorting sequences of RLZ factors. We characterize when direct factor comparisons are necessary and when they can be bypassed using RLZ specific shortcuts. We further introduce a method for extending truncated factors into right-maximal matches, enabling the recovery of matching statistics from the RLZ parse. Experimentally, RLZ sorting achieved speedups of up to 3.93x over character-based sorting. Together, these results advance the use of the RLZ format as a foundation for compressed construction.
bioinformatics2026-08-10v2Making Biorisk Measurable: A Bayesian Framework for Laboratory Risk Management
Prodanov, D.Abstract
Biosafety risk assessment traditionally relies on categorical scales embodied by the four WHO Risk Groups and biocontainment levels. Mapping such categories to quantitative metrics is an open problem for the field: the classifications are too coarse for operational decision-making, yet strictly probabilistic language remains inaccessible to most safety professionals, laboratory managers, and decision-makers. To bridge these gaps, the present work develops a quantitative Bayesian framework for laboratory risk management that combines WHO Risk Group classification as a prior with a Markov chain model of the incident--disaster escalation chain. Risk is reported on a log-risk scale that transforms multiplicative probabilities into additive quantities, mirroring the decibel scale in acoustics. The framework accommodates longitudinal updating with local incident data and quantifies the separate contributions of training, preventive maintenance, and inspection to system-level safety. Resource allocation recommendations are derived that complement existing compliance frameworks with auditable, evidence-based prioritisation. The framework is illustrated on synthetic BSL-3 scenarios and shifts the perspective of biorisk governance from static compliance assessment to dynamic risk and resource management.
bioinformatics2026-08-10v2Mapping disease traits onto spatiotemporal domain landscapes through scalable multi-sample integration of spatial transcriptomics
Zhang, S.; Luo, S.; Luo, Y.; Su, S.; Shi, Y.; Liu, L.; Li, W.; Tian, T.; Li, J.Abstract
High-resolution spatial transcriptomics (ST) is generating massive datasets involving multiple tissue sections, continuous developmental stages, or various disease conditions. Yet, their scale, biological heterogeneity, and substantial batch effects hinder data integration for further analysis, particularly the temporal interpretation of disease-associated spatial domains within tissue organization landscapes. Here, we present STUltra, a scalable hypergraph framework that integrates multi-sample ST data for precise domain detection and genome-wide association study (GWAS)-based disease trait mapping. From ST datasets, STUltra first constructs interval-sampled integrative hypergraphs, in which hyperedges capture tissue neighborhoods within the slices as well as shared biological contexts across the slices. It then combines a robust graph autoencoder with contrastive learning to learn batch-corrected, spatially informed embeddings for identifying spatial domains across the tissue sections. These embeddings are then coupled with GWAS statistics to map trait-associated signals onto the integrated tissue sub-structure landscapes spanning different sections, developmental stages, and disease conditions. STUltra substantially outperforms competing methods, especially being much faster, on diverse integration benchmarks, enabling efficient processing of datasets containing millions of spots. As a case study on a mouse embryo dataset, STUltra recovered continuous developmental programs and mapped congenital heart disease signals beyond cardiac regions to broader programs including vascular and mesenchymal domains. On a mouse Alzheimer's disease dataset, STUltra prioritized localized AD-associated genetic signals to microglia and a disease-expanded C1QA/HEXB-high astrocyte state. Together, STUltra provides a scalable framework to bridge the gap between spatiotemporal tissue structures and complex disease traits, facilitating discoveries from previously unrecognized molecular and cellular architectures across diverse contexts.
bioinformatics2026-08-10v2nf-cavalier: A Nextflow Pipeline for Rare Disease Variant Prioritization and Reporting
Munro, J. E.; Reid, J.; Bahlo, M. E.; Bennett, M. F.Abstract
nf-cavalier is a Nextflow pipeline that automates genomic variant annotation, filtering, and reporting for individuals with rare Mendelian diseases. The pipeline takes as input variant callsets for an individual, family, or rare disease cohort, together with a target gene panel or a phenotype of interest. Variants are then filtered using various customisable criteria, including predicted gene consequence, computational pathogenicity predictions, population frequency, and familial segregation. The sequencing data for candidate variants is then visualised for human review. Candidate variant results are returned in user-friendly output formats, including interactive HTML reports and PowerPoint slide decks, with embedded links to external resources that enable rapid review by clinical research teams. nf-cavalier is maintained on GitHub (bahlolab/nf-cavalier) and licensed under the permissive MIT open-source licence.
bioinformatics2026-08-10v1Biasing Conformational Sampling in AlphaFold 3 and Boltz-2 via Pair Representation Scaling
Suzuki, S.; Amagasa, T.Abstract
Deep learning has transformed protein structure prediction, yet most systems return a single dominant conformation with little control over alternative functional states. We introduce pair representation scaling, an inference-time method that biases conformational sampling in diffusion-based structure predictors by multiplying the latent pair representation by a single scalar before the Pairformer trunk, without retraining, an auxiliary model, or a second forward pass. On 86 two-state targets spanning domain motions and membrane transporters, scaling broadens the conformational ensembles of both AlphaFold 3 and Boltz-2 and recovers alternative states that default inference misses, most strongly in AlphaFold 3, where the gains extend even to targets deposited after the training cutoff. It approaches the alternative-state recovery of alignment-based sampling methods, and the benefit persists even without a multiple sequence alignment. The predicted distance distributions show that scaling shifts the encoded two-state distribution toward the experimentally observed alternative state, a directed modulation rather than arbitrary perturbation. Pair representation scaling is an interpretable, low-cost handle on the conformational ensembles of deep learning structure predictors.
bioinformatics2026-08-09v3pyfraglib: An integrated cfDNA fragmentomics platform
Schuette, D.; Godfrey, L. K.; Schneider, J.; Borchmann, S.; Heger, J.-M.; Schwarz, R. F.Abstract
Summary: Cell-free DNA (cfDNA) fragmentomics is the analysis of a diverse set of cfDNA fragment features, e.g. fragment length profiles, windowed protection scores, and end motifs. As such it requires software tooling for fragment extraction, statistical feature modeling, and cohort-level comparative analysis. In silico simulations can facilitate the development and validation of new methods by generating testing datasets with known ground truth. Existing tools address individual aspects of this workflow but none provide all necessary capabilities within a single package. Results: We present pyfraglib, a platform integrating fragment extraction from short- and long-read sequencing, statistical feature modeling (Gaussian mixture and NMF decomposition of fragment length profiles, end motif diversity, windowed protection scores), cohort-level differential testing of said features, and a simulation module. The library is exposed through a command-line interface, a Python API, and a Nextflow pipeline. We demonstrate pyfraglib in two ways. First, on two simulated 20-sample cohorts we show that pyfraglib's per-sample and cohort-level analyses recover the differences introduced by construction. Second, we apply pyfraglib to 88 cfDNA samples from a central nervous system lymphoma (CNSL) study and construct a fragmentomics score combining an NMF signature with end motif and WPS summaries via a classifier trained on cerebrospinal fluid and healthy donor plasma samples. As a proof of concept and applied to 66 baseline patient plasma samples, the score identifies a high-risk subgroup with worse failure-free survival (log-rank p=0.0247). Conclusions: pyfraglib integrates sample- and cohort-level fragmentomics analyses as well as in silico simulation within a consistently engineered Python framework. pyfraglib source code and documentation are available at https://github.com/schwarzlab-ccb/pyfraglib.
bioinformatics2026-08-09v2A Systematic Evaluation of Single-Cell Batch Integration Metrics and sBEE, A New Metric
Myradov, M.; HOUDJEDJ, A.; Tastan, O.; Kazan, H.Abstract
Background: Single-cell RNA sequencing (scRNA-seq) datasets generated across laboratories and experimental conditions often exhibit batch effects that obscure biological variation. Numerous computational methods have been developed to integrate such datasets, making robust benchmarking essential. Evaluation metrics play a central role in assessing integration quality. However, existing metrics each capture only specific aspects of batch mixing and often rely on assumptions that are violated in practice, such as every cell type being present in every batch, balanced batch composition, or simple cluster geometries. Consequently, benchmarking studies frequently report discordant rankings of integration methods, complicating interpretation and method selection. Results: We systematically evaluate widely used batch integration metrics using controlled synthetic scenarios that isolate common integration challenges, including imbalanced batch composition, partial cell-type overlap, heterogeneous cluster densities, and varying cluster geometries. Our analysis reveals distinct strengths, limitations, and failure modes, demonstrating that assessments of batch mixing can change substantially across different data characteristics. These observations motivated the development of sBEE (single-cell Batch Effect Evaluator), a metric that assesses batch mixing at the single-cell level using two complementary measures: local neighborhood composition and cross-batch distance relationships. Across diverse simulated scenarios, sBEE produces stable, interpretable evaluations and remains robust to the failure modes identified in existing metrics. We further validate these findings on multiple real-world scRNA-seq datasets. Conclusions: Our work provides a comprehensive characterization of the behavior of widely used batch integration metrics and introduces sBEE, a robust metric that combines complementary neighborhood and distance-based information to provide more reliable assessments of batch mixing. By identifying the conditions under which existing metrics succeed or fail and providing a unified alternative, sBEE enables more consistent benchmarking of single-cell integration methods.
bioinformatics2026-08-09v2Addressing challenges in agentic retrieval of structured data from biomedical databases
Halder, A.; Singh, M.; Kesarwani, R.; Mathew, B.; Bhattacharya, N.; Chikhaliya, O.; Motwani, D.; Peela, S. C. M.; Samanta, S.; Nagdev, H.; Piyush, ; Muddemmanavar, P.; Farooq, M.; Ahuja, G.; Sengupta, D.Abstract
Biomedical research increasingly relies on expert-curated databases to connect diseases, genes, variants, phenotypes, pathways and therapeutics. Incomplete or irreproducible retrieval can distort interpretation, mechanistic inference, variant assessment or therapeutic prioritization despite an unchanged evidence base. Agentic natural-language-to-SQL (NL2SQL) systems and Model Context Protocol-enabled agents can autonomously retrieve evidence from biomedical databases, including Open Targets and the Highly Confident Drug-Target Database, in response to natural-language questions. However, they may truncate large result sets, miss records when terminology differs from database vocabulary and return different outputs across runs. We systematically quantify these failures and introduce BioChirp, a goal-directed autonomous system for reliable retrieval from structured biomedical databases, built on interpretation-execution separation. A coordinated language-model layer interprets intent and resolves database-field assignments, after which entity resolution maps user terms to database terminology. A deterministic Steiner-tree planner and executor then retrieve matching records without further language-model involvement. Across curated biomedical databases spanning more than one million associations, BioChirp retrieved thousands of records for exhaustive queries, whereas baseline agents returned at most a few hundred or failed. Across 70 queries, five runs and three BioChirp database backends, BioChirp achieved a median cross-run Jaccard similarity of 1.0. Across 910 database-question evaluations derived from expert-written BioASQ questions and spanning ten databases, BioChirp retrieved at least one database record in 85.4% of cases and produced a fully correct answer in 51.1% of cases, compared with 44.6% and 17.4% for agentic NL2SQL. BioChirp is publicly available at https://biochirp.iiitd.edu.in.
bioinformatics2026-08-09v2Assessing Computational Models for Pharmacogenomic Variant Interpretation
Pucci, F.; Hermans, P.; Tsishyn, M.; Cusato, J.; Rooman, M.Abstract
Accurately predicting the effects of pharmacogenomic variants is essential for the development of personalized therapeutic strategies, as genetic variability can influence drug response differently across patients. Here, we assessed several computational approaches using a dataset of pharmacogenomic variants with either clinical annotations or functional characterization by deep mutational scanning, compiled from the literature, with an additional focus on CYP2C9, a clinically relevant drug-metabolizing enzyme. Our results show that, despite recent methodological advances, substantial room for improvement remains. In particular, current methods struggle to distinguish gain-of-function variants associated with increased drug clearance and fast-metabolizer phenotypes from neutral variants, whereas loss-of-function variants that reduce drug clearance are predicted more accurately. The integration of structural and evolutionary information appears to be a key strategy for improving performance, with the coevolution-based StructureDCA method achieving the highest accuracy compared with classical genetic variant-effect predictors and recent deep learning approaches, including the pathogenic-variant predictor AlphaMissense and general protein language model-based methods. Finally, our results indicate that computational models can complement in vitro experiments in clinical variant interpretation, as StructureDCA predictions showed better agreement with clinically annotated phenotypes than large-scale deep mutational scanning data in several cases.
bioinformatics2026-08-09v1pbcftools: parallel execution of bcftools for large variant call sets
Zhang, G.Abstract
Summary bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the results by hand. We present pbcftools, a Perl wrapper that does this automatically: it splits the genome into chunks, runs an ordinary bcftools command on each in parallel, and reassembles the outputs by a method suited to the data type. Across Linux servers, Windows/WSL2 workstations and Apple laptops, with bcftools 1.21 to 1.24, parallel output was identical to serial output for every command tested. On 1000 Genomes Phase 3 data, operations writing compressed VCF ran 10.8 to 21.1 times faster with 32 cores and up to 35.4 times with 64, those writing text 3.7 to 12.8 times, and merging 100 VCF files 19.2 times. pbcftools also runs on LSF and Slurm clusters. Availability and implementation pbcftools is written in Perl (>= 5.16) and requires bcftools; local parallel execution also requires Perl module Parallel::ForkManager. It is released under the MIT license at https://github.com/zhangge-uc/pbcftools (DOI: 10.5281/zenodo.21780361).
bioinformatics2026-08-09v1A Pseudo-Longitudinal Methylome Projection Framework Defines a Buccal PACE-like Aging-Rate Score from Cross-Sectional DNA Methylation Data
Shoji, T.; Nakaki, R.Abstract
Background: DNA methylation-based biomarkers have enabled robust estimation of biological age across tissues, and longitudinally trained measures such as DunedinPACE provide estimates of the pace of aging from blood methylomes. However, longitudinal methylation data are often unavailable, particularly for minimally invasive tissues such as buccal mucosa. Here, we developed a pseudo-longitudinal framework to estimate a buccal mucosa-derived PACE-like aging-rate score from cross-sectional methylome data. Methods: We used a buccal biological age estimator as an internal pseudo-time axis. Methylation beta-values were transformed to M-values, and CpG-specific smooth functions of biological age were fitted in cross-validation. Local derivatives of these functions were used to project each individuals buccal methylome forward by a small time step. The projected methylome was converted back to beta-values, biological age was recalculated, and the change in biological age per unit time was defined as a pseudo-aging velocity. This raw velocity was transformed to a non-negative PACE-like score centered at 1.0. We then trained cross-fitted models to predict the derived score from buccal CpG methylation profiles. Results: In 151 individuals, the proposed score was reproducibly predicted from buccal methylomes in out-of-fold analysis, with a Pearson correlation of 0.706 and Spearman correlation of 0.710 between observed and predicted PACE-like scores. Sensitivity analyses across CpG selection size and regression models showed broadly consistent performance. In contrast, the proposed buccal PACE-like score showed only modest association with measured DunedinPACE, and alternative attempts to reconstruct DunedinPACE from buccal methylomes, including supervised proxy modeling and buccal-to-blood CpG imputation, showed limited sample-level performance. Conclusions: These results support the feasibility of deriving a tissue-specific PACE-like aging-rate score from cross-sectional buccal methylome data by treating biological age as a pseudo-time axis. The proposed score should not be interpreted as a replacement for blood-derived DunedinPACE, but rather as an exploratory buccal methylome dynamics index that may capture tissue-specific aging-related variation.
bioinformatics2026-08-09v1Elucidating biosynthetic pathways related to the synthesis of small halogenated peptidic natural products in marine sponge microbiomes
Loureiro, C.; Schorn, M. A.; Alanjary, M.; Kuipers, B.; Louwen, J. J. R.; van der Oost, J.; Medema, M. H.; Sipkema, D.Abstract
Marine sponges are known sources of bioactive natural products (NPs), many of which are produced by associated bacterial symbionts via encoded biosynthetic gene clusters (BGCs). A particularly interesting subclass of sponge-derived NPs is comprised of small, brominated alkaloids, which are recovered from diverse habitats and host sponge taxonomies. Despite having been described decades ago, most of these NPs do not have an elucidated biosynthetic origin. We queried metagenomes of several sponge species by making use of a minimal set of core enzymes that we postulate to be necessary to produce these small peptidic NPs: an FADH2-dependent halogenase and an AMP-binding adenylation enzyme. This revealed a variety of novel BGC architectures, many of which showed conservation among sponge host phylogenies and were encoded in the genomes of diverse sponge-associated bacteria. Furthermore, we identified a BGC in the sponge G. barretti that is potentially linked to the production of the iconic barettins, given its enzymatic machinery and specific acidobacterial origin. The present work contributes to the challenging quest to link orphan brominated NPs to their parent BGCs in the sponge holobiont and beyond.
bioinformatics2026-08-09v1Scalable Extraction of Information on Protein-Protein Interactions using Topological Data Analysis
Mukherjee, A.; Park, B.; Malmstrom, A.; Cisewski-Kehe, J.; Van Lehn, R. C.; Zavala, V. M.Abstract
Protein-protein interactions (PPIs) govern a wide range of cellular functions. The ability to predict PPI interfaces from protein molecular surfaces is important for understanding protein function and enabling therapeutic discovery. While recent advances in structure-based learning, particularly molecular-surface geometric deep learning frameworks, have demonstrated that protein surfaces encode rich geometric and physicochemical information, such approaches often remain computationally intensive and data-hungry. Alternatively, topological data analysis (TDA) has emerged as a mathematically rigorous framework for extracting robust, multiscale shape information from complex data. In this work, we introduce a scalable TDA framework for extracting information on PPIs directly from localized protein surface patches. Our approach leverages multiscale topological descriptors, evaluated from patch-wise point cloud representations of protein mesh surfaces, combined with supervised machine learning models for interface prediction. On a full dataset of 3,362 proteins, the proposed approach substantially reduced computational cost relative to an established geometric deep learning method, MaSIF-site, decreasing preprocessing time from approximately 27 s/protein to 5-8 s/protein and total training time from approximately 6 h to 1-1.3 h. Importantly, this computational reduction is achieved while maintaining mean test area under the receiver operating characteristic curve (AUC) values of 0.76 and 0.77 for patch radii of 0.9 nm and 1.2 nm, respectively, thus approaching the MaSIF-site test AUC of 0.84. Our results suggest that topology offers a scalable and computationally efficient approach for high-throughput extraction of information from complex biomolecular interfaces.
bioinformatics2026-08-09v1pysigscore: gene signatures scoring across bulk and single-cell transcriptomics
Giacomello, T.; Mazzara, S.; Abbruzzese, G.; Barberis, A.; tangherloni, a.; Buffa, F. M.Abstract
High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures.
bioinformatics2026-08-09v1Move BeTween modAlities (MBTA) employs flow matching to predict single cell data modalities
Xu, B.; Zhang, Y.; Michor, F.Abstract
Integrating diverse molecular modalities to obtain a comprehensive view of cellular identity remains a major challenge in single-cell biology. A fundamental but underappreciated obstacle is structural mismatch-the phenomenon in which the neighborhood structure of a cell differs depending on which molecular modality is used to define it. Existing approaches typically embed modalities into a shared latent space, which actively erases the structural differences between modalities that make multimodal measurements scientifically valuable. Here we introduce Move BeTween modAlities (MBTA), the first framework explicitly designed to address structural mismatch. Rather than forcing modalities into a shared representation, MBTA maintains modality-specific latent spaces and connects them via flow matching, preserving the structural integrity of each modality while enabling accurate cross-modal translation. Across extensive benchmarks on multi-modal single-cell datasets, MBTA consistently outperformed existing methods, with the largest gains observed in datasets with pronounced structural mismatch. Applied to joint genomic and transcriptomic profiles of breast cancer patients, MBTA identified transcriptomic lineage relationships corroborated by genomic variation and outperformed state-of-the-art transcriptomics-based copy number inference methods. Extending this framework to mouse embryonic development, we reconstructed temporal trajectories jointly defined by gene expression and seven complementary epigenetic modalities. MBTA can connect any number of molecular readouts without erasing their individual character, serving as the computational foundation for assembling multi-layered portraits of cells.
bioinformatics2026-08-09v1GaugeFixer: overcoming parameter non-identifiability in models of sequence-function relationships
Marti-Gomez, C.; McCandlish, D. M.; Kinney, J. B.Abstract
Background: Mathematical models of sequence-function relationships are widely used in computational biology. A key challenge when interpreting these models is that many different choices for model parameters can encode the same sequence-function relationship. These ambiguities, called "gauge freedoms," must be removed before parameter values can be meaningfully interpreted. Doing this requires imposing additional mathematical constraints on parameter values, a process called "fixing the gauge." We recently developed mathematical methods for fixing the gauge of a large class of commonly used models. The most direct computational implementation of these methods is often impractical, however, because it requires a projection matrix whose size scales quadratically with the number of model parameters. Results: Here we introduce GaugeFixer, a Python package that exploits the mathematical structure of gauge-fixing projections to achieve log-linear scaling in computation time and linear scaling in memory usage. GaugeFixer thus requires orders of magnitude less time and memory than standard matrix multiplication, and can fix the gauge of models having millions of parameters in seconds. To demonstrate its utility, we used GaugeFixer to analyze an empirical fitness landscape for translation initiation in bacteria. This analysis revealed striking similarities (but also fine-scale differences) in ribosome-binding preferences at different positions relative to the start codon, thereby aiding the interpretation of this otherwise unwieldy landscape. Conclusions: GaugeFixer thus fills an important gap in the computational tools available for interpreting quantitative models of sequence-function relationships.
bioinformatics2026-08-07v4Beyond Bisulfite Sequencing: Resolving 5-hmC with Nanopore Sequencing Unmasks the True-5mC Methylation Entropy Landscape
Bertocchi, U.; Katz, E.; Jeffet, J.; Grunwald, A.; Gabay, N.; Deek, J.; Verma, S.; Shwartz, A.; Umschweif-Nevo, G.; Lerer, B.; Roichman, Y.; Ebenstein, Y.Abstract
DNA methylation dynamically regulates cellular function and phenotype. At the tissue level, stochastic variation in methylation patterns, measured as methylation entropy, drives plasticity, development, cancer, and aging. Demethylation is facilitated by erasure of 5-methylcytosine (5mC) via the oxidized intermediate 5-hydroxymethylcytosine (5hmC), but bisulfite sequencing cannot distinguish these modifications, classifying both as 5mC. Using nanopore sequencing with direct detection of 5mC and 5hmC, we quantified how this historical conflation affects genome-wide methylation levels and methylation entropy in kidney cancer and the mouse medial prefrontal cortex. Bisulfite-like analysis introduced systematic, tissue-specific shifts in methylation distributions, influencing biological interpretation. However, these effects were modest in the low-5hmC kidney cancer samples, where pathway-level results remained highly concordant. Our findings demonstrate that True-5mC-based methylation entropy redefines the physical mapping of epigenomes, demonstrating that, in some contexts, what was previously interpreted as stochastic maintenance failure is frequently the structured signature of distinct, mechanistically interpretable cytosine biochemistry.
bioinformatics2026-08-07v3Optimizing broadly neutralizing antibodies via all-atom interaction modeling and pre-trained language models
Song, Y.; Wu, F.; Wang, R.; Zheng, W.; He, B.; Yan, Q.; Huang, X.; Li, Y.; Chen, S.; Yuan, Q.; Rao, J.; Tang, Z.; Zhou, J.; He, H.; Zhao, J.; Yang, Y.; Yao, J.Abstract
Antibody optimization is a fundamental challenge, and the identification of antibody-antigen interactions is crucial in the optimization process. However, current methods cannot accurately predict antibody antigen interactions, providing limited functional guidance to improve the time-consuming and costly traditional optimization techniques. Here, we present InterAb and InterAb-Opt, a unified computational framework that integrates all-atom modeling with antibody language models to predict antibody antigen interactions and enable antibody optimization. Leveraging the proposed all-atom modeling approach, AtomInter, and pre-trained antibody language models, InterAb outperforms existing methods in predicting antibody specificity and antibody-antigen binding affinity. InterAb successfully identified influenza A virus-binding antibodies from an antibody library and accurately detected high-affinity antibodies in the AIntibody competition. Empowered by the robust functional insights from InterAb, InterAb-Opt was developed to optimize broadly neutralizing antibodies. For R1-32 antibody, biolayer interferometry results reveal that 85%, 80%, 90%, and 67.5% of the 40 InterAb-Opt-optimized antibodies exhibit enhanced binding affinities to wild-type SARS-CoV-2, Lambda, BQ.1.1, and EG.5.1, respectively, with a maximum improvement of up to 96-fold. For the newly emerging BA.2.86 and KP.3, 55% and 52.5% of the optimized antibodies notably transition from non-binding to binding. Neutralization assays demonstrated that the optimized antibodies exhibited enhanced neutralization activity across multiple targets, highlighting the capability of InterAb-Opt in engineering broadly neutralizing antibodies. This technology enables precise analysis of antibody-antigen interactions and optimization of broadly neutralizing antibodies, holding promise for addressing challenges in immune evasion and vaccine design.
bioinformatics2026-08-07v2Mind the gap: quantifying individual-population gap in depressive symptom dynamics through energy landscapes
Tsutsumi, M.; Kubo, T.; Kato, T. A.; Naoki, H.Abstract
People do not always feel as they appear. Someone who seems stable may struggle internally, whereas someone who appears distressed may experience it differently. This gap matters in psychiatry, where assessment relies on symptom scales and external evaluation. Here we developed mindGAP (Measuring INDividual-population GAPs in psychiatric energy landscapes), a hierarchical varia- tional Bayesian framework that uses longitudinal questionnaire data to estimate population-level and individual-level pairwise maximum entropy models (pMEMs). We applied mindGAP to time- series PHQ-9 data from 248 participants. The population landscape contained three major states, whereas individual-level landscapes often diverged from this shared structure. We quantified this gap as individual-population landscape divergence, which was associated not only with depressive sever- ity but also with modern-type depression-related traits (TACS-22) and interpersonal sensitivity-self traits (IPS-22). Thus, mindGAP opens a route to quantifying a previously unquantified gap between population-level and individual-level symptom organization.
bioinformatics2026-08-07v2FIDDL: depth-matched negative controls distinguish genuine interspecific introgression from competitive-mapping artifact
Taylor, K.; Shumaker, K. A.; Gray, S. J.; Bochman, M. L.Abstract
Interspecific introgression is routinely detected by competitively mapping reads to a concatenated multi-species reference and calling regions where a non-focal species recruits coverage. Using strains that cannot contain the ancestry being detected, we show this design generates substantial false-positive signal through two mechanisms with opposite phylogenetic-distance signatures. Standard nuclear assemblies omit the mitochondrion and 2-micron plasmid, leaving high-copy cytoplasmic reads without a legitimate target; completing the reference preferentially removes signal from the most divergent donor. Genuine cross-species sequence conservation inflates the most closely related donor. Masking chromosome ends removes its subtelomeric part but plateaus at a non-zero floor, and the interior residual traces to conserved single-copy genes where a short read carries under one base of discriminating information. The floor grows with sequencing depth (1.19% of callable positions at 50x, 2.02% at 147x, 3.85% at 393x in a pure strain), is not mitigated by long reads, and appears at sub-diploid dosage - three properties widely read as evidence of authenticity. Because the discriminating information is below single-read resolution, no read-level filter separates artifact from introgression; we show three that fail. What works is locus-level: a consensus-phylogenetic test (29/29 specificity on confirmed artifact) and an allele-fraction donor-match test, complementary and validated in both directions on independent published introgression. We package the comparative controls as FIDDL (False Introgression Detection via Depth-matched controls and Loci-recurrence), an open-source tool, withdraw two of our own analysis-ready calls, and show re-analysis of published wild isolates reduces low-confidence introgression by ~53% while leaving high-confidence signal intact.
bioinformatics2026-08-07v1SLIM: A small linear model with STRING embeddings for single-cell genetic perturbation prediction
Hu, D.; Pielies Avelli, M.; Jensen, L. J.; Rasmussen, S.Abstract
Predicting cellular responses to genetic perturbations is central to understanding gene function and prioritizing therapeutic targets, but experimental screens cannot exhaustively cover genes, cell types, and perturbation combinations. Recent benchmarks have shown that simple baselines can match or outperform substantially more complex models, suggesting that informative biological priors may be as important as model capacity. Here we present SLIM, a lightweight extension of the bilinear model of Ahlmann-Eltze et al. SLIM represents perturbations with 64-dimensional embeddings derived from the STRING protein network and predicts mean transcriptional responses through a closed-form ridge-regression estimator. It then constructs single-cell populations by retrieving training cells and rescaling each gene to match the predicted mean. We evaluated SLIM against four deep learning models and two simple baselines on four single-gene perturbation datasets and one combinatorial perturbation dataset. Across these within-dataset benchmarks, SLIM achieved competitive mean-response accuracy, ranked first in eight of twelve single-gene dataset-metric comparisons, and produced substantially lower maximum mean discrepancy values than the evaluated alternatives. The model has 640 trainable parameters and fitted each benchmark dataset in under 10 seconds on a CPU. These results show that compact biological representations can support accurate and computationally efficient perturbation prediction. Code is available at https://github.com/RasmussenLab/SLIM.
bioinformatics2026-08-07v1A confound-diagnostic toolkit for in silico perturbation with single-cell foundation models
Qiu, R.; Zhao, M. M.Abstract
Deleting a gene token from a cell's input sequence offers a convenient native strategy for in silico perturbation, but the resulting embedding delta may not represent a biological knockout response. Apparent effects can instead reflect gene identity, universal responsiveness, limited tokenization coverage, library-size contamination, or circular state scoring. Here, we present a confound-diagnostic framework combining held-out increment testing, responsiveness adjustment, coverage gating, library-size diagnostics, and de-circularized state-shift analysis, together with a numerically matched reimplementation of frozen Geneformer's perturbation engine. Across Frangieh and Replogle datasets and linear and nonlinear readouts, the native embedding delta provided no reproducible held-out improvement beyond gene identity. Signal-injection calibration showed that the test detected injected residual signal, whereas native increments remained below its detection floor. Matched controls traced apparent positives to raw-count library-size structure, broad responsiveness, and self-referential scoring, while coverage constrained perturbation applicability and estimate stability without establishing biological specificity. This model-adaptable framework helps determine when foundation-model perturbation readouts warrant biological interpretation.
bioinformatics2026-08-07v1Benchmarking single-cell foundation models in a zero-shot setting
Gaballa, Y.; Ahmed, S.; Abdelaal, T.Abstract
Single-cell foundation models have recently emerged as a promising approach for learning general-purpose representations from large-scale transcriptomic data. These models are trained on millions of cells and are designed to transfer their learned representations to a wide range of downstream tasks. However, their practical benefits compared to traditional approaches are still not fully understood. This study evaluates four foundation models, namely scGPT, SCimilarity, UCE, and Transcriptformer, across four downstream tasks: cell type annotation, human data integration, cross-species data integration, and protein expression prediction. Embeddings generated by each model were assessed using multiple public single-cell datasets and compared against conventional machine learning baselines. Performance was measured using task-specific evaluation metrics, including classification, integration, and regression metrics. The results showed that foundation model embeddings did not consistently outperform traditional approaches. In the cell type annotation task, baseline methods achieved the strongest performance across most datasets. For protein expression prediction, however, embeddings from the foundation models generally produced more accurate predictions than the baseline, with SCimilarity achieving the lowest prediction error and Transcriptformer obtaining the highest correlation scores. In the data integration task, all foundation models produced moderate results, while scVI (the baseline) achieved the strongest integration performance. Overall, the results suggest that current single-cell foundation models provide useful representations for some downstream tasks in zero-shot conditions but do not yet offer a universal replacement for task-specific methods. Their effectiveness remains dependent on the application and evaluation setting.
bioinformatics2026-08-07v1StainX: GPU-accelerated batch stain normalization for computational pathology at scale
Moustafa, S.; Zheng, Y.; Rendeiro, A. F.Abstract
Stain normalization reduces color variability in histopathology whole-slide images, but cohort-scale pipelines lack fused multi-image batch transforms for classical methods. We present StainX, a GPU-accelerated batch stain normalization framework built around a two-stage fit/transform interface. It implements histogram matching, Macenko, and Reinhard normalizers through a portable PyTorch backend and an optional CUDA backend that fuses per-pixel operations for batch throughput. On NVIDIA GPUs, the fused CUDA path outperforms the torch CPU backend by 168x, 70x, and 48x for Reinhard, histogram matching, and Macenko respectively, and exceeds the fastest GPU peers by 7-8x (Reinhard) and 2x (Macenko) at comparable accuracy. StainX also provides user-selectable precision modes, a documented Python API, continuous integration testing, and online documentation. Source code available at https://github.com/rendeirolab/stainx, and documentation at https://stainx.readthedocs.io. Implemented in Python. Runs on Linux, macOS, and Windows.
bioinformatics2026-08-07v1