Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
WITHDRAWN: Deviation Error: assessing machine learning predictions for replicate measurements in genomics and beyond
Abdulnabi, H.; Westwood, J. T.Abstract
The authors have withdrawn this manuscript because the foundational metric used throughout the study, the Deviation Error, requires substantial mathematical formalization to be established as a proper scoring rule. Completing this rigorous theoretical validation and the necessary overhaul of the accompanying synthetic case studies requires significant additional research. We intend to revisit the Deviation Error methodology in future research. Therefore, the authors do not wish this work to be cited as reference for the project. If you have any questions, please contact the corresponding author.
bioinformatics2026-07-27v3PrimeKG-Plus: a refreshed and rare-disease-enriched precision medicine knowledge graph
Nguyen, T. T. D.; Nguyen-Phuong, T.; Nguyen, Q.-H.; Abbasi, A. M.; Le Phan, H.-D.; Nguyen, L. B.-A.; Phan, N.-T.; Curabaz, N. N.; Hauser, A. S.; Tanoli, Z.; Nguyen, D. T.; Kooistra, A. J.Abstract
Biomedical knowledge evolves rapidly, yet most disease-centered knowledge graphs remain unchanged after publication. We present PrimeKG-Plus, an extension of PrimeKG that updates all 20 original data resources to releases available as of Dec 2025 and incorporates additional resources, including OpenTargets, RepurposeDrugs, and nSIDES. The graph is further expanded with relations extracted from 637 PubMed abstracts and PMC full-text articles using a large-language-model-assisted curation workflow. Extracted relations were refined through normalization, UMLS-based synonym mapping, SapBERT embedding based similarity ranking, and human expert review to ensure data quality. This literature-driven expansion focuses on rare neurological disorders, including Canavan disease, Niemann-Pick disease type C, Tay-Sachs disease and Batten disease. Network topology analyses indicate that the expanded graph improves indirect drug-disease connectivity of 3-6 hops through the graph while adding 447,288 Drug-Protein-Disease paths linking drug-disease pairs that were previously unreachable. Temporal FDA validation captured 55 new molecular entities after the June 2021 PrimeKG data cut-off as drug nodes, 46 absent from the original PrimeKG. By integrating newly incorporated and updated resources with systematically curated literature-derived associations, PrimeKG-Plus provides an up-to-date knowledge graph for network-based drug repurposing and precision medicine applications. Data and code are publicly available.
bioinformatics2026-07-27v3The practical impact of numerical variability on structural MRI measures of Parkinson's disease
Chatelain, Y. M. B.; Sokołowski, A.; Sharp, M.; Poline, J.-B.; Glatard, T.Abstract
Numerical variability is rarely quantified in neuroimaging despite many measures relying on subtle morphometric differences across individuals. We instrumented FreeSurfer 7.3.1, a widely used neuroimaging pipeline, to simulate numerical differences across computational environments, and used it to measure numerical variability in MRI analyses of Parkinson's disease patients and controls. In multiple cortical and subcortical regions, numerical variation reached nearly one-third of the population variability, altering statistical conclusions about group differences and clinical associations. To assess the impact of numerical noise in existing studies, we developed a practical tool that estimates the Numerical-Population Variability Ratio (NPVR) in a study, and propagates the resulting numerical variability to common statistics and associated p-values. By applying this framework to thirteen previously published studies reporting MRI measures in Parkinson's disease, we quantified the probability of numerically induced false positives and false negatives in the literature, highlighting a substantial impact of numerical variability on MRI measures of Parkinson's disease with an average significance-flip probability of 5% for cross-sectional studies and 10% for longitudinal studies. These results underscore the importance of systematically evaluating numerical stability in neuroimaging and provide a practical framework to do so.
bioinformatics2026-07-27v3Perturbation response decomposition enables biologically aligned generalization to unseen perturbations and cellular contexts
Molina, A.; Zhang, X.Abstract
Predicting single-cell responses to genetic perturbations could reveal the vast combinatorial space of perturbations and cellular contexts that is infeasible to measure experimentally, yet current deep learning models generalize poorly and often fail to outperform simple baselines. Here we demonstrate that generalizability in perturbation prediction requires identifying and representing distinct components of cellular response rather than on increasing model complexity alone. We introduce a decomposition framework that explicitly separates transcriptional responses into global, perturbation-specific, cell-line-specific, and perturbation-by-cell-line interaction components. Applied to four CRISPR-interference Perturb-seq screens on multiple cell lines, our framework reveals that these components have distinct structures and information requirements. The global response component is low-dimensional, reflects recurrent proliferation and stress response programs, and can be inferred from control gene expression. In contrast, the perturbation and cell-line specific components are high-dimensional and cannot be recovered from control expression. We therefore develop response-component-aligned models that map biological priors, such as gene coessentiality, onto the geometry of observed transcriptional responses. Critically, this alignment enables simple linear or multilayer perceptron (MLP) based models to outperform state-of-the-art architectures across multiple generalization settings, including unseen cell lines and combinations of unseen perturbations. Together, our framework for response decomposition and alignment provides a principled basis for evaluating and designing perturbation-prediction models, showing that generalization depends primarily on matching biological information to the response components rather than on model complexity alone.
bioinformatics2026-07-27v1PUMA: A Phenotypic Unsupervised Model of Aging Reveals Distinct Aging Dimensions
Ghorbani, F.; Nollen, E. A. A.; Guryev, V.Abstract
Aging is a multidimensional process, yet most aging models reduce it to a single score to estimate biological aging and predict health outcomes, disease risk, or mortality. Here, we introduce PUMA (Phenotypic Unsupervised Model of Aging), a framework that characterizes aging through multiple phenotypic dimensions. Applying PUMA to more than 1,000 traits spanning behavioral, psychological, social, physical, environmental, and biomedical domains in over 150,000 individuals from the Lifelines cohort, we identified seven distinct phenotypic aging dimensions. These dimensions were significantly associated with the future incidence of major age-related diseases, including cancer, diabetes, COPD, heart failure, stroke, and Parkinson's disease, demonstrating the potential of PUMA to stratify individuals by disease risk and identify phenotypic domains for targeted intervention. Notably, dimensions reflecting psychosocial factors, particularly cumulative life stress, predicted disease risk as strongly as, or more strongly than, traditional biomedical risk factors, highlighting the importance of psychological and social influences in aging and disease risk. The observation that phenotypic aging dimensions are differentially associated with the future risk of age-related diseases supports a multidimensional model of aging, indicating that aging is not a single uniform process. Because PUMA relies on accessible phenotypic data, it provides an interpretable framework for disease risk stratification and targeted preventive strategies.
bioinformatics2026-07-27v1Decoding the Oral-Cardiac Axis: FCN1 and LYN as Key Players in the Molecular Dialogue between Acute Myocardial Infarction and Periodontitis
Zhang, K.; Wang, Y.Abstract
Background: Acute myocardial infarction (AMI) and periodontitis (PD) have been epidemiologically linked, but the molecular mechanisms underlying this association remain elusive. We aimed to elucidate shared pathogenic signatures between AMI and PD using a comprehensive bioinformatics approach. Methods: We integrated transcriptomic data from multiple Gene Expression Omnibus datasets, including an AMI cohort profiled from enriched circulating endothelial cells (GSE66360) with whole-blood validation (GSE48060), and PD gingival tissue cohorts (GSE16134 and GSE10334). Weighted gene co-expression network analysis, protein-protein interaction network analysis, and LASSO-based feature selection were applied. Functional enrichment, immune deconvolution, transcription factor-gene regulatory network analysis, and single-cell RNA sequencing analysis were performed to characterize shared molecular features. Results: We identified 95 shared differentially expressed genes (DEGs) between AMI and PD. By intersecting the shared upregulated DEGs with disease-associated WGCNA modules, we obtained 46 candidate shared genes. LASSO-based feature selection further highlighted FCN1 and LYN as overlapping candidates, which showed good discriminative performance in both training and external validation cohorts in ROC analyses. Enrichment analyses suggested that the shared signature was mainly related to myeloid cell migration, phagocytosis, and neutrophil-related inflammatory pathways (e.g., neutrophil extracellular trap formation). Immune deconvolution in PD gingival tissues suggested increased plasma cells and neutrophils and decreased resting memory CD4+ T cells; immune deconvolution results in the AMI cohort were interpreted cautiously due to the CEC-enriched sample source. Single-cell analysis revealed that FCN1 and LYN were predominantly expressed in macrophage/monocyte-derived populations. Conclusions: Our study suggests a shared inflammatory and immune-mediated transcriptomic program linking AMI and PD, and identifies FCN1 and LYN as candidate shared immune markers. These findings provide a molecular rationale for the oral-cardiovascular association and are hypothesis-generating, warranting future experimental and prospective validation.
bioinformatics2026-07-27v1Spaceland: Histology-Guided Reconstruction of High-Resolution Whole-Organ 3D Molecular Atlases from Sparse Spatial Transcriptomics
Xu, F.; Zhuang, Z.; Zhu, Y.; Ying, B.; Hou, N.; Lin, W.; Wang, L.; Yang, C.; song, j.Abstract
Reconstructing whole organs in three-dimensional molecular detail is a key step toward building virtual organs for modeling tissue organization, disease progression and drug perturbation responses. However, high-resolution whole-organ spatial transcriptomic profiling remains impractical, forcing a trade-off between reconstruction fidelity and sampling density. Here, we introduce Spaceland, a morphology-guided framework that reconstructs continuous, high-resolution 3D molecular landscapes from sparsely sampled spatial transcriptomic sections and serial H&E histology. Spaceland formulates this task as learning continuous gene-expression fields within a morphology-informed histological space. It constructs a dense 3D morphological scaffold by optical-flow interpolation of foundation model-derived H&E representations and decodes sparse spot-level transcriptomic measurements onto an 8 m histology-aligned grid. Across mouse olfactory bulb, mouse hemibrain and spatiotemporal planarian regeneration, Spaceland generalized across platforms, tissue scales and biological contexts. In mouse benchmarks, Spaceland bridged 400 m molecular gaps, resolved sub-spot organization, outperformed ST-based interpolation and 2D H&E-based prediction methods, and remained robust with 160-320 m H&E intervals. In planarian regeneration, it enabled time-resolved whole-organism analysis from only four Visium sections per stage, revealing dynamic neoblast-neural spatial remodeling. Together, Spaceland shifts 3D molecular atlas construction from exhaustive experimental sampling toward data-driven virtual tissue and organ modeling, providing a scalable route to whole-organ molecular reconstruction.
bioinformatics2026-07-27v1An integrated single-cell atlas of human lung across the lifespan
Huang, L.; Huang, Z.; Feng, S.; Fang, K.; Wu, J.; Ao, Y.; Huang, L.; Zhang, J.; Sun, H.; Miao, Z.Abstract
The healthy adult lung is largely quiescent but can respond to injury and replace damaged or lost cells. However, the identities and origins of regenerative cell states remain poorly understood. Here we present the Human Developmental Lung Cell Atlas (HDLCA), a curated single-cell reference atlas of the respiratory system across the lifespan, integrating 253 datasets from 225 studies, and comprising over 18 million cells from 3,198 samples and 2,460 individuals across 34 anatomical locations in both health and disease. The HDLCA defines 142 consensus lung cell types, including 13 rare and 8 previously undescribed ones. Notably, we identify a rare intermediate alveolar epithelial progenitor (Int AP) population in normal lungs, with transcriptional signatures shared by alveolar type 2 (AT2) and alveolar type 1 (AT1) cells. Trajectory inference indicates that Int AP cells arise from both AT2 and SCGB1A1+ secretory cells and give rise to AT1 cells, with Hippo-YAP/TAZ signalling governing lineage specification. Int AP cells are associated with genetic susceptibility to idiopathic pulmonary fibrosis and chronic obstructive pulmonary disease, and their dysfunction may contribute to abnormal alveolar repair and regenerative failure in both diseases. The HDLCA provides a unified atlas-level reference for mapping respiratory cellular diversity, uncovering reparative cell dynamics and informing regenerative and cell-based therapeutic strategies.
bioinformatics2026-07-27v1Discrete Inverse Rendering: Biological Data Analysis with Integer Programming
Kirkegaard, J. B.; Zdyb, F. O.Abstract
Biological image analysis is full of discrete decisions: whether an object is present, which of several overlapping detections is real, whether two detections match across time, and whether a cell divides. Standard pipelines resolve them locally with non-max suppression, thresholding, or greedy linking, committing before all image and temporal evidence is in. We recast such problems as discrete inverse rendering: candidate renderings are generated then jointly selected to reconstruct the movie subject to temporal and biological constraints, solved to certified optimality with a modern integer-programming solver. The same formulation covers suppression of overlapping detections, selection of a structure as a path, and event-structured tracking with birth, death, and division. Applied to C. elegans splines, sperm flagella, and dividing cells, the method matches specialised state-of-the-art pipelines across three imaging modalities on a single objective, with the largest gains where per-frame segmentation is unreliable (on a low-signal fluorescence movie of Huh7 hepatoma cells, detection F1 doubles from 0.31 to 0.58).
bioinformatics2026-07-27v1Estimation of biological age using HRV data: comparison of the Klemera-Dubal method with the multiple linear regression method
Pysaruk, A.Abstract
The aim of this study was to compare three methods for estimating biological age (BA) based on heart rate variability (HRV) indices recorded in the supine position: multiple linear regression (MLR), MLR corrected for regression-to-the-mean bias using the Dubina method (1984), and the Klemera-Doubal Method (KDM, 2006). A total of 343 subjects were examined (193 women and 150 men aged 20-90 years); nine time- and frequency-domain HRV indices were analyzed (NN, lnSDNN, lnRMSSD, lnVLF, lnLF, lnHF, VLF%, LF%, HF%). The association between HRV indices and chronological age was substantially in women and men (p<0.01). KDM produced the lowest estimation error (MAE=4.34 years in women, 3.67 years in men), whereas uncorrected MLR produced the largest error (MAE=6.57 and 7.47 years, respectively); the Dubina correction substantially improved MLR accuracy (MAE=4.22 and 4.82 years) but relied directly on actual chronological age, which limits its independent diagnostic value. The strengths and limitations of each method and recommendations for their application are discussed. Keywords: biological age, heart rate variability, Klemera-Doubal method, multiple linear regression.
bioinformatics2026-07-27v1Explainable Artificial Intelligence for Cross-Dataset Generalizable Biomarker Discovery in Cardiovascular diseases (CVDs)
Abbasi, A. F.; Sajjad, M.; Vollmer, S.; Dengel, A.; Asim, M. N.Abstract
CVDs are heterogeneous, multifactorial disorders that remain the leading cause of global mor- tality from infancy to old age. It requires an early identification and treatment of risk factors to accelerate disease prevention and morbidity improvement. Advancements in transcriptomics tech- nologies gives large pool of heterogenous gene expression data. The technical heterogeneity of gene expression data reduces ability to compare multiple cross-platform datasets at once. To bridge gap, we systematically evaluate three data harmonization techniques: Shambhala-2, TDM, and UPC to align heterogeneous data into a shared expression space while preserving biological signals. Our pipeline integrates 25 independent datasets comprising 983 samples across 23 distinct CVDs phe- notypes from both RNA-seq and microarray platforms. The framework benchmarks 35 Machine learning (ML) and Deep learning (DL) classifiers, including Transformers and ResNets, across three data modalities such as RNA-seq, microarray hybridization and RNA-seq + microarray and multiple tissue types. To ensure clinical trustworthiness, we apply multiple Explainable artificial intelligence (XAI) methods, such as SHapley additive exPlanations (SHAP) and Integrated gradi- entss (IGs), and assess their reliability using quantitative metrics like Area over the perturbation curve (AOPC), Sensitivity, and Infidelity. Results indicate that Shambhala-2 provides superior harmonization by maximizing the biological signal-to-platform ratio. Evaluation of XAI methods reveals that Shapley-based approaches offer the highest stability for identifying influential genomic features in high-dimensional data. Functional enrichment and pathway analyses further confirmed the involvement of identified biomarkers in key cardiovascular processes, including inflammation, immune regulation, oxidative stress, and vascular remodeling. Collectively, this study provides a scalable and interpretable road-map that integrates XAI with cross-dataset biomarker discovery, supporting the transition toward precision cardiology.
bioinformatics2026-07-27v1Detecting the Information Flow in Proteins by Hodge Decomposition
Hacisuleyman, A.Abstract
Allosteric communication in proteins is commonly quantified as a directed or undirected coupling between residues, but such descriptors mix distinct modes of signalling into a single pattern. Here we treat the net transfer entropy flux from a dynamic Gaussian network(dGNM) model as an edge flow on the residue contact graph and apply the combinatorial Hodge decomposition, which dissects the flow orthogonally into a gradient (global source--to--sink hierarchy), a curl (local three--clique circulation) and a harmonic (cavity--scale circulation) component. Applied to the wild--type KRAS and ten oncogenic KRAS variants spanning the principal GTPase--cycle mechanism classes, partial--hydrolysis position--12 (G12D, G12C, G12S), GAP--occluding position--12 (G12V, G12R), catalytic switch II (Q61R, Q61H), fast--cycling (G13D, A146T) and a combined steric and catalytic double mutant (G12D/Q61H), on a side--chain--centroid contact network, the decomposition shows that the transfer entropy flux is overwhelmingly hierarchical: the gradient term carries 97.5--98.4% of the flux in every variant (permutation p = 0.002), and the recovered scalar potential is strongly anti--correlated with each residue's net outgoing transfer entropy (Spearman {rho} {approx} -0.92 to -0.95). The hierarchy is conserved in magnitude but relocated by mutations: the dominant information sources move from the C terminal 5/hypervariable region in wild type into the nucleotide--processing core, the switch I/II machinery and the 4/distal lobe in a way that tracks the GTPase--cycle mechanism of the substitution, while the sinks remain fixed. The method provides a parameter--free, residue--level readout of how mutations of different mechanism reposition the source of allosteric signalling in KRAS.
bioinformatics2026-07-27v1B-SMART-Former: An Explainable Transformer-Based Deep Learning Model for Predicting Drug-Drug Interactions Between Biotech and Small-Molecule Drugs
Nasiri, F.; Hooshmand, M.; Nouroozi, M.Abstract
Drug-drug interactions between biotech and small-molecule drugs play a critical role in medication safety and therapeutic efficacy. However, most existing computational DDI prediction methods focus primarily on interactions between small-molecule drugs, leaving biotech-small-molecule interactions comparatively underexplored. In this study, we propose B-SMART-Former, an explainable deep learning framework for predicting interaction types between biotech and small-molecule drugs. The proposed framework integrates ChemBERTa embeddings and Morgan molecular fingerprints for small molecules with ProtBERT embeddings for biotech drugs, eliminating the need for similarity-based features while leveraging complementary molecular representations. These multimodal features are processed by a hybrid architecture that combines Transformer-based self-attention, residual convolutional learning, and a multi-layer perceptron classifier to capture both global contextual dependencies and local discriminative patterns. The model is formulated as a multi-class classification task and evaluated using stratified 10-fold cross-validation. To improve model transparency, Integrated Gradients is employed as a post-hoc explainability method to identify the molecular features that contribute most strongly to each prediction. Experimental results demonstrate that B-SMART-Former achieves a micro-averaged AUROC of 0.9978 and an AUPR of 0.9682 while relying solely on intrinsic molecular representations, remaining competitive with similarity-based approaches. The proposed framework offers an effective and explainable solution for biotech-small-molecule DDI prediction and provides a practical foundation for future computational drug interaction studies.
bioinformatics2026-07-27v1Functional Characterization of Transcriptome-Wide Isoform Switching in Hürthle Cell Carcinoma (HCC)
Butt, R. S.; Amir, A.; Paracha, R. Z.Abstract
Hurthle cell carcinoma (HCC) is an aggressive form of thyroid cancer. While mitochondrial DNA mutations and chromosomal losses have been identified in HCC, isoform switching, and its functional consequences remain uncharacterized. This study reanalyzed NCBI GEO dataset GSE228870 (n = 32), using Salmon and IsoformSwitchAnalyzeR() to identify isoform switching. The analysis resulted in 371 switches across 335 genes showing functional consequences including loss of protein domains, shorter open reading frames (ORFs), loss of signal peptides and novel sub-cellular localizations. Most significant isoform switches (q-value < 0.05, |dIF| > 0.1) were observed in LAMA2, LSP1, MAD2L2, FBLN2 and CXCL12, implicating extracellular matrix dysregulation, DNA damage response, immune signaling and cytoskeleton regulation. These genes are expressed in normal thyroid (median TPM 20.69, 11.66, 14.79, 134.1 & 80.76). However, specific isoforms of LAMA2 and MAD2L2 are not expressed in normal thyroid, explaining tumor-specific expression in HCC. Alternative transcription termination site (ATTS) gain was significant, suggesting altered 3' end in HCC transcripts. TCGA SpliceSeq showed LSP1, FBLN2 and CXCL12 undergo alternative promoter (LSP1 exon1 PSI=94.5%, FBLN2 exon2 PSI=99.0%) and alternative termination (CXCL12 exon3.3 PSI=53.9%) in thyroid cancer, suggesting ATTS and alternative transcription start site (ATSS) as shared splicing dysregulation mechanisms. This is the first systematic characterization of isoform-level dysregulation in HCC.
bioinformatics2026-07-27v1SyntenyPair Explorer: an installation-free, browser-based tool for interactive pairwise genome synteny visualization
Gibbons, J. G.Abstract
Comparisons of genome structure between related organisms are central to understanding genome evolution, gene family dynamics, and the genomic basis of phenotypic variation. Synteny, the conserved co-localization of genes along chromosomes, is most readily interpreted visually, yet many existing synteny visualization tools require local software installation, command-line proficiency, and/or dedicated server infrastructure, and produce static images that cannot be explored interactively. Here, I present SyntenyPair Explorer, a lightweight, installation-free tool for interactive visualization of synteny between two genomes. The application runs entirely within a standard web browser as a single, self-contained HTML file with no external dependencies and no server-side component. SyntenyPair Explorer accepts standard file formats already produced by common comparative genomics workflows, including FASTA genome assemblies (used to compute optional assembly summary statistics), GFF3/GTF gene annotations, and either BLAST tabular output (outfmt 6) or MCScanX collinearity files. Syntenic relationships are resolved by gene-identifier matching between the relationship file and the gene annotations, so that each relationship corresponds to a discrete gene-to-gene link. Interactive features include continuous zoom and pan, gene search with automatic centering of the partner genome on the syntenic counterpart, synteny block coloring, extensively customizable gene highlights and annotation callouts, session saving and restoration, and publication-quality image export. I demonstrate the tool by visualizing structural differences at the alpha-amylase loci between two strains of the industrially important fungus Aspergillus oryzae. SyntenyPair Explorer lowers the technical barrier to interactive synteny visualization and is freely available under the MIT license at https://github.com/GibbonsLabGenomics/SyntenyPair-Explorer, with a live browser-based version at https://gibbonslabgenomics.github.io/SyntenyPair-Explorer/.
bioinformatics2026-07-27v1Utanos: A general-purpose shallow whole-genome sequencing analysis workflow identifies interpretable copy number signatures
Douglas, J. M.; Lynch, B. J.; Yiu, J. C. H.; Nicholson, S.; Vasquez-Rios, C.; Ma, D.; Huntsman, D. G.; Park, Y.Abstract
Summary: A modular FASTQ-to-figures solution for analyzing low-depth or shallow whole-genome sequencing (sWGS) data. Shallow WGS can be used to detect copy number (CN) aberrations, Homologous Recombination Deficiency (HRD), and to detect and create CN signatures. Growing in popularity, this sequencing type is used for neonatal diagnostics and studying cancer. One of the major benefits is the reduced cost compared with deeper sequencing modalities such as Whole Genome Sequencing (WGS). With just 15 million reads often targeted, and sample substrate options like Formalin-Fixed, Paraffin-Embedded (FFPE) blocks widely available, this approach enables an affordable study to be performed at scale. Our pipeline and R package are an end-to-end solution implemented with reusability and modularity in mind. It makes entry and exit from the ecosystem easy, providing regular standardized output formats throughout execution. The pipeline is written in the well-supported and cross-platform Nextflow framework and has been submitted for inclusion in nf-core. Additionally, a Docker image for the utanos R package has been created to improve modularity. Availability and Implementation: The latest version of all software is freely available on GitHub. For the full processing pipeline, visit: https://github.com/Huntsmanlab/swgs-processing-pipeline. For just the utanos R package, visit: https://github.com/Huntsmanlab/utanos.
bioinformatics2026-07-27v1Towards Principled Evaluation of Single-Cell Perturbation Prediction Models
Schäfer, P. S. L.; Reid, K.; Boldyga, Z.; Aksu, E. D.; Hakem, H.; Saez-Rodriguez, J.Abstract
Single-cell perturbation experiments measure how interventions alter cellular phenotypes. However, the number of possible perturbations and biological contexts far exceeds what can be tested experimentally. Motivated by this constraint, predictive models aim to extrapolate cellular responses to unseen conditions. Despite substantial efforts in model development, benchmark studies have reached inconsistent conclusions about the capabilities of current perturbation-response models. A major challenge is that evaluation protocols vary widely across studies, making results difficult to compare. Furthermore, the lack of consensus on evaluation hampers progress because it is unclear which predictive capabilities new models should prioritize. To help build consensus on evaluation principles, we develop a taxonomy that decomposes evaluation protocols into their representation, metric, score transformation, and reporting strategies. We characterize how these choices determine which aspects of prediction quality a benchmark measures and discuss criteria for selecting and assessing protocols in relation to specific benchmarking goals. We additionally provide scPertEval, a Python package with reference implementations of selected evaluation protocols, and use it to assess protocol behavior across seven publicly available single-cell perturbation datasets. By making evaluation choices and their underlying trade-offs explicit, we aim to stimulate a community discussion about developing more comparable and task-aligned evaluation protocols.
bioinformatics2026-07-27v1deepthought: the microscopy acquisition stack as an object of study
Kesavan, P. S.; Devadasan, S.; Bohra, D.Abstract
Analysis-in-the-loop microscopy has been demonstrated many times, but it is rarely used outside the laboratories that build it. Each demonstration constructs its own acquisition infrastructure, so little transfers between them, and it has remained unclear which parts of the problem are already solved. In this work, the microscopy acquisition stack was itself treated as the object of study, and was investigated by construction. A minimal stack was built end to end, and four applications were then driven through it as test conditions. These were fixed-cell high-throughput immunofluorescence, live time-lapse imaging of an unsynchronized population, fluorescence anisotropy imaging, and autonomous focus and exposure. Each element of the stack was then classified by how it varied across these applications. Device access, sequencing, data storage and viewing held constant, and mature implementations of each were adopted unchanged. Interpretation, or how an image becomes a set of entities, differed with the application and belongs behind an interface. Two elements had nothing available to adopt and were therefore built. These are a geometric representation of the sample that a plan can traverse, and a representation of a run that yields detected objects rather than images. With those two in place, feedback from analysis into acquisition was ordinary control flow. The system was applied to the DNA damage response, where 22,000 cells were acquired and analyzed without operator intervention, and cells were followed through mitosis over 24 hours in an unsynchronized population without chemical synchronization. One element, targeting, or the choice of where to observe next, varies between applications and remains unabstracted in this implementation. It is identified here as the next requirement.
bioinformatics2026-07-26v3An integrated resource for systems-level analysis of aging hallmarks and associated genes
Tiwari, R.; Balaji, M.; Chivukula, N.; Sil, P.; Samal, A.Abstract
Aging is a complex biological process involving progressive cellular dysfunction, tissue decline, and increased susceptibility to multiple chronic diseases. A systemic view of aging through its established hallmarks provides a structured framework to understand this complexity and drive therapeutic discovery. To this end, we present AgingHallmarksDB, an interactive web platform that enables systems-level analysis of hallmark-associated gene sets. Aging-related genes were first curated from seven established resources, and those present in at least two of these resources were considered as consensus aging-related genes. Using functional annotations derived from GO, KEGG, and Reactome, a total of 3111 genes were mapped to the 11 aging hallmarks, of which 2593 were supported by additional experimental or manually curated evidence, with 1089 of these forming the consensus set. Further, AgingHallmarksDB supplements hallmarks and gene annotations with tissue or cell-type class specificity, exosomal profiles, and regulatory interactions, enabling users to perform hallmark enrichment, protein-protein and regulatory interaction-based analysis. To elucidate the interconnectedness of hallmarks, network separation and proximity analyses of hallmark-associated modules were performed, revealing functional overlap between hallmarks on the human interactome. Furthermore, the utility of AgingHallmarksDB was demonstrated through network topological and regulatory interaction analyses of common hallmark genes, hallmark enrichment and network proximity analyses of eight chronic age-related diseases and seven developmental and congenital diseases, and gene set enrichment analysis of a PM2.5-associated skin transcriptome. Together, these analyses highlight the utility of AgingHallmarksDB for aging hallmark-centred systems biology research. The resource is accessible at https://cb.imsc.res.in/aginghallmarksdb/.
bioinformatics2026-07-26v2Phenotype-driven de novo molecular design from gene expression signatures
Xu, Y.; Kuang, T.; Ge, S.; Wu, H.; Wang, M.; Xu, H.; An, F.; Ma, Z.; Cheng, Q.; Ren, Z.Abstract
Target-based and structure-guided drug design remain central to modern drug discovery, but complementary strategies are needed when predefined targets or binding pockets do not fully capture disease biology. Gene-expression signatures provide scalable system-level readouts of disease and perturbation states, making them attractive inputs for phenotype-guided molecular design. However, preserving phenotypic information during molecular generation remains challenging, and chemically plausible molecules may lose connection to the intended biological response. Here, we present Tx2Mol, a transcriptome-guided framework that translates gene-expression signatures into candidate molecules while maintaining biological guidance throughout generation. We evaluated Tx2Mol across three biological settings: bulk gene perturbation, single-cell perturbation, and patient-derived disease signatures; and three validation dimensions: chemical plausibility, structural compatibility, and phenotypic preservation. Across 10 cancer-relevant bulk gene-perturbation benchmarks, Tx2Mol outperformed 9 transcriptome-guided baselines, improving maximum Tanimoto similarity to known ligands by 24.10% on average and by 50.67% on HDAC1. Structure-based analyses further supported structurally novel candidates with favorable predicted target binding. Tx2Mol also generalized to noisy single-cell perturbation profiles and preserved drug-induced transcriptional responses through in silico drug-perturbation validation. Patient-derived disease signatures further guided molecular generation toward approved-drug chemical space. Together, these results support gene-expression phenotypes as actionable guidance signals for phenotype-directed molecular design and candidate prioritization.
bioinformatics2026-07-25v2Glycine molecule radical: Predicted properties and dipeptide formation
Synak, J.; Blazewicz, J.Abstract
The formation of the first peptide bonds under prebiotic conditions remains an open question. We used density functional theory calculations with the B3LYP functional to predict a novel pathway of peptide bond formation, which could have taken place without the sophisticated catalysts used by modern biological systems, utilising only radical chemistry. To make our investigation more extensive, the properties of intermediates (glycine-derived radicals) were thoroughly analysed, using the DFT, resonance hybrid model and orbital hybridisation. These methods shed more light on the exact nature of processes which should take place, explaining why this pathway could be favoured by the system. The result is a series of reactions, which without any sophisticated catalysts and with relatively low electronic energy barrier (<20 kcal/mol) can lead to formation of dipeptides, suggesting also a possible starting point for further peptide chain extension.
bioinformatics2026-07-25v2STAR Suite: an open-source single-executable transcriptomics engine for reproducible, AI agent-assisted processing
Hung, L.-H.; Baker, D.; Flynn, B.; Huangfu, D.; Luo, R.; Robson, P.; Zhou, T.; Yeung, K. Y.Abstract
Processing sequencing data means chaining many specialized tools, a barrier for bench biologists and for the AI agents increasingly used to run analyses. Open, integrated pipelines remain incomplete: the de facto single-cell standard, Cell Ranger, is proprietary - its license bars redistribution, modification, and non-10x use - so it cannot serve as a shareable, AI-discoverable layer, and no production-ready open-source pipeline exists for 10x Flex. Here, we present STAR Suite, which extends the STAR aligner into a single executable for transcriptomics processing.Adapter handling, feature-barcode assignment, Flex probe processing, SLAM-seq analysis, sorting, and quality control are integrated directly into the 28,228-line STAR codebase with no additional external dependencies, adding 132,226 lines of C/C++ made tractable through human-directed AI software engineering. Novel methodologies include dynamic thread interleaving, fast-Hamming feature matching, alignment-validated probe hashing, and variance-based auto-trimming (SLAM-seq); in-process integrations include Y-chromosome removal, per-sample variant masking, and Variational Bayes transcript quantification (bulk RNA-seq). STAR Suite reproduces Cell Ranger 9.0.1 to gene-level Pearson 0.99-1.0 while running 3.6- to 5.7-fold faster, and 3.9-fold faster than stepwise bulk pipelines. Used across the NIH MorPhiC consortium, its source code, workflow recipes (morphic-recipes), and immutable per-run provenance records (morphic-provenance) are released under the permissive open-source MIT license.
bioinformatics2026-07-24v5Phylogenetic detection of protein sites associated with continuous traits
Duchemin, L.; Muntane, G.; Boussau, B.; Veber, P.Abstract
Comparative genomic data can be used to look for substitutions in coding sequences that are associated with the variation of a particular phenotypic trait. A few statistical methods have been proposed to do so for phenotypes represented by discrete values. For continuous traits, no such statistical approach has been proposed, and researchers have resorted to sensible but uncharacterized criteria. Here, we investigate a phylogenetic model for coding sequences where amino acid preferences at a site are given by a continuous function of a quantitative trait. This function is inferred from the amino acids and the trait values in extant species and requires inferred point estimates of ancestral values of the trait at internal nodes. For detecting sites whose evolution is associated with this trait, we use a significance test against the hypothesis that amino acid preference does not depend on the trait. This procedure is compared to simpler strategies on simulated alignments. It displays an increased recall for low false positive rates, which is of special importance for performing whole-genome scans. This comes however at a much higher computational cost, and we suggest using a simple test to filter promising candidate sites. We then revisit a dataset of alignments for 62 species of mammals, using longevity as a phenotypic trait. We apply our method to three protein families that have previously been proposed to display sites associated with variation in lifespan in mammals. Using a graphical representation extracted from the detailed phylogenetic analysis of candidate sites, we suggest that the evidence for this in the sequence data alone is weak. The proposed method has been added to our Pelican software. It is available at https://gitlab.in2p3.fr/phoogle/pelican and can now be used with both discrete and continuous phenotypes to search for sites associated with phenotypic variation, on data sets with thousands of alignments.
bioinformatics2026-07-24v5ViroSeek: a viral detection pipeline for second-generation sequencing
Berger, A.; Lefebvre, M. J. M.; Dainat, J.; Jiolle, D.; Conclois, I.; Talignani, L.; Mastriani, E.; Cornelie, S.; Berthet, N.; Paupy, C.Abstract
Arbovirus emergences represent a rising public health issue and are exacerbated by climate change and globalization. Virome analysis has become a key approach for monitoring and managing infectious diseases, yet existing tools often remain technically complex and inaccessible to non-specialists. In this context, we present ViroSeek, a reproducible and accessible bioinformatics pipeline specifically designed for the taxonomic analysis of second-generation sequencing data from target-enriched libraries. ViroSeek performs a series of automated steps: quality control, trimming, host and bacterial sequence removal, assembly, taxonomic assignment, read remapping for quantification, and PCR duplicate removal. The whole process is designed to produce a clear, usable viral taxonomy table that is suitable for diversity studies. ViroSeek was empirically validated on enriched control samples containing a known panel of viruses. All the expected viruses were correctly detected. Bacterial and host contaminant sequences were effectively removed. The pipeline is freely available and fully documented, supporting its adoption and adaptation by the research community.
bioinformatics2026-07-24v2Leveraging Uncertainty Estimates for Drug Response Prediction in Cancer Cell Lines
Iversen, P.; Renard, B. Y.; Baum, K.Abstract
Machine learning models for drug response prediction in cancer cell lines carry the potential to advance precision oncology by tailoring treatments to the molecular tumor profile. Their application is challenged by variability in prediction quality and distribution shifts between training and application. Uncertainty estimation provides more information on the predictive distribution than point estimates, enabling comprehensive decision support and downstream analysis of predictions. Yet, the most effective uncertainty estimator is domain-specific. In this work, we benchmark uncertainty-aware models for drug response prediction. We focus on epistemic uncertainty via ensemble agreement, and aleatoric uncertainty via distributional modeling, or both. We find that ensemble-based estimates are more sensitive to distribution shift and can flag out-of-distribution examples. In contrast, distributional models yield stronger prediction error reductions among high-confidence subsets. Despite higher computational cost, the combination can provide both advantages: an ensemble of neural networks that estimate a Gaussian predictive distribution can reduce the mean squared error by 64 percent when restricting predictions to the 10 percent most confident drug-cell line pairs, and reliably indicates distribution shifts and platform differences. Beyond benchmarking, probabilistic predictions can identify drugs whose uncertainty bounds overlap with therapeutically relevant ranges. We also show that uncertainty estimates enable a new dimension of model interpretability: by attributing predicted uncertainty to input features, we identify genes that signal unpredictability of drug response rather than sensitivity or resistance. We further demonstrate uncertainty-guided selection of measurements for active learning. In summary, including uncertainty in drug response prediction supports better-informed model application. The code is available at https://github.com/PascalIversen/LUDRP.
bioinformatics2026-07-24v2DOME Copilot: A resource to automate transparent reporting of artificial intelligence methods
Farrell, G.; Attafi, O. A.; Fragkouli, S.-C.; Heredia, I.; Fernandez Tobias, S.; Harrison, M.; Hermjakob, H.; Jeffryes, M.; Mehdiabadi, M.; Obregon Ruiz, M.; Pearce, M.; Pechlivanis, N.; Quaglia, F.; Lopez Garcia, A.; Psomopoulos, F.; Tosatto, S. C. E.Abstract
Artificial intelligence (AI) methods are transforming life science research and witnessing unprecedented adoption across literature. While AI is driving impactful discoveries, this rapid growth in application has inundated researchers with poorly described models and datasets which lack transparency, impeding reusability and eroding trust. The DOME Recommendations aimed to address this issue and established structured reporting guidelines to standardize AI methodology descriptions. However, the need to manually create these comprehensive disclosures imposed a major bottleneck, requiring substantial effort from authors to comply. To bridge this gap, we present DOME Copilot, an open-source, easily deployable system that automates the generation of transparent AI methodology reports supplementing publications. Evaluated against a benchmark of human-created annotations, DOME Copilot was determined to match or exceed manual annotation quality across the majority of reporting fields while reducing the creation time from hours to minutes. Ultimately, the system provides a novel and scalable solution to restore confidence in complex life science AI method publications.
bioinformatics2026-07-24v2Revisiting Logistic Regression for High-Dimensional Gene Expression Data
Souza, R. d. O.; Rodrigues, W. F.; Couto, B.; Dos Santos, M. A.Abstract
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.
bioinformatics2026-07-24v1FusedFCR: A Fused Forward Continuation-Ratio model for marker selection along cell-fate trajectories
Mattila, C.; Chakraborty, A.; Angel, P.; Cao, S.; Sonawane, K.; Hill, E.; Chung, D.; Neelon, B.; Seal, S.Abstract
Time-course single-cell RNA sequencing (scRNA-seq) data collected across ordered stages provide population-level snapshots of differentiation, disease progression, and aging. Supervised pseudotime methods use observed stage labels to reconstruct continuous progression but generally do not identify marker genes associated with changes from one stage to the next. Unsupervised pseudotime-based marker selection methods infer latent trajectories directly from expression data and identify trajectory-associated genes, but do not explicitly link these associations to the observed stages. We propose FusedFCR, a regularized forward continuation-ratio model that represents cellular progression through a sequence of conditional transitions across ordered stages. FusedFCR combines a lasso penalty for gene selection with a fusion penalty that encourages similar effects across adjacent transitions while allowing transient and direction-changing associations. The resulting transition-specific coefficients support interpretable gene selection and a continuous pseudotime-like projection anchored to the observed developmental stages. In simulations, FusedFCR accurately recovered gene-effect trajectories and improved predictive performance relative to alternative methods. Applied to mouse pancreatic beta-cell differentiation across seven time points and human extravillous trophoblast differentiation across four time points, FusedFCR identified biologically interpretable genes associated with distinct developmental transitions. Gene set enrichment analysis further revealed stage-specific pathway activity consistent with known developmental biology, while held-out stage-classification accuracy was competitive or superior across both datasets. Together, these results show that FusedFCR complements pseudotemporal ordering by identifying which molecular programs change and when those changes emerge along the developmental trajectory. An accompanying R package is available on GitHub.
bioinformatics2026-07-24v1Graph in Graph (GiG): A novel graph AI framework for integrating and interpreting whole-person healthmedical and omics data
Zhang, H.; Lu, Y.; Fang, K.; XU, Z.; Akbary Moghaddam, V.; An, P.; Jin, S.; Wojczynski, M.; Province, M.; Li, F.Abstract
Medical records and omics data are rapidly becoming standard in healthcare settings, which characterize the whole-person from dysfunctional molecules to phenotypes, and thus offer potential for precise disease diagnosis and target discovery. Whereas, it remains an open problem to systematically integrate and interprete medical record and omics data of individual patients. In this study, for the first time, we propose a novel graph AI model framework, Graph in Graph (GiG), to integrate and interpret the whole-person medical and omic datasets. Specifically, the medical record data is modeled using a person-phenotype graph, followed by omics signaling graphs of invidival patients, which enables the integration of information learned from omic-signaling graph and medical phenotype features to characterize individual patients and to prioritize important omic biomarkers and phenotypes. As an exploratory study, we applied and evaluated the GiG model to study the type 2 diabetes (T2D) and pre-T2D vs healthy using the Long Life Family Study (LLFS) cohort, which enrolls families with exceptional longevity to uncover biological mechanisms of healthy aging with medical and omics data. The evaluation results showed that GiG not only achieve a high prediction but also can interpret the prediction by ranking the essential clinical and omic biomarkers. The GiG framework can be applied to other studies by effectively integrating and interpreting medical and omic datasets for disease diagnosis and pathogenesis discovery.
bioinformatics2026-07-24v1Structural bioinformatics of three Epstein-Barr Virus (EBV) Integral Membrane Proteins and their water-soluble QTY analogs
Zhang, S.; Sun, Z.; Chen, E.Abstract
The Epstein-Barr virus (EBV) is a highly prevalent virus worldwide that is associated with several lymphoid and epithelial malignancies. However, extensive research on EBV integral membrane proteins BILF1, LMP1 and LMP2, has been scarce due to their hydrophobic transmembrane domains. Our study applies the QTY code (glutamine, threonine, tyrosine) to design water-soluble analogs of BILF1, LMP1 and LMP2 with reduced hydrophobicity, where we systematically replaced hydrophobic amino acid residues leucine (L), isoleucine (I), valine (V), and phenylalanine (F) with structurally similar polar residues glutamine (Q), threonine (T), and tyrosine (Y). We retrieved their native sequences from UniProt, identified transmembrane domains using Protter, then performed QTY design through the Protein Solubilizing Server (PSS). We then predicted native and QTY structures using in silico prediction tools AlphaFold3, ColabFold, and Boltz-2. Our analyses demonstrate that despite significant protein sequence replacements in their transmembrane domains (54.15%-61.59%) and increased intrinsic solubility, the QTY analogs exhibited minimal changes in isoelectric point (0.00-0.15 decrease) and molecular weight (0.7-1.2 kDa increase). Additionally, structural superpositions between QTY analogs and native structures using PyMOL yield low RMSD values (0.217[A] -1.202[A]). Our results demonstrate the QTY codes ability to design detergent-free analogs of BILF1, LMP1 and LMP2 with substantially reduced hydrophobicity and aggregation propensity whilst preserving native-like structures. Our results may facilitate protein characterization studies, therapeutic research on EBV, and other protocols that typically require protein solubilization.
bioinformatics2026-07-24v1CNVeil resolves haplotype-specific copy number and uncovers subclonal architecture hidden from total copy number profiling in single-cell cancergenomes
Yuan, W.; Luo, C.; Hu, Y.; Zhang, L.; Wen, Z.; Liu, Y. H.; Fan, X. M.; Zhou, X. M.Abstract
Single-cell DNA sequencing (scDNA-seq) resolves copy number variation (CNV) at single-cell resolution, revealing tumor heterogeneity and subclonal structure. Most existing methods, however, infer only total copy number. Haplotype-resolved copy number, which captures allelic imbalance and clonal evolution, remains far less developed, largely because low coverage, allelic dropout, and technical noise in scDNA-seq make phased allelic inference substantially harder than total copy number estimation. We present CNVeil, a haplotype-aware framework that infers total, allele-specific, and chromosome-scale haplotype-resolved copy number from scDNA-seq data. CNVeil first builds robust total copy number profiles through highly variable bin selection, hierarchical clustering, subclone-aware ploidy estimation, and cross-cell consensus segmentation. Using this profile as a stable scaffold, it infers allele-specific copy number with an expectation-maximization algorithm applied to heterozygous SNP allele counts, then reconstructs haplotype-specific copy number by enforcing coherent haplotype orientation across adjacent segments via dynamic programming. We benchmarked CNVeil against 12 state-of-the-art methods, including eight total copy number callers, two allele-specific callers, and two haplotype-resolved callers, across 20 simulated and real datasets spanning six experimental settings, including high-multiplexed single-nucleus sequencing, Acoustic Cell Tagmentation (ACT), and 10x Chromium. This constitutes the largest comparative evaluation of single-cell copy number inference methods to date. CNVeil consistently outperformed existing tools in segmentation accuracy, ploidy inference, subclone identification, and allele-specific copy number estimation. In a breast cancer multi-omics (wellDR-seq) cohort, CNVeil uncovered haplotype-specific subclonal diversification invisible to total copy number analysis alone and linked allele-specific copy number states to transcriptional variation. By transforming sparse single-cell allelic signals into chromosome-scale haplotype-resolved profiles, CNVeil closes a major methodological gap and provides a scalable framework for studying tumor evolution and functional genomic heterogeneity at single-cell resolution.
bioinformatics2026-07-24v1Deconvolution-derived cell-type expression targets for personal genome sequence-to-expression prediction
Sim, S.; Shen, L.Abstract
Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue--cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122--0.142 for sequence-derived approaches and 0.081--0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution--pseudobulk agreement for sequence-derived models ($r=0.35$--$0.43$ across tissue--cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially outperform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.
bioinformatics2026-07-24v1VizR: An Interactive Web Platform for End-to-End RNA-Seq Analysis and Visualization in Plant Biology
Jeon, W.-T.; Jung, H.; Shim, D.; Lee, Y.Abstract
RNA sequencing (RNA-seq) is widely used to investigate transcriptional programs in plant biology, yet the need to combine multiple specialized tools and bioinformatics expertise to convert raw sequencing reads into biologically interpretable results remains a major technical barrier for many plant biologists. Here, we present VizR (VIsualiZation of Rna seq), a web-based platform that integrates end-to-end RNA-seq analysis and visualization within a single integrated environment. VizR automates upstream processing, including quality control, adapter trimming, genome alignment, and transcript quantification, and connects the resulting expression data to downstream exploratory analyses. Its interface is designed to make expression patterns immediately searchable and interpretable: users can query genes through an equalizer-style expression-pattern interface, inspect expression profiles using inline heatmaps embedded in gene tables, and perform context-integrated gene ontology analysis throughout the workflow. VizR also supports comparative analysis through interactive Venn diagram module, allowing users to transfer gene sets directly from result tables. As a Docker-based application, VizR can be deployed locally and accessed through a standard web browser. By unifying automated RNA-seq processing, interactive visualization, and functional interpretation, VizR lowers the technical barrier to transcriptome analysis and provides a practical platform for plant biology research.
bioinformatics2026-07-24v1GraTools, an user-friendly tool for exploring and manipulating pangenome variation graphs
Ravel, S.; Marthe, N.; Carrette, C.; Mohamed, M.; Sabot, F.; Tranchant-Dubreuil, C.Abstract
Background: Pangenome variation graphs (PVGs), which represent genomic diversity through multiple genomes alignment, are powerful tools for studying genomic variations in populations. However, current tools often lack integration, efficiency, or require format conversions, to use them, hindering their usability. Results: Here, we introduce GraTools, a fast and user-friendly command-line tool for manipulating PVGs using the original GFA file. After a one-time graph import, GraTools enables rapid subgraph extraction, fasta sequence retrieval, and comprehensive analyses, including core/dispensable genome ratio calculation or group-specific segment identification. The import step results in conversion in standard data formats (BAM/BED), enabling the reuse of well-optimized existing tools, allowing an efficient storage and the querying of the PVGs large complex data structures. Scalability is ensured by a modular architecture sup- porting parallel processing and asynchronous I/O operations. GraTools supports coordinates defined on both the primary reference as well as from alternative genomes within the graph without re-importing, and its outputs can be easily visualized or manipulated using external tools. Using an Asian rice pangenome graph (13 accessions), we demonstrate its ability to easily extract subgraphs, compute depth statistics, and identify subspecies-specific segments. An intuitive command-line interface, a real-time execution feedback and a detailed logging system make this tool suitable for a wide range of applications, from population genetics to breeding and genomic medicine, for both biologists and bioinformati- cians. Conclusions: Through its unified graph manipulation interface, GraTools offers an interesting alternative to the few existing tools for manipulating PVGs, facil- itating rapid, efficient and flexible downstream analyses. It is available as an open-source tool (GNU GPLv3), with its documentation available at https: //gratools.readthedocs.io.
bioinformatics2026-07-23v3Fast and accurate taxonomic domain assignment of short metagenomic reads using BBERT
Alekhin, D.; Alon, M.; Sidi, T.; Perez Mazeh, S.; Carmi, G.; Finkel, O. M.; Erez, A.Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses.
bioinformatics2026-07-23v3Hobrac: a reference-guided workflow for genome comparison and synteny visualization
Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.Abstract
Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO rthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.
bioinformatics2026-07-23v2Statistical tests for bivariate spatial association across multi-omics data with disjoint coordinates
Hawinkel, S.; Hu, W.; Velten, B.; Maere, S.Abstract
Spatial biology has entered a new era of multimodal profiling, with multiple, high-dimensional spatial omics types being measured on consecutive tissue slices, or co-assayed on the same slice. Interest then lies in statistical testing for spatial association between the features of the different modalities, to gain insight in biological processes. One major challenge is the multitude of bivariate combinations, leading to high computational demands. Another difficulty is the difference in spatial resolution between technologies, implying no one-to-one matching between the measurement spots of the different modalities, even after alignment. As a result, common statistical measures such as joint distributions and correlations are not defined, and tests need to rely on spatial vicinity only. Moreover, we argue that many existing bivariate association tests address an inappropriate null hypothesis, or make inappropriate assumptions, both implying absence of spatial autocorrelation in any of the features and leading to misleading conclusions. As a remedy, we modify tests for the detection of spatially variable genes (Moran's I, Gaussian processes and generalized additive models) to derive bivariate spatial association tests across modalities with non-overlapping coordinate sets, and provide variance estimators that do account for spatial autocorrelation. We develop inference methods for single sections as well as for replicated experiments with multiple sections, and compare their performance in nonparametric and parametric simulations. Finally, we apply the newly developed methods to two co-assayed spatial transcriptomics and metabolomics datasets from mouse and human. The full suite of tests is available from github.com/sthawinke/sbivar as the R-package sbivar.
bioinformatics2026-07-23v2GenoSim: A Forward-Time Genotype Simulator for Clinical and Population Genetics with Population Stratification
Bakar, A.; Gul, R.; Haq, W. u.; Afghani, T.Abstract
Motivation: Next-generation sequencing studies in clinical genetics are often limited by the scarcity of human genotype data, which stems from ethical, regulatory, and economic barriers. The shortfall is sharpest in consanguineous populations, which are common in South Asia and the Middle East, where family-based designs need large pedigrees that are rarely sequenced in full. Existing simulators do not combine pedigree-aware propagation, realistic population stratification, and clinical export formats in one tool. Results: We present GenoSim, an R package for forward-time simulation of diploid SNP genotypes. It runs in two modes: a population mode implementing inbreeding-adjusted Hardy-Weinberg sampling, Wright-Fisher drift, directional selection, recurrent mutation, and Haldane recombination across multiple generations; and a pedigree-constrained mode that ingests real family VCFs and a pedigree, reconstructs phase where the pedigree makes it identifiable, propagates genotypes through the observed family structure, and appends synthetic generations. Version 1.1.1 adds population stratification through the Balding-Nichols model parameterised by gnomAD v3.1 fixation indices (F_ST) for eight ancestry groups (AFR, AMR, EAS, EUR, FIN, MID, SAS, ASJ), empirical allele-frequency loading from external reference panels, and admixed-cohort simulation. Analysis functions cover Hardy-Weinberg testing, linkage disequilibrium, runs of homozygosity, principal component analysis, founder-referenced and between-generation F-statistics, and Nei gene diversity. Availability and implementation: GenoSim is available as an R package at https://github.com/malikbak/GenoSim under the MIT licence. It requires R [≥] 4.0.0 and depends only on base R packages (stats, utils, graphics, grDevices, tools).
bioinformatics2026-07-23v2GeneKnow: An auditable AI framework for source-grounded biological evidence synthesis
Zhang, H.; Sittipongpittaya, B.; Zang, C.Abstract
Biological function is context dependent, yet synthesizing evidence for a gene's function in a defined cell type, disease, or perturbation remains labor-intensive. Here we present GeneKnow, an auditable artificial intelligence (AI) framework for source-grounded biological evidence synthesis. GeneKnow separates evidence-critical operations, including literature retrieval, passage selection, provenance tracking, and bibliography construction, from generative steps that are constrained to semantic analysis and literature synthesis. GeneKnow supports multi-paper discovery and single-paper inspection while preserving links from synthesized claims to source passages, generating trustworthy syntheses without fabricated citations and minimizing hallucinations. Systematic benchmarking showed that GeneKnow achieved higher claim support and citation fidelity than leading general-purpose and scientific AI systems. These results demonstrate that an AI system with controlled division of labor between deterministic and generative components can substantially improve the fidelity and auditability of biomedical literature synthesis.
bioinformatics2026-07-23v2EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-07-23v2A pan-cancer analysis of microRNA tissue specificity and its association with dysregulation
Poptsova, M.; Ismailov, A.; Belogurov, A.; Evpak, A.Abstract
MicroRNAs are frequently dysregulated in cancer, yet how their tissue-specificity is remodeled during malignant transformation remains poorly characterized. Here we systematically quantified the tissue-specificity of miRNAs across normal (GTEx) and tumor (TCGA) tissues using the Tau index, and compared its distribution between healthy and cancerous states. To robustly define dysregulation, we combined two independent analyses: a binomial test over per-project differential expression across 17 matched normal tissues within TCGA cohort, and a TCGA-GTEx pan-tissue comparison of mean expression. The change in specificity ({Delta}Tau) separated up- from down-regulated miRNAs, showing moderate agreement with the binomial signal and a strong correlation with the expression-based contrast. Finally, we identified 6 miRNAs that lose tissue-specificity upon transformation while remaining consistently upregulated (miR-519a-5p, miR-512-3p, miR-522-3p, miR-105-5p, miR-935, miR-1269a). Functional analysis of experimentally validated targets showed significant enrichment for converging on core oncogenic programs for miR-512-3p, miR-105-5p and miR-935, such as apoptosis and cellular-stress regulation, TP53, FoxO, PI3K-Akt/mTOR signaling, immune modulation. Collectively, integrating specificity dynamics with dysregulation evidence pinpoints candidate miRNAs with coordinated, cancer-relevant regulatory roles and highlights those with favorable tissue specificity profiles for therapeutic targeting.
bioinformatics2026-07-23v2A transparent multicriteria and fuzzy classification approach for genome-based probiotic candidate prioritisation
Ounissi, N. E.; Gomri, M. A.; El Hadef El Okki, M.Abstract
The identification of novel probiotic candidates with potential health-promoting properties remains a major challenge in food biotechnology and increasingly relies on in silico screening of genomic information. However, probiogenomic markers are heterogeneous, and safety, survival-colonisation, and functional-benefit traits do not contribute equally to probiotic potential. This study developed the Structured Probiotic Potential Index (SPPI), a fuzzy multicriteria system for genome-based probiotic candidate prioritisation. A hierarchical evaluation structure was established from probiogenomic evidence and organised into three main pillars and fourteen subcriteria. Expert judgements were collected using the Analytic Hierarchy Process, followed by consistency-based curation and weight aggregation. A curated dataset of 48 complete bacterial genomes, distributed into probiotic, potentially probiotic, neutral, and pathogenic groups, was taxonomically validated and analysed using an automated probiogenomic screening pipeline. Genome-wide screening generated 3,218 binary genomic features, from which curated probiogenomic markers were mapped to the scoring hierarchy. The resulting index was formulated as a normalised expert-weighted equation integrating safety, survival-colonisation, and functional-benefit components. SPPI prioritised genomes according to weighted probiogenomic profiles and separated pathogenic genomes from favourable probiotic and potentially probiotic profiles within the analysed dataset. Downstream Fuzzy Comprehensive Evaluation transformed the continuous score into probiotic/potentially probiotic, neutral, and pathogenic classes. Under internal leave-one-out evaluation, all 48 genomes were assigned to their expected reference classes, while confidence analysis distinguished high-confidence from borderline assignments. Post-classification comparison with ProbML showed concordant behaviour for most genomes and discordant predictions for selected cases. These results support SPPI as a transparent genome-based decision-support system for early probiotic candidate prioritisation before experimental validation.
bioinformatics2026-07-23v1PhytoFam: A Nextflow Pipeline for Genome-Wide Analysis of Plant Gene Families
Parajuli, S.; Adhikari, B.; Fennell, A.; Nepal, M. P.Abstract
Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complicates reproducibility and parameter tracking, and may require additional steps to remove redundant protein isoforms. To address these challenges, we developed PhytoFam, a Nextflow-based workflow that automates gene family identification from proteome input through phylogenetic reconstruction. The pipeline integrates HMMER for candidate sequence identification, isoform-aware deduplication, InterProScan for domain confirmation, BLAST reciprocal best hit (RBH) analysis for orthology assignment, MUSCLE for multiple sequence alignment with optional outgroup incorporation, TrimAl for alignment trimming, and IQ-TREE3 for phylogenetic reconstruction. PhytoFam is portable across local workstations and high-performance computing environments and supports deployment through Conda, Docker, and Singularity. We validated the workflow using the Morus alba MADS-box gene family, where the complete analysis finished in 1 h 10 min (9 CPU h). IQ-TREE3 accounted for most of the execution time, whereas InterProScan showed the highest memory requirement with a peak resident set size of 4.5 GB. PhytoFam provides a reproducible, automated, and scalable solution for plant gene family identification and phylogenetic analysis. The pipeline is freely available at https://github.com/sanamparajuli/PhytoFam.
bioinformatics2026-07-23v1An openly licensed benchmark and per-gene calibration map for missense pathogenicity predictors on activating cancer drivers
Lee, S.-G.Abstract
Missense pathogenicity predictors such as AlphaMissense are increasingly used in clinical variant interpretation, yet they are trained on germline labels dominated by loss-of-function (LOF) variants. Using an openly licensed, reproducible benchmark of 768 Cancer Gene Census genes scored with 49 predictors (labels from CIViC, COSMIC, cancerhotspots, ClinVar and gnomAD), we show that 42 of 49 tools (86%) score oncogene, gain-of-function (GOF) variants worse than tumour-suppressor variants. This under-scoring is mechanistically characterized: missed drivers occupy low-conservation, solvent-exposed, non-destabilizing positions (phyloP 2.51 versus 7.89; relative solvent accessibility 0.671 versus 0.185; gene-clustered p = 4.8x10-20 and 2.3x10-35), and, counter-intuitively, the unsupervised and protein-language models now entering clinical use are the most affected. Per-gene oncogenic thresholds span 0.07-0.99, so a single global cut-off is mis-calibrated for most genes; we provide a per-gene calibration map. A cancer-calibrated stack (OncoCal) modestly improves discrimination over the best single tool (AUROC {approx} 0.93 versus 0.87), rescues drivers such as JAK2 V617F (0.334[->]0.57), and generalizes to independent deep mutational scanning data. We provide an openly licensed framework to interpret and recalibrate these tools in the somatic setting rather than a replacement predictor.
bioinformatics2026-07-23v1pHaseMD4AI: Phase-Space Dynamics Dataset with Chemical and pH Perturbations for Physically and Kinetically Consistent Biomolecular AI
Song, T.; Guo, Y.; He, J.; Liu, Z.; Low, M.; Wang, K.; Zhang, Y.; Li, Z.; Huang, Y.; Wang, Y.Abstract
Protein function emerges from dynamic conformational ensembles and transitions that are challenging to characterize experimentally and computationally. Recent advances in generative AI have created new opportunities for learning molecular thermodynamics, kinetics, and conformational evolution directly from simulation data, but progress is limited by the availability of large-scale datasets that combine rigorous sampling, complete phase-space information, and diverse physicochemical perturbations. Here, we present pHaseMD4AI, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins. The peptide branch includes a complete set of canonical tripeptide and tetrapeptide systems together with post-translationally modified (PTM) and protonation-state datasets, providing synchronized atomic coordinates (R), velocities (V), forces (F), and Markov state model-based kinetic annotations. An accompanying web portal (https://isb.zju.edu.cn/md4ai/) enables users to browse, visualize, and download trajectories, annotations, and metadata. As an example application, we demonstrate a sequence-based model that can predict residue-level equilibrium dihedral distributions from sequence. pHaseMD4AI provides a resource for developing and benchmarking molecular machine learning methods while supporting broader studies of biomolecular dynamics under sequence, post-translational modification, and protonation-state perturbations.
bioinformatics2026-07-23v1scSAID: A Comprehensive Cross-Species Single-Cell Skin Atlas Reveals Species-Specific Responses to Psoriasis
Ren, Y.; Shen, Y.; Jin, L.; Huang, Y.; Deng, Y.; Xiao, Y.; Wang, C.Abstract
Single-cell RNA sequencing has rapidly expanded the scale of skin transcriptomic data, yet these datasets remain fragmented across studies spanning different species, diseases and experimental manipulations. An up-to-date, comprehensive, and queryable single-cell cross-species repository for skin is still lacking. Here, we present scSAID (skin-scsaid.com), a single-cell database with an interactive web portal offering a broad suite of in-depth analyses for human and mouse skin. It integrates more than 1.2 million high-quality cells collected from 252 samples, establishing a unified reference for cell-type annotation, cross-species comparison and pathological studies. Using psoriasis as a case study, we demonstrate how scSAID can be used to evaluate how faithfully mouse models reproduce human pathology. Systematic comparison with the imiquimod-induced mouse model revealed numerous species-specific molecular signatures of psoriasis, including human-specific NFKB1 activation and STAT1 involvement, indicating that the current mouse model captures only limited aspects of the disease. We further introduce psoSpotter, a disease-biomarker-selection algorithm that, coupled with in silico perturbation using scSAID data, uncovers PPIA as a novel psoriasis drug target, illustrating the potential of scSAID for identifying therapeutic approaches. Overall, scSAID delivers a large-scale, cross-species skin single-cell resource and analysis platform, opening new opportunities for the discovery of disease-relevant targets in skin diseases.
bioinformatics2026-07-23v1Integrated Bioinformatics Analysis of PALB2 Reveals Expression Patterns, Molecular Interactions, and Prognostic Significance in Breast Cancer
Bithi, A. J.; Rahat, M. H.Abstract
Background: Partner and Localizer of BRCA2 (PALB2) is a key tumor suppressor gene involved in homologous recombination mediated DNA repair through its interactions with BRCA1 and BRCA2. Germline alterations in PALB2 have been associated with hereditary breast cancer risk; however, its broader molecular role in breast cancer progression and prognosis requires further investigation. Methods: A comprehensive in silico analysis of PALB2 was performed using publicly available databases and bioinformatics platforms. Differential expression of PALB2 in breast cancer were evaluated using GEPIA2. Prognostic significance was assessed through Kaplan Meier analyses for overall survival (OS) and disease free survival (DFS). Protein protein interaction (PPI) networks were constructed using STRING. Functional enrichment analyses of PALB2-associated genes were conducted using g. Mutational profiling of PALB2 in breast cancer was performed using cBioPortal with data from TCGA breast cancer cohorts. Results: PALB2 expression was elevated in breast tumor tissues compared with normal breast tissues. Survival analyses revealed no statistically significant association between PALB2 expression and either overall survival (HR = 0.88, p = 0.44) or disease free survival (HR = 0.74, p = 0.11). Protein interaction analysis revealed strong interactions between PALB2 and major DNA repair proteins including BRCA1, BRCA2, RAD51, RAD51C, FANCD2, and BRIP1. Functional enrichment analysis showed limited significant pathway enrichment, with only marginal transcription factor motif enrichment observed. Mutational analysis demonstrated diverse genomic alterations including missense mutations, truncating mutations, copy number gains, and shallow deletions. Conclusion: The findings support the biological relevance of PALB2 in breast cancer through its elevated expression and strong connectivity within DNA repair pathways. However, PALB2 expression alone does not appear to serve as an independent prognostic indicator. Further studies integrating genomic, transcriptomic, and clinical parameters are required to clarify its role in breast cancer progression and therapeutic response. Keywords: PALB2; Breast Cancer; Bioinformatics; Gene Expression Analysis; Protein Protein Interaction; Survival Analysis; Mutation Profiling; Homologous Recombination; TCGA; GEPIA2
bioinformatics2026-07-23v1Transcriptional landscape of direct reprogramming toward the hematopoietic lineage
Cwycyshyn, J.; Stansbury, C.; Golts, S.; Lee, H.; Pickard, J.; Meixner, W.; Rajapakse, I.; Muir, L. A.Abstract
Direct reprogramming of human fibroblasts into hematopoietic stem cells (HSCs) offers a promising strategy for generating autologous cells to treat blood and immune disorders. Current protocols are limited by low efficiency and insufficient tools for evaluating reprogramming outcomes. Although functional assays are the standard for confirming cell identity, they require fully reprogrammed cells, limiting their utility during protocol development. To address this, we assembled a single-cell transcriptomic reference atlas of hematopoietic reprogramming and tested an algorithmically-predicted transcription factor recipe for HSC induction. Long-read single-cell RNA sequencing of CD34+ reprogrammed cells revealed progressive loss of fibroblast identity alongside induction of early hematopoietic and endothelial programs, with reference-atlas benchmarking placing reprogrammed cells in an intermediate transcriptomic state between fibroblasts, endothelial cells, and HSCs. Isoform-level analysis further revealed transcriptional remodeling not captured by gene-level analyses. This experimental-computational framework offers a generalizable strategy for characterizing partially reprogrammed states and guiding optimization of reprogramming protocols.
bioinformatics2026-07-22v5Benchmarking the Impact of Data Leakage on the Performance of Knowledge Graph Embedding Models for Biomedical Link Prediction
BRIERE, G.; STOSSKOPF, T.; LOIRE, B.; BAUDOT, A.Abstract
Motivation: Knowledge Graphs (KGs) are increasingly used to organize complex biomedical knowledge into structured representations of entities and relations. Knowledge Graph Embedding (KGE) models facilitate efficient exploration of KGs by learning compact representations, and are widely applied to biomedical link prediction, for instance to uncover new therapeutic uses for existing drugs. Despite extensive work on KGE models, current evaluations often overlook data leakage, which can artificially inflate performance and undermine benchmark validity. Data leakage can arise when (1) there is redundancy between training and test sets, (2) the model leverages illegitimate features, or (3) the test set does not reflect real-world inference scenarios. Results: We assess the impact of data leakage on KGE-based link prediction across three biomedical KGs, using both decoder-only and GNN-based models. We first demonstrate the impact of train-test redundancies and implement a systematic procedure to detect and remove them. Using permutation experiments, we investigate whether node degree acts as an illegitimate predictive feature, and find no evidence that predictions are driven by degree alone. Finally, we evaluate how well common test set sampling strategies reflect real-world inference in drug repurposing. Comparing random and cold-start splits with an independent Orphanet-derived test set, we observe a substantial performance drop on the latter, indicating that current practices may overestimate how well KGE models generalize. Overall, our findings highlight the importance of rigorous benchmark design and careful evaluation of the generalization ability of KGE models for biomedical link prediction. Availability and Implementation: All code and results are openly available on GitHub at https://github.com/galadrielbriere/data_leakage_kge_benchmark.git.
bioinformatics2026-07-22v3Graph Lens Lite: A browser-based tool for interactive visualization and exploration of biological networks
Ley, M.; Keska-Izworska, K.; Fillinger, L.; Walter, S. M.; Baumgärtel, F.; Bono, E.; Galou, L.; Andorfer, P.; Hauser, P.; Leierer, J.; Kratochwill, K.; Perco, P.Abstract
Biological network visualization together with graph-based analyses are key techniques in systems biology and network medicine to detect patterns and generate hypotheses regarding disease pathobiology, drug target identification, biomarker prioritization, and digital drug discovery. Network representations provide an intuitive way to communicate and share research findings. We have developed Graph Lens Lite, a browser-based tool that combines rich visualization with a streamlined interface for exploring and sharing biological networks. It offers an expressive query language, topological network analysis, interactive filtering, visual grouping, customizable layouts, a data editor, fine-grained property-based styling, animated edge-flow visualization, community detection, and a context-aware locally powered AI assistant particularly suited for exploring molecular models of disease pathobiology or drug mechanism of action. We demonstrate its utility on a curated network model of autosomal dominant polycystic kidney disease. Graph Lens Lite is open source, with a live web version available at https://delta4ai.github.io/GraphLensLite/.
bioinformatics2026-07-22v3