Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Adding layers of information to scRNA-seq data using pre-trained language models
Krissmer, S. M.; Menger, J.; Rollin, J.; Vogel, T. M.; Binder, H.; Hackenberg, M.Abstract
Pre-trained language models promise to enrich single-cell analyses with contextual information from large biomedical text corpora, but it remains unclear how to optimally align this knowledge with quantitative scRNA-seq data. To address this, we construct text-based training datasets from both scRNA-seq data and biomedical literature targeted to the experimental setting at hand. We then fine-tune lightweight encoder-only biomedical language models to learn a shared, literature-enriched representation. Controlled evaluations across immune and developmental datasets show that this representation preserves cell identity while adding robust and interpretable contextual layers of functional, disease-associated, and developmental information to single-cell analysis workflows.
bioinformatics2026-09-04v3Hierarchical Breakdown of RNA Structure Prediction in CASP16: From Reliable Local Helices to Speculative Multimer Assembly
Nithin, C.; Pilla, S. P.; Kmiecik, S.Abstract
CASP16 provided a community-wide benchmark for assessing RNA structure prediction, including the first large-scale blind assessment of RNA-RNA multimer prediction. CASP16 results showed that accurate three-dimensional modeling, especially for RNA-RNA multimers, remains a major challenge across the field. In this work, we use the submissions of our group (LCBio) as a diagnostic case study to examine the current limits of RNA structure prediction. In the official CASP16 best-of-submitted-models analysis, our workflow ranked first in the RNA-RNA multimer category and remained competitive for monomers. This makes the submitted model set useful for examining why high-ranking multimer predictions can still deviate substantially from experimental structures. We combine hierarchical analysis with representative case studies to connect this field-wide limitation to specific structural failure modes, showing that prediction accuracy decreases from relatively reliable canonical base-pairing and local helical organization to less reliable non-canonical interactions, stacking geometry, tertiary motifs, and assembly-level features. In RNA-RNA multimers, errors in monomer structure can combine with uncertainty in interface geometry and model selection, reducing the accuracy of the assembled complexes. These findings point to monomer structure accuracy, interface modeling, and model selection as key areas for improving RNA-RNA multimer prediction.
bioinformatics2026-09-04v3RNA structure conservation in plastids across plant evolution
Mehta, D.; Xiao, C.; Hua, J.; Siqueira Reis, R.Abstract
Plastid genomes are deeply evolutionary conserved. RNA structures within the primary or mature transcript play central role in plastid regulation of RNA processing, stability, and translation. However, the identity and conservation of RNA structures selected in plastid's evolution are still largely elusive. Here, we developed a stringent, covariation-based pipeline that perform an unbiased screen for conserved RNA secondary structures across entire plastid genomes. We analysed ~14,000 plastid genomes and identified a repertoire of 57 high-confidence conserved structures. We recovered known functional classes, e.g., 16S rRNA, tRNA, group II intron, and 3' end stem-loop, evidencing that our genome-wide analysis is reliable. We further uncovered novel putative cis-acting structures within the UTRs and introns of key photosynthetic genes, including psbN, clpP, and atpF, as well as putative trans-acting antisense RNAs to petB and psbT, suggesting uncharacterized elements with major regulatory function. Experimental in vivo RNA probing demonstrated that nearly half of the conserved structures adopt the predicted conformation in Arabidopsis plastid. Our comprehensive, yet stringent atlas of conserved plastid RNA structures provides the foundations for new regulatory discoveries in plastid biology.
bioinformatics2026-09-04v3Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach
Kalkhoven, E.; Baak, R. E.; Hooiveld, G. J. E. J.; Schrauwen, P.; Hoeks, J.; Raymakers, R.; van der Stolpe, A.; Kersten, S.Abstract
Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.
bioinformatics2026-09-04v2Integrating Spatially Adjusted Protein Summaries for Survival Prediction in Spatial Proteomics
Ahn, S.; Oh, E. J.; Prada, D.; Shojaie, A.Abstract
Recent advances in spatial proteomics, particularly imaging mass cytometry, enable the measurement of protein expression at the single-cell level while preserving a spatial context. Conventional survival analyses, however, typically rely on patient-level averages of protein intensities and therefore overlook spatial heterogeneity and tissue architecture. To address this limitation, we introduce a framework that incorporates spatial information into survival modeling by generating spatially adjusted protein summaries (SAPS). In this approach, cell-level protein intensities within each patient are modeled using spatial spline regression to capture spatial trends. From these models, we extract two complementary features: a spatially adjusted mean expression and a residual variance that reflects cell-to-cell variability unexplained by spatial effects. These summaries are then incorporated into Cox proportional hazards models in combination with clinical covariates. We further show that our estimator is asymptotically equivalent to an oracle estimator under mild regularity conditions. In simulation studies, our proposed framework achieved improved predictive performance compared to other alternative methods. The application of the method to breast cancer imaging mass cytometry data indicate that spatially adjusted summaries may enhance survival prediction and reveal biologically interpretable spatial protein patterns, suggesting high translational potential. This methodology offers an efficient means of translating complex spatial proteomics data into patient-level features, providing both improved survival prediction and new insights into the role of spatial heterogeneity in cancer outcomes. R package is available on the Comprehensive R Archive Network repository at https://cran.r-project.org/web/packages/SurvSPro/index.html
bioinformatics2026-09-04v2Poly Pipeline: A Polyvalent Spatial Transcriptomics Workflow Validated Across Polyploid and Diploid Organisms
Carvalho, P. C.; Millsteed, T.; Henry, R. J.Abstract
Spatial transcriptomics (ST) has emerged as a transformative approach for visualizing tissue landscapes, yet it faces significant challenges regarding data standardization, sparsity, and the analysis of complex genomes, particularly polyploid plants. To address these limitations, we introduce Poly Pipeline, a robust and universal bioinformatic workflow designed to streamline analysis across diverse plant and animal genomes. The pipeline integrates a comprehensive converter for proprietary formats, clustering algorithms, and hdWGCNA co-expression networks, which indirectly preserves the expression signatures of low-expressed duplicated genes. Benchmarking across datasets from wheat, rice, Arabidopsis, and mouse demonstrated the broad applicability of the pipeline in identifying relevant clusters, showing effectiveness across diverse organisms and data types. By providing a unified and reproducible framework, Poly Pipeline addresses a critical gap in analyzing genomic redundancy, especially that related to polyploidy, and promotes FAIR data principles for the broader scientific community.
bioinformatics2026-09-04v1Restrictome-EVOLVE: population-resolved haplotype architecture of human antiviral restriction-factor loci
Maghembe, R. S.; Bahati, S. Y.; Makaranga, A.Abstract
Human antiviral restriction factors act across multiple stages of viral replication, but whether their population-resolved haplotype architecture differs systematically from comparable genomic regions is unclear. We tested this using phased public human genomic data from 660 individuals in seven African and African-diaspora populations, representing 30 canonical restriction-factor units and 436 target windows. Each canonical unit was compared with 80 exact matched genomic controls, yielding 2,400 frozen controls and 4,429,760 target-control endpoint comparisons across 19 retained haplotype endpoints. All 30 canonical units showed lower differentiation effects and lower robust population-private haplotype effects than their matched controls. Within-population diversity effects were higher in 19 of 30 units, whereas dominant-haplotype concentration effects were lower in 22 of 30. Nineteen units occupied a deconcentrated/high-diversity state, eight a concentrated/low-diversity state, and three a lower-diversity/lower-concentration state. Of 127 global endpoint/context summaries, 88 reached a global false-discovery-rate q value below 0.05; 76 were lower in restriction-factor targets and 12 were higher. Directional sign-test inference detected widespread repeated displacement relative to matched controls, whereas no matched-cell empirical-rank test reached global false-discovery-rate significance. These results show that human antiviral restriction-factor loci occupy a reproducible matched-control haplotype architecture characterized by attenuated population partitioning and reduced robust private structure, together with substantial locus-specific variation in within-population diversity and haplotype concentration. The comparative framework separates population-genomic structure from claims of functional or adaptive causality.
bioinformatics2026-09-04v1When DL-Based Prescreening Meets Synthon-Based Docking: Target-Adapting PharmacoNet via MEL-Steered Correction
Liu, W.; Hong, Y.; Ku, T.; Lee, W.; Nguyen, E.; Xu, A.; Katritch, V.Abstract
As chemical libraries expand into the trillions of molecules, Virtual SYNthon Hierarchical Enumeration Screening (V-SYNTHES) has emerged as a leading strategy for making gigascale virtual screening computationally tractable. In V-SYNTHES, a Minimal Enumeration Library (MEL) of chemical fragments is docked against a target first, and only the top-scoring fragments are expanded into full ligands for large-scale docking. However, among the large number of comparably well-docked fragments, only a small fraction can be expanded under a fixed docking budget, leaving most similarly promising fragments unexplored. General-purpose prescreening tools can be adopted to address this constraint, reallocating the same docking budget across a larger pool of fragments' enumerated full ligands by their proxy score. However, such tools are applied without accounting for target-specific pocket environments. One such method, PharmacoNet, predicts interaction hotspots from a protein structure and ranks candidates via graph matching against a fixed set of interaction-type weights. We recognize that V-SYNTHES's initial fragment-docking step, ordinarily used only for selection of best fragments for expansion, already reveals which of these hotspots and interaction types a given pocket actually favors, and we can recover this signal to fine-tune PharmacoNet accordingly. We introduce MEL-Steered PharmacoNet, a parameter-efficient adaptation framework that specializes PharmacoNet to a given target through two composable mechanisms: (i) empirical density-map steering of predicted pharmacophore hotspots, and (ii) empirical fine-tuning of interaction-type scoring weights. Across three structurally distinct GPCR targets (CB2, GPR91, 5-HT2AR), MEL-Steered PharmacoNet achieves substantial enrichment factor (EF100) gains over a random baseline, and improves EF100 over PharmacoNet by 8.94x, 6.87x, and 1.69x, respectively. The fitted per-target weights further reveal distinct, chemically interpretable interaction profiles that PharmacoNet's generic fixed weights fail to capture. These results show that fragment-docking data already generated by the standard V-SYNTHES pipeline can adapt a general-purpose pharmacophore prescreening method to an individual target, significantly improving its performance while retaining its ultra-fast screening ability, with no additional experimental data or model retraining.
bioinformatics2026-09-04v1Evaluating performance bias in face-to-BMI vision transformer models across diverse human populations
Hoffman, J.; Gurven, M.; Kaplan, H.; Stieglitz, J.; Trumble, B. C.; Beheim, B.; Hooper, P. L.; Lee, R. B.; Phelps, J. R.; Hill, K.; Codding, B. F.; Brewer, S.; Lim, Y. A. L.; Lea, A. J.; Wallace, I. J.; Venkataraman, V. V.; Kraft, T. S.Abstract
Computer vision models that estimate body mass index (BMI) from facial features offer a non-invasive, low-cost alternative to physical measurement, with uses in telemedicine, emergency care where a scale or measuring tools arent available, automated self-monitoring, and large-scale epidemiological research. Most of these models, however, are trained on government records, social media images, and celebrity photographs, sources that introduce dataset biases and fail to represent the general public. This study tests how well a face-to-BMI machine learning model generalizes across populations, specifically how morphological diversity and population-specific training data affect cross-cultural accuracy. We trained and evaluated Vision Transformer (ViT-H/14) models on paired BMI measurements and facial photographs from four Indigenous populations: the Orang Asli of Malaysia, the Ju/hoansi of Southern Africa, the Sama residing in the Philippines, and the Tsimane of Bolivia. To evaluate how training data composition affects predictions, we compared four training strategies, from single-population models (focal models) to models trained on the full combined global dataset (global models). In-distribution training always produced the best performance. Models exposed to a target populations morphology, whether focal or global, consistently predicted BMI most accurately for that population. But when a target population differed from the training sample, adding more cross-cultural variation to training improved out-of-distribution predictions. Therefore, training on a populations own data works best when that data exists, and training on data spanning a wide range of human morphology is the strongest fallback when it doesnt. These findings suggest that while target population training data produces the most accurate results, training on datasets that capture global morphological variation substantially improves performance in unrepresented populations. Broader diversity in training data is essential for developing machine learning health tools that generalize reliably across human populations.
bioinformatics2026-09-04v1siProGenA: Generative siRNA Candidate Construction via Position Proposal and Guide Generation
Ma, Z.; Zhou, J.; Wang, R.; Deng, Z.; Wu, Z.; Zheng, Y.Abstract
Small interfering RNAs (siRNAs) are short guide RNAs that recruit the RNA-induced silencing complex (RISC) to complementary target sites on messenger RNAs (mRNAs), triggering Ago2-mediated cleavage and gene silencing. siRNA design requires compact candidate sets that cover a target while preserving efficacy, specificity, and practical sequence constraints. Existing pipelines usually enumerate candidate windows, assign a canonical guide to each window, and then rank preconstructed siRNA--mRNA pairs. This has produced strong pairwise efficacy predictors, but leaves a candidate-construction gap: candidate positions and guide sequences are fixed before the model begins to rank them. We address this gap by decomposing siRNA candidate construction into two generative decisions: where to place candidates within an mRNA segment, and what constrained guide variants to consider at a candidate position. We instantiate this framework as siProGenA, using a Discrete Denoising Diffusion Probabilistic Model (D3PM) for mRNA-conditioned position proposal and a Bayesian Flow Network (BFN) for temperature-controlled guide generation. On 62 positive test segments, the diversity-aware final library reaches Hit@1 = 0.790 and Hit@5 = 0.903. In a measured-site controlled Stage~2 evaluation, seed- and cleavage-preserving variants outscore the canonical complement for 89.8% of measured sites, with supporting gains across additional computational scorers, random-mismatch controls, and biophysical diagnostics. Together, the results support a modular proposal--generation view of siRNA candidate construction for prioritizing compact candidate sets.
bioinformatics2026-09-04v1BROOQS: Spectral Methods Resolve Level-1 Hybridization Cycles without Tests of Symmetry
Arasti, S.; Mirarab, S.Abstract
Modern phylogenomic analyses often seek to reconstruct both vertical and reticulate evolutionary histories. While the prevalence of non-vertical evolution is increasingly appreciated, inferring networks remains conceptually challenging and computationally demanding. Following the success of quartet-based methods for handling gene tree discordance, several quartet-based network inference methods have been developed. A key insight of these methods is that level-1 networks can be constructed by first building a multifurcating tree called tree-of-blobs and then resolving each polytomy into a cycle. This two-step approach makes the problem easier both conceptually and computationally. However, these quartet-based methods often rely on noisy statistical tests of asymmetry in quartet frequencies. Moreover, they either enumerate all quartets, losing some scalability, or subsample them, losing information. We introduce BROOQS, a quartet-based method for resolving trees of blobs into a level-1 phylogenetic network. BROOQS efficiently aggregates information from all quartets around a blob without enumerating them, builds a pairwise similarity matrix, and uses robust spectral ordering algorithms to recover the cyclic ordering without relying on individual quartet symmetry tests. We prove theoretically that our spectral method is consistent under the network multi-species coalescent (NMSC) model. Across simulated and empirical datasets, BROOQS consistently improves accuracy and scalability compared to existing methods and extends to thousands of taxa.
bioinformatics2026-09-04v1On doubting image quality assessment metrics for microscopy virtual staining
Li, W.-s.; Way, G. P.Abstract
Pairing label-free microscopy with virtual staining could reduce the cost and experimental burden of fluorescence microscopy, but its impact is conditional on generalizable inference. Most virtual staining studies assess performance using image quality assessment (IQA) metrics developed for natural images, yet how well these metrics translate to microscopy remains unknown. Here, we examined the behavior of seven commonly-used full-reference training objectives and metrics, MAE, PSNR, SSIM, foreground PSNR and SSIM, LPIPS, and DISTS, under controlled image degradation and realistic out-of-distribution virtual staining. We applied graded intensity, textural, and morphological transformations to Cell Painting images spanning 18 cell lines, seeding densities, and fluorescence channels. Channel, cell line identity and seeding density explained substantial metric variation after controlling for degradation magnitude. DISTS and foreground metrics showed more favorable balance between degradation sensitivity and biological invariance, although no metric reported performance independent of biological context. Incrementally degrading images and evaluating concomitant metric degradation further revealed that most metrics used only a small fraction of their nominal numerical ranges and frequently plateaued while image degradation visibly continued. We next trained three popular virtual staining model architectures (UNet, WGAN-GP, UNeXt) on five U2-OS seeding densities separately, and computed metrics on model predictions across 17 unseen cell lines. We observed that architecture and training U2-OS seeding density together explain less than 2% of metric variation. Visual inspection suggested comparable scores across cell lines correspond to qualitatively distinct errors, such as differences in cell morphology and marker intensity. These findings show that conventional IQA metrics do not effectively translate to virtual staining applications. Selection or optimization of virtual staining models against real application such as in label-free high content drug screening should instead be approached in an application-oriented fashion.
bioinformatics2026-09-04v1Gene function prediction from bulk coexpression is bounded by cell-type-level signal
Adrian-Hamazaki, A.; Pavlidis, P.Abstract
It is widely accepted in genomics that coexpression of RNA transcripts suggests a commonality of function. This intuition is explicitly leveraged in machine learning methods that predict gene function, where it is often combined with other features such as protein interactions and sequence similarity. For example, including coexpression data from human tissue expression boosts performance for predicting Gene Ontology annotations. However, the biological underpinnings of this observation have not been well-investigated. Building on earlier results from our group, in this work we show that gene function is predictable from coexpression substantially because it reflects differences in expression between cell types, and these differences are also intrinsic to the ground truth labels. Using simulations and analyses of real data, we show that variance in the cellular composition of bulk samples impacts function learnability and attribute this to cell type marker gene content in the GO terms. We further show that cell type profiles, where the relationship between gene expression and cell type is made transparent, are effective for predicting gene function while increasing interpretability. These results indicate that function prediction models trained on bulk coexpression are largely limited to cell-type-level resolution rather than fine-grained biochemical function, with direct consequences for how such predictions should be interpreted.
bioinformatics2026-09-04v1Modelling interpretable patient-level representationsfrom structured and simple multimodal data
Oksza-Orzechowski, K.; Lazecka, M.; Koperski, L.; Wojtowicz, D.; Mozejko, M.; Schulz, D.; Liechti, R.; Marzetta, F.; Morfouace, M.; Hong, H. S.; Tissot, S.; Bodenmiller, B.; Staub, E.; Szczurek, E.Abstract
Patient cohort profiling increasingly includes structured views for multiple modalities, such as single-cell RNA sequencing, spatial transcriptomics or proteomics, and histology, each providing multiple subobservations per patient, including single cells, spatial spots or patches. To model such data along with simple patient-level views, current multimodal integration methods typically rely on separately precomputed summaries and fail to fully leverage information in structured views. Here we present FACTMx, a variational framework that jointly models structured and simple views to learn interpretable patient-level representations. FACTMx couples latent patient factors with subobservation clustering and per-patient component proportions, enabling direct interpretation and downstream association analyses. The framework supports different structured-view mixture assumptions, including topic- and Gaussian-structured data, while retaining modular encoder-decoder parameterisations. In simulations spanning sparse and dense dependencies and multiple noise regimes, FACTMx improved reconstruction, integration and recovery of structured components relative to previous methods. Applied to non-small cell lung cancer cohorts, FACTMx captured survival-associated latent signals linked to immune microenvironments, gene expression pathways and spatially coherent histological patterns. In a longitudinal coronary syndrome cohort, FACTMx highlighted an outcome-associated axis connected to ejection-fraction change, immune cell states, soluble mediators and cardiac injury markers. These results support joint structured-simple modelling for interpretable multimodal patient stratification.
bioinformatics2026-09-04v1QuickSeg: A fast, versatile and accurate algorithm for genomic copy number segmentation using dynamic programming
Schlotmann, B.; Favero, F.; Locallo, A.; Weischenfeldt, J. L.Abstract
Copy number alterations are among the most common genomic aberrations in cancer and their accurate identification relies on robust segmentation of sequencing read-depth signals. Existing segmentation methods typically balance computational efficiency against segmentation accuracy and remain sensitive to technical artifacts present in sequencing data. Here, we present QuickSeg, a fast and versatile methodology that uses an exact dynamic programming algorithm to detect copy number segments using median-based error function. Motivated by the observation that sequencing depth distributions contain a small but pervasive population of outlying observations, this approach provides increased robustness to technical noise while simultaneously reducing the computational complexity of the segmentation problem. Across whole-genome sequencing of cancer cohorts, using breakpoint-supported somatic copy number alterations, we demonstrate improved segmentation precision over two widely used baseline methods, Circular Binary Segmentation (CBS) and Piecewise Constant Fitting (PCF), across a broad range of sensitivity thresholds. QuickSeg also consistently outperformed both methods with respect to runtime and memory usage. Collectively, our results show that robust median-based optimization provides both biological and computational advantages for copy number segmentation, enabling accurate analysis of large sequencing cohorts with minimal computational requirements.
bioinformatics2026-09-04v1AltraFlowSOM: A Semi-Supervised Framework for Imaging Mass Cytometry Phenotyping
ANILKUMAR REKHA, A.; Bettacchioli, E.; Le Dantec, C.; Hemon, P.; Jouve, P. E.; Hillion, S.Abstract
Imaging Mass Cytometry (IMC) enables the simultaneous quantification of 40+ protein markers at single cell resolution in tissue, however biologically faithful phenotyping at scale remains a critical bottleneck. Unsupervised clustering fragments coherent populations or conversely merges biologically incoherent ones into a single cluster, supervised classifiers impose a closed vocabulary, and the presence of rare subsets (encoding clinically relevant biology) in conjunction with abundant subsets may be detrimental to detection performances. We present AltraFlowSOM, a semi-supervised extension of FlowSOM that embeds partial expert annotations directly into self-organizing map training via a two-layer SuperSOM architecture, balancing label-guided topology anchoring with unsupervised discovery. By anchoring the map to biologically labelled reference points, AltraFlowSOM circumvents the canonical dependency between batch correction and clustering. Evaluated under Leave-one-out cross validation on two independent IMC cohorts, Lupus Nephritis (n=22 ROIs) and Sjogren syndrome (n=10 ROIs), AltraFlowSOM outperformed all unsupervised and supervised baseline on Adjusted Rand Index, F1 scores (macro and weighted), weighted purity and in the identification of rare populations. The median Treg cell recovery exceeded that of all comparator methods. AltraFlowSOM resolves the scalability-alignment-discovery trilemma, by establishing a semi-supervised SOM as a generalizable method for high dimensional IMC phenotyping.
bioinformatics2026-09-04v1Experimental and In Silico Analysis of the Structural Dynamics of Dengue NS2B-NS3 Protease
Muthuvel, s. k.; BALAKRISHNAN, S. S.; MANJINI, S.; DHAL, K.; DAS, S.; S, D.Abstract
The Dengue virus (DENV), a major global health concern, causes dengue fever, predominantly affecting tropical and subtropical regions. The NS2B-NS3 protease complex of DENV is critical for viral replication, making it a promising target for antiviral drug development. This study investigates the structural dynamics of the NS2B-NS3 protease under varying pH conditions, integrating experimental and computational approaches. The recombinant NS2B-NS3 protease was expressed in E. coli, purified using affinity chromatography, and analysed for purity via SDS-PAGE. Circular dichroism (CD) spectroscopy revealed pH-dependent secondary structural changes, indicating stability at neutral to slightly basic pH and destabilization under acidic and highly basic conditions. Dynamic light scattering (DLS) analysis demonstrated protein aggregation and structural heterogeneity under extreme pH levels. Complementary in silico techniques, including homology modelling and molecular dynamics (MD) simulations, provided detailed insights into the conformational changes of the protease. The modelled structure, validated and refined through computational tools, revealed structural stability at physiological pH, with notable disruptions at pH extremes. DSSP (Dictionary of Secondary Structure of Proteins) and principal component analyses highlighted significant secondary structural transitions, especially at acidic pH where -helices and {beta}-sheets transformed into random coils Docking and MD simulation was carried out with the small molecule for therapeutic analysis between Dengue NS2bNS3 and small molecule. This study emphasizes the pH-dependent conformational plasticity of the NS2B-NS3 protease, contributing to the understanding of its functional mechanisms and providing a foundation for the rational design of pH-specific inhibitors. These findings underscore the importance of structural biology in advancing therapeutic strategies against dengue fever.
bioinformatics2026-09-04v1A pan-cohort transcriptional landscape of breast cancer maps subtype and microenvironmental programs
Arora, S.; Suresh, R.; Holland, N.; Glatzer, G.; Jensen, M.; Konnick, E. Q.; Pritchard, C.; Li, Y.; Parsons, H. A.; Hurvitz, S. A.; Holland, E. C.Abstract
Breast cancer comprises heterogeneous transcriptional states that are incompletely captured by discrete clinical or molecular subtype labels. To visualize this heterogeneity in a unified framework, we integrated bulk RNA-seq data from 2,284 patient samples across 13 studies using 18,089 protein coding genes, a harmonized processing pipeline, batch correction, consensus clustering and PaCMAP dimensionality reduction to construct an interactive breast cancer transcriptional landscape. Consensus clustering identified five major regions, which were annotated using PAM50 scores calculated for each sample: Luminal A, Luminal B, HER2 enriched, and two basal associated clusters. The basal clusters separated into an immune rich region marked by T cell-inflamed, tumor-associated macrophages (TAM), and low-purity signatures, and a cell-cycle driven region enriched for proliferation and DNA replication programs. Overlay of marker genes, pathways, kinases, neuronal like signaling programs, cancer associated fibroblasts (CAF) states, and TAM programs revealed spatially organized subtype biology and microenvironmental heterogeneity. Finally, projection of therapy associated resistance signatures identified landscape regions linked to predicted resistance to HER2-targeted therapy and hormone receptor directed endocrine therapies. By enabling interactive exploration of transcriptional states, marker genes, pathways, and therapeutic response programs, this resource provides a community framework for biomarker discovery in breast cancer.
bioinformatics2026-09-04v1Creating DNAm Algorithms Using the Illumina Methylation Screening Array (MSA)
Seale, K.; Hassouneh, S.; Giosan, I.; Sugden, K.; Balague-Dobon, L.; Dwaraka, V.; Lasky-Su, J. A. B.; Mallin, M.; Caspi, A.; Moffitt, T.; Smith, R.; Carreras-Gallo, N.Abstract
Most established DNA methylation (DNAm) biomarkers were developed on legacy Illumina EPIC arrays. The Infinium Methylation Screening Array (MSA) offers a lower-cost, higher-throughput alternative with reduced probe content, but EPIC-trained algorithms cannot be assumed to transfer directly. Here we present a reproducibility-based framework for developing and transferring DNAm algorithms on the MSA. Using paired biological replicates profiled on EPICv1 and MSA (1,764 EPICv1-MSA sample pairs, plus within-array MSA replicates on the same and different beadchips), we quantified probe-level agreement using mean absolute error (MAE) and intraclass correlation coefficients (ICC). Of 140,150 CpG sites shared between EPICv1 and MSA, 40,786 (29.1%) met both stability criteria (MAE < 0.05 and ICC(2,k) > 0.6). This stable feature space supported two modelling streams. First, we trained 134 epigenetic biomarker proxies (EBPs) natively on MSA, with and without kernel principal component analysis (kPCA) for sample-level harmonisation. All 134 reached same-beadchip ICC(2,1) >= 0.80 (median 0.97) and 96.3% reached different-beadchip ICC(2,1) >= 0.60 (median 0.81), with a median Spearman correlation of 0.48 against observed values. Among the 72 kPCA-selected models with a comparable stable-probe baseline, 70 (97%) showed higher cross-beadchip ICC (median improvement +0.18). Second, we transferred three established clocks using model-specific strategies: OMICmAge and SystemsAge were retrained to estimate their EPICv1-derived values (held-out test-set rho = 0.944 and 0.912-0.949), whereas DunedinPACE required stable-probe normalisation and robust linear calibration, which raised cross-array ICC(2,1) from 0.784-0.810 to 0.891-0.925 and reduced MAE from 0.085-0.089 to 0.041-0.050 across three sample sets. Reduced probe content does not preclude reproducible DNAm biomarker measurement, and transfer strategy must be matched to model architecture.
bioinformatics2026-09-04v1Modeling Patient-Reported Pain Trajectories with Frequent Minimum and Maximum Scores
Liu, Y.; Harris, R. E.; Clauw, D.; Bayman, E.; Leroux, A.; Lindquist, M. A.Abstract
Chronic pain is a widespread public health issue that imposes substantial health, emotional, and economic burdens on individuals and communities. Because pain is subjective and lacks objective biomarkers, it is typically measured using patient-reported scores, often on a numerical scale from zero to ten. Increasingly, pain studies use ecological momentary assessment, with multiple daily assessments over days and across study phases (e.g., a series of baseline and post-intervention assessments). These data frequently show many ratings at the extremes (i.e., at minimum or maximum pain scores), commonly referred to as zero- and one-inflation in the statistical literature, along with considerable within-person variability both within and across days. These phenomena present challenges for statistical analyses, as they violate assumptions of most commonly used statistical techniques (e.g., the normality assumption of linear mixed models). We propose a Bayesian beta-binomial mixed-effects model for modeling potential zero- or one-inflated pain scores while accounting for variability using random effects on the mean and variance parameters across subjects. A simulation study demonstrates that the method accurately estimates model parameters across realistic sample sizes, time points, and zero- and one-inflation levels. An application to data from two longitudinal pain studies demonstrates that the model fits the data better and, when correctly specified, yields accurate uncertainty intervals for longitudinal changes in pain compared to existing models, especially for zero- and one-inflated outcomes. Additionally, the model directly estimates the probability of clinically meaningful pain events. The proposed method provides a powerful statistical framework for studying the patient-reported pain trajectories.
bioinformatics2026-09-03v4kamino: fast proteome-wide variant calling for amino acid phylogenomics
Derelle, R.; Lees, J. A.; Chindelevitch, L.Abstract
Amino acid-based phylogenetics usually relies on first clustering and aligning orthologous proteins. This approach is powerful but computationally demanding. Here, we present kamino, a reference- and alignment-free method that rapidly builds amino acid phylogenomic alignments directly from proteomes. As with similar algorithms, homologous regions are identified through shared sequences flanking variable regions. The method uses local changes in recoded k-mer occupancy to efficiently identify variable positions within these homologous regions and extract the corresponding pseudo-aligned sequences. It generates phylogenetically informative alignments across diverse prokaryotic and eukaryotic datasets. Phylogenetic analyses show that it accurately recovers Mycobacterium tuberculosis lineages, most curated GTDB taxa, and relationships consistent with published Drosophila and mammalian phylogenies, while producing signals broadly similar to BUSCO-based approaches. Runtimes are comparable to genome-based alignment-free methods and several orders of magnitude faster than classical marker-based pipelines, with moderate memory requirements. The method performs well across a broad range of divergence levels, from within-species comparisons to family-level prokaryotic and phylum-level eukaryotic datasets. kamino therefore provides a fast and simple route from proteomes to phylogenomic alignments across a broad range of evolutionary scales. The program is implemented in Rust and freely available at https://github.com/rderelle/kamino.
bioinformatics2026-09-03v2CoTRA: an integrated R/Shiny framework for transparent bulk and single-cell RNA-seq analysis
Seemab, U.; Vainionpaa, K.; Tanoli, Z.; Leinonen, H. O.Abstract
Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.
bioinformatics2026-09-03v2AI semantics for biomedical data integration
McLaughlin, J.; Puig-Barbe, A.; Ibrahim, A.; Pava, D.; Pendlington, Z. M.; Matentzoglu, N.; Sollis, E.; Foreman, A.; Wilson, R.; Lopez Gomez, F.; Harris, L.; Adeleye, Y.; Kaur, S.; Meldal, B.; Smedley, D.; Parkinson, H.Abstract
Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge. Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries.
bioinformatics2026-09-03v2DAQplugin: Interactive Deep Learning-Based Validation of Cryo-EM Protein Models in ChimeraX
Terashi, G.; Zhu, H.; Kihara, D.Abstract
Although an increasing number of protein structures are determined by cryogenic electron microscopy (cryo-EM), structure modeling frequently suffers from residue misassignments and sequence register shifts, particularly in regions with ambiguous density. Here, we present DAQplugin, a ChimeraX plugin for real-time evaluation of protein models against cryo-EM density maps using the deep-learning-based residue-wise model quality (DAQ) score. Unlike existing validation tools that are typically applied after model construction, DAQplugin enables interactive validation during model building and refinement. DAQ has been shown to accurately identify residue assignment errors, including sequence register shifts, as well as local conformational modeling errors. DAQplugin also provides guidance for correcting sequence register shifts by suggesting alternative residue placements along the backbone. The plugin is computationally efficient and runs on standard CPUs without requiring GPU hardware, enabling deep-learning-based validation on ordinary laptops during interactive model building, model-map fitting, and refinement. DAQplugin facilitates more accurate interpretation of cryo-EM density maps and improve the reliability assessment of protein structure models.
bioinformatics2026-09-03v2MetaClaw: an auditable AI agent for end-to-end, multi-directional metagenomic and multi-omics analysis
Zhang, H.; Li, Z.; Lagniton, P. N. P.; Wang, Z.; Zhao, L.; Li, W.; Duan, p.; Jiang, X.; Ning, K.Abstract
End-to-end omics analysis requires more than selecting tools: a usable agent must bind data correctly, execute long workflows without blocking, preserve provenance and recover the biological conclusions that motivate an analysis. Existing LLM-driven bioinformatics agents automate parts of this process, but their operational dependencies and conclusion-level validity are often unclear. Here we present MetaClaw, an auditable agent that maps a user request to a registered workflow, executes standardized upstream processing on FlowHub, and runs study-specific downstream analyses in network-isolated OpenClaw containers. A YAML registry and an explicit plan-submit-poll-finalise lifecycle record file bindings, parameters, scripts, environments and outputs in per-job bundles. Across the full cohorts of four published studies (769 metagenomic profiles), MetaClaw recovered 4/4 sorghum marker groups, 3/3 RRMS features, 4/5 canonical CRC markers among the top 20 classifier features and 5/5 permafrost marker groups. In 45 model-by-prompt runs, upstream completion was consistent whereas downstream validity depended on the backend and instruction detail; three decoy-tested endpoints showed no significant differences. In 48 ablation sessions, removing the registry, planning loop or manifest caused distinct losses, with registry removal increasing time, tool calls and token cost. MetaClaw therefore connects standardized upstream execution, local analytical flexibility and conclusion-level validation in a rerunnable framework for metagenomic and microbiome multi-omics analysis.
bioinformatics2026-09-03v2Architectonic Spandrels in the Origin of Enzymatic Function
Poley-Gil, M.; Fernandez-Martin, M.; Banka, A.; Heinzinger, M.; Rost, B.; Valencia, A.; Parra, R. G.Abstract
How new molecular functions emerge during protein evolution remains a fundamental question in molecular biology. The energy landscape theory states that proteins are minimally frustrated, i.e. they have minimized their internal conflicts, to allow robust folding. Yet, not all energetic conflicts are eliminated, with functional regions such as catalytic residues and ligand-binding sites being often enriched in frustrated interactions, trading localized stability for biological activity. However, it is still uncertain whether this functional frustration is an evolutionary adaptation, positively selected despite its energetic cost or an inevitable physical byproduct of the fold architecture. Here, we combine reverse folding, structure prediction, and sequence analysis with local frustration profiling to address this long-standing question. Unexpectedly, we found that reverse folding algorithms are unable to energetically minimize evolutionary conserved frustration at specific residues, even when detrimental to overall structural stability. We propose that these frustration hotspots act as architectural spandrels, inherent physical constraints of the fold that evolution subsequently co-opts for function. Our findings connect biophysical constraints and evolutionary selection, providing a new framework to understand how functional specificity emerges in protein landscapes.
bioinformatics2026-09-03v2AnnFlux: object-conditioned neural stochastic differential equations for single-cell perturbation dynamics
Choi, H.; Byeon, G.; Park, H.; Park, J.; Lim, S.; An, J.-Y.Abstract
Single-cell perturbation profiling measures responses to genetic and chemical interventions, yet most models learn a static map, ignoring how populations move over time and how perturbations combine. AnnFlux, an object-conditioned stochastic differential equation, learns a drift field in latent cell-state space. Conditioning on the perturbing object makes the field queryable one object at a time, yielding per-object drifts comparable across genes and drugs. By learning a drift field tailored to each perturbation context, it interpolates a held-out timepoint in an epithelial-mesenchymal transition time course and predicts unseen perturbations. Beyond point estimates, AnnFlux improves distributional fidelity and predicts responses to held-out perturbation combinations. An IFN-response signature predicted by AnnFlux was associated with TLS proximity in an independent pan-cancer spatial atlas. This framework maps perturbation-driven cell-state evolution as continuous trajectories and represents unseen perturbations using prior-knowledge embeddings.
bioinformatics2026-09-03v1An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.
Villalba, M. P. C.; Bustamante, F. E.Abstract
Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.
bioinformatics2026-09-03v1FibrilNet maps conserved and tissue-specific molecular environments across systemic amyloidoses
Guzzi, P. H.; Ugo, L. H.; Carbonari, V.; Lio', P.; Veltri, P.Abstract
Systemic amyloidoses are initiated by distinct amyloidogenic precursor proteins but frequently contain recurrent extracellular, complement, lipid-transport and matrix-remodelling components. Whether these recurrent proteins form a conserved systems-level environment across amyloid diseases, and how strongly that environment depends on precursor and tissue context, remains unresolved. We developed FibrilNet, a network framework that integrates experimentally defined amyloid proteomes with a human protein protein interaction graph and Gene Ontology derived semantic information. FibrilNet compares topology-only random walk with restart (RWR) with ontology aware semantic RWR in frozen leave-one-out module reconstruction and precursor-seeded prioritization tasks. The human graph contains 17,997 proteins and 925,977 physical interactions, with a 9-dimensional semantic representation of interaction context. In expanded cardiac transthyretin amyloidosis (ATTR), semantic-RWR increased mean reciprocal rank (MRR) from 0.00167 to 0.05015 and Recall@100 from 0.0199 to 0.3377, improving 132 of 151 held-out targets. Significant semantic gains were also observed in renal serum amyloid A amyloidosis (AA) and leukocyte chemotactic factor 2 amyloidosis (ALECT2). Across compact ATTR, light-chain amyloidosis (AL), AA and ALECT2 modules, APCS, VTN and TIMP3 formed a direct four-disease recurrent core, while APOE occurred in three of four modules. A tissue-aware ATTR analysis showed limited overlap between cardiac and neurologic modules (19 shared proteins; Jaccard 0.0569). In the hTTR-A97S peripheral-nerve model, semantic-RWR significantly improved reconstruction of the 202-protein mapped neurologic module, with the strongest evidence concentrated in the downregulated proteomic program. TTR-seeded propagation improved with semantic information but remained weak in absolute terms, separating precursor identity from the distributed downstream molecular environment. These results support a multilayer model in which a restricted conserved amyloid environment coexists with precursor-, tissue- and disease-specific organization
bioinformatics2026-09-03v1Joint ancestry inference reveals the landscape of archaic introgression in admixed populations
Medina Tretmanis, J.; Anorve-Garibay, V.; Peede, D.; Banuelos, M. M.; Avila Arcos, M. C.; Jay, F.; Huerta-Sanchez, E.Abstract
Studying the evolutionary history of archaic segments in recently admixed individuals requires inferring both continental and archaic ancestry in admixed genomes. Here, we present TRACTINATOR, the first deep-learning method for simultaneous inference of continental and archaic ancestry in admixed human genomes. The model combines SNP sequences, population allele-frequency information, and S* statistics to improve both inference tasks. By learning relationships between haplotypes and population allele frequencies, TRACTINATOR can generalize across genomic regions and even across different genomic datasets. We train our model using both real and synthetic data, and show that augmenting with synthetic data improves accuracy for both continental and archaic ancestry inference. Finally, we apply TRACTINATOR to admixed Latin American populations from the 1,000 Genomes Project, revealing how archaic ancestry is distributed within chromosomal segments of African, European and Indigenous American ancestry in Latin American individuals. For candidates of adaptive introgression, we also infer whether the archaic haplotype was introduced via European or Indigenous American ancestors.
bioinformatics2026-09-03v1Atlantis: An integrative database for human proteome structural and functional sites
De Oliveira Rosa, N.; Ferronato, P.; Varisco, M.; Matic, M.; Ruscio, M.; Miglionico, P.; Raimondi, F.Abstract
Understanding protein mechanisms in health and disease requires characterizing the functional roles of individual amino acid residues. To explore the role of residues and their mutations, we have developed Atlantis, a database that integrates structural and functional information at the human proteome residue level. A graph database enables complex queries and the retrieval of integrated information for multiple functional analysis of protein systems. A Model Context Protocol (MCP) connector allows the interrogation of the resource through Large Language Models (LLMs) or agentic frameworks for biomedical research. Atlantis annotates over 11M residues across 20k human proteins, identifying hundreds thousands intra- and inter-protein contacts in PDB as well as AlphaFoldDB structures. We also provide the possibility to analyze and integrate predicted 3D complexes inputted by the user, and we showcased these features on hundreds of AlphaFold-multimer complexes of GPCRs and LRRK2 interaction networks. The tool is freely accessible at https://atlantis.bioinfolab.sns.it/.
bioinformatics2026-09-03v1Comprehensive characterization of genomic, transcriptomic and epigenomic artifacts introduced in formalin-fixed, paraffin-embedded tissues.
Zmuda, E. J.; Covington, K. R.; Kandoth, C.; Ng, C. K. Y.; Weisenberger, D. J.; Bowlby, R.; Chu, A.; Corbett, R.; Brooks, D.; Cherniack, A. D.; Murray, B.; Ling, S.; Zhao, W.; Shinbrot, E.; Lim, R. S.; Bootwalla, M. S.; Hinoue, T.; Hu, J.; Kakkar, N.; Lawrence, M.; McLellan, M.; Miller, C.; Morton, D.; Mungall, A. J.; Muzny, D.; Sougnez, C.; Stojanov, P.; Xi, L.; Schultz, N.; Akbani, R.; Curtis, C.; Fulton, B.; Gibbs, R. A.; Getz, G.; Spellman, P.; Marra, M.; Ladanyi, M.; Tarnuzzer, R.; Sofia, H.; The Cancer Genome Atlas Research Network, ; Hutter, C.; Wheeler, D. A.; Gastier-Foster, J.; LairAbstract
Genomic, transcriptomic and epigenomic characterization has accelerated the discovery of clinically-relevant alterations in cancer, predominantly using fresh frozen (FF) specimens. However, clinical molecular pathology laboratories prefer formalin-fixed paraffin-embedded (FFPE) methods, known to introduce artifacts at the nucleic acid level, over fresh frozen methods. Extending the multi-platform analysis to FFPE specimens for comprehensive clinical molecular diagnosis requires a thorough understanding of the consequence of formalin-fixation. We present a detailed multi-platform characterization of FFPE preservation using paired FF specimens as the 'gold standard'. DNA and RNA were obtained from 38 patients across 6 cancer types using a FFPE optimized co-isolation. The impact of FFPE on exome sequencing was dependent on filtering, where a minimum coverage or supporting read filter can mitigate FFPE-specific false positives. Copy number alterations, MSI assessment, mutational signatures, and DNA methylation were comparable between FFPE and FF. FFPE biases in RNA expression can be overcome when using biology-relevant genes and we describe a novel consequence of FFPE on miRNA species diversity. Collectively, this data provides a broad view of FFPE artifact and offers best practices for overcome these biases.
bioinformatics2026-09-03v1From neuropeptide and receptor annotation to ligand-receptor pairing: a sequence- and structure-based framework for mapping the neuropeptide-receptor interactome in Gryllus bimaculatus
Margaretha, F.; Sakamoto, M.; Seike, H.; Nagata, S.; Nakamura, Y.; Mochizuki, T.Abstract
Neuropeptides and their G protein-coupled receptors (GPCRs) control much of insect physiology and behaviour, but in Gryllus bimaculatus, an emerging model and edible insect, receptor sequence similarity hinders the mapping of which peptide each GPCR activates. We re-annotated a chromosome-scale genome (BUSCO 95.3%, from 86.7% on insecta_odb12) with comprehensive curation of 48 neuropeptide precursor families (51 loci, including seven not previously identified) and 134 candidate GPCRs (66 rhodopsin-class, 68 secretin-class), providing a near complete neuropeptide-receptor interactome catalogue. We modelled all 15,946 peptide-receptor pairs with AlphaFold3 and Boltz-2 and scored each interface with pLDDT and ipSAE. Ranking these scores, and cross-checking the top candidate for each family against a receptor phylogeny of known ligand specificity, gave a confident, phylogenetically related receptor for 27 of 35 curated receptor groups. These matches confirm the structural scorings with existing deorphanization data and propose receptors for peptides with no prior functional evidence. The annotation, curated peptide and receptor sets, and ranked complexes are available through CricketBase (https://cricket.annotation.jp), a genome browser with a structure viewer of peptide-receptor complexes, providing a resource for G. bimaculatus endocrinology and a workflow to deorphanize GPCRs in other non-model insects.
bioinformatics2026-09-03v1Bravais Lattice Sampling: Geometry-Guided Sparse Probing for Connected-Component Detection in 3D Discretized Spaces
Carrascoza, F.Abstract
We introduce Bravais Lattice Sampling (BLS), a two-phase method for detecting connected high-density regions in three-dimensional space. BLS places probe sites on a Bravais lattice scaled to the expected nearest-neighbour distance dNN of the target structures, then recovers cluster boundaries by depth-first expansion seeded only from occupied probes, replacing the exhaustive raster scan that conventional connected-component labelling uses to discover seeds. The spacing between probe sites is set from the covering radius of the lattice, which is what allows the method to state in advance the size below which a cluster may escape detection. The second phase, an expansion refinement activated only on probes that return an occupied voxel, verifies every edge, so the components returned are true connected components. BLS versatility allows for selection of different Bravais lattice unit cells to match the target structure; for amorphous, non-crystalline shapes, BLS can default to a simple face-centred cubic unit cell, where the expected minimum cluster size is the only parameter that needs to be set. The current BLS implementation has been developed as a post-processing tool for molecular dynamics trajectories, and was tested for searching water ice clusters of different morphologies. BLS returns component counts and maximum cluster sizes identical to exhaustive-labeller algorithms, with 100% recall; it runs at about 0.94 times the cost of depth-first search, and at 0.84 to 0.90 times the cost of the fastest other labeller in our benchmark set. This algorithm, although implemented by us for molecular dynamics applications, could be of interest in other domain areas where searching for high-density elements in 3D space is relevant.
bioinformatics2026-09-03v1Cell-type separability predicts annotation accuracy and outweighs algorithm choice: a factorial benchmark across seven scRNA paradigms
Wardhana, O.; Zeng, Z.; Lu, X.Abstract
Automated cell-type annotation is a prerequisite for most single-cell RNA-sequencing (scRNA-seq) analyses, but the rapid proliferation of methods spanning marker-based, correlation-based, classical machine-learning, deep-learning, semi-supervised, large-language-model (LLM), and transformer foundation-model paradigms has outpaced head-to-head evaluation. Existing benchmarks rely on convenience samples of real datasets in which cell count, class imbalance, cell-type number, and differential-expression strength co-vary uncontrollably, precluding causal attribution of performance to any dataset property. To resolve this, we benchmarked 63 tools across seven paradigms using a Taguchi L9(34) orthogonal array that varies four dataset properties independently, progressively reconfiguring experimental control across five phases: fully controlled simulation, within-platform and cross-platform real-data validation, database-connected and LLM-based annotation under ontology-aware scoring, and fine-tuned foundation models. Using standardized oracle inputs and Cohen's {kappa}, we found that, within the ranges tested, the major paradigms achieved comparable accuracy. Accuracy was predicted near-linearly by the separability of cell types in a shared expression embedding, measured as k-nearest-neighbor (kNN) purity, a relationship that held across sequencing platforms and in fine-tuned foundation models. We attributed the vast majority of {kappa} variance to dataset structure and only a small share to tool identity. Computational cost traded against workflow accessibility rather than accuracy: accessible correlation-based and LLM-based approaches performed competitively, while foundation models matched them only after fine-tuning. Because our oracle design isolates algorithmic capability from upstream noise, these results reframe how methods should be selected: the field's near-term gains lie in strengthening infrastructure--prioritizing tool accessibility, standardized evaluation, and robustness to pipeline variation.
bioinformatics2026-09-03v1BioIMA: a one-click desktop tool for standardized extraction of phenotypic traits from biological images
Qumu, X.; Dan, X.; Feng, J.; Cui, Y.; Gong, Y.; Hou, Y.; Lai, Q.; Wang, Z.; Zhang, Y.; Zhu, Y.; Yu, Y.; Zhang, F.; Todesco, M.; Wang, J.Abstract
Standardized extraction of quantitative phenotypes from images is increasingly important across plant biology, from ecological and evolutionary studies to genetics, breeding, and functional genomics. However, as large image datasets are increasingly used for trait analysis, many biologically relevant traits, including size, shape, color, and spatial patterning, are still measured manually or using fragmented semi-automated workflows. These limitations reduce throughput, reproducibility, and accessibility, especially for researchers without computational expertise. Here, we present BioIMA, an open-source desktop tool for rapid and standardized phenotyping from biological images. BioIMA integrates foundation model-based segmentation with automated trait computation, allowing users to extract quantitative measurements from images through an intuitive graphical interface and without model training. To validate its performance, we quantified a set of knot morphological traits in two Populus species, as these measurements are typically time-consuming to perform manually. Automatic measurements showed strong agreement with manual ImageJ-based measurements (R2 > 0.95), while reducing per-image processing time by approximately 75% (from ~15 s to ~4 s). BioIMA was further applied to diverse plant datasets, including Helianthus and Rhododendron images with varying morphologies and background conditions. Although developed for plant phenotyping, BioIMA may also be extended to other biological samples where region-based size, shape, or color traits are of interest. By combining accessibility and standardization in a lightweight local application, BioIMA provides a practical community resource for image-based phenotyping in ecological and evolutionary studies.
bioinformatics2026-09-03v1Towards Sparse Causal Features for Zero-shot Mutation Effect Prediction in a Protein Language Model
Mohanty, S.; Phutela, M.; Green, A. G.Abstract
Protein language models (pLMs) such as ESM-2 achieve strong zero-shot mutation-effect prediction, yet the internal computations supporting these predictions remain poorly understood. We introduce a sparse feature circuit framework that combines sparse autoencoders, integrated-gradients attribution, and activation patching to identify the latent features that causally mediate zero-shot mutation effect prediction in ESM-2 650M. We evaluate this framework over 67 mutations ranging from strongly deleterious to weakly deleterious in the DNAJA1 J-domain, where ESM-2 predictions agree strongly with deep mutational scanning measurements. We find that circuits selected by indirect effect recover the model's predictions more efficiently and provide more informative biological explanations than those selected by raw activation changes, showing that activation magnitude does not necessarily reflect causal importance. We find that related substitutions reuse substantial portions of their recovered circuits, ranging from 40% to 75%, and that the shared features often represent residues in three-dimensional contact with the mutation site. To our knowledge, our work provides the first causal, feature-level account of zero-shot mutation effect prediction in a pLM.
bioinformatics2026-09-03v1Beyond benchmark accuracy: machine-learning turnover-number predictors require system-level validation
Rimon Martinez, M. J.; Lottermoser, J.; Bouillon, A. T. C.; Vranken, W. F.; Zehetner, L.; Zanghellini, J.Abstract
Enzyme turnover numbers (kcat) are essential for kinetic models and enzyme-constrained genome-scale metabolic models (ecGEMs), but measured values are sparse and therefore increasingly estimated using machine learning (ML). Although these predictors are commonly evaluated by global regression metrics, their practical utility depends on how errors propagate through downstream models. We benchmarked six current kcat predictors on a curated BRENDA-derived dataset and five of them on EnzyExtract. To assess the influence of training-set proximity, we compared each benchmark dataset with the available training data for each predictor. We then used the predicted kcat values to parameterize ecGEMs of Saccharomyces cerevisiae and evaluated growth predictions across 19 conditions. We find that benchmark accuracy is moderate even on the BRENDA-derived dataset and drops sharply on EnzyExtract, where all predictors achieve R2 values of 0.20 or lower. This decline is accompanied by substantially lower overlap between the benchmark and training datasets, with exact sequence matches ranging from 24% to 78% for BRENDA, compared with 9% to 26% for EnzyExtract. However, that overlap alone does not explain differences in generalization across predictors. Moreover, downstream performance is also not explained by benchmark ranking. Across 19 conditions, none of the tool-specific ecGEMs consistently reproduces the experimentally observed variation in growth. In glucose minimal medium, the weakest benchmark performer yields the most accurate growth prediction in the downstream ecGEMs, whereas higher-ranked predictors produce larger deviations in growth. We trace this mismatch to localized errors at high-leverage positions in yeast's metabolic network, where underpredicted mitochondrial ADP/ATP carrier turnover numbers restrict adenine nucleotide exchange and impose an apparent limitation on cytosolic ATP supply. Relaxing this constraint shifts predicted growth toward the experimental reference. Thus, ML-derived kcat values can affect not only quantitative growth predictions but also the phenotype a mechanistic model appears to identify. These results argue for application-driven validation of biological parameter predictors in the downstream systems they are intended to support.
bioinformatics2026-09-03v1Non-Covalent Poly(ADP-ribose) Signaling Organizes a Circadian E3 Ligase Network in the Brain
Harikrishnan, S.; Kang, S.-U.Abstract
Although PAR biology has traditionally been studied through covalent PARylation, non-covalent PAR-binding proteins provide an additional mechanism for interpreting transient PAR signals and converting them into downstream regulatory programs. Among these effectors, E3 ubiquitin ligases are uniquely positioned to couple PAR sensing to selective ubiquitination, thereby integrating stress signaling with proteostatic control. Because circadian systems depend heavily on temporally coordinated protein turnover, we hypothesized that PAR-binding E3 ligases may form a circadian-structured regulatory layer within the brain. To test this, we integrated GTEx v10 brain transcriptomics, GWAS Catalog gene-mapped associations, CIRCA circadian phase annotations, and Human Protein Atlas single-cell transcriptomic resources to characterize the organization of PAR-binding E3 ubiquitin ligases across neural tissues. Across the brain, the E3 ligase repertoire was broadly deployed yet regionally structured, with cerebellar and cortical enrichment patterns preserved within the PAR-binding subset. Representative ligases spanning circadian regulation, DNA repair, and neurodegeneration-relevant pathways displayed distinct abundance and regional-variability archetypes across GTEx brain regions. Human genetic analyses demonstrated that E3 ligases associated with cognition-, neurodegeneration-, and sleep/circadian-related phenotypes were disproportionately PAR-binding, supporting convergence between PAR-responsive ubiquitin regulation and disease-relevant biology. Circadian phase analyses further revealed that PAR-binding ligases occupy structured, non-random circadian windows within the broader E3 background, including distinct co-phasing relationships with BMAL1 and CRY1. Finally, cell-type enrichment analyses identified microglia as the dominant compartment for circadian-linked and PAR-binding circadian E3 weighting within the brain E3 program. Together, these findings support a systems-level framework in which non-covalent PAR-binding E3 ubiquitin ligases constitute a brain-deployed, circadian-organized regulatory layer that couples PAR signaling to time-dependent ubiquitin control in neural systems.
bioinformatics2026-09-03v1Morphologic intratumoral heterogeneity from routine whole-slide histopathology is prognostic for survival in primary central nervous system lymphoma: development in the LOC Network and international external validation
Rincon de la Rosa, L.; Barillot, N.; Hernandez-Verdin, I.; Velasco, R.; Mathon, B.; Eimer, S.; Vignes, J. R.; Rousseau, A.; Paillassa, J.; Ahle, G.; Lerintiu, F.; Uro-Coste, E.; Oberic, L.; Tabouret, E.; Appay, R.; Gauchotte, G.; Taillandier, L.; Marolleau, J.-P.; Adam, C.; Ursu, R.; Cuzzubbo, S.; Schmitt, A.; Nichelli, L.; Pons-Escoda, A.; Charlotte, F.; Davi, F.; Le Garff-Tavernier, M.; Choquet, S.; Soussain, C.; Vidal, N.; Gonzalez-Barca, E.; Climent, F.; Lopez, P.; Drieux, F.; Veresezan, E.-L.; Heming, M.; Meyer zu Hörste, G.; Grauer, O.; Jardin, F.; Mokhtari, K.; Houillier, C.; Hoang-XuaAbstract
Background: Clinical scores incompletely capture outcomes in primary central nervous system lymphoma (PCNSL). We quantified morphologic heterogeneity in pretreatment hematoxylin and eosin (H\&E) whole slides. Patients and methods: Three independent cohorts of immunocompetent, HIV- and EBV-negative patients treated recently were analyzed: LOC 2023 (122 slides), phase III BLOCAGE-01 (245 slides; NCT02313389), and external Barcelona (BCN; 41 slides). UNI embeddings, prototype learning, spatial metrics, and elastic-net Cox regression defined ITH-C. Results: Models achieved bootstrap-corrected concordance of 0.797--0.834. Age-, sex-, and KPS-adjusted ITH-C HRs were 1.29 (95\% CI 1.01--1.64), 1.27 (1.07--1.51), and 2.13 (1.35--3.37), respectively. Adding ITH-C increased MSKCC C-index from 0.671 to 0.717, 0.560 to 0.593, and 0.588 to 0.706. Spatial transcriptomics linked ITH-C to immune programs. Conclusions: Routine H\&E encodes prognostic spatial heterogeneity in PCNSL. ITH-C complements clinical scores, supporting prospective risk stratification.
bioinformatics2026-09-03v1An alignment-last approach enables rapid transcriptomic biomarker discovery in large cohorts
Narmanli, E.; Lanau, A.; Neacsu, M.; Koshkina, M. K.; Fumeron, P.; Martin, P.; Servant, N.; Perrin-Gilbert, N.; Waterfall, J. J.Abstract
Canonical transcriptomic analysis requires committing from the outset to a reference genome or transcriptome, which imposes a predefined feature set, usually annotated genes or isoforms. Alignment and annotation dilute the signal through feature-level aggregation, discard any sequence absent from the reference, and require reprocessing the entire dataset for each new question (mutations, fusions, transposable elements). Here, we introduce the alignment-last paradigm, in which the read becomes the unit of comparison across samples, and alignment is deferred to annotate only the relevant sequences. Querying the merome, a reference-free cohort k-mer index, with just a handful of reads (about 0.01% of a sample's) reveals the cohort's transcriptomic structure in bulk and single-cell data. At single-cell resolution, these reads outperform genes for cell classification and rediscover, without supervision, a transposable-element signature (VL30) of exhausted T cells. Finally, unsupervised read-level differential analysis recovers established lncRNA biomarkers; uncovers new prognostic transposable-element reads in adrenocortical carcinoma and sarcomas; and extracts signals even from reads that fail to align.
bioinformatics2026-09-03v1Macrophage signature-based prediction of cancer treatment response using MIL-attention
Madgwick, M.; Witham, S.; Occhetta, M.; Haneklaus, M.; Camanzi, B.; Smyrnakis, M.; Gardiner, L.-J.Abstract
Predicting immunotherapy response from single-cell data remains difficult due to patient-level labels, extreme class imbalance, and highly heterogeneous macrophage states. We present a Multiple Instance Learning (MIL) framework that treats each patient as a bag of macrophage embeddings derived from a single-cell RNA foundation model. The architecture incorporates an attention-based pooling mechanism with reduced model complexity, dropout-enhanced regularization and explicit attention penalties to improve stability in small-sample regimes. To address imbalanced clinical datasets, MIL outputs are optimized with a combined focal loss and supervised contrastive objective that simultaneously sharpens class boundaries and improves representation clustering. Across three cancer datasets, this approach outperforms pseudobulk aggregation, embedding baselines and standard MIL variants. Attention-weighted attribution and transcriptional regulatory analysis reveal distinct macrophage programs, interferon and antigen-presentation networks in responders versus hypoxia-linked regulatory modules in non-responders. This shows the potential of MIL to uncover predictive and mechanistically interpretable immune states.
bioinformatics2026-09-03v1DAG-HEART: Directed Acyclic Graph-Guided Health Equity-Aware Representation Transfer Learning Framework for Breast Cancer
Baek, M.; Wang, J.; Wan, S.Abstract
Breast cancer outcome prediction remains challenging for underrepresented populations because genomic datasets are demographically imbalanced and conventional multi-omics integration largely relies on undirected molecular similarity. We developed DAG-HEART, a directed acyclic graph-guided multi-omics transfer-learning framework that extends our previous transfer learning strategy with data augmentation. Using TCGA-BRCA mRNA, miRNA, and DNA-methylation data, DAG-HEART was evaluated for progression-free interval prediction in a data-minority group. DAG-guided nonlinear integration consistently improved predictive performance relative to direction-agnostic and correlation-based representations, while biologically motivated directional constraints generally outperformed reversed or unconstrained structures. Recurrently selected features converged on extracellular-matrix and regulatory pathways and supported clinically meaningful risk stratification. DAG-HEART provides an interpretable strategy for combining directed multi-omics structure with transfer learning under data imbalance across racial groups.
bioinformatics2026-09-03v1Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach
Kalkhoven, E.; Baak, R. E.; Hooiveld, G. J. E. J.; Schrauwen, P.; Hoeks, J.; Raymakers, R.; van der Stolpe, A.; Kersten, S.Abstract
Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.
bioinformatics2026-09-03v1spatialMET: an open and scalable framework for spatial metabolomics analysis
Mekonnen, Y. A.; Ospina, O. E.; Rubio, V.; Welsh, E.; Uddin, R.; Ackerman, H. D.; Soupir, A.; Cox, J. E.; Fridley, B. L.; Flores, E. R.; Koomen, J.; Stewart, P. A.Abstract
Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.
bioinformatics2026-09-03v1CROWN: Curated Repository Of Well-resolved Noncovalent interactions
Poelmans, R.; Van Eynde, W.; Bruncsics, B.; Bruncsics, B.; Arany, A.; Moreau, Y.; Voet, A. R.Abstract
The development of machine-learning models for protein-ligand interactions is constrained by the quality and diversity of the available structural data. Existing resources force researchers into a trade-off: carefully curated collections such as PDBBind and HiQBind offer high structural reliability but cover only a narrow slice of the Protein Data Bank (PDB), whereas large-scale resources such as PLINDER provide broad coverage with minimal quality control. We present CROWN (Curated Repository Of Well-resolved Non-covalent interactions), a machine-learning-ready dataset that reconciles scale and rigor through a fully automated preprocessing pipeline. Starting from the PDB database, CROWN applies a series of interleaved quality filters and processing stages that address crystallographic resolution, ligand identity, pocket completeness, structural repair, interaction quality, and protonation at physiological pH. The pipeline finishes with a constrained energy-minimization step built on custom flat-bottomed restraints - a step absent from all existing protein-ligand datasets - that balances crystallographic evidence against the relaxation of intramolecular strain. By reconciling the heterogeneous refinement practices of different depositions without distorting the experimentally observed binding geometry, this step yields a structurally uniform collection of 178,263 complexes, representing a roughly four-fold increase in protein diversity over PDBBind and HiQBind. Rather than organizing the data around sparsely available, bias-prone binding affinities, CROWN adopts a geometry-centric design philosophy that treats the three-dimensional arrangement of atoms at the binding interface as a self-consistent source of information. To demonstrate its value as a training resource, we trained two knowledge-based scoring functions on CROWN and benchmarked them on CASF-2016: relative to HiQBind-trained counterparts, CROWN-trained models showed markedly improved ranking power (mean Spearman correlation rising from 0.509 to 0.637) and docking power (top-1 near-native pose recovery of 0.785 versus 0.724). Because CROWN imposes no requirement for affinity labels, it can in principle support any model that learns from or is evaluated against protein-ligand complex structures. We anticipate that it will serve as a broadly useful resource for tasks such as the training of binder generation, protein design or protein folding models conditioned on bound ligands, the development of scoring functions or benchmarking of interaction-prediction methods.
bioinformatics2026-09-02v2DESPOT: Direction-Enhanced Scoring POTentials
Poelmans, R.; Bruncsics, B.; Arany, A.; Van Eynde, W.; Shemy, A.; Moreau, Y.; Voet, A. R.Abstract
Knowledge-based potentials (KBPs) remain among the most reliable and interpretable scoring functions for protein-ligand interactions, yet most share two structural limitations. They assume that the space around each protein atom is isotropic, and their interaction-conditioned reference state cannot represent regions of space that are preferentially left empty. We introduce DESPOT (Direction-Enhanced Scoring POTentials), an all-atom anisotropic KBP that overcomes both. DESPOT classifies atoms into isotropic, axially symmetric, and fully anisotropic symmetry classes from their hybridization and bonding environment, and discretizes the surrounding interaction space using the according symmetry. By adopting a positionally averaged reference state and using a void ligand atom type, it learns, for every point around a protein atom, the probability that the point is occupied by a given ligand atom type or preferentially left empty - a ligand-independent description that naturally encodes steric exclusion. This occupancy-conditioned potential captures the precise, atom-level placement of ligand atoms; we pair it with a complementary geometry-conditioned, residue-level formulation (DESPOT-screen, in the spirit of KORP-PL) and combine the two inverse-Boltzmann scores into a consensus score, DESPOT-combo. Derived from 110,943 curated, energy-minimized complexes drawn from the CROWN database and evaluated on the CASF-2016 benchmark, DESPOT achieves competitive scoring power (Pearson r = 0.61), while DESPOT-combo attains best-in-class docking power (89.5% top-1 success); all anisotropic DESPOT variants significantly outperform isotropic KBPs and established empirical scoring functions in virtual screening. Anisotropy is decisive for rejecting geometrically implausible poses, and uniting the atom-level precision of DESPOT with the implicit flexibility tolerance of the residue-level score yields the most consistent performance across tasks. Because the same occupancy-conditioned potentials can be evaluated over an empty grid, DESPOT generates molecular interaction fields as well, unifying pose scoring with direction-aware binding-site characterization within a single interpretable model.
bioinformatics2026-09-02v2Accessible and reproducible deployment reveals the practical boundaries of single-cell foundation models
Hou, S.; Yang, P.; Ma, W.; Xiang, J.; Wang, J. X.; Wan, H.; Ma, Y.; Zhou, X.Abstract
Single-cell foundation models (scFMs) have been widely promoted as a unifying paradigm for transcriptomic analysis, yet whether large-scale pretraining translates into reproducible biological advantages remains unclear. Their adoption is further hindered by heterogeneous implementations, preprocessing requirements, and computational environments. Here we develop a unified, automated, and reproducible framework for standardized deployment and controlled evaluation of scFMs across datasets, computational environments, training regimes, and downstream analyses, substantially lowering the technical barriers to their use. Leveraging this framework, we systematically investigate thirteen scFMs alongside established methods across nearly one hundred datasets spanning diverse biological contexts. Our analyses reveal clear practical boundaries to scFM utility. First, increased model scale, architectural complexity, pretraining corpus size, or input encoding does not consistently translate into superior downstream performance. Instead, measurable properties of embedding geometry provide a model-agnostic, representation-level explanation for differences in zero-shot performance across diverse model families. Second, the benefits of pretrained representations depend strongly on the biological and supervision regime: scFMs provide their clearest advantages under extremely limited supervision, particularly for rare-cell annotation and open-set detection of source-absent cell states, whereas established methods remain competitive or preferable in most other settings. Task-matched analyses further show that scFM representations transfer inconsistently to spatial-domain recovery, while their gene embeddings capture broad functional relatedness without reliably recovering context-specific regulatory relationships. Together, these results establish that scFM utility is neither universal nor determined simply by model scale alone, but varies with learned representation geometry, biological context, and supervision. By combining reproducible deployment with large-scale empirical and mechanistic investigation, our framework provides a principled foundation for determining when foundation-model pretraining offers genuine practical value and when simpler approaches remain sufficient.
bioinformatics2026-09-02v2Interactive downstream proteomics analysis with MiraProt using Mueller cell proteomes from equine recurrent uveitis
Schmalen, A.; Fleischer, A. B.; Riedel, B. M.; Deeg, C. A.Abstract
Mass spectrometry-based proteomics requires downstream analysis of processed protein abundance data, including data inspection, filtering, statistical testing, functional enrichment, protein set comparison, network analysis, and visualization. MiraProt was developed as a modular, metadata-aware R Shiny platform that integrates these steps in a single interactive workflow for processed protein-level proteomics data. Its metadata-aware design enables identifiers, sample information, experimental conditions, transformations, and derived data columns to be defined during data preparation and reused consistently across downstream analyses. To demonstrate its use, we reanalyzed a previously published label-free proteomic dataset of primary retinal Mueller cells from healthy horses and horses with equine recurrent uveitis (ERU). ERU is a naturally occurring autoimmune eye disease of horses characterized by recurrent intraocular inflammation triggered by autoreactive T-cells. Mueller cells are specialized retinal macroglia with various functions such as maintaining retinal ion homeostasis and supporting retinal neuron metabolism. Of 193 proteins with an adjusted p-value [≤] 0.05, 187 also showed at least a twofold abundance difference between ERU-derived and control Mueller cells. Functional enrichment highlighted nuclear RNA processing, chromatin-associated structures, DNA and RNA binding, interferon responses, and cell-cycle-associated programs. Gene set enrichment analysis identified positive enrichment of Interferon Alpha Response, Interferon Gamma Response, and MYC-, E2F-, and G2M-associated gene sets. Network analysis of shared proteins further linked this signature to DNA replication, mitotic checkpoint control, and RNA processing. ERU-derived Mueller cells also showed increased abundance of MHC class II-associated proteins. Together, these findings identified an interferon-responsive, cell-cycle-associated, and MHC class II-associated Mueller cell protein signature in ERU and generated experimentally testable hypotheses for further mechanistic studies. MiraProt provides an accessible, metadata-aware framework for reproducible downstream exploration of processed proteomic datasets and prioritization of candidate proteins and pathways for experimental follow-up.
bioinformatics2026-09-02v1High-Resolution Subtyping of Pediatric Low-Grade Glioma Using an Integrated Meta-Clustering Framework
Tuerhanbayi, B.; Wang, J.; Wan, S.Abstract
Pediatric low-grade glioma (pLGG) is the most common type of brain tumor in children, accounting for approximately 30% of all central nervous system tumors in children. pLGG has multiple molecular subtypes that differ in disease progression, recurrence patterns, and treatment responses. Conventional wet lab approaches including molecular profiling and histopathological studies for pLGG characterization are time consuming, costly, and laborious. Recently, methods based on artificial intelligence (AI) or machine learning (ML) have been widely used for pLGG molecular categorization, but most of them can only identify two or three pLGG subtypes. To more comprehensively characterize the molecular subtypes of pLGG and their potential biological and therapeutic significance, we develop an integrated meta-clustering approach, namely Meta-pLGG, that can explore high resolution molecular subtypes and their transcriptional heterogeneity for pLGG. Specifically, we first performed multiple rounds of random projection (RP) to generate dimension-reduced feature vectors from pLGG transcriptomics data, each of which was subsequently clustered by different clustering algorithms including hierarchical clustering, K-means, Self-Organizing Maps (SOM), Non-negative Matrix Factorization (NMF), Gaussian Mixture Model (GMM), and Spectral Clustering, as base clustering methods. Then, to yield robust clustering performance, we integrated the clustering results of these RP based individual clustering algorithms by adopting a weighted meta-clustering (wMetaC) approach. Results based on 532 pLGG patients suggested that our proposed approach demonstrated superior stability and discriminative powers for higher resolution pLGG subtyping compared to conventional approaches. Based on consensus matrix analysis, we identified two major pLGG mega-subtypes, with one further subdivided into three subgroups and the other into two. Then, we performed cluster specific differential gene expression analysis, molecular pathway analysis, and gene-drug-disease association analysis. The results showed that the identified five subgroups exhibited significant subtype-specific transcriptomic heterogeneity. In summary, our meta-clustering approach demonstrated much higher performance and robustness in identifying higher resolution molecular subtypes of pLGG, revealing the molecular heterogeneity within pLGG and potentially providing new insights for more precise molecular subtyping and precision therapy.
bioinformatics2026-09-02v1