Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
Target Preference Maps: A machine learning model generalizing transferable drug-receptor interactions and guiding drug discovery
Menezes, F.; Wahida, A.; Froehlich, T.; Grass, P.; Zaucha, J.; Napolitano, V.; Siebenmorgen, T.; Pustelny, K.; Barzowska-Gogola, A.; Rioton, S.; Didi, K.; Bronstein, M.; Czarna, A.; Hochhaus, A.; Plettenburg, O.; Sattler, M.; Nissen-Meyer, J.; Conrad, M.; Kurzrock, R.; Popowicz, G. M.Abstract
Modern AI models can decode the genomic landscape and protein structure world. Yet, they fail to generalize to one of the most important fields: small-molecule drug discovery. Since the late 1970s, the advent of macromolecular crystallography inspired the notion that structural knowledge alone could enable a lock-and-key approach to drug design. However, drug discovery continues to depend on costly, resource-intensive, and largely serendipitous screening campaigns that probe only an infinitesimal fraction of the drug-like chemical space. Despite some successful cases, our understanding of, and reasoning from, non-bonded interaction chemistry remains limited for general applicability. Furthermore, though structural databases contain hundreds of thousands of entries, a strong historical bias pervades protein-drug structures, hindering reliable advances through AI scaling. Here, we present a machine-learning framework that learns atom-type-specific spatial preference maps from local protein microenvironments in protein-ligand structures. By excluding ligand topology from the model input and learning from local atom-level environments, the framework is designed to reduce dependence on whole-ligand memorization and to capture transferable interaction preferences. The resulting maps recover chemically meaningful interaction patterns, including cases involving bridging waters and metal-dependent environments. The model was validated using retrospective and prospective real-world data in drug optimization when targeting a challenging protein-protein interface. This shows that the method can provide interpretable workflows to guide molecule optimization and provide input for downstream generative or docking workflows.
bioinformatics2026-07-21v11LeafRank: A phylodynamic framework for inferring relative fitness from single-cell phylogenies in chromosomally unstable tumors
Wu, C.; Leder, K.; Wang, Z.; Sun, R.Abstract
Tumors contain cancer cells with diverse growth potentials that shape evolutionary trajectories, yet this fitness diversity remains difficult to quantify in cases of whole-genome duplication (WGD) and chromosomal instability. We present LeafRank, a mathematical framework that leverages single-cell DNA-seq phylogenies to infer the relative fitness of individual cells. Using a multi-type branching process model, LeafRank integrates full tree topology, including branch lengths and bifurcation patterns, to estimate marginal fitness probabilities under punctuated evolutionary regimes driven by rare driver events. To account for elevated aberration rates following WGD, we introduce a tree-rescaling strategy that adjusts for lineage-specific genomic instability. Unlike methods focused on predefined subclones, LeafRank ranks all sampled cells, enabling flexible assessment of growth heterogeneity. Simulations demonstrate high accuracy across spatial and non-spatial virtual tumors. Applied to ovarian cancer, LeafRank reveals directional and parallel selection in WGD tumors and identifies recurrent copy number events enriched in high-fitness lineages. WGD lineages do not show immediate growth advantages but acquire fitness through subsequent alterations.
bioinformatics2026-07-21v4RNAStabFormer: Region-Aware Multi-Task Hybrid Learning for RNA Stability Prediction from Pulse-Chase Transcriptomics
Wang, S.; Zhang, C.Abstract
RNA stability is a major post-transcriptional regulator of gene expression, yet sequence-based prediction from pulse-chase transcriptomics remains difficult because labels depend on the time window, quantification region, and replicate quality. We present RNAStabFormer, a controlled RNA stability framework centered on a Region-Aware Multi-Task Hybrid Transformer (RAMHT). RAMHT encodes the nucleotide context of the 5-prime UTR, coding sequence (CDS), and 3-prime UTR; incorporates an additional CDS codon stream; upgrades engineered sequence features into a tabular interaction branch; and uses gated multi-task regression to predict four ENCODE BrU-seq/BruChase-seq RNA stability proxies, with Exon 6 h/0 h as the primary task. Across 26 outer data splits, including 23 chromosome holdout tests, a heterogeneous three-member RAMHT ensemble achieves a mean Pearson correlation of 0.773 on the primary task, statistically matching an engineered-feature XGBoost baseline, which also achieves 0.773. The mean paired difference is +0.000004, with a bootstrap 95 percent confidence interval from -0.003845 to +0.004077 and a Wilcoxon p-value of 0.8613. The ensemble improves upon the strongest individual RAMHT member, increasing the mean Pearson correlation from 0.768 to 0.773 and outperforming it on 23 of the 26 data splits. A strictly nested XGBoost-RAMHT blended model further increases the correlation to 0.775. Evaluations conducted on identical data splits also show that the ensemble outperforms frozen full-length mRNA language-model embeddings, which achieve a correlation of 0.760, and public LAMAR-DR transfer learning, which achieves a correlation of 0.180. Gate analysis, ablation experiments, and sequence recoding analyses indicate that engineered sequence grammar remains the dominant source of predictive information, whereas the nucleotide and codon branches provide complementary signals localized primarily within the CDS. RNAStabFormer narrows the performance gap between neural RNA sequence models and strong tabular baselines while retaining an extensible architecture for model interpretation and biological data integration.
bioinformatics2026-07-21v2Panomap: Unbiased Nanopore Signal Mapping with Pangenome Variation Graphs
Shih, P. J.; Sanghani, Z.; Guarracino, A.; Gamaarachchi, H.; Batten, C.Abstract
Motivation: Signal-space nanopore mappers enable real-time mapping and filtering decisions directly from raw nanopore signals. However, existing signal-space mappers are built around linear references, and using a single representative reference can introduce reference bias when the sample diverges from that reference. Pangenome reference collections can reduce this bias by representing diversity across related reference sequences, but linear-reference signal mappers must treat each sequence as a separate target, redundantly storing shared sequences. Pangenome variation graphs provide a more compact representation by storing shared sequences once and encoding variants as alternative paths through the graph. Although sequence-to-graph mapping is well established for basecalled reads, existing signal-space methods do not directly use pangenome variation graphs. Results: We present Panomap, the first signal-space mapper that operates on pangenome variation graphs. Panomap maps raw nanopore signals to graph references, allowing signal-space mapping to use pangenome diversity while representing shared sequences once. We evaluate Panomap in three settings. First, when a single reference already maps the sample well, Panomap preserves mapping accuracy as additional reference sequences are added to the reference collection, while state-of-the-art signal-space tools regress. Second, when the exact sample strain is absent from the reference collection, Panomap benefits from adding related assemblies from the same species to the pangenome reference. Third, using a highly polymorphic locus, we show that Panomap can map reads from alleles not represented in the reference collection by using related alleles in the pangenome, with the largest gains for more divergent alleles and for decisions made from short prefixes of the read signal. In addition, Panomap's graph index scales sublinearly with pangenome collection size. Together, these results show that Panomap brings population-aware reference representation into signal-space mapping. Availability and Implementation: Panomap is open source and available at https://github.com/cornell-brg/panomap.
bioinformatics2026-07-21v2Genomic, Transcriptomic, and Regulomic Analyses Do Not Support Profound Autism as a Distinct Biological Category
Eicher, T. D.; Ne'eman, A.; Quackenbush, J. D.Abstract
The Lancet Commission on the Future of Care and Clinical Research in Autism proposed the construct of "profound autism" as a recognizable subtype of autism. Supporters argue that this classification is necessary to ensure that autistic persons with severe impairment receive appropriate research attention and policy support, whereas critics contend that the construct lacks scientific validity and may reflect social or political considerations more than biological distinction. To inform this debate, we evaluate whether the proposed "profound autism" category represents a distinct genetic phenotype using multiple molecular data types collected in a large cohort. Across genomic, transcriptomic, and regulatory analyses, we find no evidence supporting "profound autism" as a biologically distinct phenotypic group. Instead, differences emerge primarily in inferred gene regulatory networks distinguishing nonspeaking from speaking autistic children, suggesting potential regulatory mechanisms contributing to speech ability. These findings suggest that future research into severe impairment may be more productive if focused on specific traits -- such as speech impairment -- rather than attempting to define a distinct biological subtype within the multidimensional phenomenon of autism.
bioinformatics2026-07-21v2SEEK-VEC: Augmenting topic modeling with spectral ensemble learning
Danning, R.; Ke, Z. T.; Ma, R.; Lin, X.Abstract
Count data are ubiquitous across many applications in which understanding latent patterns is of interest. Topic modeling is a powerful tool for detecting latent structure in count data. However, standard topic modeling methods are often constrained by their restrictive assumptions, susceptible to noise, and sensitive to misspecification of the number of topics. Here, we introduce SEEK-VEC (Spectral Ensembling of topic models with Eigenscore for K-agnostic Vocabulary Embedding and Classification), an ensemble topic modeling framework that integrates insights from multiple candidate topic models through a spectral ensembling procedure. SEEK-VEC produces a meta-structure matrix containing prioritization scores and grouping scores that enable variable classification, interactive pattern discovery, and model diagnostics. Through simulations, we demonstrate that SEEK-VEC augments the performance of standard topic models for identifying important vocabulary words and understanding the relationships among them, particularly when signal strength is weak. We apply SEEK-VEC to the MADStat dataset of statistical abstracts and demonstrate its utility for evaluating the proposed interpretation of a topic model.
bioinformatics2026-07-21v2Molecular Mimicry in Inflammatory Bowel Disease: Multi-layered Functional and Sequence-level Analysis of Gut Microbial Proteins Mimicking the Human Proteome
Anand, A. A.; Mishra, P.; Srivathsa, V. S.; Yadav, V.; Samanta, S. K.Abstract
Inflammatory bowel disease (IBD), encompassing Crohn's disease (CD) and ulcerative colitis (UC), is a chronic inflammatory disorder whose pathogenesis involves intricate host-microbiome interactions. Molecular mimicry (the structural or functional resemblance between microbial and host proteins) represents a plausible mechanism by which gut microbiota may trigger or perpetuate autoimmune responses. Here we present a comprehensive, multi-layered molecular mimicry in silico pipeline (MMIP) analysis of 39 baseline shotgun metagenomic samples from Human Microbiome Project 2 (HMP2/IBDMDB). Using DIAMOND-based homology search against the SwissProt followed by UniProt Retrieve/ID Mapping (URIM), we characterized microbial protein functional space through three complementary frameworks: (i) normalized GO term frequency comparison across diagnostic groups, (ii) protein family (PFAM) domain enrichment analysis, and (iii) sequence-level mimicry analysis identifying microbial proteins with direct homology to human proteins. Taxonomic profiling using MetaPhlAn3 pre-computed abundance profiles provided further biological context. CD microbiomes exhibited greater enrichment of immune-relevant biological processes and a higher per-sample sequence-level mimicry rate than UC. CD and UC also showed distinct pathobiont profiles, with CD enriched for oral-origin taxa including Haemophilus parainfluenzae and UC enriched for Fusobacterium nucleatum and Prevotella species. Notably, healthy gut microbiome maintains coordinated mimicry of host neuronal, signal-recognition-particle-associated, and antimicrobial peptide machinery, a repertoire dismantled in IBD and replaced by disease-subtype-specific signatures, alongside a candidate mimicry link between CD-enriched bacteria and NOD2. These results represent the first metagenome-wide characterization of molecular mimicry across IBD subtypes using shotgun metagenomic data, offering new mechanistic insight into how microbial dysbiosis may contribute to immune dysregulation in IBD. Keywords: Inflammatory bowel disease (IBD), Crohn's disease (CD), Ulcerative colitis (UC), Gastroenteritis, Molecular mimicry
bioinformatics2026-07-21v2COSMOS: A FAIR-aligned infrastructure for clinical trial data validation, warehousing, and interactive discovery
Roberts, C. A.; Nilsson-Takeuchi, A.; Stuart, C.; Chivers, M.; Soares, P.; Foure, V.; Griffiths, G.; Niazi, U.Abstract
The Clinical Omics System for Metadata and Outcome Storage (COSMOS) is an open-source, FAIR- and GCP aligned clinical trial unit (CTU) infrastructure designed to streamline the transition of academic clinical and multi-omic trial datasets into curated, analysis-ready repositories within Secure Data Environments (SDEs). By integrating an automated Data Quality and Data Validation (DQ&DV) "Trust Layer" with a relational Structured Query Language (SQL) schema, COSMOS enables programmatic and interactive data access via R Shiny applications. Dynamic integration of omics and clinical data is achieved through analytical data structures (e.g. ExpressionSet objects) linked via relational database identifiers. This lowers technical barriers for researchers and promotes governed data reuse and secondary discovery.
bioinformatics2026-07-21v1Bayesian Factor Analysis for Binary and Ordinal Phenotypes with Missingness
Shashaank, N.; Knowles, D. A.Abstract
Binary and ordinal phenotypes are common in clinical screening and self-reported questionnaires, but many factor analysis and matrix factorization methods are only applicable for quantitative phenotypes with real-valued and/or continuous data distributions. To address this, we propose FABOr (Factor Analysis of Binary and Ordinal data), a Bayesian framework for matrix factorization in which the low-rank matrices are latent variables with continuous priors while the phenotypes are observed variables modeled with appropriate binary/ordinal likelihoods. We also develop missing not at random (MNAR) extensions of FABOr for analyzing data with structured missingness. In experiments with simulated phenotypes, we found that FABOr performs similarly to the best-performing benchmark methods on binary data at imputation and exceeds the performance of all tested benchmark methods on ordinal data. We then applied FABOr to analyze real-world binary and ordinal phenotypes from the Simons Foundation SPARK dataset on autism spectrum disorder (ASD) and found that it improved imputation accuracy by up to 5% on binary data and up to 23% on ordinal data relative to the benchmark methods.
bioinformatics2026-07-21v1A new set of DNA methylation variants in the human genome show predominant tissue specificity and sensitivity to reprogramming with a potential for disease susceptibility.
Anne, A.; Kumar, L.; Singh, M.; Choudhury, S.; Das, S.; Zimmer-Bensch, G.; Bandyopadhyay, D.; K, N. M.Abstract
Analyses of 3,370 normal human tissues of ectodermal, endodermal and mesodermal origins identified 12,587 regions averaging ~585 bp with significant differences in DNA methylation levels within identical tissues. These methylation variants (MeVars) occurred in 8,037 genes enriched in neurological disorders and cancers of which, majority were tissue-specific rather than being systemic. This somatic variation was reduced by reprogramming in vitro into iPSCs and in vivo during spermatogenesis. Analysis of prefrontal cortices showed a higher incidence of MeVars in the candidate genes in controls than schizophrenia patients wherein a subset showed significantly altered transcript levels. Similar effects were observed for oral tissues and skin fibroblast cells. MeVars showed significant association with SINE1, simple and low complexity repeats, H3K27me3, H3k9me3 and H3K4me1 modifications and EZH2, SUZ12 and REST binding sites. Collectively, MeVars have postzygotic origins with an ability to reset during reprogramming, adding a new dimension in the form of epigenetic diversity and its relevance to disease susceptibility in humans.
bioinformatics2026-07-21v1ProtSyntax: a protein large language model for decoding post-translational modification syntax and function
Lin, Y.Abstract
Post-translational modifications (PTMs) regulate protein function through dependencies among residue chemistry, sequence context, three-dimensional microenvironments and modification states, yet most predictors model sites independently and do not connect modification propensity to functional consequences. Here we present ProtSyntax, a PTM-centered protein language model trained on 4.25 million examples spanning 40 PTM classes and supervised for kinase specificity, PTM crosstalk and enzyme kinetics. ProtSyntax integrates bidirectional long-range modeling with geometry-gated attention in a sparse mixture-of-experts architecture and uses adaptive multi-objective learning to couple residue-level PTM syntax to protein-level function. Across 40 PTM-site benchmarks, ProtSyntax improved mean MCC and AP by 12.7% and 10.7%, respectively, relative to the best-performing baselines. It also distinguished authentic sites from structurally incompatible motif decoys, transferred to rare PTMs, recovered crosstalk, linked PTM perturbations to enzyme-kinetic changes and identified disease-associated PTM disruptions. Together, ProtSyntax provides an interpretable framework for decoding PTM regulation across the proteome.
bioinformatics2026-07-21v1PepCL: A replay-based continual learning framework for updating peptide-MHC models
Chati, P. M.; Lashkari, V. D.; Salhotra, A.; Bruno, P. M.; Ntranos, V.Abstract
Understanding peptide-major histocompatibility complex (MHC) class I binding is critical for effective vaccine and immunotherapy design but is a combinatorially complex challenge for which prediction models have become essential. MHC ligands are typically identified at scale via untargeted mass spectrometry (MS), and this has built a strong base for peptide-MHC model training. However, MS incompletely captures the vast peptide-MHC space due to technical, sampling, and biological biases. Although recently developed experimental assays have queried such blind spots yielding complementary information, existing peptide-MHC predictors have not yet incorporated these orthogonal data and are not designed to be updated as new data are generated. Here, we introduce PepCL (Peptide-MHC Continual Learning), a continual learning framework for updating peptide-MHC predictors with new assay data while explicitly preserving prior MS knowledge. To enable PepCL, we also develop MHCPrime, a new state-of-the-art pan-allelic peptide-MHC prediction model, trained on publicly available MS data, that can be effectively updated under our framework. We demonstrate that PepCL allows MHCPrime to learn previously unseen, assay-specific information while preventing catastrophic forgetting that is typically observed with conventional fine-tuning. We evaluate PepCL and MHCPrime in a variety of biological contexts, including infectious disease and cancer, and show improved peptide-MHC prediction that transfers across alleles for broader applicability in clinical settings. Overall, our results establish PepCL as a flexible framework for extending the utility of peptide-MHC models by improving their predictive performance as immunopeptidomics assays continue to evolve and new data become available.
bioinformatics2026-07-21v1An automated platform for spatial functional modeling and fingerprint analysis of tissue molecular landscapes
Hajihosseini, M.; Patino-Martinez, E.; Ghosal, R.; Kaplan, M. J.; Pyne, S.Abstract
Spatial transcriptomics (ST) enables high resolution molecular profiling while preserving tissue architecture, creating new opportunities to investigate how disease-associated pathways are organized within tissues. However, existing analytical approaches largely focus on individual pathways or cell types and do not provide a unified framework for modeling spatially varying pathway interactions across tissue sections and anatomical planes. Here, we present an integrative framework, Spatial Fingerprints Analytics (SFinx), that combines reference-free deconvolution, pathway activity reconstruction, and Spatial Functional Data Analysis (SFDA) to map localized pathway activity and pathway phenotype interactions in complex tissues. Applying SFinx to10x Visium ST datasets from murine lupus nephritis, we reconstructed continuous spatial landscapes of pathway activity and disease-associated phenotypes across kidney sections. This approach identified anatomically restricted inflammatory domains characterized by coordinated activation of immune pathways and revealed substantial spatial heterogeneity in pathway crosstalk across renal compartments. Using generalized additive models with tensor-product splines, we quantified spatially varying associations between lupus nephritis and neutrophil activation pathways across tissue sections, uncovering regions with both positive and negative relationships that would be obscured by conventional bulk analyses. Multi-slice integration further demonstrated reproducible spatial interaction patterns while accounting for section-specific variability. Together, SFinx transforms mixed-spot transcriptomic measurements into interpretable spatial pathway landscapes and interaction maps, providing a general framework for identifying localized disease mechanisms. This approach reveals previously unrecognized spatial organization of inflammatory signaling in lupus nephritis and presents a broadly applicable strategy for studying spatially coordinated biological processes in autoimmune disorders, cancer, and neurodegenerative diseases.
bioinformatics2026-07-21v1Stratified Immune Profiling Uncovers Prognostic Heterogeneity Beyond MYCN Amplification and Age in Neuroblastoma
Magno, J. M.; Muzzi, J. C. D.; Resende, J. S. S.; Querne, L. B. P.; Alvarenga, L. M.; Cavalli, L. R.; Figueiredo, B. C.; Castro, M. A. A.Abstract
Neuroblastoma is the most common extracranial solid tumor in children, presenting remarkable clinical heterogeneity with survival outcomes ranging from spontaneous regression to aggressive progression. MYCN oncogene amplification and age at diagnosis are established prognostic factors that are typically treated as independent covariates in risk stratification, yet their joint influence on the tumor immune microenvironment remains poorly understood. Here we show that stratifying patients by both variables simultaneously reveals six reproducible immune subtypes with distinct transcriptional programs and prognostic significance. Consensus clustering of immunomodulatory gene expression profiles from 149 patients in the TARGET-NBL cohort identified subtypes whose survival trajectories differ significantly within clinical strata defined by MYCN status and age at diagnosis. A linear Support Vector Machine classifier trained on these subtypes, using immunomodulatory gene expression combined with MYCN amplification status and age at diagnosis as predictive features, achieved 97.2% accuracy and a Cohen's Kappa of 0.963 under 10-fold cross-validation, and generalized to an independent cohort of 493 patients (GSE62564). Kaplan-Meier analysis revealed significant survival differences across subtypes in both cohorts (TARGET-NBL: log-rank p = 0.0018; GSE62564: log-rank p < 0.0001). Single-sample gene set enrichment analysis identified differential activation of proliferative and immune response pathways across subtypes, consistent between both cohorts. These findings suggest that integrating MYCN amplification status and age at diagnosis as joint determinants of immune organization may reveal prognostic heterogeneity that is not fully captured when these factors are considered independently.
bioinformatics2026-07-21v1GeneAutomate: A Browser-Based, Integer-Indexed Platform for Dual-Gene-List Functional Annotation and Interactive Network Visualization
Singh, R. P.; Kumar, A.Abstract
Comparative interpretation of two gene lists, for example, two treatment arms, two tissues, or a discovery and a validation cohort, is a routine task in functional genomics. While several tools offer dual-list comparison (e.g., EnrichmentMap, RRHO packages), they typically require local software installation, R/Bioconductor, or manual reconciliation of separate single-list outputs. Most widely used web-based enrichment tools (DAVID, g:Profiler, Enrichr, ShinyGO, WebGestalt) are built around the analysis of a single gene list at a time, and those that support comparison often lack interactive, publication-ready visualization or depend on server-side query latency. Here we present GeneAutomate, a browser-based tool purpose-built for side-by-side comparison of two gene lists. GeneAutomate performs Over-Representation Analysis (ORA) against Gene Ontology (GO) and Reactome using an exact hypergeometric test with Benjamini-Hochberg false discovery rate correction, and Gene Set Enrichment Analysis (GSEA) when ranked (log2 fold-change) input is supplied, alongside Protein-Protein Interaction (PPI) subgraph extraction from BioGRID physical interactions. All reference data (Gene Ontology, Reactome, BioGRID, and NCBI/Ensembl identifier cross-references) are pre-compiled offline into a single integer-indexed database of approximately 32 MB for Homo sapiens, in which every gene identifier Ensembl ID, Entrez ID, official symbol, or alias is resolved to one canonical integer prior to any user query. This design removes live database round-trips from the runtime path, enabling fast, at-your-desk enrichment without installation or a server-side per-query bottleneck. The tool renders thirteen interactive, D3.js- and Cytoscape.js-based comparative visualizations, including a Rank-Rank Hypergeometric Overlap (RRHO) heatmap, a GO-slim "Radar/Spider" functional fingerprint, and chord/edge-bundled cross-talk diagrams that are, to our knowledge, not offered as an integrated set by any existing academic or commercial ORA/GSEA platform. GeneAutomate is an unfunded, individual student project developed with feedback from a professor, and is in its final stage of development. It requires no installation or login. We describe the tool's architecture, statistical methods, and comparative feature set relative to established academic tools (DAVID, ShinyGO, g:Profiler, Enrichr, WebGestalt, STRING, PANTHER, GeneMANIA, Cytoscape, clusterProfiler, GSEA, Metascape) and commercial platforms (IPA, MetaCore, Pathway Studio, iPathwayGuide, Partek Pathway), and we state candidly the current version's limitations, which are planned to be the added in next version: single-species (human-only) coverage, no upstream regulator analysis, and comparison currently limited to two (occasionally three) concurrent lists. GeneAutomate is available at https://geneautomate.tech/.
bioinformatics2026-07-21v1Extended t-cores for the de novo identification of transposable elements and other inexact repeats from short read RNAseq data
Darmon, S.; Mary, A.; Lacroix, V.Abstract
Transcribed repeats represent a major challenge in the de novo assembly of transcriptomes from short RNA-seq reads. Young transposable elements (TEs) and other inexact repeats create dense and ambiguous regions in the assembly graph, preventing the correct assembly of transcripts. In this paper, we introduce a fully de novo method based on the discovery of dense regions in the compacted De Bruijn graph (DBG) to identify such repeats directly from short reads RNA-seq data, without requiring a reference genome or repeat database. Our approach defines the extended t-cores, subgraphs of the DBG that capture the complex topology induced by highly expressed inexact repeats appearing in RNA-seq reads. Independently of its interest for transcriptome assembly, the proposed method appears to be effective for the de novo identification of repeats in transcriptomes. After classifying cores using sequence-based motifs to distinguish simple repeats from potential TEs, we demonstrate its potential for the de novo discovery of transposable elements. We validate the approach on a Mus musculus dataset using expressed TE consensus sequences, showing that extended t-cores correspond to known expressed TE families. We also illustrate its de novo discovery potential on a non-model species, Canis lupus familiaris, where the method was also able to recover known transposable elements.
bioinformatics2026-07-20v2Leveraging multiplicity in biologically informed neural networks to uncover disease heterogeneity
Gankin, D.; Beltrao, P.Abstract
Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map to named biological entities. Yet BINNs have been hard to train at biobank scale, and the reliability of their interpretations remains largely untested. Here we present a fast BINN implementation trained on UK Biobank genotype and plasma proteomics data from about 500,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph topology, because node degree and layer position influence the scores. Normalization reduces this bias but can weaken enrichment for known disease genes. Second, BINNs show substantial predictive multiplicity, that is, independently trained models with identical architecture and data reach similarly accurate solutions while prioritizing different genes and pathways. Although this multiplicity makes single-model explanations unstable, the range of interpretations can itself reveal disease biology. Across 100 replicate BINNs for type 2 diabetes, we find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways, mirroring known disease heterogeneity. Thus, analyzing the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms.
bioinformatics2026-07-20v1FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations.
Clemente-Larramendi, I.; Hillion, S.; Cornec, D.; Jamin, C.; Foulquier, N.Abstract
Single-cell RNA sequencing (scRNA-seq) enables detailed characterization of cellular heterogeneity, yet understanding the full cellular and regulatory environment of complex tissues remains challenging. In the era of large single-cell atlases, this technology has become increasingly accessible, and datasets have grown in scale and statistical power. As a result, sample representation methods have emerged as a promising strategy to summarize patient-level biological variation. However, most existing approaches rely on unsupervised learning frameworks with ambiguous biological interpretability. Here we present a Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method. FloREN models single-cell data as a heterogeneous network integrating cells and genes together with gene regulatory and cell-cell communication relationships. Through condition-aware embeddings and interpretable attention networks, FloREN enables improved sample stratification and biomarker discovery. In addition, the framework supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
bioinformatics2026-07-20v1Metagenomic profiling of bacterial endosymbionts in wild mutant and permethrin-susceptible head lice shows an expanded microbiota in the resistant strain
Mohammadi, J.; Alipour, H.; Azizi, K.; Kalantari, M.; Moemenbellah-Fard, M. D.Abstract
Background Human head lice (Pediculus humanus capitis de Geer) are known to harbor diverse maternally inherited bacterial symbionts. These endosymbiotic bacteria may contribute to insecticide degradation, potentially helping lice withstand particular environmental pressures. Using next-generation sequencing (NGS), this study investigated the bacterial symbionts present in wild head-louse populations and characterized their phylogenetic relationships in mutant and permethrin-susceptible strains. Methods Head lice specimens were collected from 10 locations across Fars Province, Iran. Following DNA extraction, the samples were analyzed using polymerase chain reaction (PCR), and the resulting amplicons were sequenced to detect mutations in specimens from each location. Lice were then classified according to the presence or absence of mutations and subjected to NGS to characterize the symbiotic bacterial communities in mutant and putatively permethrin-susceptible strains. Results Mutant strains were detected at only three sampling stations. Bioinformatic analysis of the nucleic acid sequences revealed three exons and two introns, with an expected amplicon length of 582 bp. NGS analysis showed that Candidatus Riesia pediculicola (Arsenophonus), belonging to the phylum Proteobacteria, was the predominant bacterial genus. Actinobacteria and Firmicutes were the second- and third-most abundant phyla, respectively. Most of the remaining 25 bacterial taxa were associated with the mutant strain. Additionally, two previously unreported bacterial genera were deposited in GenBank. Conclusions The distinct distribution of Arsenophonus species between susceptible and mutant head-lice strains, along with the greater abundance of Escherichia, Shigella, Lawsonella, and Megamonas in mutant strains, highlights the need for advanced metagenomic analyses to determine how these endosymbionts may help their hosts withstand specific environmental disturbances.
bioinformatics2026-07-20v1A Label-Free Multi-Metric Pipeline for Benchmarking Single-Cell RNA-Sequencing Clustering and Testing the Reproducibility of Cell-Type Heterogeneity
Yousef, Z.; Simone, J.; Klein, D.; Cho, H.; Bhuyan, A.; Bhatt, P.; Wu, J.; Wang, H.; Cai, L.Abstract
A discovered sub-population from single-cell transcriptomic data is only meaningful if it is reproducible, yet clustering is usually done with one method on one embedding and rarely tested. We present a label-free, multi-metric pipeline that reframes clustering as an auditable, methods-blind decision and separates two notions of stability that are commonly conflated: reproducibility under cell resampling (bootstrap) and reproducibility under re-embedding (retraining the representation). The pipeline evaluates seven clustering configurations across cluster counts using five non-redundant quality metrics. As a whole-dataset control on a mouse retinal atlas, it recovers an eight-cell-type annotation at 96.3% accuracy (adjusted Rand index, ARI = 0.91) without labels. We then validate the discovery mode on two cell types with opposite ground truth. On bipolar cells, which have well-established subtypes, the pipeline accepts the sub-structure: across-embedding reproducibility rises with cluster number to a high plateau (mean pairwise ARI ~0.93 near the ~15 known bipolar subtypes), with quality metrics improving in parallel. On rod photoreceptors, treated as homogeneous, it rejects over-clustering: the metric-selected partition passes a bootstrap-stability check but is not reproducible when the embedding is retrained (mean pairwise ARI = 0.69), and the metrics do not improve with cluster number. On synthetic data, the test recovers real structure down to a 5% subpopulation while rejecting null data (high sensitivity and specificity). Bootstrap stability alone is therefore insufficient evidence for sub-population; the across-embedding test discriminates real sub-structure from over-clustering and applies to any cell type as a reproducible alternative to single-method, single-embedding clustering.
bioinformatics2026-07-20v1GDTR: Layer-wise Settling Depth Reveals Biological Grammar in Genomic Foundation Models
Cho, Y.; Kang, J.; Park, S.; Kim, S.Abstract
Genomic foundation models capture sequence regularities, yet existing interpretability tools rarely ask where in the layer stack a biological grammar becomes stable. We introduce GDTR, the Genomic Deep-Thinking Ratio, a training-free residual-stream lens that assigns each nucleotide token a settling depth : the first layer at which its representation stabilizes against the post-final-norm reference. On Evo 2 7B, splice donor and acceptor sites settle approximately two layers earlier than intronic contexts; enhancer-like cCREs show a smaller but measurable shift; and a chromosome 22 calibration transfers to held-out chromosome 17. Perturbing canonical splice donors shows that the signal is bidirectional: disrupting the central GT motif deepens settling, whereas shuffling the flanking grammar makes the preserved motif settle earlier. Differential GDTR further reveals consequence-associated peak-disruption depths across ClinVar variants, with synonymous substitutions peaking deepest but showing broad class overlap. GDTR therefore provides a layer-wise interpretability axis for genomic foundation models, complementary to existing prediction and variant-scoring tools.
bioinformatics2026-07-20v1EBD-DTI: Episodic Bridge Diffusion for Zero-Shot Cold-Start Drug-Target Interaction Prediction
Liu, J.; Le, J.; Wei, C.; Liu, M.; Yin, Z.Abstract
Predicting drug-target interactions (DTI) for entirely unseen drugs or proteins---the cold-start problem---remains a critical challenge in computational drug discovery. While sequence-based methods naturally support zero-shot generalization, they often ignore relational topology, and existing graph-based approaches either rely on global diffusion that blurs the boundary between inductive and transductive evaluation or require a few known interaction samples at test time (few-shot). We present EBD-DTI, a framework that enables zero-shot inference in graph-based DTI models without requiring any known interactions for unseen entities. The key innovation is episodic cold-start training: at each epoch, a random subset of training entities is masked and treated as pseudo-cold, forcing the model to learn cold-start inference with explicit gradient supervision. A bridge-conditioned local subgraph, together with multi-hop diffusion, provides cold entities with relational context from their nearest observed neighbors. Experiments on three benchmarks (BioSNAP, BindingDB, and DrugBank) demonstrate that EBD-DTI achieves competitive or superior performance compared to state-of-the-art methods under strict zero-shot evaluation, with episodic training improving AUC by up to 12%.
bioinformatics2026-07-20v1Learnable Graph Network Model (LGNM): A Physics Constrained Graph Neural Network with Quantum Hamiltonian Learning
Sharma, B.; Sarkar, C.Abstract
Elastic Network Models (ENMs), particularly the Gaussian Network Model (GNM) and its distance-weighted variant (mENM), predict per-residue protein flexibility from C contact graphs at low computational cost. Their central limitation is the assumption of uniform spring constants, which ignores the chemical identity, burial depth, and evolutionary conservation of individual residue contacts. We introduce the Learnable Graph Network Model (LGNM), a heterogeneous ENM in which per-edge spring constants are parameterised by per-residue flexibility coefficients predicted by a physics-constrained Graph Neural Network (GNN). The GNN is trained on molecular dynamics (MD)-derived root-mean-square fluctuation (RMSF) profiles from 413 proteins in the ATLAS database, using fold-disjoint CATH superfamily splits. The learning objective is an instance of the Quantum Neural PDE (QNPDE) Hamiltonian learning framework, with K = 3 operator types enabling an O(K) quantum gradient. On 91 held-out test proteins, LGNM achieves mean per-protein Pearson correlation r = 0.8549 , versus r = 0.8024 for mENM. The implementation of this methodology is available at https://lgnm.compbiosysnbu.in/ allowing researchers to evaluate flexibility and downstream processes. Keywords: Protein flexibility; Elastic Network Model; Physics-constrained Graph Neural Network; Residue fluctuation.
bioinformatics2026-07-20v1From vehicles to wildlife: transferable deep learning for trajectory generation
Patras, J.; Fablet, R.; Brunel, A.; Roy, A.; Bugoni, L.; Tavares Nunes, G.; Barbraud, C.; Jacoby, J.; Benboudjema, S.; Passuni, G.; Delord, K.; Lanco, S.Abstract
Realistic simulation of animal movement is fundamental to conservation, habitat modeling, and ecological scenario evaluation. Traditional approaches struggle to capture multi-scale trajectory dynamics, while generative deep learning for complete trajectory simulation remains largely unexplored in ecology due to data scarcity. We show that a diffusion model pre-trained on millions of vehicle GPS trajectories can be fine-tuned on hundreds of seabird central-place foraging trips to simultaneously generate ecologically realistic trajectories for five species and six breeding colonies. The fine-tuned model consistently outperforms four state-of-the-art baselines (GAN, VAE, HMM, iSSF) across movement metrics spanning step dynamics, spatial distribution, and behavioral temporality, with comparable or shorter computation times. Domain transfer reaches full performance in 35 minutes versus 7 hours from scratch, and conditioning on species and colony enables generalization to unseen combinations from as few as 10 trajectories. These results establish cross-domain transfer learning as a new paradigm for data-efficient generative animal movement modeling.
bioinformatics2026-07-20v1Interpretable Prediction of Phase Separation and Disease Variant Effects in Intrinsically Disordered Regions
Zhao, M.; Kumar, S.Abstract
Coding mutations within intrinsically disordered regions (IDRs) of proteins are increasingly implicated in human diseases yet remain poorly interpreted by conventional variant-effect predictors that rely on structural stability and conservation-based metrics. Quantifying disruption of IDR-mediated liquid-liquid phase separation (LLPS) offers a biophysically principled approach to interpreting the pathogenic impact of such variants. However, existing LLPS predictors suffer from training biases toward self-separating proteins, show limited performance on partner-dependent phase separation, and often lack interpretability for variant prioritization. We present an interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent. Our two-step classifiers outperform existing methods on independent benchmark datasets, with the largest gains for partner-dependent LLPS proteins. Beyond classification, our framework identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS. Applied to disease-associated variant databases, we found that pathogenic mutations are enriched in predicted phase-separating regions and frequently perturb LLPS propensity scores, implicating mutation-induced LLPS dysregulation as a potential pathogenic mechanism for numerous diseases. Overall, our framework provides an accurate, interpretable approach for identifying phase-separating proteins and linking aberrant phase-separation behavior to disease pathogenesis.
bioinformatics2026-07-20v1DNAS-Bench: Deterministic Nucleic Acid Screener Benchmarking
Wong, H. C.; Kohno, T.; Nivala, J.Abstract
The rapid growth of biotechnology manufacturing for synthetic DNA and proteins has raised concerns that adversaries could exploit commercial synthesis pipelines to create biological weapons. Without effective safeguards, an attacker could seek regulated genetic sequences from synthesis providers; while synthetic DNA is not itself a pathogen or toxin, access to such sequences can lower barriers to downstream misuse, motivating robust order-time screening. To mitigate this risk, Biosecurity Screening Software (BSS) systems have been developed to flag potentially malicious synthesis orders. Here, we propose one of the first deterministic benchmarks for evaluating the robustness of Biosecurity Screening Software. Our framework enables systematic testing of BSS behaviors and potential on specific nucleic-acid sequences and on targeted regions of malicious genomes. Our framework allows for insights into what is being flagged as malicious in BSSs, leading to potential discussions if specific BSS is fit for a specific manufacturing pipeline. We additionally introduce a dataset of manipulated genomes derived from the HHS and USDA Select Agents and Toxins List. When evaluated on this dataset, SeqScreen flags 42% of the sequences as malicious, while Commec flags 10.2%. Across a range of manipulation strategies, we find that simple manipulations, such as padding sequences by adding a repeated nucleotides at 1.5 times the original length, perform nearly as well as more targeted methods, such as embedding malicious sequences within benign genomic context. Padding-based methods trail embedding-based methods by only 0.75 percentage points in average detection rate. Consistent with prior reports from BSS developers and studies, we observe a sharp drop in detection rate when input sequence length falls below a critical threshold, typically between 50 and 100 base pairs (bp). Under our threat model, this implies that an adversary can bypass most existing safeguards by splitting a target genome into fragments shorter than ~50 bp. Fragment-level analysis further reveals that some toxin regions evade detection entirely by SeqScreen, while other malicious genomes remain detectable even when fragmented into 30-50 base-pair segments. We open-source this benchmark to support reproducible evaluation of BSS robustness and to inform the development of next-generation biosecurity screening tools (https://github.com/HenryCWong/DNAS-Bench). For ethical concerns we only open-source the framework while the data is available upon request.
bioinformatics2026-07-20v1multiScaleAC: Cell-Cell interaction with Moran's I as a function of kernel bandwidth
Soupir, A. C.; Hayes, M. T.; Manley, B. J.; Wang, X.; Wrobel, J.; Peres, L. C.; Fridley, B. L.Abstract
Over the last decade, spatial transcriptomic technology has transformed our understanding of tissue architecture including cell-cell interactions within the tumor immune microenvironment. A specific use-case of increasing interest is leveraging the spatial statistical relationship of genes whose protein products are known to be involved in ligand-receptor interactions. One methodological limitation of this approach has been the requirement to choose one radius around a cell as a parameter that can come with selection biases. Rather, interactions between cells vary in strength across a range of spatial scales that single-radius choice may miss. To fill this gap we developed `multiScaleAC` to extended Moran's I, a correlation measure that accounts for locations of values, by employing a Gaussian kernel applied to locations and varying the bandwidth parameter h. The resulting Moranis I\left(h\right) then can be compared between samples using functional data analysis. In the current study, we used simulations to show that our framework has well controlled Type I error due to the use of permutations for assessing significant interactions. We also demonstrate that `multiScaleAC` has high statistical power to identify a significant interaction when a true interaction is simulated (1.00 at bandwidths greater than 5) and increasing power as bandwidth increases when negative interaction is simulated. We found `multiScaleAC` largely captures similar significant ligand-receptor profiles in 8 Visium samples of colon tissue using the same bandwidth as `spatialDM` but without removing low-weight spots from the weight matrix (73.7% - 87.2%). Applying the `multiScaleAC` framework to our previous single-cell spatial transcriptomics data (COL4A1-ITGAV in the stromal compartment of clear cell renal cell carcinoma) followed by functional principal component analysis, we found functional principal component 1 to represent global interaction elevation/depression. Associating functional principal component 1 scores with immunotherapy exposure showed significantly higher scores in stromal tissues exposed to immunotherapy than those naive to immunotherapy, indicating an overall higher interaction of cell expressing COL4A1-ITGAV. These findings recapitulate our previous study while reducing bias in neighbor selections. We believe this is the first study to apply a functional extension of Moran's I in combination with functional data analysis to understand cell-cell interaction over spatial scales.
bioinformatics2026-07-20v1Interpretable Peripheral Blood Cell Classification via Vision-Language Concept Bottleneck and Soft Decision Tree
Chen, K.; Hu, T.Abstract
Motivation: Deep learning classifiers for medical image analysis typically function as black boxes, disclosing neither the image features underlying their predictions nor the reasoning by which individual decisions are reached. Peripheral blood cell classification exemplifies this challenge: experienced laboratory professionals identify cell types through structured morphological criteria---nucleus shape, chromatin texture, nucleus-to-cytoplasm ratio, granularity, and staining properties---yet existing automated systems cannot express their reasoning in these same terms, impeding clinical audit and verification. Results: We present a two-stage interpretable pipeline that addresses both levels of opacity. In the first stage, a frozen domain-adapted vision-language model (ConceptCLIP) projects each cell image onto a 70-dimensional vector of morphological concept scores via zero-shot cosine similarity, eliminating the need for per-image concept annotations. In the second stage, a Soft Decision Tree (SDT) classifies cells solely on these concept scores, producing a deterministic, concept-based decision path for each prediction. On BloodMNIST (eight cell types, 3,421 test images), the full pipeline achieves 94.86% test accuracy---approximately 3 percentage points below the black-box ceiling---while providing fully traceable decision logic. Post-training histological annotation confirms that the learned routing logic aligns with established hematological morphology criteria and reveals an emergent separation of immature granulocyte subtypes (promyelocyte versus metamyelocyte) without subtype supervision, demonstrating that concept-based decision trees can recover clinically meaningful distinctions beyond the granularity of the training labels. Availability and implementation: The source code, trained SDT weights, precomputed concept score data, and inference scripts are publicly available at https://github.com/aquamarineaqua/CLIP-CBM-SoftDecisionTree.
bioinformatics2026-07-20v1DPCGS: a computational framework for linking GWAS to single-cell transcriptomics in complex traits and diseases
Liu, C.; Shen, B.; Li, J.; Zhu, R.; Yang, P.; Wu, B.; Xuan, Y.; Yang, S.; Yuan, B.; Yang, N.; Ma, L.; Liu, Q.; Dai, S.; Zhang, Y.Abstract
Complex traits and diseases arise from the interplay between genetic variation and cellular heterogeneity, making it essential to understand how genetic risk manifests at the cellular level. However, connecting genome-wide association studies (GWAS) to specific cell populations remains challenging due to cellular complexity and the prevalence of noncoding variants. Here, we present DPCGS, a computational framework that systematically integrates GWAS summary statistics with single-cell RNA-sequencing (scRNA-seq) data to identify trait-associated cell populations, genes, and regulatory programs. DPCGS is based on the principle that GWAS-prioritized genes should exhibit elevated expression in relevant cells compared with matched controls. Benchmark analyses with simulated datasets showed that DPCGS consistently outperforms existing methods, achieving higher accuracy and sensitivity in detecting trait-relevant cells. Applications to diverse scRNA-seq datasets further validated its robustness, revealing oligodendrocytes and astrocytes as key subpopulations in Alzheimer's disease and macrophages and B cells in asthma. These analyses also highlighted potential molecular regulators, including CD74, FOS, FLI1, and AP-1 transcription factors. Together, these findings establish DPCGS as a versatile framework for dissecting the cellular and molecular basis of complex traits and diseases, with broad implications for biomarker discovery and therapeutic development.
bioinformatics2026-07-20v1Mapping Tumor-Microenvironment dependencies with TMEformer: A spatial foundation framework enabling in silico perturbation
Chen, S.; Zhu, G.; Yang, L.; Wei, X.; Li, S.; Liu, P.; Chen, Q.; Zhang, Z.; Liu, D.; Tang, Y.; Xu, G.; Zhou, M.; Luo, J.; Huang, L.; Chen, B.; Ou, S.; Jiang, J.Abstract
Despite the fundamental role of spatial context in driving tumor progression, most current computational models for virtual perturbation have largely overlooked its importance. Here, we introduce TMEformer, a tumor microenvironment-aware deep learning framework that leverages high-resolution spatial transcriptomics to jointly model intrinsic tumor cell programs and local microenvironmental signals by explicitly incorporating spatial architecture. Validated across diverse tumor spatial transcriptomic cohorts, TMEformer enables virtual perturbations that capture functional dependencies within local cellular ecosystems. Despite being trained on cancer-specific spatial datasets, TMEformer outperforms baseline models pretrained on large-scale corpora in capturing key tumor transitions, including lineage plasticity and the emergence of therapy resistance. Systematic perturbation analyses prioritize tumor-intrinsic transcription factors and TME-derived ligands that drive disease progression, recovering established regulators and revealing novel candidates. Furthermore, TME-derived embeddings improve the spatial stratification of tumor cells and align more closely with pathological architecture. Together, TMEformer establishes a general framework for modeling tumors as spatially coupled, perturbable ecosystems.
bioinformatics2026-07-17v2MAJEC: unified gene, isoform, and locus-level transposable element quantification from RNA-seq
Lim, T.-Y.; Firestone, A. J.Abstract
Background: The study of transposable elements (TEs) has become increasingly central to fields such as cancer biology, immunology, and aging. Accurately quantifying disease- or laboratory-mediated perturbations in these elements is critical to support this expanding research, yet current RNA-seq pipelines struggle with the pervasive overlap between TEs and protein-coding genes. Existing tools either aggregate to the subfamily level with no locus resolution (TEtranscripts), or provide locus-level quantification without modeling gene overlap (Telescope), with the latter attributing over 40% of TE signal to the 1.1% of loci that overlap gene exons. Results: We present MAJEC (Momentum Accelerated Junction Enhanced Counting), a unified Expectation-Maximization (EM) framework that jointly quantifies genes, transcript isoforms, and individual TE loci from BAM alignments in a single pass. Splice junction evidence informs transcript-level priors, enabling MAJEC to probabilistically distinguish genic from TE-derived reads. This approach was independently validated against Salmon and RSEM on isoform quantification benchmarks. The joint feature space reduces exon-overlap contamination of locus-level TE estimates from 43% of total signal (Telescope) to 5% (MAJEC), while preserving subfamily-level accuracy (differential expression r = 0.987 vs TEtranscripts). Using paired biological vignettes, we demonstrate that MAJEC correctly resolves both the false TE reactivation artifacts endemic to TE-only models, and the false gene upregulation artifacts that occur when heuristic rules misassign genuine intragenic TE transcription. Conclusion: MAJEC simultaneously produces the isoform and locus-level resolution that TEtranscripts lacks, with greater accuracy than Telescope, and runs faster than either.
bioinformatics2026-07-17v2Beyond Bisulfite Sequencing: Resolving 5-hmC with Nanopore Sequencing Unmasks the True-5mC Methylation Entropy Landscape
Bertocchi, U.; Katz, E.; Jeffet, J.; Grunwald, A.; Gabay, N.; Deek, J.; Verma, S.; Shwartz, A.; Umschweif-Nevo, G.; Lerer, B.; Roichman, Y.; Ebenstein, Y.Abstract
DNA methylation dynamically regulates cellular function and phenotype. At the tissue level, stochastic variation in methylation patterns, measured as methylation entropy, drives plasticity, development, cancer, and aging. Demethylation is facilitated by erasure of 5-methylcytosine (5mC) via the oxidized intermediate 5-hydroxymethylcytosine (5hmC), but bisulfite sequencing cannot distinguish these modifications, classifying both as 5mC. Using nanopore sequencing with direct detection of 5mC and 5hmC, we quantified how this historical conflation affects genome-wide methylation levels and methylation entropy in kidney cancer and the mouse medial prefrontal cortex. Bisulfite-like analysis introduced systematic, tissue-specific shifts in methylation distributions, influencing biological interpretation. However, these effects were modest in the low-5hmC kidney cancer samples, where pathway-level results remained highly concordant. Our findings demonstrate that True-5mC-based methylation entropy redefines the physical mapping of epigenomes, demonstrating that, in some contexts, what was previously interpreted as stochastic maintenance failure is frequently the structured signature of distinct, mechanistically interpretable cytosine biochemistry.
bioinformatics2026-07-17v2Nextstrain automates real-time phylogenetic analysis of open data for endemic and emerging pathogens
Andrews, K. R.; Chang, J.; Roemer, C.; Hadfield, J.; Lin, V.; Brito, A. F.; Daodu, R.; Joia, I. A.; Kistler, K.; Li, A. W.; Moncla, L. H.; Paredes, M. I.; Kuhnert, D.; Torres, L. M.; Voitl, L.; Aksamentov, I.; Hodcroft, E. B.; Huddleston, J.; McCrone, J. T.; Anderson, J. S.; Sibley, T. R.; Lee, J.; Neher, R. A.; Bedford, T.Abstract
Motivation: Genome sequencing provides an exceptional window into the evolutionary and epidemiological dynamics of endemic and emerging pathogens, and thus allows for better, more targeted, public health interventions. Online genomic surveillance platforms can provide near real-time insight into these dynamics. Results: Nextstrain provides continually updated real-time genomic surveillance for 21 viruses and the bacterial pathogen Mycobacterium tuberculosis, with most analyses relying solely on open sequence data. Each pathogen includes steps to fetch and curate open data, classify sequences using established nomenclature systems, perform phylogenetic analyses, and share the results publicly. These analyses are automated, with most running daily to provide continually updated snapshots of pathogen evolution. Availability and Implementation: All source code is available at https://github.com/nextstrain. Phylogenetic results can be visualized and downloaded at https://nextstrain.org/pathogens, and open sequence data and curated metadata are available at https://nextstrain.org/pathogens/files.
bioinformatics2026-07-17v2SoftHybrid: A Hybrid Imputation Algorithm Optimised for Single-Cell Proteomics Data
Shi, Y.; Davis, S.; Charles, P. D.; Taylor, S.; Dombi, E.; Berridge, G.; Ebner, D.; Fischer, R.Abstract
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive Missing-Not-At-Random (MNAR) sparsity. Existing imputation methods typically target either Missing-At-Random (MAR) or MNAR mechanisms, resulting in a trade-off between replicate consistency and preservation of biological variation, and are largely designed for bulk data. Here, we introduce SoftHybrid, a data-driven imputation framework that jointly models missingness and protein abundance to estimate the probability of MNAR, enabling continuous weighting between MAR- and MNAR-oriented strategies. SoftHybrid requires no external priors (cell type labels, group annotations, predefined missingness assumptions, etc.), enabling fully unsupervised applications. Across ground truth benchmarks and real single-cell proteomics datasets, SoftHybrid outperforms existing methods at low input and matches or exceeds their performance at the mini-bulk level. By preserving proteomic structure and abundance accuracy, it enhances the recovery of biologically meaningful signals. SoftHybrid is implemented as an R package and is freely available on GitHub.
bioinformatics2026-07-17v2Retention, not flux: endpoint confounding caps computational prediction of peptide skin penetration, with a delivery-aware reframing
Komianos, N.; Prakash, P.Abstract
Bioactive peptides are now central to cosmetic and dermatological actives, yet predicting whether a given sequence will reach its site of action in skin remains unsolved. We contend that the dominant framing, predicting a single binary "skin permeability" label from sequence, is ill-posed, and that this, rather than a shortage of modelling power, explains the field's stalled predictive performance. The scope of the claim is narrow: barrier-crossing propensity is a legitimate, learnable function of molecular structure, whereas the vehicle- and endpoint-agnostic binary label that the literature supplies is not. We support this with a first-principles analysis and a study of public-source data. First, the experimental endpoint most commonly reported, transdermal flux into a diffusion-cell receptor compartment (OECD Test Guideline 428), conflates two opposite outcomes (genuine deep delivery and undesired systemic transport) and is, for a cosmetic active, frequently a failure signal rather than a success signal. That receptor flux is an imperfect measure of cutaneous bioavailability is long established in dermatopharmacokinetics; our contribution is to show that the same confound, inherited through scraped labels, is what caps machine learning from sequence. Second, reported "permeability" is a property of the sequence x delivery-vehicle x measurement-compartment triad, two terms of which are usually unrecorded. Third, on public-source data, a physicochemical intrinsic-permeability estimate (Potts-Guy) carries no positive predictive signal for scraped penetration labels (grouped AUC 0.45, 95% CI 0.40-0.51); sequence-only classifiers plateau in the mid-0.70s with diminishing returns as labels accumulate (AUC 0.70-0.77); and the same descriptor pipeline on a clean single-endpoint membrane dataset scores materially higher (AUC 0.83, non-overlapping CI). Our proposed reframing separates barrier-crossing (data-driven, sequence-level) from depth-and-retention (physics-driven, delivery-aware) and treats intrinsic transdermal flux as a regulatory risk axis; we close by proposing a triad-annotated reporting schema and a seed benchmark.
bioinformatics2026-07-17v2PFM: perturbed flow matching for structure-based drug design
Yu, Y.; Xu, G.; Xie, Z.; Yang, Y.; Jiang, Y.; Zhou, X.; Li, K.Abstract
Generating 3D molecules that bind to specific protein targets via generative models has shown great promise in structure-based drug design. Recently, diffusion-based methods have achieved promising results, but their reliance on high sampling steps poses risks of slowing the drug discovery process due to increased time and computational costs. In this work, we propose a novel method named Perturbed Flow Matching (PFM), which significantly reduces sampling steps by leveraging a Flow Matching framework. PFM introduces a unique perturbed conditional probability path design that incorporates pocket binding site information and atom type-coordinate coupled information to enhance molecular generation performance. Experiments on CrossDocked2020 dataset demonstrate that PFM generates molecules with competitive 3D structures and state-of-the-art (SOTA) binding affinities towards the protein targets, achieving an Avg. of -7.12. Additionally, PFM accelerates the generation of valid molecules by a factor of 21.3, while demonstrating potential for further improvement. The code is available at https://github.com/kurisu92725/PFM.
bioinformatics2026-07-17v1SST-MAE: Learning Spectral-Spatio-Temporal Representations from Plant Hyperspectral Time Series to Discover Complex Genotype-Phenotype Relations
Okyere, F. G. G.; Mehrem, S. L.; Snoek, B. L.; Van den Ackerveken, G.; Abeln, S.Abstract
Understanding the link between genetic variation and observable traits is key to crop breeding. Hyperspectral imaging captures physiological and biochemical profiles, but current supervised methods require costly trait annotations and treat each observation as a static snapshot, ignoring the temporal dynamics of plant development. We introduce SST-MAE, a self-supervised framework that learns genotype-discriminative representations from plant hyperspectral developmental trajectories, without requiring phenotypic labels. The model learns to reconstruct masked information, capturing multiple growth trajectories. Validated on 194 field-grown lettuce genotypes across eight time points, the frozen encoder serves as a feature extractor for downstream genotype classification. SST-MAE outperforms raw spectral and linear baselines, achieving AUROC > 0.89 for anthocyanin pigmentation SNPs and 0.77 for leaf serration. The learned features are highly label-efficient, attaining near-full performance with only 30-50% of labeled data, offering a scalable pathway toward high-throughput genetic screening from image-based phenotypes. Keywords: Genotype prediction, Hyperspectral imaging, Masked autoencoder, Plant phenotyping, Self-supervised learning
bioinformatics2026-07-17v1Cell-Hub: a graphical interface for end-to-end single-cell RNA sequencing analysis
macaux, g.; Di Gallo, M.; Taglietti, V.; Amthor, H.; Maire, P.Abstract
Abstract Single-cell and single-nucleus RNA sequencing have become increasingly widespread, creating a significant demand for accessible analysis tools in research laboratories. Despite this need, the bioinformatics expertise required for such analyses remains rare. Cell-Hub addresses this gap by enabling single-cell data analysis for all researchers, regardless of computational background. Cell-Hub is a comprehensive, free, and open-source framework built on R/Shiny and distributed as a Docker image, integrating Seurat 5, CellChat 2, and Monocle 3 within a unified graphical interface. It supports all essential steps of single-cell RNA-seq analysis: data loading, quality control, normalization, clustering, multi-dataset integration, differential expression, and biomarker detection. Cell-Hub further incorporates ligand-receptor interaction inference powered by GaspouDB, a consolidated database of 11,563 mouse and 9,604 human interactions derived from CellChat, CellPhoneDB, CellTalkDB, and MultiNicheNet as well as trajectory inference via Monocle 3. All analyses produce publication-ready visualizations with flexible export options. By integrating these analytical frameworks into a single, intuitive interface requiring no programming expertise, Cell-Hub represents a significant step toward democratizing single-cell genomics for the broader research community.
bioinformatics2026-07-17v1Systematic evaluation and benchmarking of text summarization methods for biomedical literature: From word-frequency methods to language models
Baumgärtel, F.; Bono, E.; Fillinger, L.; Galou, L.; Keska-Izworska, K.; Walter, S.; Andorfer, P.; Kratochwill, K.; Perco, P.; Ley, M.Abstract
The rapid expansion of biomedical literature demands automated summarization tools that can reliably condense research articles into concise, accurate overviews. We benchmarked 62 text summarization methods - ranging from frequency-based and TextRank extractors to modern encoder-decoder models (EDMs) and large language models (LLMs) - on a set of 1,000 biomedical abstracts for which author-generated highlights sections were available as reference summaries. Models were evaluated using a composite suite of metrics covering lexical overlap (ROUGE-1/2/L, BLEU, METEOR), embedding-based semantic similarity (RoBERTa, DeBERTa, all-mpnet-base-v2), and factual consistency (AlignScore). Our results indicate that general-purpose language models (LMs) achieve the highest overall scores across both lexical and semantic metrics, outperforming both reasoning-oriented and domain-specific models. Within the general-purpose group, medium-sized models, typically runnable on a single node, often outperform frontier-scale counterparts, suggesting an optimal balance between model capacity and computational efficiency. Statistical extractive methods lag behind all neural approaches. These findings provide a systematic reference for selecting summarization tools in biomedical research and highlight that broad pretraining remains more effective than narrow domain adaptation for generating high-quality scientific summaries.
bioinformatics2026-07-16v4DIOPT: the DRSC Integrative Ortholog Prediction Tool, 2026 update
Hu, Y.; Comjean, A.; Gao, C.; Yamamoto, S.; Mohr, S.; Perrimon, N.Abstract
Mapping orthologous proteins is a critical step for cross-species literature mining, data integration, experimental design, and more, making the ability to quickly predict orthologs across species a key tool for functional genomic studies. The DRSC Integrative Ortholog Prediction Tool (DIOPT) was initially developed in 2011 to provide a centralized portal for identifying predicted orthologs among major model organisms. By integrating results from multiple ortholog prediction algorithms, DIOPT allows users to compare predictions across methods and prioritize high-confidence ortholog relationships. Over the years, we regularly updated the underlying genome annotations and refreshed predictions from each integrated algorithm. In addition, both the number of supported species and the number of ortholog prediction algorithms incorporated into the platform have grown. The web portal has also been enhanced with new features designed to improve usability, facilitate data exploration, and support a broader range of research applications. We also developed a sister version of DIOPT tailored specifically for arthropod species; this enables researchers working with a diverse set of insects and related organisms to perform ortholog mapping and comparative analyses more effectively. Together, these developments ensure that DIOPT remains a robust and broadly useful resource for functional genomics research.
bioinformatics2026-07-16v3M6AFormer Prioritizes Unannotated Functional m6A Candidate Sites in the Human m6A Epitranscriptome
Niu, Z.; Liu, C.; Gu, L.Abstract
N6-methyladenosine (m6A) is a pervasive RNA modification with critical roles in post-transcriptional regulation, yet accurate transcriptome-wide identification of functional m6A sites remains challenging. Here, we present M6AFormer, a hybrid deep-learning framework that combines convolutional feature extraction with a lightweight Transformer to capture both local sequence motifs and broader contextual dependencies. M6AFormer consistently outperformed representative m6A predictors, including MST-M6A, CLSM6A and deepSRAMP. Transcriptome-wide scanning revealed a large repertoire of previously unannotated candidate m6A sites that retained hallmark m6A features, including canonical motif enrichment, characteristic spatial distribution and preferential overlap with m6A writer and reader binding regions. Importantly, M6AFormer-predicted sites were broadly associated with genetic and disease-relevant features, including SNPs, sequence variants and GWAS-linked loci, suggesting their potential contribution to human disease mechanisms. Finally, experimental validation confirmed a previously unreported m6A site in NEU4 mRNA and demonstrated its functional impact on cancer cell migration. Together, M6AFormer provides an accurate, interpretable and biologically informative framework for m6A site discovery.
bioinformatics2026-07-16v2JanusX: an integrated and high-performance platform for scalable genome-wide association studies and genomic selection
Fu, J.; Jia, A.; Wang, H.; Liu, H.-J.Abstract
As genomic datasets expand in both sample size and marker density, genome-wide association studies (GWAS) and genomic selection (GS) require workflows that remain statistically rigorous, computationally efficient, and reproducible across the full analysis path, from genotype matrix to decision-relevant outputs. Here we present JanusX, an integrated high-performance framework that provides a streamlined, user-oriented workflow for GWAS and GS by unifying data handling, model execution, and visualization. Across simulated and real datasets, JanusX maintained high concordance with established baselines while substantially reducing runtime and memory usage. In GWAS, JanusX achieved up to a 19-fold speedup over GEMMA in linear mixed model (LMM) inference, and implemented additional LMM inference based on a sparse genomic relationship matrix with GRAMMAR-Gamma calibration, alleviating computational and memory bottlenecks in large-scale cohorts. JanusX also provides a FarmCPU implementation within its GWAS module, achieving a median 11.4-fold runtime improvement and reducing peak memory usage by 84.9% relative to rMVP. In GS, JanusX integrates an optimized best linear unbiased prediction (BLUP) backend that adaptively selects sample- and SNP-space solvers and incorporates a Preconditioned Conjugate Gradient (PCG) solver. This implementation efficiently completes five-fold cross-validation of 500k individuals x 500k single-nucleotide polymorphisms (SNPs) in 35.1 minutes with only 14.3 gibibyte (GiB) of peak memory. Beyond BLUP, JanusX integrates Bayesian and machine-learning predictors under a single interface with compact automatic tuning to ensure robust cross-model performance. JanusX therefore enables efficient locus discovery and genomic prediction under consistent analytical assumptions, even in large-scale cohorts.
bioinformatics2026-07-16v2Biological Continued Pretraining Reshapes the Capability Profile of a Foundation Model Without Catastrophic Forgetting
Wang, L.Abstract
It is widely assumed that continued pretraining (CPT) on a narrow, out-of-distribution corpus such as raw biological sequence must trade away a general-purpose model's broad competence --- the "alignment tax" or catastrophic-forgetting intuition. We test this directly, without any new training, by re-analyzing three checkpoints from a single lineage of a 26B-parameter Mixture-of-Experts model (Gemma-4-26B-A4B): the instruction-tuned base, the same model after biological CPT (8.7B tokens of DNA, protein, and biomedical text), and after subsequent supervised fine-tuning (SFT). Across three independent capability axes --- general knowledge/reasoning (MMLU, ARC, HellaSwag), code generation (MBPP), and biomedical knowledge (BixBench) --- we find that biological CPT does not degrade the model; it lifts it: MMLU +13 points, MBPP pass@1 nearly doubles (0.33 to 0.63), and BixBench discrimination rises sharply (MCC 0.23 to 0.92). The single measured regression is truthfulness (TruthfulQA -8.8 points), a small and interpretable domain drift. A clean vocabulary-expansion ablation (<0.4 pt on every general metric) confirms the gains are attributable to CPT, not tokenizer changes. Crucially, subsequent SFT narrows the model back: all three axes fall to near-base levels, revealing a consistent division of labor --- CPT re-organizes and lifts the shared capability substrate; SFT cashes it out onto target tasks. We argue this reframes biological sequence not as a competitor for a foundation model's capacity but as a form of structured scientific data that reshapes its capability profile, and that CPT and SFT should be budgeted as complementary rather than substitutable stages. All checkpoints, evaluation code, and per-example outputs are public.
bioinformatics2026-07-16v1Evaluating the use of non-linear models in data-driven rescoring of peptide-spectrum matches
Nameni, A.; Declercq, A.; Gabriels, R.; Degroeve, S.; Martens, L.; Bouwmeester, R.Abstract
In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly supports this task through peptide-spectrum match (PSM) rescoring, in which a classifier, typically a linear semi-supervised model, refines the initial matching score. However, Mokapot allows the user to choose among different machine learning algorithms of increasing complexity, from the default linear support vector machine (LSVM) to random forest and XGBoost. Here, we use an entrapment approach to assess the effect of this increasing complexity on PSM identification and the accuracy of the estimated false discovery rate (FDR). We show that, while more complex models increase the number of identified PSMs at a fixed FDR threshold, this gain reflects a bias towards random matches from the target proteome database rather than genuine identifications. Indeed, for the most complex model, the entrapment FDR reaches 6.3% instead of the estimated 1% decoy FDR. This bias thus yields overly optimistic FDR estimates, indicating that model complexity in PSM rescoring must be carefully balanced against this overfitting risk.
bioinformatics2026-07-16v1Reference Regulatory Element-Guided Gene Expression Analysis for Mechanistic Inference of Gene Regulatory Networks
Ren, L.; Debnath, I.; Duren, Z.Abstract
Regulatory genomics faces a depth-breadth gap: deep multi-omics provides regulatory detail but is difficult to scale, whereas broad expression datasets often lack the regulatory structure needed for mechanistic Gene Regulatory Network (GRN) analysis. We developed Regulatory Elements Guided Analysis (REGA), an interpretable framework that uses reference Regulatory Element (RE) catalogs to infer transcription factor (TF)-RE-gene programs from gene expression data. Across ChIP-seq, knockdown, Hi-C, cis- and trans-eQTL benchmarks, REGA prioritized functional REs, improved RE-gene and TF-gene inference over existing baselines, including methods using more data, and recovered coherent regulatory modules. In PsychENCODE snRNA-seq, REGA identified disease-associated modules and TF activities, linked regulatory dysregulation to genetic risk, and detected cross-cell-type neuronal-glial programs. In spatial transcriptomics, REGA linked cell-intrinsic regulatory programs with intercellular ligand-receptor communication; in Perturb-seq, it mapped perturbation responses to trait-associated regulatory architectures. REGA enables scalable, interpretable GRN analysis across expression datasets.
bioinformatics2026-07-16v1Differential multi-omics analysis of pulmonary arterial hypertension microvascular endothelial cells for differential drug response
Hiort, P.; Weiss, A.; Krentz, J.; Schermuly, R. T.; Bogaard, H.-J.; Conrad, T.; Szulcek, R.; Baum, K.Abstract
Pulmonary arterial hypertension (PAH) represents a heterogeneous group of disorders that involves complex molecular dysregulations, which are not fully captured by single-omics analyses. We apply our network-based multi-omics analysis framework, DrDimont, to transcriptomic, proteomic, phosphoproteomic, and kinase screening data from lung microvascular endothelial cells of PAH patients and controls. Thereby, we extend the functionality of DrDimont to incorporate kinase-kinase interactions during the construction of condition-specific multi-omics networks. Kinase interactions are inferred from phosphorylations of screened substrates that are weighted by kinase-substrate predictions. Differential interaction scores from the network-based analysis between PAH and control uncover alterations centered on kinases, in particular top hits relating to MAPK signaling, such as MAPK13, MAP2K, or upstream IRAK1, and other MAPK/MAP2K family members. Further highly differential nodes were ACADSB, GPX7, DSE (for proteins), and AIM1, LY96, CHSY3 (for mRNAs). When prioritizing drug candidates by mapping drug targets onto the differential network, we find high scores for the drug tacrolimus (FK506) and several anti-neoplastic MAPK inhibitors (e.g., selumetinib, trametinib), as well as agents acting on general proliferation via (mitochondrial) DNA transcription (e.g., epirubicin, topotecan). Integrating kinase activity screens into our explainable multi-omics network-based analyses reveals kinase-centered alterations and therapeutic hypotheses in PAH that complement single layer classical differential expression analyses.
bioinformatics2026-07-16v1CurateMake: an auditable workflow for multi-source ITS reference database harmonisation and phylogenetic validation
Gardette, A.; Belda, E.; Prifti, E.; Zucker, J.-D.Abstract
1. Reference databases shape the taxonomic resolution, uncertainty, and reproducibility of metabarcoding analyses. For ITS barcodes, public references are distributed across repositories with different taxonomic conventions, geographic coverage, and annotation practices, creating conflicts, missing ranks, and misannotations when databases are merged or compared. 2. We introduce CurateMake, a reproducible Snakemake workflow for ITS reference database construction, harmonisation, and validation. It integrates four public sources (UNITE, BOLD, PLANiTS, and CALeDNA) and user-supplied databases, combines Catalogue of Life name harmonisation with ITSx-based region standardisation, MSA/HMM-based alignment grouping, and SATIVA phylogenetic validation. Raw, CoL-harmonised, and SATIVA-validated annotation layers are retained throughout to compare curation effects while preserving flagged records for review. 3. We evaluated CurateMake on 3.58 million ingested sequences and controlled error-injection simulations. ITSx expanded the final harmonised database to 5.19 million barcode-resolved entries by recovering ITS1 and ITS2 sub-regions from full-length ITS records. Across the full dataset, normalised intra-cluster entropy decreased from Raw to CoL-harmonised to SATIVA-validated annotations, consistent with improved taxonomic coherence. In simulations, CurateMake achieved the highest correction rate across 1%-50% corruption and, at 15% corruption, corrected 42% +/- 1% of introduced errors, compared with 28% +/- 1% for CoL alone and 0% for SATIVA without the workflow's alignment infrastructure. 4. These results show that nomenclatural harmonisation and phylogeny-informed validation address complementary error classes, with phylogenetic validation contributing measurably only within taxon-coherent alignments in this benchmark. CurateMake therefore provides a reproducible, provenance-tracked framework for auditable ITS reference database curation in metabarcoding workflows.
bioinformatics2026-07-16v1SHINE: Decoding transcriptional-metabolic microenvironments through higher-order spatial integration
Du, B.; Wong, J. W. H.; Huang, Y.Abstract
Spatial omics technologies are expanding to co-profile transcriptomics and metabolomics on the same tissue slide, providing complementary views of gene expression and biochemical activity to reveal molecular programs within native tissue microenvironments. However, integrating the transcriptome and metabolome remains technically challenging due to spatial misalignment, resolution disparity, and higher-order cross-modality interactions. Here, we present SHINE, a hypergraph-based computational framework for the joint analysis of spatial gene expression and metabolic networks derived from the co-profiling slide, focusing on representation learning and cross-modality interaction. Across multiple datasets, SHINE consistently outperformed existing methods for domain segmentation and biomarker co-localization and provided interpretable insights into metabolic-transcriptional microenvironments. Specifically, in Parkinson's disease mouse models, SHINE accurately delineates dopaminergic neuron-depleted regions and reconstructs coherent dopamine-associated axes. In human lung and breast cancers, SHINE resolves tumor-associated spatial regions and identifies spatially organized gene-metabolite programs associated with the tumor microenvironment. SHINE enables scalable spatial multi-omics integration across diverse biological systems.
bioinformatics2026-07-16v1Integrating suboptimal secondary structures, AI-assisted genomic synteny, and evolutionary conservation to identify bacterial ncRNA homologs beyond sequence similarity
Panek, J.Abstract
A bioinformatic approach for genome-wide identification of homologs of bacterial non-coding RNAs (ncRNAs) integrating structural similarity, genomic synteny, and evolutionary conservation is presented. The structural similarity is detected using an algorithm for genome-wide identification of loci in genomic intergenic regions (IGRs) containing sequences capable of adopting secondary structures similar to that of the query ncRNA. The algorithm scans IGR sequences using a sliding window with a predefined step. For each window, suboptimal secondary structures are predicted and compared with the template structure to compute structural similarity scores. These scores are evaluated statistically on a genome-wide scale to infer homology of the RNAs represented by the predicted structures. Loci encoding statistically significant structures are further filtered using genomic synteny of the query ncRNAs inferred from genomic annotations. ChatGPT was used to assist in identifying literature-supported biological relationships between genes with distinct functional annotations. Syntenic loci with the structures are then examined for homologs in related species, as evolutionary conservation among related species is a common feature of ncRNAs Using this approach, we predicted novel homologs of the spot42 RNA-encoding spf gene in Glaciecola and Pseudoalteromonas genomes, and ms1 RNA genes in Frankia and Bifidobacterium genomes, where previous homology searches had failed.
bioinformatics2026-07-16v1Unraveling the Comedone Switch through Single-Cell Resolution of Human Acne Lesions
Duez, T.; Rolka, T.; Torocsik, D.; Reuter, H.; Al, B.; Gallinat, S.; Baumbach, J.; Holzscheck, N.Abstract
Acne vulgaris is one of the most prevalent inflammatory skin diseases worldwide, yet the molecular events initiating comedogenesis remain poorly understood. The comedone switch hypothesis proposes that acne originates from an imbalance in lineage commitment within the junctional zone of the pilosebaceous unit, promoting infundibular differentiation at the expense of sebaceous gland maintenance. However, direct evidence from human acne tissue at single-cell resolution has been lacking. Here, we integrated single-cell transcriptomic datasets from healthy skin, non-lesional skin of acne patients, and lesional acne tissue to reconstruct the earliest stages of comedogenesis. We identified a previously uncharacterized cell population in non-lesional skin with transcriptomic features consistent with a microcomedone and mapped this population across independent datasets to reconstruct the transcriptional comedone architecture. Comedonal remodeling was characterized by enhanced keratinization and inflammatory programs. Quantitative analyses supported a shift from sebaceous toward infundibular cell fate, providing first data-driven evidence for the comedone switch hypothesis in human acne. Beyond the pilosebaceous unit, we identified broader epithelial alterations, including loss of POSTN and ERRFI1 expression in basal interfollicular epidermal keratinocytes. Together, these findings provide a cell-resolved framework for human comedogenesis and identify candidate mechanisms linking genetic susceptibility, environmental triggers, and lineage imbalance within the upper hair follicle.
bioinformatics2026-07-16v1