Latest bioRxiv papers
Category: bioinformatics — Showing 50 items
STAR Suite: an open-source single-executable transcriptomics engine for reproducible, AI agent-assisted processing
Hung, L.-H.; Baker, D.; Flynn, B.; Huangfu, D.; Luo, R.; Robson, P.; Zhou, T.; Yeung, K. Y.Abstract
Processing sequencing data means chaining many specialized tools, a barrier for bench biologists and for the AI agents increasingly used to run analyses. Open, integrated pipelines remain incomplete: the de facto single-cell standard, Cell Ranger, is proprietary - its license bars redistribution, modification, and non-10x use - so it cannot serve as a shareable, AI-discoverable layer, and no production-ready open-source pipeline exists for 10x Flex. Here, we present STAR Suite, which extends the STAR aligner into a single executable for transcriptomics processing.Adapter handling, feature-barcode assignment, Flex probe processing, SLAM-seq analysis, sorting, and quality control are integrated directly into the 28,228-line STAR codebase with no additional external dependencies, adding 132,226 lines of C/C++ made tractable through human-directed AI software engineering. Novel methodologies include dynamic thread interleaving, fast-Hamming feature matching, alignment-validated probe hashing, and variance-based auto-trimming (SLAM-seq); in-process integrations include Y-chromosome removal, per-sample variant masking, and Variational Bayes transcript quantification (bulk RNA-seq). STAR Suite reproduces Cell Ranger 9.0.1 to gene-level Pearson 0.99-1.0 while running 3.6- to 5.7-fold faster, and 3.9-fold faster than stepwise bulk pipelines. Used across the NIH MorPhiC consortium, its source code, workflow recipes (morphic-recipes), and immutable per-run provenance records (morphic-provenance) are released under the permissive open-source MIT license.
bioinformatics2026-07-24v5Phylogenetic detection of protein sites associated with continuous traits
Duchemin, L.; Muntane, G.; Boussau, B.; Veber, P.Abstract
Comparative genomic data can be used to look for substitutions in coding sequences that are associated with the variation of a particular phenotypic trait. A few statistical methods have been proposed to do so for phenotypes represented by discrete values. For continuous traits, no such statistical approach has been proposed, and researchers have resorted to sensible but uncharacterized criteria. Here, we investigate a phylogenetic model for coding sequences where amino acid preferences at a site are given by a continuous function of a quantitative trait. This function is inferred from the amino acids and the trait values in extant species and requires inferred point estimates of ancestral values of the trait at internal nodes. For detecting sites whose evolution is associated with this trait, we use a significance test against the hypothesis that amino acid preference does not depend on the trait. This procedure is compared to simpler strategies on simulated alignments. It displays an increased recall for low false positive rates, which is of special importance for performing whole-genome scans. This comes however at a much higher computational cost, and we suggest using a simple test to filter promising candidate sites. We then revisit a dataset of alignments for 62 species of mammals, using longevity as a phenotypic trait. We apply our method to three protein families that have previously been proposed to display sites associated with variation in lifespan in mammals. Using a graphical representation extracted from the detailed phylogenetic analysis of candidate sites, we suggest that the evidence for this in the sequence data alone is weak. The proposed method has been added to our Pelican software. It is available at https://gitlab.in2p3.fr/phoogle/pelican and can now be used with both discrete and continuous phenotypes to search for sites associated with phenotypic variation, on data sets with thousands of alignments.
bioinformatics2026-07-24v5Leveraging Uncertainty Estimates for Drug Response Prediction in Cancer Cell Lines
Iversen, P.; Renard, B. Y.; Baum, K.Abstract
Machine learning models for drug response prediction in cancer cell lines carry the potential to advance precision oncology by tailoring treatments to the molecular tumor profile. Their application is challenged by variability in prediction quality and distribution shifts between training and application. Uncertainty estimation provides more information on the predictive distribution than point estimates, enabling comprehensive decision support and downstream analysis of predictions. Yet, the most effective uncertainty estimator is domain-specific. In this work, we benchmark uncertainty-aware models for drug response prediction. We focus on epistemic uncertainty via ensemble agreement, and aleatoric uncertainty via distributional modeling, or both. We find that ensemble-based estimates are more sensitive to distribution shift and can flag out-of-distribution examples. In contrast, distributional models yield stronger prediction error reductions among high-confidence subsets. Despite higher computational cost, the combination can provide both advantages: an ensemble of neural networks that estimate a Gaussian predictive distribution can reduce the mean squared error by 64 percent when restricting predictions to the 10 percent most confident drug-cell line pairs, and reliably indicates distribution shifts and platform differences. Beyond benchmarking, probabilistic predictions can identify drugs whose uncertainty bounds overlap with therapeutically relevant ranges. We also show that uncertainty estimates enable a new dimension of model interpretability: by attributing predicted uncertainty to input features, we identify genes that signal unpredictability of drug response rather than sensitivity or resistance. We further demonstrate uncertainty-guided selection of measurements for active learning. In summary, including uncertainty in drug response prediction supports better-informed model application. The code is available at https://github.com/PascalIversen/LUDRP.
bioinformatics2026-07-24v2DOME Copilot: A resource to automate transparent reporting of artificial intelligence methods
Farrell, G.; Attafi, O. A.; Fragkouli, S.-C.; Heredia, I.; Fernandez Tobias, S.; Harrison, M.; Hermjakob, H.; Jeffryes, M.; Mehdiabadi, M.; Obregon Ruiz, M.; Pearce, M.; Pechlivanis, N.; Quaglia, F.; Lopez Garcia, A.; Psomopoulos, F.; Tosatto, S. C. E.Abstract
Artificial intelligence (AI) methods are transforming life science research and witnessing unprecedented adoption across literature. While AI is driving impactful discoveries, this rapid growth in application has inundated researchers with poorly described models and datasets which lack transparency, impeding reusability and eroding trust. The DOME Recommendations aimed to address this issue and established structured reporting guidelines to standardize AI methodology descriptions. However, the need to manually create these comprehensive disclosures imposed a major bottleneck, requiring substantial effort from authors to comply. To bridge this gap, we present DOME Copilot, an open-source, easily deployable system that automates the generation of transparent AI methodology reports supplementing publications. Evaluated against a benchmark of human-created annotations, DOME Copilot was determined to match or exceed manual annotation quality across the majority of reporting fields while reducing the creation time from hours to minutes. Ultimately, the system provides a novel and scalable solution to restore confidence in complex life science AI method publications.
bioinformatics2026-07-24v2ViroSeek: a viral detection pipeline for second-generation sequencing
Berger, A.; Lefebvre, M. J. M.; Dainat, J.; Jiolle, D.; Conclois, I.; Talignani, L.; Mastriani, E.; Cornelie, S.; Berthet, N.; Paupy, C.Abstract
Arbovirus emergences represent a rising public health issue and are exacerbated by climate change and globalization. Virome analysis has become a key approach for monitoring and managing infectious diseases, yet existing tools often remain technically complex and inaccessible to non-specialists. In this context, we present ViroSeek, a reproducible and accessible bioinformatics pipeline specifically designed for the taxonomic analysis of second-generation sequencing data from target-enriched libraries. ViroSeek performs a series of automated steps: quality control, trimming, host and bacterial sequence removal, assembly, taxonomic assignment, read remapping for quantification, and PCR duplicate removal. The whole process is designed to produce a clear, usable viral taxonomy table that is suitable for diversity studies. ViroSeek was empirically validated on enriched control samples containing a known panel of viruses. All the expected viruses were correctly detected. Bacterial and host contaminant sequences were effectively removed. The pipeline is freely available and fully documented, supporting its adoption and adaptation by the research community.
bioinformatics2026-07-24v2Graph in Graph (GiG): A novel graph AI framework for integrating and interpreting whole-person healthmedical and omics data
Zhang, H.; Lu, Y.; Fang, K.; XU, Z.; Akbary Moghaddam, V.; An, P.; Jin, S.; Wojczynski, M.; Province, M.; Li, F.Abstract
Medical records and omics data are rapidly becoming standard in healthcare settings, which characterize the whole-person from dysfunctional molecules to phenotypes, and thus offer potential for precise disease diagnosis and target discovery. Whereas, it remains an open problem to systematically integrate and interprete medical record and omics data of individual patients. In this study, for the first time, we propose a novel graph AI model framework, Graph in Graph (GiG), to integrate and interpret the whole-person medical and omic datasets. Specifically, the medical record data is modeled using a person-phenotype graph, followed by omics signaling graphs of invidival patients, which enables the integration of information learned from omic-signaling graph and medical phenotype features to characterize individual patients and to prioritize important omic biomarkers and phenotypes. As an exploratory study, we applied and evaluated the GiG model to study the type 2 diabetes (T2D) and pre-T2D vs healthy using the Long Life Family Study (LLFS) cohort, which enrolls families with exceptional longevity to uncover biological mechanisms of healthy aging with medical and omics data. The evaluation results showed that GiG not only achieve a high prediction but also can interpret the prediction by ranking the essential clinical and omic biomarkers. The GiG framework can be applied to other studies by effectively integrating and interpreting medical and omic datasets for disease diagnosis and pathogenesis discovery.
bioinformatics2026-07-24v1Revisiting Logistic Regression for High-Dimensional Gene Expression Data
Souza, R. d. O.; Rodrigues, W. F.; Couto, B.; Dos Santos, M. A.Abstract
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.
bioinformatics2026-07-24v1FusedFCR: A Fused Forward Continuation-Ratio model for marker selection along cell-fate trajectories
Mattila, C.; Chakraborty, A.; Angel, P.; Cao, S.; Sonawane, K.; Hill, E.; Chung, D.; Neelon, B.; Seal, S.Abstract
Time-course single-cell RNA sequencing (scRNA-seq) data collected across ordered stages provide population-level snapshots of differentiation, disease progression, and aging. Supervised pseudotime methods use observed stage labels to reconstruct continuous progression but generally do not identify marker genes associated with changes from one stage to the next. Unsupervised pseudotime-based marker selection methods infer latent trajectories directly from expression data and identify trajectory-associated genes, but do not explicitly link these associations to the observed stages. We propose FusedFCR, a regularized forward continuation-ratio model that represents cellular progression through a sequence of conditional transitions across ordered stages. FusedFCR combines a lasso penalty for gene selection with a fusion penalty that encourages similar effects across adjacent transitions while allowing transient and direction-changing associations. The resulting transition-specific coefficients support interpretable gene selection and a continuous pseudotime-like projection anchored to the observed developmental stages. In simulations, FusedFCR accurately recovered gene-effect trajectories and improved predictive performance relative to alternative methods. Applied to mouse pancreatic beta-cell differentiation across seven time points and human extravillous trophoblast differentiation across four time points, FusedFCR identified biologically interpretable genes associated with distinct developmental transitions. Gene set enrichment analysis further revealed stage-specific pathway activity consistent with known developmental biology, while held-out stage-classification accuracy was competitive or superior across both datasets. Together, these results show that FusedFCR complements pseudotemporal ordering by identifying which molecular programs change and when those changes emerge along the developmental trajectory. An accompanying R package is available on GitHub.
bioinformatics2026-07-24v1Structural bioinformatics of three Epstein-Barr Virus (EBV) Integral Membrane Proteins and their water-soluble QTY analogs
Zhang, S.; Sun, Z.; Chen, E.Abstract
The Epstein-Barr virus (EBV) is a highly prevalent virus worldwide that is associated with several lymphoid and epithelial malignancies. However, extensive research on EBV integral membrane proteins BILF1, LMP1 and LMP2, has been scarce due to their hydrophobic transmembrane domains. Our study applies the QTY code (glutamine, threonine, tyrosine) to design water-soluble analogs of BILF1, LMP1 and LMP2 with reduced hydrophobicity, where we systematically replaced hydrophobic amino acid residues leucine (L), isoleucine (I), valine (V), and phenylalanine (F) with structurally similar polar residues glutamine (Q), threonine (T), and tyrosine (Y). We retrieved their native sequences from UniProt, identified transmembrane domains using Protter, then performed QTY design through the Protein Solubilizing Server (PSS). We then predicted native and QTY structures using in silico prediction tools AlphaFold3, ColabFold, and Boltz-2. Our analyses demonstrate that despite significant protein sequence replacements in their transmembrane domains (54.15%-61.59%) and increased intrinsic solubility, the QTY analogs exhibited minimal changes in isoelectric point (0.00-0.15 decrease) and molecular weight (0.7-1.2 kDa increase). Additionally, structural superpositions between QTY analogs and native structures using PyMOL yield low RMSD values (0.217[A] -1.202[A]). Our results demonstrate the QTY codes ability to design detergent-free analogs of BILF1, LMP1 and LMP2 with substantially reduced hydrophobicity and aggregation propensity whilst preserving native-like structures. Our results may facilitate protein characterization studies, therapeutic research on EBV, and other protocols that typically require protein solubilization.
bioinformatics2026-07-24v1CNVeil resolves haplotype-specific copy number and uncovers subclonal architecture hidden from total copy number profiling in single-cell cancergenomes
Yuan, W.; Luo, C.; Hu, Y.; Zhang, L.; Wen, Z.; Liu, Y. H.; Fan, X. M.; Zhou, X. M.Abstract
Single-cell DNA sequencing (scDNA-seq) resolves copy number variation (CNV) at single-cell resolution, revealing tumor heterogeneity and subclonal structure. Most existing methods, however, infer only total copy number. Haplotype-resolved copy number, which captures allelic imbalance and clonal evolution, remains far less developed, largely because low coverage, allelic dropout, and technical noise in scDNA-seq make phased allelic inference substantially harder than total copy number estimation. We present CNVeil, a haplotype-aware framework that infers total, allele-specific, and chromosome-scale haplotype-resolved copy number from scDNA-seq data. CNVeil first builds robust total copy number profiles through highly variable bin selection, hierarchical clustering, subclone-aware ploidy estimation, and cross-cell consensus segmentation. Using this profile as a stable scaffold, it infers allele-specific copy number with an expectation-maximization algorithm applied to heterozygous SNP allele counts, then reconstructs haplotype-specific copy number by enforcing coherent haplotype orientation across adjacent segments via dynamic programming. We benchmarked CNVeil against 12 state-of-the-art methods, including eight total copy number callers, two allele-specific callers, and two haplotype-resolved callers, across 20 simulated and real datasets spanning six experimental settings, including high-multiplexed single-nucleus sequencing, Acoustic Cell Tagmentation (ACT), and 10x Chromium. This constitutes the largest comparative evaluation of single-cell copy number inference methods to date. CNVeil consistently outperformed existing tools in segmentation accuracy, ploidy inference, subclone identification, and allele-specific copy number estimation. In a breast cancer multi-omics (wellDR-seq) cohort, CNVeil uncovered haplotype-specific subclonal diversification invisible to total copy number analysis alone and linked allele-specific copy number states to transcriptional variation. By transforming sparse single-cell allelic signals into chromosome-scale haplotype-resolved profiles, CNVeil closes a major methodological gap and provides a scalable framework for studying tumor evolution and functional genomic heterogeneity at single-cell resolution.
bioinformatics2026-07-24v1Deconvolution-derived cell-type expression targets for personal genome sequence-to-expression prediction
Sim, S.; Shen, L.Abstract
Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue--cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122--0.142 for sequence-derived approaches and 0.081--0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution--pseudobulk agreement for sequence-derived models ($r=0.35$--$0.43$ across tissue--cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially outperform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.
bioinformatics2026-07-24v1VizR: An Interactive Web Platform for End-to-End RNA-Seq Analysis and Visualization in Plant Biology
Jeon, W.-T.; Jung, H.; Shim, D.; Lee, Y.Abstract
RNA sequencing (RNA-seq) is widely used to investigate transcriptional programs in plant biology, yet the need to combine multiple specialized tools and bioinformatics expertise to convert raw sequencing reads into biologically interpretable results remains a major technical barrier for many plant biologists. Here, we present VizR (VIsualiZation of Rna seq), a web-based platform that integrates end-to-end RNA-seq analysis and visualization within a single integrated environment. VizR automates upstream processing, including quality control, adapter trimming, genome alignment, and transcript quantification, and connects the resulting expression data to downstream exploratory analyses. Its interface is designed to make expression patterns immediately searchable and interpretable: users can query genes through an equalizer-style expression-pattern interface, inspect expression profiles using inline heatmaps embedded in gene tables, and perform context-integrated gene ontology analysis throughout the workflow. VizR also supports comparative analysis through interactive Venn diagram module, allowing users to transfer gene sets directly from result tables. As a Docker-based application, VizR can be deployed locally and accessed through a standard web browser. By unifying automated RNA-seq processing, interactive visualization, and functional interpretation, VizR lowers the technical barrier to transcriptome analysis and provides a practical platform for plant biology research.
bioinformatics2026-07-24v1Fast and accurate taxonomic domain assignment of short metagenomic reads using BBERT
Alekhin, D.; Alon, M.; Sidi, T.; Perez Mazeh, S.; Carmi, G.; Finkel, O. M.; Erez, A.Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses.
bioinformatics2026-07-23v3GraTools, an user-friendly tool for exploring and manipulating pangenome variation graphs
Ravel, S.; Marthe, N.; Carrette, C.; Mohamed, M.; Sabot, F.; Tranchant-Dubreuil, C.Abstract
Background: Pangenome variation graphs (PVGs), which represent genomic diversity through multiple genomes alignment, are powerful tools for studying genomic variations in populations. However, current tools often lack integration, efficiency, or require format conversions, to use them, hindering their usability. Results: Here, we introduce GraTools, a fast and user-friendly command-line tool for manipulating PVGs using the original GFA file. After a one-time graph import, GraTools enables rapid subgraph extraction, fasta sequence retrieval, and comprehensive analyses, including core/dispensable genome ratio calculation or group-specific segment identification. The import step results in conversion in standard data formats (BAM/BED), enabling the reuse of well-optimized existing tools, allowing an efficient storage and the querying of the PVGs large complex data structures. Scalability is ensured by a modular architecture sup- porting parallel processing and asynchronous I/O operations. GraTools supports coordinates defined on both the primary reference as well as from alternative genomes within the graph without re-importing, and its outputs can be easily visualized or manipulated using external tools. Using an Asian rice pangenome graph (13 accessions), we demonstrate its ability to easily extract subgraphs, compute depth statistics, and identify subspecies-specific segments. An intuitive command-line interface, a real-time execution feedback and a detailed logging system make this tool suitable for a wide range of applications, from population genetics to breeding and genomic medicine, for both biologists and bioinformati- cians. Conclusions: Through its unified graph manipulation interface, GraTools offers an interesting alternative to the few existing tools for manipulating PVGs, facil- itating rapid, efficient and flexible downstream analyses. It is available as an open-source tool (GNU GPLv3), with its documentation available at https: //gratools.readthedocs.io.
bioinformatics2026-07-23v3A pan-cancer analysis of microRNA tissue specificity and its association with dysregulation
Poptsova, M.; Ismailov, A.; Belogurov, A.; Evpak, A.Abstract
MicroRNAs are frequently dysregulated in cancer, yet how their tissue-specificity is remodeled during malignant transformation remains poorly characterized. Here we systematically quantified the tissue-specificity of miRNAs across normal (GTEx) and tumor (TCGA) tissues using the Tau index, and compared its distribution between healthy and cancerous states. To robustly define dysregulation, we combined two independent analyses: a binomial test over per-project differential expression across 17 matched normal tissues within TCGA cohort, and a TCGA-GTEx pan-tissue comparison of mean expression. The change in specificity ({Delta}Tau) separated up- from down-regulated miRNAs, showing moderate agreement with the binomial signal and a strong correlation with the expression-based contrast. Finally, we identified 6 miRNAs that lose tissue-specificity upon transformation while remaining consistently upregulated (miR-519a-5p, miR-512-3p, miR-522-3p, miR-105-5p, miR-935, miR-1269a). Functional analysis of experimentally validated targets showed significant enrichment for converging on core oncogenic programs for miR-512-3p, miR-105-5p and miR-935, such as apoptosis and cellular-stress regulation, TP53, FoxO, PI3K-Akt/mTOR signaling, immune modulation. Collectively, integrating specificity dynamics with dysregulation evidence pinpoints candidate miRNAs with coordinated, cancer-relevant regulatory roles and highlights those with favorable tissue specificity profiles for therapeutic targeting.
bioinformatics2026-07-23v2GeneKnow: An auditable AI framework for source-grounded biological evidence synthesis
Zhang, H.; Sittipongpittaya, B.; Zang, C.Abstract
Biological function is context dependent, yet synthesizing evidence for a gene's function in a defined cell type, disease, or perturbation remains labor-intensive. Here we present GeneKnow, an auditable artificial intelligence (AI) framework for source-grounded biological evidence synthesis. GeneKnow separates evidence-critical operations, including literature retrieval, passage selection, provenance tracking, and bibliography construction, from generative steps that are constrained to semantic analysis and literature synthesis. GeneKnow supports multi-paper discovery and single-paper inspection while preserving links from synthesized claims to source passages, generating trustworthy syntheses without fabricated citations and minimizing hallucinations. Systematic benchmarking showed that GeneKnow achieved higher claim support and citation fidelity than leading general-purpose and scientific AI systems. These results demonstrate that an AI system with controlled division of labor between deterministic and generative components can substantially improve the fidelity and auditability of biomedical literature synthesis.
bioinformatics2026-07-23v2Hobrac: a reference-guided workflow for genome comparison and synteny visualization
Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.Abstract
Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO rthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.
bioinformatics2026-07-23v2GenoSim: A Forward-Time Genotype Simulator for Clinical and Population Genetics with Population Stratification
Bakar, A.; Gul, R.; Haq, W. u.; Afghani, T.Abstract
Motivation: Next-generation sequencing studies in clinical genetics are often limited by the scarcity of human genotype data, which stems from ethical, regulatory, and economic barriers. The shortfall is sharpest in consanguineous populations, which are common in South Asia and the Middle East, where family-based designs need large pedigrees that are rarely sequenced in full. Existing simulators do not combine pedigree-aware propagation, realistic population stratification, and clinical export formats in one tool. Results: We present GenoSim, an R package for forward-time simulation of diploid SNP genotypes. It runs in two modes: a population mode implementing inbreeding-adjusted Hardy-Weinberg sampling, Wright-Fisher drift, directional selection, recurrent mutation, and Haldane recombination across multiple generations; and a pedigree-constrained mode that ingests real family VCFs and a pedigree, reconstructs phase where the pedigree makes it identifiable, propagates genotypes through the observed family structure, and appends synthetic generations. Version 1.1.1 adds population stratification through the Balding-Nichols model parameterised by gnomAD v3.1 fixation indices (F_ST) for eight ancestry groups (AFR, AMR, EAS, EUR, FIN, MID, SAS, ASJ), empirical allele-frequency loading from external reference panels, and admixed-cohort simulation. Analysis functions cover Hardy-Weinberg testing, linkage disequilibrium, runs of homozygosity, principal component analysis, founder-referenced and between-generation F-statistics, and Nei gene diversity. Availability and implementation: GenoSim is available as an R package at https://github.com/malikbak/GenoSim under the MIT licence. It requires R [≥] 4.0.0 and depends only on base R packages (stats, utils, graphics, grDevices, tools).
bioinformatics2026-07-23v2Statistical tests for bivariate spatial association across multi-omics data with disjoint coordinates
Hawinkel, S.; Hu, W.; Velten, B.; Maere, S.Abstract
Spatial biology has entered a new era of multimodal profiling, with multiple, high-dimensional spatial omics types being measured on consecutive tissue slices, or co-assayed on the same slice. Interest then lies in statistical testing for spatial association between the features of the different modalities, to gain insight in biological processes. One major challenge is the multitude of bivariate combinations, leading to high computational demands. Another difficulty is the difference in spatial resolution between technologies, implying no one-to-one matching between the measurement spots of the different modalities, even after alignment. As a result, common statistical measures such as joint distributions and correlations are not defined, and tests need to rely on spatial vicinity only. Moreover, we argue that many existing bivariate association tests address an inappropriate null hypothesis, or make inappropriate assumptions, both implying absence of spatial autocorrelation in any of the features and leading to misleading conclusions. As a remedy, we modify tests for the detection of spatially variable genes (Moran's I, Gaussian processes and generalized additive models) to derive bivariate spatial association tests across modalities with non-overlapping coordinate sets, and provide variance estimators that do account for spatial autocorrelation. We develop inference methods for single sections as well as for replicated experiments with multiple sections, and compare their performance in nonparametric and parametric simulations. Finally, we apply the newly developed methods to two co-assayed spatial transcriptomics and metabolomics datasets from mouse and human. The full suite of tests is available from github.com/sthawinke/sbivar as the R-package sbivar.
bioinformatics2026-07-23v2EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery
Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.Abstract
Motivation: As biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. Result: EcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAI's multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimer's Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research.
bioinformatics2026-07-23v2scSAID: A Comprehensive Cross-Species Single-Cell Skin Atlas Reveals Species-Specific Responses to Psoriasis
Ren, Y.; Shen, Y.; Jin, L.; Huang, Y.; Deng, Y.; Xiao, Y.; Wang, C.Abstract
Single-cell RNA sequencing has rapidly expanded the scale of skin transcriptomic data, yet these datasets remain fragmented across studies spanning different species, diseases and experimental manipulations. An up-to-date, comprehensive, and queryable single-cell cross-species repository for skin is still lacking. Here, we present scSAID (skin-scsaid.com), a single-cell database with an interactive web portal offering a broad suite of in-depth analyses for human and mouse skin. It integrates more than 1.2 million high-quality cells collected from 252 samples, establishing a unified reference for cell-type annotation, cross-species comparison and pathological studies. Using psoriasis as a case study, we demonstrate how scSAID can be used to evaluate how faithfully mouse models reproduce human pathology. Systematic comparison with the imiquimod-induced mouse model revealed numerous species-specific molecular signatures of psoriasis, including human-specific NFKB1 activation and STAT1 involvement, indicating that the current mouse model captures only limited aspects of the disease. We further introduce psoSpotter, a disease-biomarker-selection algorithm that, coupled with in silico perturbation using scSAID data, uncovers PPIA as a novel psoriasis drug target, illustrating the potential of scSAID for identifying therapeutic approaches. Overall, scSAID delivers a large-scale, cross-species skin single-cell resource and analysis platform, opening new opportunities for the discovery of disease-relevant targets in skin diseases.
bioinformatics2026-07-23v1Integrated Bioinformatics Analysis of PALB2 Reveals Expression Patterns, Molecular Interactions, and Prognostic Significance in Breast Cancer
Bithi, A. J.; Rahat, M. H.Abstract
Background: Partner and Localizer of BRCA2 (PALB2) is a key tumor suppressor gene involved in homologous recombination mediated DNA repair through its interactions with BRCA1 and BRCA2. Germline alterations in PALB2 have been associated with hereditary breast cancer risk; however, its broader molecular role in breast cancer progression and prognosis requires further investigation. Methods: A comprehensive in silico analysis of PALB2 was performed using publicly available databases and bioinformatics platforms. Differential expression of PALB2 in breast cancer were evaluated using GEPIA2. Prognostic significance was assessed through Kaplan Meier analyses for overall survival (OS) and disease free survival (DFS). Protein protein interaction (PPI) networks were constructed using STRING. Functional enrichment analyses of PALB2-associated genes were conducted using g. Mutational profiling of PALB2 in breast cancer was performed using cBioPortal with data from TCGA breast cancer cohorts. Results: PALB2 expression was elevated in breast tumor tissues compared with normal breast tissues. Survival analyses revealed no statistically significant association between PALB2 expression and either overall survival (HR = 0.88, p = 0.44) or disease free survival (HR = 0.74, p = 0.11). Protein interaction analysis revealed strong interactions between PALB2 and major DNA repair proteins including BRCA1, BRCA2, RAD51, RAD51C, FANCD2, and BRIP1. Functional enrichment analysis showed limited significant pathway enrichment, with only marginal transcription factor motif enrichment observed. Mutational analysis demonstrated diverse genomic alterations including missense mutations, truncating mutations, copy number gains, and shallow deletions. Conclusion: The findings support the biological relevance of PALB2 in breast cancer through its elevated expression and strong connectivity within DNA repair pathways. However, PALB2 expression alone does not appear to serve as an independent prognostic indicator. Further studies integrating genomic, transcriptomic, and clinical parameters are required to clarify its role in breast cancer progression and therapeutic response. Keywords: PALB2; Breast Cancer; Bioinformatics; Gene Expression Analysis; Protein Protein Interaction; Survival Analysis; Mutation Profiling; Homologous Recombination; TCGA; GEPIA2
bioinformatics2026-07-23v1A transparent multicriteria and fuzzy classification approach for genome-based probiotic candidate prioritisation
Ounissi, N. E.; Gomri, M. A.; El Hadef El Okki, M.Abstract
The identification of novel probiotic candidates with potential health-promoting properties remains a major challenge in food biotechnology and increasingly relies on in silico screening of genomic information. However, probiogenomic markers are heterogeneous, and safety, survival-colonisation, and functional-benefit traits do not contribute equally to probiotic potential. This study developed the Structured Probiotic Potential Index (SPPI), a fuzzy multicriteria system for genome-based probiotic candidate prioritisation. A hierarchical evaluation structure was established from probiogenomic evidence and organised into three main pillars and fourteen subcriteria. Expert judgements were collected using the Analytic Hierarchy Process, followed by consistency-based curation and weight aggregation. A curated dataset of 48 complete bacterial genomes, distributed into probiotic, potentially probiotic, neutral, and pathogenic groups, was taxonomically validated and analysed using an automated probiogenomic screening pipeline. Genome-wide screening generated 3,218 binary genomic features, from which curated probiogenomic markers were mapped to the scoring hierarchy. The resulting index was formulated as a normalised expert-weighted equation integrating safety, survival-colonisation, and functional-benefit components. SPPI prioritised genomes according to weighted probiogenomic profiles and separated pathogenic genomes from favourable probiotic and potentially probiotic profiles within the analysed dataset. Downstream Fuzzy Comprehensive Evaluation transformed the continuous score into probiotic/potentially probiotic, neutral, and pathogenic classes. Under internal leave-one-out evaluation, all 48 genomes were assigned to their expected reference classes, while confidence analysis distinguished high-confidence from borderline assignments. Post-classification comparison with ProbML showed concordant behaviour for most genomes and discordant predictions for selected cases. These results support SPPI as a transparent genome-based decision-support system for early probiotic candidate prioritisation before experimental validation.
bioinformatics2026-07-23v1PhytoFam: A Nextflow Pipeline for Genome-Wide Analysis of Plant Gene Families
Parajuli, S.; Adhikari, B.; Fennell, A.; Nepal, M. P.Abstract
Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complicates reproducibility and parameter tracking, and may require additional steps to remove redundant protein isoforms. To address these challenges, we developed PhytoFam, a Nextflow-based workflow that automates gene family identification from proteome input through phylogenetic reconstruction. The pipeline integrates HMMER for candidate sequence identification, isoform-aware deduplication, InterProScan for domain confirmation, BLAST reciprocal best hit (RBH) analysis for orthology assignment, MUSCLE for multiple sequence alignment with optional outgroup incorporation, TrimAl for alignment trimming, and IQ-TREE3 for phylogenetic reconstruction. PhytoFam is portable across local workstations and high-performance computing environments and supports deployment through Conda, Docker, and Singularity. We validated the workflow using the Morus alba MADS-box gene family, where the complete analysis finished in 1 h 10 min (9 CPU h). IQ-TREE3 accounted for most of the execution time, whereas InterProScan showed the highest memory requirement with a peak resident set size of 4.5 GB. PhytoFam provides a reproducible, automated, and scalable solution for plant gene family identification and phylogenetic analysis. The pipeline is freely available at https://github.com/sanamparajuli/PhytoFam.
bioinformatics2026-07-23v1An openly licensed benchmark and per-gene calibration map for missense pathogenicity predictors on activating cancer drivers
Lee, S.-G.Abstract
Missense pathogenicity predictors such as AlphaMissense are increasingly used in clinical variant interpretation, yet they are trained on germline labels dominated by loss-of-function (LOF) variants. Using an openly licensed, reproducible benchmark of 768 Cancer Gene Census genes scored with 49 predictors (labels from CIViC, COSMIC, cancerhotspots, ClinVar and gnomAD), we show that 42 of 49 tools (86%) score oncogene, gain-of-function (GOF) variants worse than tumour-suppressor variants. This under-scoring is mechanistically characterized: missed drivers occupy low-conservation, solvent-exposed, non-destabilizing positions (phyloP 2.51 versus 7.89; relative solvent accessibility 0.671 versus 0.185; gene-clustered p = 4.8x10-20 and 2.3x10-35), and, counter-intuitively, the unsupervised and protein-language models now entering clinical use are the most affected. Per-gene oncogenic thresholds span 0.07-0.99, so a single global cut-off is mis-calibrated for most genes; we provide a per-gene calibration map. A cancer-calibrated stack (OncoCal) modestly improves discrimination over the best single tool (AUROC {approx} 0.93 versus 0.87), rescues drivers such as JAK2 V617F (0.334[->]0.57), and generalizes to independent deep mutational scanning data. We provide an openly licensed framework to interpret and recalibrate these tools in the somatic setting rather than a replacement predictor.
bioinformatics2026-07-23v1pHaseMD4AI: Phase-Space Dynamics Dataset with Chemical and pH Perturbations for Physically and Kinetically Consistent Biomolecular AI
Song, T.; Guo, Y.; He, J.; Liu, Z.; Low, M.; Wang, K.; Zhang, Y.; Li, Z.; Huang, Y.; Wang, Y.Abstract
Protein function emerges from dynamic conformational ensembles and transitions that are challenging to characterize experimentally and computationally. Recent advances in generative AI have created new opportunities for learning molecular thermodynamics, kinetics, and conformational evolution directly from simulation data, but progress is limited by the availability of large-scale datasets that combine rigorous sampling, complete phase-space information, and diverse physicochemical perturbations. Here, we present pHaseMD4AI, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins. The peptide branch includes a complete set of canonical tripeptide and tetrapeptide systems together with post-translationally modified (PTM) and protonation-state datasets, providing synchronized atomic coordinates (R), velocities (V), forces (F), and Markov state model-based kinetic annotations. An accompanying web portal (https://isb.zju.edu.cn/md4ai/) enables users to browse, visualize, and download trajectories, annotations, and metadata. As an example application, we demonstrate a sequence-based model that can predict residue-level equilibrium dihedral distributions from sequence. pHaseMD4AI provides a resource for developing and benchmarking molecular machine learning methods while supporting broader studies of biomolecular dynamics under sequence, post-translational modification, and protonation-state perturbations.
bioinformatics2026-07-23v1Transcriptional landscape of direct reprogramming toward the hematopoietic lineage
Cwycyshyn, J.; Stansbury, C.; Golts, S.; Lee, H.; Pickard, J.; Meixner, W.; Rajapakse, I.; Muir, L. A.Abstract
Direct reprogramming of human fibroblasts into hematopoietic stem cells (HSCs) offers a promising strategy for generating autologous cells to treat blood and immune disorders. Current protocols are limited by low efficiency and insufficient tools for evaluating reprogramming outcomes. Although functional assays are the standard for confirming cell identity, they require fully reprogrammed cells, limiting their utility during protocol development. To address this, we assembled a single-cell transcriptomic reference atlas of hematopoietic reprogramming and tested an algorithmically-predicted transcription factor recipe for HSC induction. Long-read single-cell RNA sequencing of CD34+ reprogrammed cells revealed progressive loss of fibroblast identity alongside induction of early hematopoietic and endothelial programs, with reference-atlas benchmarking placing reprogrammed cells in an intermediate transcriptomic state between fibroblasts, endothelial cells, and HSCs. Isoform-level analysis further revealed transcriptional remodeling not captured by gene-level analyses. This experimental-computational framework offers a generalizable strategy for characterizing partially reprogrammed states and guiding optimization of reprogramming protocols.
bioinformatics2026-07-22v5Benchmarking the Impact of Data Leakage on the Performance of Knowledge Graph Embedding Models for Biomedical Link Prediction
BRIERE, G.; STOSSKOPF, T.; LOIRE, B.; BAUDOT, A.Abstract
Motivation: Knowledge Graphs (KGs) are increasingly used to organize complex biomedical knowledge into structured representations of entities and relations. Knowledge Graph Embedding (KGE) models facilitate efficient exploration of KGs by learning compact representations, and are widely applied to biomedical link prediction, for instance to uncover new therapeutic uses for existing drugs. Despite extensive work on KGE models, current evaluations often overlook data leakage, which can artificially inflate performance and undermine benchmark validity. Data leakage can arise when (1) there is redundancy between training and test sets, (2) the model leverages illegitimate features, or (3) the test set does not reflect real-world inference scenarios. Results: We assess the impact of data leakage on KGE-based link prediction across three biomedical KGs, using both decoder-only and GNN-based models. We first demonstrate the impact of train-test redundancies and implement a systematic procedure to detect and remove them. Using permutation experiments, we investigate whether node degree acts as an illegitimate predictive feature, and find no evidence that predictions are driven by degree alone. Finally, we evaluate how well common test set sampling strategies reflect real-world inference in drug repurposing. Comparing random and cold-start splits with an independent Orphanet-derived test set, we observe a substantial performance drop on the latter, indicating that current practices may overestimate how well KGE models generalize. Overall, our findings highlight the importance of rigorous benchmark design and careful evaluation of the generalization ability of KGE models for biomedical link prediction. Availability and Implementation: All code and results are openly available on GitHub at https://github.com/galadrielbriere/data_leakage_kge_benchmark.git.
bioinformatics2026-07-22v3Graph Lens Lite: A browser-based tool for interactive visualization and exploration of biological networks
Ley, M.; Keska-Izworska, K.; Fillinger, L.; Walter, S. M.; Baumgärtel, F.; Bono, E.; Galou, L.; Andorfer, P.; Hauser, P.; Leierer, J.; Kratochwill, K.; Perco, P.Abstract
Biological network visualization together with graph-based analyses are key techniques in systems biology and network medicine to detect patterns and generate hypotheses regarding disease pathobiology, drug target identification, biomarker prioritization, and digital drug discovery. Network representations provide an intuitive way to communicate and share research findings. We have developed Graph Lens Lite, a browser-based tool that combines rich visualization with a streamlined interface for exploring and sharing biological networks. It offers an expressive query language, topological network analysis, interactive filtering, visual grouping, customizable layouts, a data editor, fine-grained property-based styling, animated edge-flow visualization, community detection, and a context-aware locally powered AI assistant particularly suited for exploring molecular models of disease pathobiology or drug mechanism of action. We demonstrate its utility on a curated network model of autosomal dominant polycystic kidney disease. Graph Lens Lite is open source, with a live web version available at https://delta4ai.github.io/GraphLensLite/.
bioinformatics2026-07-22v3CESAR: A R Package for High-Sensitivity Detection of Copy Number Variations in ctDNA Using Segmentation and Anchor Recalibration
Lu, J.; Ni, S.; Wang, L.; Wu, N.; Jiang, X.Abstract
Background: Detecting copy number variations (CNVs) in circulating tumor DNA (ctDNA) is crucial for the companion diagnosis and resistance monitoring of various solid tumors (e.g., NSCLC, Glioblastoma). However, when tumor-derived DNA fractions are extremely low (often <1%), traditional depth-based methods frequently fail due to non-linear sequencing depth fluctuations and probe-specific capture biases inherent to targeted Next-Generation Sequencing (NGS). Methods: We developed CESAR (CNV Estimation with Segmentation and Anchor Recalibration), a novel computational tool optimized for ultra-sensitive, tumor-only CNV detection in targeted NGS panels. CESAR utilizes Circular Binary Segmentation (CBS) to re-partition target regions based on relative capture efficiency. It then introduces a dynamic "anchor" selection algorithm that identifies a personalized set of genomic segments mirroring the non-linear coverage behavior of each target gene. By minimizing the Coefficient of Variation (CV) through iterative anchor selection, CESAR effectively recalibrates the baseline to suppress technical noise. Results: Validation using standard DNA reference materials demonstrated that CESAR successfully identified both amplifications (e.g., MET, ERBB2, EGFR) and relative copy number deletions at ultra-low tumor fractions. Notably, CESAR achieved stable detection of focal alterations as subtle as 2.18 copies (a mere 1.09x fold change relative to the diploid baseline), while maintaining zero false positives in control regions. Evaluation across distinct clinical biofluids, 36 clinical plasma samples and 41 glioma cerebrospinal fluid (CSF) samples, identified critical, previously undetected CNV events, including subtle ERBB2 gains and distinct MET deletions. Furthermore, comprehensive benchmarking revealed that CESAR consistently outperformed the widely used CNVkit, particularly in suppressing technical variance and resolving ultra-low-level copy number gains that CNVkit failed to distinguish from background noise. Conclusions: CESAR provides a highly stable and sensitive algorithmic framework for tumor-only CNV calling in liquid biopsies, facilitating precise therapeutic decision-making in precision oncology.
bioinformatics2026-07-22v2Integrative structure determination of a human mitochondrial contact site and cristae organizing system (MICOS) sub-assembly
Jindal, M.; Mahato, R.; Das, S.; Guha, A.; Majila, K.; Arvindekar, S.; Vaidya, A. T.; Viswanath, S.Abstract
The Mitochondrial contact site and Cristae Organizing System (MICOS) complex is an inner mitochondrial membrane (IMM) assembly present at the cristae junction. It is responsible for regulating cristae formation and remodeling. However, its structure is not known. We applied Bayesian integrative structure determination to characterize the structure of the Mic60, Mic19, Mic10, and Mic13-containing MICOS complex combining AlphaFold predictions with data from crosslinking mass spectrometry, biochemical assays, electron tomography, homology modeling, and sequence alignments. The integrative structure revealed novel mutual interfaces among Mic10N,C, Mic60LBS1,LBS2,mitofilin, and Mic13central,C, which were experimentally validated. Several likely-pathogenic missense mutations also localize to these novel interfaces, highlighting their importance. Our results indicate that Mic13 likely facilitates MICOS assembly by binding Mic10 in the IMM-proximal region and Mic60 in the intermembrane space. Taken together, our integrative approach sheds light on the structure and assembly of the MICOS complex.
bioinformatics2026-07-22v1IntraTalker - Modelling Intracellular Signalling in1 Cellular Crosstalk
Kloker, V.; Nagai, J. S.; Feng, Z.; Mavrommatis, L.; Hermanns, L.; Moscoso, J. M. J.; Ruiz, M.; Kuppe, C.; Costa, I. G.Abstract
Single-cell sequencing has advanced the study of cell-cell communication, yet most methods focus on intercellular ligand-receptor interactions while neglecting downstream intracellular signalling cascades and the possibility that downstream target genes themselves encode ligands, thereby propagating communication across multiple cells. We present IntraTalker+CrossTalkeR that combines intracellular (IntraTalker) and intercellular (CrossTalkeR) signalling from multimodal single-cell data. IntraTalker infers cell-type-specific transcription factor activities and constructs receptomes that link receptors to downstream target genes, which are then integrated with ligand-receptor predictions in CrossTalkeR. To prioritize signalling receptors, the framework performs in silico receptor perturbation. In a murine bone marrow dataset, this recovered the known function of the Il1r1 receptor in driving myeloid progenitor cell-state changes. In a human myocardial infarction dataset, it predicted a novel role for IGF1R signalling in driving the differentiation of fibroblasts towards progenitor fibroblast states.
bioinformatics2026-07-22v1A Standardized Methodology for FAIRness Assessment and Multi-Dimensional Scoring in Agrosystem Research Data Infrastructures
Haleem, A. U.; Arend, D.; Etukala, J. R.; Mazon, E. R.; Schmidt, M.; Jung, J.; Martini, D.; Usadel, B.; Neidiger, C.; Ulrich, R.; Lange, M.Abstract
The NFDI-consortium FAIRagro has established a systematic framework for evaluating the FAIRness of Research Data Infrastructures (RDIs) within the German agrosystem research landscape. While FAIR principles are widely accepted, their practical implementation by research data infrastructures (RDI) remains challenging. By operationalizing the FAIR principles into a reproducible multi-dimentional scoring methodology, this initiative addresses the critical need for a transparent and citable benchmark of RDIs that moves beyond simple compliance. This paper details the underlying assessment criteria, comprising 20 aggregated core metrics, the iterative community-driven validation process, and the integration of these metrics into the FAIRagro Search Hub. This framework evaluates RDIs, like repositories or databases, instead of sampling hosted data sets., across the four distinct categories of FAIR independently, yielding granular, pillar-specific ratings. Unlike aggregate scoring models, which can inadvertently mask technical deficiencies by averaging performance across categories, this multi-dimentional approach ensures that a repositorys distinct strengths and bottlenecks remain fully visible. Our findings demonstrate that standardized scoring not only clarifies data accessibility for users but also highlights specific operational gaps, allowing repository providers to identify precisely where the service implementation can be enhanced. By establishing this data-driven service in the agronomy domain, we provide a scalable template for the broader NFDI and EOSC ecosystems to foster a culture of excellence in research data stewardship.
bioinformatics2026-07-22v1Hobrac: a reference-guided workflow for genome comparison and synteny visualization
Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.Abstract
Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO rthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.
bioinformatics2026-07-22v1scLEMBAS: Context-Aware Modeling of Signaling Pathway Activity at Single-Cell Resolution
Baghdassarian, H. M.; Meimetis, N.; Nordenstorm, O.; Joughin, B.; Nilsson, A.; Lauffenburger, D.Abstract
Cells sense and integrate extracellular cues through intracellular signaling networks that reshape transcription factor activity to dictate cellular responses. Signaling activity is difficult to decipher: it is non-linear, and it contains extensive feedback and crosstalk. Furthermore, the same perturbation can elicit markedly different responses depending on context (e.g., cell type, disease state, and tissue microenvironment) such that identical stimuli produce diverse responses in multicellular populations. Consequently, there is a vast combinatorial space of complex interactions and context-dependent responses that necessitate computational models. Computational models of single-cell perturbation responses are demonstrated to predict cellular responses, but are often limited in mechanistic insight. Prior knowledge networks offer a route to bridge predictive capability and interpretability. Here we present scLEMBAS, a context-aware, gray-box neural network that models signaling pathway activity at single-cell resolution while preserving mechanistic grounding. scLEMBAS encodes a prior-knowledge network of protein-protein interactions as a recurrent neural network whose learnable edge weights correspond to signaling interaction strengths. It also captures context and individual cell variance through compositional bias terms. An adversarial approach allows the model to answer a single-cell counterfactual - what a given cell's TF activity would be under a different perturbation or context - while involving mechanistic rather than simply relational information. Across two scRNA-seq datasets spanning single- and multi-perturbation settings, scLEMBAS accurately predicts out-of-distribution combinations of perturbation and context. Capturing population variance across individual cells enables the model to predict cell subtype specific perturbation responses, despite being agnostic to such labels. Beyond prediction, scLEMBAS learned parameters are biologically interpretable: learned edge weights carry information beyond network topology and "self-prune" spurious interactions, while the categorical bias nominates proteins associated with cell-type-specific perturbation states. Overall, scLEMBAS enables quantitative dissection of how signaling pathway activity is reshaped by perturbation within specific cellular contexts.
bioinformatics2026-07-22v1Learning Minimal Gene Programs for Disease-Aligned Representations
Madduri, A.; Patel, C. J.Abstract
Identifying small, interpretable gene sets that robustly capture disease-associated variation in single-cell transcriptomic data remains a central challenge for biological interpretation and experimental follow-up. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-held-out evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data.
bioinformatics2026-07-22v1Tangerine: A Python framework for dynamic gene regulation analysis from transcriptomic time series
Narendra, T.; Schweikert, G.Abstract
Motivation: Time-series single-cell transcriptomics enables the study of dynamic gene regulation. However, standard computational tools frequently aggregate temporal data into static, dense topologies, obscuring the precise regulatory rewiring that drives developmental transitions. Further, navigating the inherent noise of statistical inference without losing biological interpretability remains an important bottleneck. Results: We present Tangerine, a Python framework for the dynamic reconstruction and interactive exploration of time-varying gene regulatory networks. Tangerine integrates time-constrained metacell aggregation with regularized linear modelling and non-parametric correlation to infer dynamic topologies. To solve the interpretability gap, it features a browser-based visual analytics engine. Tangerine empowers researchers to track macroscopic gene module evolution, interactively filter effect sizes, and link topological rewiring directly to raw transcriptomic evidence. Availability and implementation: Tangerine is implemented in Python and Plotly Dash. The code is available on Github at https://github.com/ntanmayee/tangerine.
bioinformatics2026-07-22v1BioPhasor: Decoding Cellular State Tensors from Multi-Omics Phasor Dynamics for Quantum Ready Systems Biology
Sigdel, D.; Panday, N.Abstract
Integrating multi-omics data - transcriptomics, proteomics, metabolomics, single-cell - remains a fundamental challenge in systems biology. We present BioPhasor, a framework that encodes each measurement as a complex phasor z = e^(i{varphi}) on the compact N-torus T^N, modelling the cell as phase-coupled oscillatory programs whose dissipative dynamics generate limit cycles and an attractor landscape. From this geometry we derive the Cell State Tensor (CST), a rank-3 tensor whose axes we root in measured multi-omics quantities: a pathway/module atlas on the regulatory axis and a directional central-dogma modality axis. Across nine scenarios on open public data (GEO, CPTAC), loaded through one unmodified data layer, we report verdicts honestly: four reproduce, three are partial, two do not. A data-driven cell-cycle axis lifts agreement with a reference method from 0.34 to 0.69; an explicit circadian origin cuts peak-time error from 10.6 to 1.4 h; and central-dogma coupling mRNA phase organising protein amplitude clears a surrogate null and is tumour-specific. Grounding the quantum-ready claim, the CST maps to a density-matrix formalism whose coherence and entropy match quantum-information counterparts, and the phasor circuit transpiles gate-for-gate to a variational quantum circuit, though no empirical advantage emerges. One loader regenerates every reported number, and the code is released.
bioinformatics2026-07-22v1AI4Life Open Calls and Public Challenges: why, how, and what we have learned.
Galinova, V.; Seifi, M.; Serrano Solano, B.; Lidayova, K.; Dalle Nogare, D.; Corbat, A. A.; Talks, J.; Giacomello, E.; Gomez-de-Mariscal, E.; Ferreira, M. G.; Fuster-Barcelo, C.; Battagliotti, J. M.; Garcia-Lopez-de-Haro, C.; Salmon, B.; Croft, M.; Yie, S. Y.; Rey-Paniagua, G.; Hu, X.; Cho, S.; Sheth, A.; Porwal, C.; Li, X.; AI4Life Consortium, ; Henriques, R.; Li, X.; Krull, A.; Klemm, A.; Munoz Barrutia, A.; Kreshuk, A.; Ouyang, W.; Jug, F.; Deschamps, J.Abstract
Within AI4Life, we ran three Open Calls and three Public Challenges (2023-2025), supporting 22 bioimage analysis projects from 151 applications and engaging 225 challenge participants, with the aim of applying FAIR deep learning in the life sciences. Our experience offers a view of the current state of bioimage analysis, the landscape of available tools, as well as the existing gaps between method developers, tool producers and potential users. It highlights that even after careful selection for AI-ready projects, most still require substantial effort to apply deep learning, and that the field still relies heavily on established, well-rounded methods to solve common problems. We come to the conclusion that for scientific AI in biology, the rate-limiting step is not methods and models but data, annotations, and shared infrastructure underneath them.
bioinformatics2026-07-22v1TrioNsight: Building a meta-predictor to evaluate the clinical impact of TrioN-like Dbl-homology domain variants
Taciroglu, A.; Aydin Son, Y.; Martin, A. C. R.; Orengo, C.Abstract
TRIO is a member of the Rho family of guanine nucleotide exchange factors (Rho GEFs), which promote the exchange of GDP for GTP to activate Rho GTPases and serve as key regulators of cellular signalling pathways. Mutations in TRIO are associated with neurodevelopmental disorders, including intellectual disability and autism spectrum disorders. TRIO contains two GEF units: one N-terminal and one C-terminal, each of which contains a Dbl-homology (DH) domain that drives its GEF activity to activate Rho GTPases. While the human proteome contains 70 highly conserved DH domains, the N-terminal DH domain of TRIO (TrioN) contains one-third of all reported DH domain pathogenic variants and has many variants of unknown significance. Numerous variant impact prediction tools exist, but most lack gene-specific considerations. Here, we describe TrioNsight, a meta-predictor designed to predict mutation impacts for TrioN and 12 highly similar human DH domains, including TrioC. TrioNsight exploits the naive-Bayes algorithm and leverages structural, evolutionary, and physiochemical features of approximately 1500 highly similar DH domains (TrioN-like DH domains) from 294 species. TrioNsight surpasses all available predictors, including AlphaMissense, achieving a Matthews' Correlation Coefficient of 0.890. Additionally, we provide a variant impact map that details the impacts of mutations at each position in the DH domain of these proteins, which can be valuable for clinical assessments. Furthermore, our approach establishes a standardised workflow adaptable for creating domain-specific variant predictors for other protein families, offering a template for improved variant interpretation.
bioinformatics2026-07-22v1IOBRpy enables agentic multi-omics decoding of anti-tumor immunity
Huang, H.; Li, X.; Liu, L.; Gu, W.; Wang, G.; Zeng, D.Abstract
Decoding the tumor immunity is pivotal for cancer immunotherapy, yet transcriptomic pipelines remain bottlenecked by fragmented tools and biased interpretations. Here we present IOBRpy, a Python toolkit driven by an innovative AI dual-agent layer for automated, highly standardized immuno-oncology workflows. Moving beyond conventional expression profiling, IOBRpy enables agentic multi-omics decoding. From raw FASTQ or TPM matrices, it seamlessly integrates upstream quality control, transcript quantification, and downstream TME parsing, encompassing signature scoring, ligand-receptor crosstalk, and cellular deconvolution. Crucially, IOBRpy expands data dimensions by incorporating complementary immunogenomic layers, empowering concurrent high-resolution SpecHLA typing and TRUST4-based TCR/BCR repertoire reconstruction from sequencing data. Deployed across two large-scale cohorts (IMvigor210 and OAKPOPLAR), IOBRpy successfully captured multi-dimensional prognostic insights. While broad HLA-I heterozygosity showed negligible impact, it precisely unmasked treatment-stratified, allele-specific survival associations (e.g., HLA-A*01 and HLA-DPA1*02) tightly coupled with distinct immunosuppressive ligand-receptor networks (such as HMGB1-THBD and EFNB2-EPHB6) and dynamic TCR clonal diversity shifts. Empowering this lifecycle is a paired agent framework: a workflow agent automatically audits project states to execute validated, path-aware commands, while a result agent evaluates tool provenance, handles method-aware adaptive visualizations, and organizes findings into evidence-constrained biological hypotheses. Collectively, IOBRpy provides a reproducible, scalable, and intelligence-augmented Python gateway to transform raw sequencing data into multi-omic, interpretation-ready discoveries for cohort-scale precision immunotherapy (https://iobr.github.io/IOBRpy/).
bioinformatics2026-07-22v1Target Preference Maps: A machine learning model generalizing transferable drug-receptor interactions and guiding drug discovery
Menezes, F.; Wahida, A.; Froehlich, T.; Grass, P.; Zaucha, J.; Napolitano, V.; Siebenmorgen, T.; Pustelny, K.; Barzowska-Gogola, A.; Rioton, S.; Didi, K.; Bronstein, M.; Czarna, A.; Hochhaus, A.; Plettenburg, O.; Sattler, M.; Nissen-Meyer, J.; Conrad, M.; Kurzrock, R.; Popowicz, G. M.Abstract
Modern AI models can decode the genomic landscape and protein structure world. Yet, they fail to generalize to one of the most important fields: small-molecule drug discovery. Since the late 1970s, the advent of macromolecular crystallography inspired the notion that structural knowledge alone could enable a lock-and-key approach to drug design. However, drug discovery continues to depend on costly, resource-intensive, and largely serendipitous screening campaigns that probe only an infinitesimal fraction of the drug-like chemical space. Despite some successful cases, our understanding of, and reasoning from, non-bonded interaction chemistry remains limited for general applicability. Furthermore, though structural databases contain hundreds of thousands of entries, a strong historical bias pervades protein-drug structures, hindering reliable advances through AI scaling. Here, we present a machine-learning framework that learns atom-type-specific spatial preference maps from local protein microenvironments in protein-ligand structures. By excluding ligand topology from the model input and learning from local atom-level environments, the framework is designed to reduce dependence on whole-ligand memorization and to capture transferable interaction preferences. The resulting maps recover chemically meaningful interaction patterns, including cases involving bridging waters and metal-dependent environments. The model was validated using retrospective and prospective real-world data in drug optimization when targeting a challenging protein-protein interface. This shows that the method can provide interpretable workflows to guide molecule optimization and provide input for downstream generative or docking workflows.
bioinformatics2026-07-21v11LeafRank: A phylodynamic framework for inferring relative fitness from single-cell phylogenies in chromosomally unstable tumors
Wu, C.; Leder, K.; Wang, Z.; Sun, R.Abstract
Tumors contain cancer cells with diverse growth potentials that shape evolutionary trajectories, yet this fitness diversity remains difficult to quantify in cases of whole-genome duplication (WGD) and chromosomal instability. We present LeafRank, a mathematical framework that leverages single-cell DNA-seq phylogenies to infer the relative fitness of individual cells. Using a multi-type branching process model, LeafRank integrates full tree topology, including branch lengths and bifurcation patterns, to estimate marginal fitness probabilities under punctuated evolutionary regimes driven by rare driver events. To account for elevated aberration rates following WGD, we introduce a tree-rescaling strategy that adjusts for lineage-specific genomic instability. Unlike methods focused on predefined subclones, LeafRank ranks all sampled cells, enabling flexible assessment of growth heterogeneity. Simulations demonstrate high accuracy across spatial and non-spatial virtual tumors. Applied to ovarian cancer, LeafRank reveals directional and parallel selection in WGD tumors and identifies recurrent copy number events enriched in high-fitness lineages. WGD lineages do not show immediate growth advantages but acquire fitness through subsequent alterations.
bioinformatics2026-07-21v4RNAStabFormer: Region-Aware Multi-Task Hybrid Learning for RNA Stability Prediction from Pulse-Chase Transcriptomics
Wang, S.; Zhang, C.Abstract
RNA stability is a major post-transcriptional regulator of gene expression, yet sequence-based prediction from pulse-chase transcriptomics remains difficult because labels depend on the time window, quantification region, and replicate quality. We present RNAStabFormer, a controlled RNA stability framework centered on a Region-Aware Multi-Task Hybrid Transformer (RAMHT). RAMHT encodes the nucleotide context of the 5-prime UTR, coding sequence (CDS), and 3-prime UTR; incorporates an additional CDS codon stream; upgrades engineered sequence features into a tabular interaction branch; and uses gated multi-task regression to predict four ENCODE BrU-seq/BruChase-seq RNA stability proxies, with Exon 6 h/0 h as the primary task. Across 26 outer data splits, including 23 chromosome holdout tests, a heterogeneous three-member RAMHT ensemble achieves a mean Pearson correlation of 0.773 on the primary task, statistically matching an engineered-feature XGBoost baseline, which also achieves 0.773. The mean paired difference is +0.000004, with a bootstrap 95 percent confidence interval from -0.003845 to +0.004077 and a Wilcoxon p-value of 0.8613. The ensemble improves upon the strongest individual RAMHT member, increasing the mean Pearson correlation from 0.768 to 0.773 and outperforming it on 23 of the 26 data splits. A strictly nested XGBoost-RAMHT blended model further increases the correlation to 0.775. Evaluations conducted on identical data splits also show that the ensemble outperforms frozen full-length mRNA language-model embeddings, which achieve a correlation of 0.760, and public LAMAR-DR transfer learning, which achieves a correlation of 0.180. Gate analysis, ablation experiments, and sequence recoding analyses indicate that engineered sequence grammar remains the dominant source of predictive information, whereas the nucleotide and codon branches provide complementary signals localized primarily within the CDS. RNAStabFormer narrows the performance gap between neural RNA sequence models and strong tabular baselines while retaining an extensible architecture for model interpretation and biological data integration.
bioinformatics2026-07-21v2Panomap: Unbiased Nanopore Signal Mapping with Pangenome Variation Graphs
Shih, P. J.; Sanghani, Z.; Guarracino, A.; Gamaarachchi, H.; Batten, C.Abstract
Motivation: Signal-space nanopore mappers enable real-time mapping and filtering decisions directly from raw nanopore signals. However, existing signal-space mappers are built around linear references, and using a single representative reference can introduce reference bias when the sample diverges from that reference. Pangenome reference collections can reduce this bias by representing diversity across related reference sequences, but linear-reference signal mappers must treat each sequence as a separate target, redundantly storing shared sequences. Pangenome variation graphs provide a more compact representation by storing shared sequences once and encoding variants as alternative paths through the graph. Although sequence-to-graph mapping is well established for basecalled reads, existing signal-space methods do not directly use pangenome variation graphs. Results: We present Panomap, the first signal-space mapper that operates on pangenome variation graphs. Panomap maps raw nanopore signals to graph references, allowing signal-space mapping to use pangenome diversity while representing shared sequences once. We evaluate Panomap in three settings. First, when a single reference already maps the sample well, Panomap preserves mapping accuracy as additional reference sequences are added to the reference collection, while state-of-the-art signal-space tools regress. Second, when the exact sample strain is absent from the reference collection, Panomap benefits from adding related assemblies from the same species to the pangenome reference. Third, using a highly polymorphic locus, we show that Panomap can map reads from alleles not represented in the reference collection by using related alleles in the pangenome, with the largest gains for more divergent alleles and for decisions made from short prefixes of the read signal. In addition, Panomap's graph index scales sublinearly with pangenome collection size. Together, these results show that Panomap brings population-aware reference representation into signal-space mapping. Availability and Implementation: Panomap is open source and available at https://github.com/cornell-brg/panomap.
bioinformatics2026-07-21v2Genomic, Transcriptomic, and Regulomic Analyses Do Not Support Profound Autism as a Distinct Biological Category
Eicher, T. D.; Ne'eman, A.; Quackenbush, J. D.Abstract
The Lancet Commission on the Future of Care and Clinical Research in Autism proposed the construct of "profound autism" as a recognizable subtype of autism. Supporters argue that this classification is necessary to ensure that autistic persons with severe impairment receive appropriate research attention and policy support, whereas critics contend that the construct lacks scientific validity and may reflect social or political considerations more than biological distinction. To inform this debate, we evaluate whether the proposed "profound autism" category represents a distinct genetic phenotype using multiple molecular data types collected in a large cohort. Across genomic, transcriptomic, and regulatory analyses, we find no evidence supporting "profound autism" as a biologically distinct phenotypic group. Instead, differences emerge primarily in inferred gene regulatory networks distinguishing nonspeaking from speaking autistic children, suggesting potential regulatory mechanisms contributing to speech ability. These findings suggest that future research into severe impairment may be more productive if focused on specific traits -- such as speech impairment -- rather than attempting to define a distinct biological subtype within the multidimensional phenomenon of autism.
bioinformatics2026-07-21v2SEEK-VEC: Augmenting topic modeling with spectral ensemble learning
Danning, R.; Ke, Z. T.; Ma, R.; Lin, X.Abstract
Count data are ubiquitous across many applications in which understanding latent patterns is of interest. Topic modeling is a powerful tool for detecting latent structure in count data. However, standard topic modeling methods are often constrained by their restrictive assumptions, susceptible to noise, and sensitive to misspecification of the number of topics. Here, we introduce SEEK-VEC (Spectral Ensembling of topic models with Eigenscore for K-agnostic Vocabulary Embedding and Classification), an ensemble topic modeling framework that integrates insights from multiple candidate topic models through a spectral ensembling procedure. SEEK-VEC produces a meta-structure matrix containing prioritization scores and grouping scores that enable variable classification, interactive pattern discovery, and model diagnostics. Through simulations, we demonstrate that SEEK-VEC augments the performance of standard topic models for identifying important vocabulary words and understanding the relationships among them, particularly when signal strength is weak. We apply SEEK-VEC to the MADStat dataset of statistical abstracts and demonstrate its utility for evaluating the proposed interpretation of a topic model.
bioinformatics2026-07-21v2Molecular Mimicry in Inflammatory Bowel Disease: Multi-layered Functional and Sequence-level Analysis of Gut Microbial Proteins Mimicking the Human Proteome
Anand, A. A.; Mishra, P.; Srivathsa, V. S.; Yadav, V.; Samanta, S. K.Abstract
Inflammatory bowel disease (IBD), encompassing Crohn's disease (CD) and ulcerative colitis (UC), is a chronic inflammatory disorder whose pathogenesis involves intricate host-microbiome interactions. Molecular mimicry (the structural or functional resemblance between microbial and host proteins) represents a plausible mechanism by which gut microbiota may trigger or perpetuate autoimmune responses. Here we present a comprehensive, multi-layered molecular mimicry in silico pipeline (MMIP) analysis of 39 baseline shotgun metagenomic samples from Human Microbiome Project 2 (HMP2/IBDMDB). Using DIAMOND-based homology search against the SwissProt followed by UniProt Retrieve/ID Mapping (URIM), we characterized microbial protein functional space through three complementary frameworks: (i) normalized GO term frequency comparison across diagnostic groups, (ii) protein family (PFAM) domain enrichment analysis, and (iii) sequence-level mimicry analysis identifying microbial proteins with direct homology to human proteins. Taxonomic profiling using MetaPhlAn3 pre-computed abundance profiles provided further biological context. CD microbiomes exhibited greater enrichment of immune-relevant biological processes and a higher per-sample sequence-level mimicry rate than UC. CD and UC also showed distinct pathobiont profiles, with CD enriched for oral-origin taxa including Haemophilus parainfluenzae and UC enriched for Fusobacterium nucleatum and Prevotella species. Notably, healthy gut microbiome maintains coordinated mimicry of host neuronal, signal-recognition-particle-associated, and antimicrobial peptide machinery, a repertoire dismantled in IBD and replaced by disease-subtype-specific signatures, alongside a candidate mimicry link between CD-enriched bacteria and NOD2. These results represent the first metagenome-wide characterization of molecular mimicry across IBD subtypes using shotgun metagenomic data, offering new mechanistic insight into how microbial dysbiosis may contribute to immune dysregulation in IBD. Keywords: Inflammatory bowel disease (IBD), Crohn's disease (CD), Ulcerative colitis (UC), Gastroenteritis, Molecular mimicry
bioinformatics2026-07-21v2COSMOS: A FAIR-aligned infrastructure for clinical trial data validation, warehousing, and interactive discovery
Roberts, C. A.; Nilsson-Takeuchi, A.; Stuart, C.; Chivers, M.; Soares, P.; Foure, V.; Griffiths, G.; Niazi, U.Abstract
The Clinical Omics System for Metadata and Outcome Storage (COSMOS) is an open-source, FAIR- and GCP aligned clinical trial unit (CTU) infrastructure designed to streamline the transition of academic clinical and multi-omic trial datasets into curated, analysis-ready repositories within Secure Data Environments (SDEs). By integrating an automated Data Quality and Data Validation (DQ&DV) "Trust Layer" with a relational Structured Query Language (SQL) schema, COSMOS enables programmatic and interactive data access via R Shiny applications. Dynamic integration of omics and clinical data is achieved through analytical data structures (e.g. ExpressionSet objects) linked via relational database identifiers. This lowers technical barriers for researchers and promotes governed data reuse and secondary discovery.
bioinformatics2026-07-21v1Bayesian Factor Analysis for Binary and Ordinal Phenotypes with Missingness
Shashaank, N.; Knowles, D. A.Abstract
Binary and ordinal phenotypes are common in clinical screening and self-reported questionnaires, but many factor analysis and matrix factorization methods are only applicable for quantitative phenotypes with real-valued and/or continuous data distributions. To address this, we propose FABOr (Factor Analysis of Binary and Ordinal data), a Bayesian framework for matrix factorization in which the low-rank matrices are latent variables with continuous priors while the phenotypes are observed variables modeled with appropriate binary/ordinal likelihoods. We also develop missing not at random (MNAR) extensions of FABOr for analyzing data with structured missingness. In experiments with simulated phenotypes, we found that FABOr performs similarly to the best-performing benchmark methods on binary data at imputation and exceeds the performance of all tested benchmark methods on ordinal data. We then applied FABOr to analyze real-world binary and ordinal phenotypes from the Simons Foundation SPARK dataset on autism spectrum disorder (ASD) and found that it improved imputation accuracy by up to 5% on binary data and up to 23% on ordinal data relative to the benchmark methods.
bioinformatics2026-07-21v1