SimBu: Bias-Aware simulation of bulk RNA-seq data with variable cell-Type composition

Alexander Dietrich, Gregor Sturm, Lorenzo Merotto, Federico Marini, Francesca Finotello, Markus List

Publikation: Beitrag in FachzeitschriftArtikelBegutachtung

7 Zitate (Scopus)

Abstract

Motivation: As complex tissues are typically composed of various cell types, deconvolution tools have been developed to computationally infer their cellular composition from bulk RNA sequencing (RNA-seq) data. To comprehensively assess deconvolution performance, gold-standard datasets are indispensable. Gold-standard, experimental techniques like flow cytometry or immunohistochemistry are resource-intensive and cannot be systematically applied to the numerous cell types and tissues profiled with high-Throughput transcriptomics. The simulation of 'pseudo-bulk' data, generated by aggregating single-cell RNA-seq expression profiles in pre-defined proportions, offers a scalable and cost-effective alternative. This makes it feasible to create in silico gold standards that allow fine-grained control of cell-Type fractions not conceivable in an experimental setup. However, at present, no simulation software for generating pseudo-bulk RNA-seq data exists. Results: We developed SimBu, an R package capable of simulating pseudo-bulk samples based on various simulation scenarios, designed to test specific features of deconvolution methods. A unique feature of SimBu is the modeling of cell-Type-specific mRNA bias using experimentally derived or data-driven scaling factors. Here, we show that SimBu can generate realistic pseudo-bulk data, recapitulating the biological and statistical features of real RNA-seq data. Finally, we illustrate the impact of mRNA bias on the evaluation of deconvolution tools and provide recommendations for the selection of suitable methods for estimating mRNA content. SimBu is a user-friendly and flexible tool for simulating realistic pseudo-bulk RNA-seq datasets serving as in silico gold-standard for assessing cell-Type deconvolution methods.

OriginalspracheEnglisch
Seiten (von - bis)II141-II147
FachzeitschriftBioinformatics
Jahrgang38
DOIs
PublikationsstatusVeröffentlicht - 1 Sept. 2022

Fingerprint

Untersuchen Sie die Forschungsthemen von „SimBu: Bias-Aware simulation of bulk RNA-seq data with variable cell-Type composition“. Zusammen bilden sie einen einzigartigen Fingerprint.

Dieses zitieren