# Pseudocount TPM Normalization Gene Expression

**URL:** <https://forum.depmap.org/t/pseudocount-tpm-normalization-gene-expression/547>\
**Category:** Q&A\
**Created:** [April 8, 2021, 9:41pm UTC](https://forum.depmap.org/t/pseudocount-tpm-normalization-gene-expression/547 "2021-04-08T21:41:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![mavergara](https://avatars.discourse-cdn.com/v4/letter/m/85f322/32.png) [@mavergara](https://forum.depmap.org/u/mavergara)\
**Post date:** [April 8, 2021, 9:41pm UTC](https://forum.depmap.org/t/pseudocount-tpm-normalization-gene-expression/547/1 "2021-04-08T21:41:19Z")

</div>

Dear all,  
I have doubt about the normalization applied for RNAseq gene expression data. In the download portal says for gene expression: **" Log2 transformed, using a pseudo-count of 1 " (CCLE)**

On the other hand, looking on the database of “Genomics of Drug Sensitivity in Cancer” (GDSC) using the web engine tool “Orcestra” ([ORCESTRA](https://www.orcestra.ca/pset/10.5281/zenodo.3905481)) they say that

_**Gene TPM Values: After estimation by the tool detailed above, gene TPM values are transformed by log2(x + 0.001).**_

Therefore, we can conclude that the only difference its just that you guys for the RNAseq expression data you just add 1 instead of 0.001 like the guys GDSC?

Thanks!

---

<div class="post-metadata">

**Author:** ![jnoorbak](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.depmap.org/jnoorbak/32/12_2.png) [@jnoorbak](https://forum.depmap.org/u/jnoorbak)\
**Post date:** [April 9, 2021, 3:45pm UTC](https://forum.depmap.org/t/pseudocount-tpm-normalization-gene-expression/547/2 "2021-04-09T15:45:21Z")

</div>

Hi, that is correct. The only difference as you mentioned is in the pseudocount value being 1 in our data reports. This does not have any significant analysis value and was mainly chosen for historical reasons and to reduce the dynamic range of expression while avoiding negative and/or infinite values. -thanks
