<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/43088?offset=360</link>
	<atom:link href="https://bioinformaticsonline.com/related/43088?offset=360" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/44713/understanding-rna-seq-normalization-methods-tpm-vs-fpkm-vs-cpm</guid>
	<pubDate>Wed, 11 Dec 2024 00:59:15 -0600</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/44713/understanding-rna-seq-normalization-methods-tpm-vs-fpkm-vs-cpm</link>
	<title><![CDATA[Understanding RNA-Seq Normalization Methods: TPM vs. FPKM vs. CPM]]></title>
	<description><![CDATA[<p>RNA sequencing (RNA-Seq) is a powerful technology used to study transcriptomes, providing insights into gene expression levels. However, raw RNA-Seq data requires normalization to account for sequencing depth and gene length, enabling accurate comparisons between genes and samples. Among the most widely used normalization methods are TPM (Transcripts Per Million), FPKM (Fragments Per Kilobase Million), and CPM (Counts Per Million). Each method has its unique principles and applications, which we&rsquo;ll explore in this blog.</p><h2>Why Normalize RNA-Seq Data?</h2><p>Normalization is a crucial step in RNA-Seq analysis for the following reasons:</p><ul>
<li>
<p><strong>Sequencing depth:</strong> Different RNA-Seq experiments produce varying numbers of reads, making direct comparisons between samples misleading.</p>
</li>
<li>
<p><strong>Gene length:</strong> Longer genes inherently generate more reads, irrespective of their actual expression level.</p>
</li>
<li>
<p><strong>Bias reduction:</strong> Normalization mitigates technical biases, enabling meaningful biological interpretation.</p>
</li>
</ul><h2>TPM (Transcripts Per Million)</h2><p>TPM measures the proportion of reads mapped to a transcript, normalized by transcript length and sequencing depth. It is calculated as:</p><h3>Key Features:</h3><ol>
<li>
<p><strong>Proportionality:</strong> TPM values sum to 1,000,000 across all transcripts in a sample, making it easier to compare between samples.</p>
</li>
<li>
<p><strong>Intuitive interpretation:</strong> TPM values directly represent the abundance of transcripts in a sample.</p>
</li>
<li>
<p><strong>Preferred for comparisons:</strong> TPM facilitates between-sample comparisons better than FPKM.</p>
</li>
</ol><h2>FPKM (Fragments Per Kilobase Million)</h2><p>FPKM normalizes read counts by transcript length and sequencing depth, but without enforcing proportionality like TPM. It is defined as:</p><h3>Key Features:</h3><ol>
<li>
<p><strong>Historical significance:</strong> FPKM was one of the first normalization methods used for RNA-Seq.</p>
</li>
<li>
<p><strong>Single-end vs. paired-end:</strong> In paired-end sequencing, FPKM becomes RPKM (Reads Per Kilobase Million).</p>
</li>
<li>
<p><strong>Limited utility:</strong> FPKM values are not as robust as TPM for cross-sample comparisons due to lack of proportionality.</p>
</li>
</ol><h2>CPM (Counts Per Million)</h2><p>CPM normalizes raw read counts by sequencing depth, without considering gene length. It is expressed as:</p><h3>Key Features:</h3><ol>
<li>
<p><strong>Simplicity:</strong> CPM is straightforward and computationally less intensive.</p>
</li>
<li>
<p><strong>Application:</strong> Suitable for non-length-dependent analyses, such as comparing total expression levels or differential expression analysis.</p>
</li>
<li>
<p><strong>Gene length agnostic:</strong> CPM does not correct for gene length, making it less ideal for measuring expression levels.</p>
</li>
</ol><h2>When to Use Each Method</h2><ul>
<li>
<p><strong>TPM:</strong> Best for comparing expression levels between samples, especially when transcript length and sequencing depth vary.</p>
</li>
<li>
<p><strong>FPKM:</strong> Useful for historical consistency but generally replaced by TPM.</p>
</li>
<li>
<p><strong>CPM:</strong> Ideal for differential expression analysis when gene length normalization is unnecessary.</p>
</li>
</ul><h2>Conclusion</h2><p>Choosing the right normalization method depends on the specific objectives of your RNA-Seq analysis. TPM&rsquo;s proportionality and robustness make it the preferred choice for most applications, while CPM serves well for differential expression studies. Although FPKM paved the way for RNA-Seq normalization, it has largely been supplanted by TPM in modern workflows. Understanding these methods and their nuances ensures accurate and meaningful interpretations of RNA-Seq data.</p><h3>References:</h3><ol>
<li>
<p>Li, B., &amp; Dewey, C. N. (2011). RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. <em>BMC Bioinformatics.</em></p>
</li>
<li>
<p>Trapnell, C., et al. (2010). Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. <em>Nature Biotechnology.</em></p>
</li>
<li>
<p>Law, C. W., et al. (2014). voom: precision weights unlock linear model analysis tools for RNA-seq read counts. <em>Genome Biology.</em></p>
</li>
</ol>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/1471/24-mb-genome-size-for-worlds-biggest-virus</guid>
	<pubDate>Thu, 08 Aug 2013 10:05:37 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/1471/24-mb-genome-size-for-worlds-biggest-virus</link>
	<title><![CDATA[2.4 Mb Genome Size for World's Biggest Virus]]></title>
	<description><![CDATA[<p>The genome size of new discovered Pandoraviruses have roughly twice the size of the record-holding Megavirus genomic code. Interestingly only 6 percent of its genes resembled the genes other organisms. It is assume that it may come from a different origin.</p><p>For detail : http://www.sciencemag.org/content/341/6143/281</p><p>http://www.npr.org/blogs/health/2013/07/18/203298244/worlds-biggest-virus-may-have-ancient-roots</p><p>&nbsp;</p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/13852/ebola-virus-disease-evdor-ebola-haemorrhagic-fever</guid>
	<pubDate>Sun, 10 Aug 2014 13:08:13 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/13852/ebola-virus-disease-evdor-ebola-haemorrhagic-fever</link>
	<title><![CDATA[Ebola virus disease (EVD)or Ebola haemorrhagic fever !!!]]></title>
	<description><![CDATA[<p>Ebola virus disease (EVD)or Ebola haemorrhagic fever is a severe and often deadly illness in humans, caused by the Ebola virus. The disease has high mortality rate, killing upto 90% of people who are infected.</p><p><img src="http://s4.reutersmedia.net/resources/r/?m=02&amp;d=20140808&amp;t=2&amp;i=959839176&amp;w=580&amp;fh=&amp;fw=&amp;ll=&amp;pl=&amp;r=LYNXMPEA770BX" width="580" height="452" alt="image" style="border: 0px;"></p><p><br />The ongoing 2014 West Africa Ebola outbreak is considered to be the largest and longest outbreak ever recorded of Ebola, killing at least 932 people and infecting more than 1,700 till date since March in Sierra Leone, Guinea, Nigeria and Liberia.<br /><br />Hence, the World Health Organisation (WHO) on 8 August, 2014 declared the killer Ebola epidemic ravaging parts of West Africa an international health emergency.<br /><br />Causes<br /><br />EVD is caused by infection with a virus of the family Filoviridae, genus Ebolavirus. While there are five identified sub-species of Ebolavirus, four viruses cause disease in humans. They are Bundibugyo virus (BDBV), Ebola virus (EBOV), Sudan virus (SUDV), Ta&iuml; Forest virus (TAFV).<br /><br />The fifth virus, Reston virus (RESTV), is not considered to be disease-causing in humans.<br /><br />According to WHO, EVD first appeared in 1976 in two simultaneous outbreaks, in Nzara, Sudan, and in Yambuku, Democratic Republic of Congo. The latter was in a village situated near the Ebola River from which the disease takes its name.</p><p>How does it spread?<br /><br />It is still unclear how Ebola spreads. However, it is believed that the first pateint becomes infected through contact with an infected animal's body fluids.<br /><br />Human-to-human transmission can occur through direct contact with blood, organs or other body fluids of infected people or exposure to objects such as needles and syringes that have been contaminated with infected secretions.<br /><br />Ebola can also be transmitted from men who have recovered from the disease through semen as it is infectious for up to 7 weeks.<br /><br />Infected dead bodies can spread Ebola as they are still infectious. So mourners who have direct contact with the body of deceased person can also get the disease.<br /><br />Who is most at risk?<br /><br />Health-care workers who do not wear appropriate protective clothing and family members who are in close contact with infected people or deceased patients.<br /><br />Signs and symptoms:<br /><br />Symptoms may occur between 2 and 21 days after contracting the infection. Common signs of Ebola include:</p><p><img src="https://scontent-b-sin.xx.fbcdn.net/hphotos-xap1/t1.0-9/p720x720/10494629_873450929332827_3274653669306581755_n.jpg" width="720" height="720" alt="image" style="border: 0px;"></p><p>Fever<br /><br />Headache<br /><br />Muscle, abdominal and joint pain<br /><br />Sore throat<br /><br />Weakness<br /><br />Diarrhea<br /><br />Vomit or cough up blood<br /><br />Chest pain<br /><br />Difficulty in breathing and swallowing<br /><br />Rash<br /><br />Hiccups<br /><br />Bleeding inside and outside the body<br /><br />Prevention<br /><br />Currently there is no vaccine available for humans. But the infection can be controlled through the use of recommended protective measures such as:<br /><br />Avoid contacting infected blood or secretions, including from those who are dead .<br /><br />Using standard precautions for all patients in the healthcare setting.<br /><br />Sterilizing equipment, and wearing protective clothing including masks, gloves, gowns and goggles.<br /><br />Washing your hands with soaps or detergents.<br /><br />Disinfecting your surroundings.<br /><br />Isolate people who have Ebola symptoms.<br /><br />Culling of infected animals, with close supervision of burial or incineration of carcasses.<br /><br />Yet, not travelling to the areas or countries where the virus is found is the best way to avoid Ebola.</p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/43940/langya-virus-update</guid>
	<pubDate>Fri, 12 Aug 2022 05:31:10 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/43940/langya-virus-update</link>
	<title><![CDATA[Langya Virus Update !]]></title>
	<description><![CDATA[<p>https://www.ncbi.nlm.nih.gov/nuccore/OM101125,OM101126,OM101127,OM101128,OM101129,OM101130?</p><p>Zoonotic Henipavirus</p><p>https://pubmed.ncbi.nlm.nih.gov/35921459/</p><p>https://www.ncbi.nlm.nih.gov/nuccore/OM069646,,OM069567,OM069568,OM069569,OM069570,OM069571,OM069572,OM069573,OM069574,OM069575,OM069576,OM069577,OM069578,OM069579,OM069580,OM069581,OM069582,OM069583,OM069584,OM069585,OM069586,OM069587,OM069588,OM069589,OM069590,OM069591,OM069592,OM069593,OM069594,OM069595,OM069596,OM069597,OM069598,OM069599,OM069600,OM069601,OM069602,OM069603,OM069604,OM069605,OM069606,OM069607,OM069608,OM069609,OM069610,OM069611,OM069612,OM069613,OM069614,OM069615,OM069616,OM069617,OM069618,OM069619,OM069620,OM069621,OM069622,OM069623,OM069624,OM069625,OM069626,OM069627,OM069628,OM069629,OM069630,OM069631,OM069632,OM069633,OM069634,OM069635,OM069636,OM069637,OM069638,OM069639,OM069640,OM069641,OM069642,OM069643,OM069644,OM069645,OM069646</p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45286/viralqc-checking-if-a-viral-genome-tells-the-whole-story</guid>
	<pubDate>Tue, 08 Sep 2026 10:35:04 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45286/viralqc-checking-if-a-viral-genome-tells-the-whole-story</link>
	<title><![CDATA[ViralQC: Checking If a Viral Genome Tells the Whole Story]]></title>
	<description><![CDATA[<p>Imagine finding a mysterious piece of a puzzle and being told it belongs to a virus. Before studying it, you would want to know two things: Is the piece really viral, and how much of the puzzle is missing?</p><p>That is the problem ViralQC (https://github.com/ChengPENG-wolf/ViralQC) aims to solve.</p><p>Viral sequences recovered from metagenomic data can be incomplete or contaminated with microbial DNA. ViralQC uses information from both DNA sequences and predicted proteins to detect contamination and estimate how complete a viral contig is.</p><p>The authors compared ViralQC with CheckV and found that ViralQC performed particularly well for longer viral contigs, improving contamination detection and completeness estimation in several test cases.</p><p>Why does this matter? Because discovering a viral sequence is only the first step. If the sequence is contaminated or incomplete, downstream analyses&mdash;such as identifying viral functions or studying evolution&mdash;can be misleading.</p><p>ViralQC provides a useful quality check before researchers trust the viral genome they have discovered.</p><p>In a world where metagenomics is uncovering enormous numbers of unknown viruses, tools like ViralQC help us separate the real viral story from an incomplete or mixed-up one.</p><p>More at&nbsp;https://academic.oup.com/bioinformatics/article/42/Supplement_2/btag463/8767288</p>]]></description>
	<dc:creator>LEGE</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/41886/coronavirus-sars-cov-2</guid>
	<pubDate>Wed, 17 Jun 2020 11:18:24 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/41886/coronavirus-sars-cov-2</link>
	<title><![CDATA[Coronavirus SARS-CoV-2]]></title>
	<description><![CDATA[<p><span>Used Nanographics Vj, our real-time molecular visualization and animation software, to create this video showing the structure of the virus. In the video, you can see the latest theory on how the RNA is organized inside of the virus particle.</span></p>
<p><span><span>On this page, you can download&nbsp;</span><a href="https://nanographics.at/projects/sars-cov-2/sars-cov-2-renders.zip">high resolution images</a><span>&nbsp;of our renderings. We made them with transparent background, so that you can use it in your work. As the research progresses, we will keep updating the model as well as the images on this page, so stay tuned!</span></span></p>
<p>&nbsp;</p><p>Address of the bookmark: <a href="https://nanographics.at/projects/sars-cov-2/" rel="nofollow">https://nanographics.at/projects/sars-cov-2/</a></p>]]></description>
	<dc:creator>Shruti Paniwala</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/34246/unicycler-hybrid-assembly-pipeline-for-bacterial-genomes</guid>
	<pubDate>Fri, 10 Nov 2017 03:58:27 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/34246/unicycler-hybrid-assembly-pipeline-for-bacterial-genomes</link>
	<title><![CDATA[Unicycler: Hybrid assembly pipeline for bacterial genomes]]></title>
	<description><![CDATA[<p><span>Unicycler is an assembly pipeline for bacterial genomes. It can assemble&nbsp;</span><a href="http://www.illumina.com/">Illumina</a><span>-only read sets where it functions as a&nbsp;</span><a href="http://cab.spbu.ru/software/spades/">SPAdes</a><span>-optimiser. It can also assembly long-read-only sets (</span><a href="http://www.pacb.com/">PacBio</a><span>&nbsp;or&nbsp;</span><a href="https://nanoporetech.com/">Nanopore</a><span>) where it runs a&nbsp;</span><a href="https://github.com/lh3/miniasm">miniasm</a><span>+</span><a href="https://github.com/isovic/racon">Racon</a><span>&nbsp;pipeline. For the best possible assemblies, give it both Illumina reads&nbsp;</span><em>and</em><span>&nbsp;long reads, and it will conduct a hybrid assembly.</span></p><p>Address of the bookmark: <a href="https://github.com/rrwick/Unicycler" rel="nofollow">https://github.com/rrwick/Unicycler</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36621/hapcut2-robust-and-accurate-haplotype-assembly-for-diverse-sequencing-technologies</guid>
	<pubDate>Tue, 15 May 2018 07:35:26 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36621/hapcut2-robust-and-accurate-haplotype-assembly-for-diverse-sequencing-technologies</link>
	<title><![CDATA[HapCUT2: robust and accurate haplotype assembly for diverse sequencing technologies]]></title>
	<description><![CDATA[HapCUT2 is a maximum-likelihood-based tool for assembling haplotypes from DNA sequence reads, designed to "just work" with excellent speed and accuracy. We found that previously described haplotype assembly methods are specialized for specific read technologies or protocols, with slow or inaccurate performance on others. With this in mind, HapCUT2 is designed for speed and accuracy across diverse sequencing technologies, including but not limited to:

NGS short reads (Illumina HiSeq)
clone-based sequencing (Fosmid or BAC clones)
SMRT reads (PacBio)
Oxford Nanopore reads
10X Genomics Linked-Reads
proximity-ligation (Hi-C) reads
high-coverage sequencing (&gt;40x coverage-per-SNP) using above technologies
combinations of the above technologies (e.g. scaffold long reads with Hi-C reads)
See below for specific examples of command line options and best practices for some of these technologies.

NOTE: At this time HapCUT2 is for diploid organisms only. VCF input should contain diploid variants.

If you use HapCUT2 in your research, please cite:

Edge, P., Bafna, V. &amp; Bansal, V. HapCUT2: robust and accurate haplotype assembly for diverse sequencing technologies. Genome Res. gr.213462.116 (2016). doi:10.1101/gr.213462.116<p>Address of the bookmark: <a href="https://github.com/vibansal/HapCUT2" rel="nofollow">https://github.com/vibansal/HapCUT2</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37291/transrate-understanding-your-transcriptome-assembly</guid>
	<pubDate>Fri, 13 Jul 2018 07:49:26 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37291/transrate-understanding-your-transcriptome-assembly</link>
	<title><![CDATA[transrate: Understanding your transcriptome assembly]]></title>
	<description><![CDATA[<p><span>Transrate is software for&nbsp;</span><em>de-novo</em><span>&nbsp;transcriptome assembly quality analysis. It examines your assembly in detail and compares it to experimental evidence such as the sequencing reads, reporting quality scores for contigs and assemblies. This allows you to choose between assemblers and parameters, filter out the bad contigs from an assembly, and help decide when to stop trying to improve the assembly.</span></p><p>Address of the bookmark: <a href="http://hibberdlab.com/transrate/index.html" rel="nofollow">http://hibberdlab.com/transrate/index.html</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/38505/allhic-phasing-and-scaffolding-polyploid-genomes-based-on-hi-c-data</guid>
	<pubDate>Thu, 20 Dec 2018 12:03:32 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/38505/allhic-phasing-and-scaffolding-polyploid-genomes-based-on-hi-c-data</link>
	<title><![CDATA[ALLHiC: Phasing and scaffolding polyploid genomes based on Hi-C data]]></title>
	<description><![CDATA[<p><span>The major problem of scaffolding polyploid genome is that Hi-C signals are frequently detected between allelic haplotypes and any existing stat of art Hi-C scaffolding program links the allelic haplotypes together. To solve the problem, we developed a new Hi-C scaffolding pipeline, called ALLHIC, specifically tailored to the polyploid genomes. ALLHIC pipeline contains a total of 5 steps:&nbsp;</span><em>prune</em><span>,&nbsp;</span><em>partition</em><span>,&nbsp;</span><em>rescue</em><span>,&nbsp;</span><em>optimize</em><span>&nbsp;and&nbsp;</span><em>build</em><span>.</span></p><p>Address of the bookmark: <a href="https://github.com/tangerzhang/ALLHiC/wiki" rel="nofollow">https://github.com/tangerzhang/ALLHiC/wiki</a></p>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>

</channel>
</rss>