<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/37241?offset=150</link>
	<atom:link href="https://bioinformaticsonline.com/related/37241?offset=150" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/44229/common-steps-for-reads-mapping</guid>
	<pubDate>Thu, 09 Mar 2023 02:48:02 -0600</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/44229/common-steps-for-reads-mapping</link>
	<title><![CDATA[Common steps for reads mapping !]]></title>
	<description><![CDATA[<div><div><div><div><div><div><div><div><div><div><p>Mapping reads to a reference genome is an essential step in many types of genomic analysis, such as variant calling and gene expression analysis. Here are some general steps to follow for mapping reads to a genome:</p><ol>
<li>
<p>Choose a read mapper: There are many read mappers available, such as BWA, Bowtie, and HISAT2. Choose a mapper that is appropriate for your type of data and research question.</p>
</li>
<li>
<p>Index the reference genome: Before mapping reads, the reference genome needs to be indexed. This involves creating an index of the genome sequence that allows the mapper to quickly find matches to the reads. Most mappers have their own indexing tools.</p>
</li>
<li>
<p>Prepare the read data: The reads should be in a format that is compatible with the mapper. Most mappers accept FASTQ or BAM files. Depending on the quality of the data, it may need to be filtered or trimmed before mapping.</p>
</li>
<li>
<p>Run the mapper: The mapper is run with the command-line interface or using a graphical user interface. The specific command depends on the mapper being used, but typically involves specifying the input data, reference genome, and output file format.</p>
</li>
<li>
<p>Evaluate the mapping results: After the mapping is complete, the results should be evaluated. This includes assessing the quality of the mapping, such as the mapping rate, the number of mapped reads, and the mapping quality score.</p>
</li>
<li>
<p>Post-processing: Depending on the analysis being performed, post-processing of the mapped reads may be necessary. This can include filtering reads based on quality, removing duplicate reads, and calling variants.</p>
</li>
</ol><p>Overall, mapping reads to a reference genome is a complex process that requires careful consideration of the type of data, the research question, and the specific mapper being used.</p></div></div></div></div></div></div></div></div></div></div>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45343/simulating-the-unseen-new-tools-reshaping-metagenomic-research</guid>
	<pubDate>Wed, 23 Sep 2026 09:42:30 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45343/simulating-the-unseen-new-tools-reshaping-metagenomic-research</link>
	<title><![CDATA[Simulating the Unseen: New Tools Reshaping Metagenomic Research]]></title>
	<description><![CDATA[<p>When benchmarking a new metagenomic pipeline the first step is generally to face a simple problem. A bioinformatician might have a good deal of real sequencing data, but there's a drawback in that the actual biological composition underlying those reads is almost never known. What organisms really were there? In what quantities? How many sequencing errors occurred? Can the pipeline correctly reconstruct the community? Since the true answers are unknown, it is difficult to deal with these questions.</p><p>Metagenomic simulators are a way of dealing with this issue; they produce artificial sequencing datasets in which both the composition of the microorganisms and the expected results are known. Researchers can then test a pipeline on such a controlled dataset and see precisely how well it functions. However, as metagenomic studies have become larger, more diverse, and more complex, simply generating artificial reads is no longer sufficient. The simulators now have to replicate community variation, sequencing errors, different sequencing platforms, massive data volumes, and even laboratory-based biases.</p><p>MIDASim (2024) was designed with a particular emphasis on microbial community dynamics. Unlike previous tools which regarded each simulated sample as individual, MIDASim is able to efficiently generate realistic multi-sample microbiome datasets, covering the cases where changes over time are being tracked. This feature is particularly useful when investigating how microbial communities vary between people, different experimental setups, or at various times. By overcoming the computational limitations of earlier simulation methods, MIDASim has made it easier to carry out large and complex microbiome simulation studies.</p><p>A major challenge is scale. Since modern metagenomic projects can involve millions or even billions of sequencing reads, generating such large synthetic datasets can become a significant computational problem. In order to tackle this issue, Izzy (2025) was developed. As a high-throughput metagenomic read simulator, Izzy is designed with speed in mind while still reproducing the typical patterns of sequencing errors. The people who developed it stated that it can be as much as 60 times faster than earlier tools, which makes large-scale simulations far more practical. Now, researchers can begin to ask not how long a simulation will take but rather what size dataset they want to test.</p><p>At the same time, the sequencing technology became more diverse. Researchers were not any longer restricted to using Illumina short reads since Oxford Nanopore and PacBio long-read sequencing also became important in the field of metagenomic research. In order to address these new requirements, MeSS (Metagenomic Sequence Simulator, 2025) was developed. The programme is based on a Snakemake framework and is therefore compatible with the Illumina, Oxford Nanopore and PacBio platforms. MeSS has also been designed to operate efficiently and to use less memory than earlier tools; it provides ready-made templates for the microbiomes of different sites in the human body, thus giving researchers a simple means of beginning to generate realistic simulated communities.</p><p>The simulation problem became even more specialized. Suppose that sequencing is not based on uniform sampling of DNA? In the case of targeted sequencing and hybridization-capture experiments, particular sequences are deliberately enriched, and the probability of capturing a target depends on both the probe and the target sequence. A standard model based on uniform sampling might fail to take into account this important feature of real experiments. That is why RAmpSim has been designed specifically to handle simulations for capture-based and targeted sequencing. Instead of depending solely on the assumption of uniform sampling, it includes a thermodynamic nearest-neighbor energy model to represent probe&ndash;target interactions and sequence-dependent enrichment. As a result, it is especially suitable for applications such as hybridization capture in metagenomics, where the efficiency of recovering a sequence can depend strongly on how it relates to the capture probes. RAmpSim thus reflects a wider trend towards simulations that aim not only to reproduce sequencing output but also key aspects of the experimental process itself.</p><p>MIDASim, Izzy, MeSS and RAmpSim all demonstrate how fast metagenomic simulation is evolving. MIDASim is able to deal with realistic variation within a large number of and changing microbial communities. Izzy overcomes the problem of generating very large datasets. MeSS provides support for a number of sequencing technologies in an efficient and repeatable manner. RAmpSim increases the experimental realism for capture-based sequencing. Although each of these tools addresses a separate issue, they all contribute to advancing the field.</p><p>Modern metagenomic simulation therefore aims not only at producing artificial FASTQ files. Researchers nowadays desire simulated datasets that reflect community changes, incorporate the characteristics of the sequencing platform, include real error patterns, take into account computational scale, and account for experimental biases. As metagenomic workflows become more advanced, the simulators have to keep up with this development. The tools are now going beyond that of simple data generators; they produce controlled versions of complex metagenomic experiments and thus provide researchers with something that ordinary sequencing data seldom does: a dataset in which the answer is known before any analysis takes place.</p>]]></description>
	<dc:creator>LEGE</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/31205/yasra-reference-based-assembler</guid>
	<pubDate>Wed, 01 Mar 2017 08:32:45 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/31205/yasra-reference-based-assembler</link>
	<title><![CDATA[YASRA: Reference based assembler]]></title>
	<description><![CDATA[<p>YASRA (Yet Another Short Read Assembler) performs comparative assembly of short reads using a reference genome, which can differ substantially from the genome being sequenced. Mapping reads to reference genomes makes use of LASTZ (Harris et al), a pairwise sequence aligner compatible with BLASTZ. Special scoring sets were derived to improve the performance, both in runtime and quality for 454 and Illumina sequence reads.</p>
<p>YASRA uses LASTZ (<a href="http://bx.psu.edu/miller_lab">http://bx.psu.edu/miller_lab</a> for released version and <a href="http://www.bx.psu.edu/%7Ersharris/lastz/newer">http://www.bx.psu.edu/~rsharris/lastz/newer</a> for newer version) for aligning the sequences to the reference genome. Please install LASTZ (the newest version on <a href="http://www.bx.psu.edu/%7Ersharris/lastz/newer">http://www.bx.psu.edu/~rsharris/lastz/newer</a>) and add the LASTZ binary in your executable/binary search path before installing YASRA.</p><p>Address of the bookmark: <a href="https://github.com/aakrosh/YASRA" rel="nofollow">https://github.com/aakrosh/YASRA</a></p>]]></description>
	<dc:creator>Abhimanyu Singh</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/41158/carefully-opt-for-human-reference-genome</guid>
	<pubDate>Tue, 18 Feb 2020 07:43:32 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/41158/carefully-opt-for-human-reference-genome</link>
	<title><![CDATA[Carefully opt for human reference genome]]></title>
	<description><![CDATA[<p><a href="http://lh3.github.io/2017/11/13/which-human-reference-genome-to-use" target="_blank">Heng Li posted several issues with the human reference genomes given in these resources</a> and suggests the following compressed FASTA file to be used as hg38/GRCh38 human reference genome.</p>
<p>if you map reads to GRCh38 or hg38, use the following:</p>
<div>
<div>
<pre><code>ftp://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz
</code></pre>
</div>
</div>
<p>There are several other versions of GRCh37/GRCh38. What&rsquo;s wrong with them? Here are a collection of potential issues:</p>
<p>More at http://lh3.github.io/2017/11/13/which-human-reference-genome-to-use</p><p>Address of the bookmark: <a href="http://lh3.github.io/2017/11/13/which-human-reference-genome-to-use" rel="nofollow">http://lh3.github.io/2017/11/13/which-human-reference-genome-to-use</a></p>]]></description>
	<dc:creator>biogeek</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/41230/curated-set-of-ribosomal-rna-rrna-reference-sequences-targeted-loci-with-verifiable-organism</guid>
	<pubDate>Sun, 23 Feb 2020 02:17:30 -0600</pubDate>
	<link>https://bioinformaticsonline.com/news/view/41230/curated-set-of-ribosomal-rna-rrna-reference-sequences-targeted-loci-with-verifiable-organism</link>
	<title><![CDATA[Curated set of ribosomal RNA (rRNA) reference sequences (targeted loci) with verifiable organism]]></title>
	<description><![CDATA[<p>MCBI have a curated set of ribosomal RNA (rRNA) reference sequences (targeted loci) with verifiable organism sources and current names. This set is critical for correctly identifying and classifying prokaryotic (bacteria and archaea) and fungal samples. To provide easy access to these sequences, we recently added a separate rRNA/ITS databases section on the nucleotide BLAST page for these targeted sequences that makes it convenient to quickly identify source organisms. The new databases are: </p><p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *16S ribosomal RNA (Bacteria and Archaea)</p><p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *18S ribosomal RNA sequences (SSU) from Fungi type and reference material&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p><p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *28S ribosomal RNA sequences (LSU) from Fungi type and reference material</p><p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *Internal transcribed spacer region (ITS) from Fungi type and reference material</p><p>You can also download these from the BLAST db FTP area.&nbsp; See the <a href="https://go.usa.gov/xdEBX" target="_blank">NCBI Insights post</a> for more detail. </p><p>Useful links</p><p>-----------------</p><p><a href="https://go.usa.gov/xdEj5" target="_blank">BLAST form with rRNA/ITS databases</a></p><p><a href="https://ftp.ncbi.nlm.nih.gov/blast/db/" target="_blank">BLAST db download</a></p><p><a href="https://www.ncbi.nlm.nih.gov/refseq/targetedloci/" target="_blank">Targeted loci</a></p><p><span style="color: black;">If you have any questions or concerns, please contact <a href="mailto:blast-help@ncbi.nlm.nih.gov" target="_blank" title="Follow link">blast-help@ncbi.nlm.nih.gov<sup><span style="color: black; text-decoration: none;"><img src="https://mail.google.com/mail/u/0?ui=2&amp;ik=024a8aa0b9&amp;attid=0.1&amp;permmsgid=msg-f:1659255165855446848&amp;th=1706dbc8408bb740&amp;view=fimg&amp;sz=s0-l75-ft&amp;attbid=ANGjdJ_drW2ArYDNLoHrQh36gm6rp2Std8ZUSplCzP6bYQSQYBsQfZ_85vOujXOdTRdaLxrR7QeEBVUbyACPBJHhFUeIglX8G7Ew7TcclzhvO7fJhiz7sIdkkDgZ7QA&amp;disp=emb" alt="https://jira.ncbi.nlm.nih.gov/images/icons/mail_small.gif" width="13" height="12" style="border: 0px;"></span></sup></a></span></p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37512/purecn-copy-number-calling-and-snv-classification-using-targeted-short-read-sequencing</guid>
	<pubDate>Thu, 09 Aug 2018 04:09:37 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37512/purecn-copy-number-calling-and-snv-classification-using-targeted-short-read-sequencing</link>
	<title><![CDATA[PureCN: copy number calling and SNV classification using targeted short read sequencing]]></title>
	<description><![CDATA[<p>This package estimates tumor purity, copy number, and loss of heterozygosity (LOH), and classifies single nucleotide variants (SNVs) by somatic status and clonality. PureCN is designed for targeted short read sequencing data, integrates well with standard somatic variant detection and copy number pipelines, and has support for tumor samples without matching normal samples.</p>
<p>Author: Markus Riester [aut, cre], Angad P. Singh [aut]</p>
<p>Maintainer: Markus Riester &lt;markus.riester at novartis.com&gt;</p>
<div id="bioc_citation_outer">
<p>Citation (from within R, enter&nbsp;<code>citation("PureCN")</code>):</p>
<div id="bioc_citation">
<p>Riester M, Singh A, Brannon A, Yu K, Campbell C, Chiang D, Morrissey M (2016). &ldquo;PureCN: Copy number calling and SNV classification using targeted short read sequencing.&rdquo;&nbsp;<em>Source Code for Biology and Medicine</em>,&nbsp;<strong>11</strong>, 13. doi:&nbsp;<a href="http://doi.org/10.1186/s13029-016-0060-z">10.1186/s13029-016-0060-z</a>.</p>
</div>
</div><p>Address of the bookmark: <a href="http://bioconductor.org/packages/release/bioc/html/PureCN.html" rel="nofollow">http://bioconductor.org/packages/release/bioc/html/PureCN.html</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37221/asplice-a-scalable-and-memory-efficient-algorithm-for-de-novo-transcriptome-assembly</guid>
	<pubDate>Tue, 03 Jul 2018 04:09:46 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37221/asplice-a-scalable-and-memory-efficient-algorithm-for-de-novo-transcriptome-assembly</link>
	<title><![CDATA[ASplice: a scalable and memory-efficient algorithm for de novo transcriptome assembly]]></title>
	<description><![CDATA[With increased availability of de novo assembly algorithms, it is feasible to study entire transcriptomes of non-model organisms. While algorithms are available that are specifically designed for performing transcriptome assembly from high-throughput sequencing data, they are very memory-intensive, limiting their applications to small data sets with few libraries.

Texas A&amp;M University researchers develop a transcriptome assembly algorithm that recovers alternatively spliced isoforms and expression levels while utilizing as many RNA-Seq libraries as possible that contain hundreds of gigabases of data. New techniques are developed so that computations can be performed on a computing cluster with moderate amount of physical memory.

Availability – A software program that implements the algorithm is available at: http://faculty.cse.tamu.edu/shsze/asplice.

Sze SH, Pimsler ML, Tomberlin JK, Jones CD, Tarone AM. (2017) A scalable and memory-efficient algorithm for de novo transcriptome assembly of non-model organisms. BMC Genomics 18(Suppl 4):387.<p>Address of the bookmark: <a href="http://faculty.cse.tamu.edu/shsze/asplice/" rel="nofollow">http://faculty.cse.tamu.edu/shsze/asplice/</a></p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/39640/flas-fast-and-high-throughput-algorithm-for-pacbio-long-read-self-correction</guid>
	<pubDate>Sat, 22 Jun 2019 12:16:39 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/39640/flas-fast-and-high-throughput-algorithm-for-pacbio-long-read-self-correction</link>
	<title><![CDATA[FLAS: fast and high throughput algorithm for PacBio long read self-correction.]]></title>
	<description><![CDATA[<p><span>FLAS, a wrapper algorithm of MECAT, to achieve high throughput long read self-correction while keeping MECAT's fast speed. FLAS finds additional alignments from MECAT prealigned long reads to improve the correction throughput, and removes misalignments for accuracy.</span></p><p>Address of the bookmark: <a href="https://github.com/baoe/flas" rel="nofollow">https://github.com/baoe/flas</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/33479/novelseq-novel-sequence-insertion-detection</guid>
	<pubDate>Fri, 09 Jun 2017 04:31:30 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/33479/novelseq-novel-sequence-insertion-detection</link>
	<title><![CDATA[NovelSeq: Novel Sequence Insertion Detection]]></title>
	<description><![CDATA[<p><span>The NovelSeq framework is designed to detect novel sequence insertions using high throughput paired-end whole genome sequencing data.</span></p>
<p>http://novelseq.sourceforge.net/Home</p>
<p>Paper at&nbsp;https://www.ncbi.nlm.nih.gov/pubmed/20385726</p><p>Address of the bookmark: <a href="http://novelseq.sourceforge.net/Home" rel="nofollow">http://novelseq.sourceforge.net/Home</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/43698/mimilook-a-phylogenetic-workflow-for-detection-of-gene-acquisition-in-major-orthologous-groups-of-megavirales</guid>
	<pubDate>Mon, 10 Jan 2022 06:32:22 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/43698/mimilook-a-phylogenetic-workflow-for-detection-of-gene-acquisition-in-major-orthologous-groups-of-megavirales</link>
	<title><![CDATA[MimiLook: A Phylogenetic Workflow for Detection of Gene Acquisition in Major Orthologous Groups of Megavirales]]></title>
	<description><![CDATA[<p><span>This tool detects statistically validated events of gene acquisitions with the help of the T-REX algorithm by comparing individual gene tree with NCBI species tree. In between the steps, the workflow decides about handling paralogs, filtering outputs, identifying Megavirale specific OGs, detection of HGTs, along with retrieval of information about those OGs that are monophyletic with organisms from cellular domains of life.&nbsp;</span></p>
<p>https://www.readcube.com/articles/10.3390%2Fv9040072</p><p>Address of the bookmark: <a href="https://pubmed.ncbi.nlm.nih.gov/28387730/" rel="nofollow">https://pubmed.ncbi.nlm.nih.gov/28387730/</a></p>]]></description>
	<dc:creator>Abhi</dc:creator>
</item>

</channel>
</rss>