<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/38886?offset=420</link>
	<atom:link href="https://bioinformaticsonline.com/related/38886?offset=420" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/42477/hifiasm-a-haplotype-resolved-assembler-for-accurate-hifi-reads</guid>
	<pubDate>Thu, 24 Dec 2020 10:03:36 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/42477/hifiasm-a-haplotype-resolved-assembler-for-accurate-hifi-reads</link>
	<title><![CDATA[Hifiasm: a haplotype-resolved assembler for accurate Hifi reads]]></title>
	<description><![CDATA[<p><span>Hifiasm is a fast haplotype-resolved de novo assembler for PacBio Hifi reads. It can assemble a human genome in several hours and works with the California redwood genome, one of the most complex genomes sequenced so far. Hifiasm can produce primary/alternate assemblies of quality competitive with the best assemblers. It also introduces a new graph binning algorithm and achieves the best haplotype-resolved assembly given trio data.</span></p><p>Address of the bookmark: <a href="https://github.com/chhylp123/hifiasm" rel="nofollow">https://github.com/chhylp123/hifiasm</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/44234/steps-to-find-palindrome-in-genomes</guid>
	<pubDate>Thu, 09 Mar 2023 02:56:54 -0600</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/44234/steps-to-find-palindrome-in-genomes</link>
	<title><![CDATA[Steps to find palindrome in genomes !]]></title>
	<description><![CDATA[<div><div><div><div><div><div><div><div><div><div><p>Palindromes are sequences of nucleotides that read the same backward as forward. They can be present in genomes and have various biological functions. Here are some methods for discovering palindromes in genomes:</p><ol>
<li>
<p>Direct sequence search: One of the simplest ways to discover palindromes is to search the genome sequence directly for palindromic sequences using pattern matching tools, such as regular expressions or string algorithms. This approach can be useful for discovering simple palindromes, but may miss more complex palindromic structures.</p>
</li>
<li>
<p>Dot plot analysis: Dot plot analysis is a graphical method that can be used to identify palindromic regions in a genome. It involves plotting the genome sequence against itself and examining the diagonal patterns that emerge. Palindromic regions will appear as symmetrical patterns along the diagonal.</p>
</li>
<li>
<p>Restriction enzyme analysis: Some restriction enzymes, such as EcoRI and HindIII, recognize palindromic sequences and cleave DNA at these sites. By digesting the genome with these enzymes and examining the resulting fragments, palindromic regions can be identified.</p>
</li>
<li>
<p>Next-generation sequencing: High-throughput sequencing technologies, such as PacBio and Oxford Nanopore, can generate long reads that can span entire palindromic regions. By mapping these reads to the genome, palindromic regions can be identified and characterized.</p>
</li>
<li>
<p>Comparative genomics: Comparing the genomes of related species can also reveal palindromic regions that are conserved across evolutionarily divergent lineages. This approach can help identify functional palindromes that are under selective pressure.</p>
</li>
</ol><p>Overall, the discovery of palindromic sequences in genomes can be accomplished using a variety of methods, each with their own advantages and limitations. A combination of these methods can provide a comprehensive understanding of the palindromic landscape of a genome.</p></div></div></div></div></div></div></div></div></div></div>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/27080/mrfast-micro-read-fast-alignment-search-tool</guid>
	<pubDate>Tue, 26 Apr 2016 03:50:06 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/27080/mrfast-micro-read-fast-alignment-search-tool</link>
	<title><![CDATA[mrFAST:  Micro Read Fast Alignment Search Tool]]></title>
	<description><![CDATA[<p><span>mrFAST is a read mapper that is designed to map short reads to reference genome with a special emphasis on the discovery of structural variation and segmental duplications. mrFAST maps short reads with respect to user defined error threshold, including indels up to 4+4 bp. This manual, describes how to choose the parameters and tune mrFAST with respect to the library settings. mrFAST is designed to find&nbsp;</span><strong><span style="text-decoration: underline;">'all'</span></strong><span>&nbsp; mappings for a given set of reads, however it can return one "best" map location if the relevant parameter is invoked.</span></p>
<p><span>More at&nbsp;http://mrfast.sourceforge.net/manual.html</span></p><p>Address of the bookmark: <a href="http://mrfast.sourceforge.net/manual.html" rel="nofollow">http://mrfast.sourceforge.net/manual.html</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/27839/lorma-a-tool-for-correcting-sequencing-errors-in-long-reads-such-those-produced-by-pacific-biosciences-sequencing-machines</guid>
	<pubDate>Wed, 15 Jun 2016 17:18:36 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/27839/lorma-a-tool-for-correcting-sequencing-errors-in-long-reads-such-those-produced-by-pacific-biosciences-sequencing-machines</link>
	<title><![CDATA[LoRMA: a tool for correcting sequencing errors in long reads such those produced by Pacific Biosciences sequencing machines]]></title>
	<description><![CDATA[<p>LoRMA is a tool for correcting sequencing errors in long reads such those produced by Pacific Biosciences sequencing machines.</p>
<p>Publication:</p>
<ul>
<li>L. Salmela, R. Walve, E. Rivals, and E. Ukkonen: Accurate selfcorrection of errors in long reads using de Bruijn graphs. Accepted to RECOMB-Seq 2016.</li>
</ul>
<p>Download:</p>
<ul>
<li><a href="https://www.cs.helsinki.fi/u/lmsalmel/LoRMA/LoRMA-0.3.tar.gz">LoRMA 0.3 source files</a></li>
<li><a href="https://www.cs.helsinki.fi/u/lmsalmel/LoRMA/README.txt">README</a></li>
</ul><p>Address of the bookmark: <a href="https://www.cs.helsinki.fi/u/lmsalmel/LoRMA/" rel="nofollow">https://www.cs.helsinki.fi/u/lmsalmel/LoRMA/</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/30555/yaha</guid>
	<pubDate>Fri, 20 Jan 2017 05:38:05 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/30555/yaha</link>
	<title><![CDATA[YAHA]]></title>
	<description><![CDATA[<p>YAHA, a fast and flexible hash-based aligner. YAHA is as fast and accurate as BWA-SW at finding the single best alignment per query and is dramatically faster and more sensitive than both SSAHA2 and MegaBLAST at finding all possible alignments. Unlike other aligners that report all, or one, alignment per query, or that use simple heuristics to select alignments, YAHA uses a directed acyclic graph to find the optimal set of alignments that cover a query using a biologically relevant breakpoint penalty. YAHA can also report multiple mappings per defined segment of the query. We show that YAHA detects more breakpoints in less time than BWA-SW across all SV classes, and especially excels at complex SVs comprising multiple breakpoints.</p>
<p><strong>Availability:</strong> YAHA is currently supported on 64-bit Linux systems. Binaries and sample data are freely available for download from <a href="http://faculty.virginia.edu/irahall/YAHA" target="pmc_ext">http://faculty.virginia.edu/irahall/YAHA</a>.</p>
<p><strong>Contact:</strong></p>
<p>http://genome.wustl.edu/people/groups/detail/hall-lab/</p><p>Address of the bookmark: <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3463118/" rel="nofollow">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3463118/</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/35621/bbtools-for-bioinformatician</guid>
	<pubDate>Thu, 15 Feb 2018 16:45:52 -0600</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/35621/bbtools-for-bioinformatician</link>
	<title><![CDATA[BBTools for bioinformatician !]]></title>
	<description><![CDATA[<p><span></span><br /><strong>BBMap.sh</strong><br /><br /></p><ul>
<li><strong>Mapping Nanopore reads</strong></li>
</ul><p><br /><span>BBMap.sh has a length cap of 6kbp. Reads longer than this will be broken into 6kbp pieces and mapped independently.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ mapPacBio.sh -Xmx20g k=7 in=reads.fastq ref=reference.fa maxlen=1000 minlen=200 idtag ow int=f qin=33 out=mapped1.sam minratio=0.15 ignorequality slow ordered maxindel1=40 maxindel2=400</pre></div><p><br /><span>The "maxlen" flag shreds them to a max length of 1000; you can set that up to 6000. But I found 1000 gave a higher mapping rate.&nbsp;&nbsp;</span><br /><br /></p><ul>
<li><strong>Using Paired-end and single-end reads at the same time</strong></li>
</ul><p><br /><span>BBMap itself can only run single-ended or paired-ended in a single run, but it has a wrapper that can accomplish it, like this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ bbwrap.sh in1=read1.fq,singletons.fq in2=read2.fq,null out=mapped.sam append</pre></div><p><span>This will write all the reads to the same output file but only print the headers once. I have not tried that for bam output, only sam output</span><br /><br /><span>Note about alignment stats: For paired reads, you can find the total percent mapped by adding the read 1 percent (where it says "mapped: N%") and read 2 percent, then dividing by 2. The different columns tell you the count/percent of each event. Considering the cigar strings from alignment, "Match Rate" is the number of symbols indicating a reference match (=) and error rate is the number indicating substitution, insertion, or deletion (X, I, D).</span><br /><br /></p><ul>
<li><strong>Exact matches when mapping small reads (e.g. miRNA)</strong></li>
</ul><p><br /><span>When mapping small RNA's with BBMap use the following flags to report only perfect matches.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">ambig=all vslow perfectmode maxsites=1000</pre></div><p><span>It should be very fast in that mode (despite the vslow flag). Vslow mainly removes masking of low-complexity repetitive kmers, which is not usually a problem but can be with extremely short sequences like microRNAs.</span></p><ul>
<li><strong>Important note about BBMap alignments</strong></li>
</ul><p><br /><span>BBMap is always nondeterministic when run in paired-end mode with multiple threads, because the insert-size average is calculated on a per-thread basis, which affects mapping; and which reads are assigned to which thread is nondeterministic. The only way to avoid that would be to restrict it to a single thread (threads=1), or map the reads as single-ended and then fix pairing afterward:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">bbmap.sh in=reads.fq outu=unmapped.fq int=f
repair.sh in=unmapped.fq out=paired.fq fint outs=singletons.fq</pre></div><p><span>In this case you'd want to only keep the paired output.&nbsp;</span><br /><br /><span>BBSplit is based on BBMap, so it is also nondeterministic in paired mode with multiple threads. BBDuk and Seal (which can be used similarly to BBSplit) are always deterministic.&nbsp;</span><br /><br /><span>--------------------------------------------------------</span><br /><br /><strong>Reformat.sh</strong></p><ul>
<li><strong>Count k-mers/find unknown primers</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.fq out=trimmed.fq ftr=19</pre></div><p><span>This will trim all but the first 20 bases (all bases after position 19, zero-based).</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ kmercountexact.sh in=trimmed.fq out=counts.txt fastadump=f mincount=10 k=20 rcomp=f</pre></div><p><span>This will generate a file containing the counts of all 20-mers that occurred at least 10 times, in a 2-column format that is easy to sort in Excel.&nbsp;</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">ACCGTTACCGTTACCGTTAC	100
AAATTTTTTTCCCCCCCCCC	85</pre></div><p><span>...etc. If the primers are 20bp long, they should be pretty obvious.&nbsp;&nbsp;</span></p><ul>
<li><strong>Convert SAM format from 1.4 to 1.3 (required for many programs)</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.sam out=out.sam sam=1.3</pre></div><ul>
<li><strong>Removing N basecalls</strong></li>
</ul><p><br /><span>You can use BBDuk or Reformat with "qtrim=rl trimq=1". That will only trim trailing and leading bases with Q-score below 1, which means Q0, which means N (in either fasta or fastq format). The BBMap package automatically changes q-scores of Ns that are above 0 to 0 and called bases with q-scores below 2 to 2, since occasionally some Illumina software versions produces odd things like a handful of Q0 called bases or Ns with Q&gt;0, neither of which make any sense in the Phred scale.</span></p><ul>
<li><strong>Sampling reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.fq out=sampled.fq sample=3000</pre></div><div><div>Code:</div><pre dir="ltr">To sample 10% of the reads:
reformat.sh in1=reads1.fq in2=reads2.fq out1=sampled1.fq out2=sampled2.fq samplerate=0.1

or more concisely:
reformat.sh in=reads#.fq out=sampled#.fq samplerate=0.1

and for exact sampling:
reformat.sh in=reads#.fq out=sampled#.fq samplereadstarget=100k</pre></div><ul>
<li><strong>Changing fasta headers</strong></li>
</ul><p><br /><span>Remove anything after the first space in fasta header.&nbsp;</span><br /><br /></p><div><div>Code:</div><pre dir="ltr"> reformat.sh in=sequences.fasta out=renamed.fasta trd</pre></div><p><span>"trd" stands for "trim read description" and will truncate everything after the first whitespace.</span></p><ul>
<li><strong>Extract reads from a sam file</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.sam out=reads.fastq</pre></div><ul>
<li><strong>Verify pairing and optionally de-interleave the reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.fastq verifypairing</pre></div><ul>
<li><strong>Verify pairing if the reads are in separate files</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in1=r1.fq in2=r2.fq vpair</pre></div><p><span>If that completes successfully and says the reads were correctly paired, then you can simply de-interleave reads into two files like this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.fastq out1=r1.fastq out2=r2.fastq</pre></div><ul>
<li><strong>Base quality histograms</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=reads.fq qchist=qchist.txt</pre></div><p><span>That stands for "quality count histogram".&nbsp;</span></p><ul>
<li><strong>Filter SAM/BAM file by read length</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=x.sam out=y.sam minlength=50 maxlength=200</pre></div><ul>
<li><strong>Filter SAM/BAM file to detect/filter spliced reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=mapped.bam out=filtered.bam maxdellen=50</pre></div><p><span>You can set "maxdellen" to whatever length deletion event you consider the minimum to signify splicing, which depends on the organism.</span><br /><span>-------------------------------------------------------------</span><br /><strong>Repair.sh</strong></p><ul>
<li><strong>"Re-pair" out-of-order reads from paired-end data files</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ repair.sh in1=r1.fq.gz in2=r2.fq.gz out1=fixed1.fq.gz out2=fixed2.fq.gz outsingle=singletons.fq.gz</pre></div><p><span>--------------------------------------------------------------</span><br /><strong>BBMerge.sh</strong><br /><br /><span>BBMerge now has a new flag - "outa" or "outadapter". This allows you to automatically detect the adapter sequence of reads with short insert sizes, in case you don't know what adapters were used. It works like this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ bbmerge.sh in=reads.fq outa=adapters.fa reads=1m</pre></div><p><span>Of course, it will only work for paired reads! The output fasta file will look like this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">&gt;Read1_adapter
GATCGGAAGAGCACACGTCTGAACTCCAGTCACATCACGATCTCGTATGCCGTCTTCTGCTTG
&gt;Read2_adapter
GATCGGAAGAGCACACGTCTGAACTCCAGTCACCGATGTATCTCGTATGCCGTCTTCTGCTTG</pre></div><p><span>If you have multiplexed things with different barcodes in the adapters, the part with the barcode will show up as Ns, like this:</span><br /><br /><span>GATCGGAAGAGCACACGTCTGAACTCCAGTCACNNNNNNATCTCGTATGCCGTCTTCTGCTTG&nbsp;&nbsp;</span><br /><br /><span>Note: For BBMerge with micro-RNA, you need to add the flag&nbsp;</span><strong>mininsert=17</strong><span>. The default is 35, which is too long for micro-RNA libraries.&nbsp;</span></p><ul>
<li><strong>Identifying adapters</strong></li>
</ul><p><span>If you have paired reads, and enough of the reads have inserts shorter than read length, you can identify adapter sequences with BBMerge, like this (they will be printed to adapters.fa):</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ bbmerge.sh in1=r1.fq in2=r2.fq outa=adapters.fa</pre></div><p><br /><span>-----------------------------------------------------------------</span><br /><br /><strong>BBDuk.sh</strong><br /><br /><span>Note: BBDuk is strictly deterministic on a per-read basis, however it does by default reorder the reads when run multithreaded. You can add the flag "ordered" to keep output reads in the same order as input reads</span></p><ul>
<li><strong>Finding reads with a specific sequence at the beginning of read</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ bbduk.sh -Xmx1g in=reads.fq outm=matched.fq outu=unmatched.fq restrictleft=25 k=25 literal=AAAAACCCCCTTTTTGGGGGAAAAA</pre></div><p><span>In this case, all reads starting with "AAAAACCCCCTTTTTGGGGGAAAAA" will end up in "matched.fq" and all other reads will end up in "unmatched.fq". Specifically, the command means "look for 25-mers in the leftmost 25 bp of the read", which will require an exact prefix match, though you can relax that if you want.</span><br /><br /><span>So you could bin all the reads with your known sequence, then look at the remaining reads to see what they have in common. You can do the same thing with the tail of the read using "restrictright" instead, though you can't use both restrictions at the same time.&nbsp;&nbsp;</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ bbduk.sh in=reads.fq outm=matched.fq literal=NNNNNNCCCCGGGGGTTTTTAAAAA k=25 copyundefined</pre></div><p><span>With the "copyundefined" flag, a copy of each reference sequence will be made representing every valid combination of defined letter. So instead of increasing memory or time use by 6^75, it only increases them by 4^6 or 4096 which is completely reasonable, but it only allows substitutions at predefined locations. You can use the "copyundefined", "hdist", and "qhdist" flags together for a lot of flexibility - for example, hdist=2 qhdist=1 and 3 Ns in the reference would allow a hamming distance of 6 with much lower resource requirements than hdist=6. Just be sure to give BBDuk as much memory as possible.</span></p><ul>
<li><strong>Removing illumina adapters (if exact adapters not known)</strong></li>
</ul><p><br /><span>If you're not sure which adapters are used, you can add "ref=truseq.fa.gz,truseq_rna.fa.gz,nextera.fa.gz" and get them all (this will increase the amount of overtrimming, though it should still be negligible).&nbsp;</span></p><ul>
<li><strong>Removing illumina control sequences/phiX reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">bbduk.sh in=trimmed.fq.gz out=filtered.fq.gz k=31 ref=artifacts,phix ordered cardinality</pre></div><ul>
<li><strong>Identify certain reads that contain a specific sequence</strong></li>
</ul><div><div>Code:</div><pre dir="ltr">$ bbduk.sh in=reads.fq out=unmatched.fq outm=matched.fq literal=ACGTACGTACGTACGTAC k=18 mm=f hdist=2</pre></div><p><span>Make sure "k" is set to the exact length of the sequence. "hdist" controls the number of substitutions allowed. "outm" gets the reads that match. By default this also looks for the reverse-complement; you can disable that with "rcomp=f".&nbsp;&nbsp;</span></p><ul>
<li><strong>Extract sequences that share kmers with your sequences with BBDuk</strong></li>
</ul><div><div>Code:</div><pre dir="ltr">$ bbduk.sh in=a.fa ref=b.fa out=c.fa mkf=1 mm=f k=31</pre></div><p><span>This will print to C all the sequences in A that share 100% of their 31-mers with sequences in B.&nbsp;</span><br /><br /></p><ul>
<li><strong>Extract sequences that contain N's with BBDuk</strong></li>
</ul><div><div>Code:</div><pre dir="ltr">bbduk.sh in=reads.fq out=readsWithoutNs.fq outm=readsWithNs.fq maxns=0</pre></div><p><span>If you have, say, 100bp reads and only want to separate reads containing all 100 Ns, change that to "maxns=99".</span><br /><br /><strong>General notes for BBDuk.sh</strong><span>&nbsp;</span><br /><br /><span>BBDuk can operate in one of 4 kmer-matching modes:</span><br /><span>Right-trimming (ktrim=r), left-trimming (ktrim=l), masking (ktrim=n), and filtering (default). But it can only do one at a time because all kmers are stored in a single table. It can still do non-kmer-based operations such as quality trimming at the same time.</span><br /><br /><span>BBDuk2 can do all 4 kmer operations at once and is designed for integration into automated pipelines where you do contaminant removal and adapter-trimming in a single pass to minimize filesystem I/O. Personally, I never use BBDuk2 from the command line. Both have identical capabilities and functionality otherwise, but the syntax is different.</span><br /><br /><span>------------------------------------------------------------------</span><br /><br /><strong>Randomreads.sh</strong></p><ul>
<li><strong>Generate random reads in various formats</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ randomreads.sh ref=genome.fasta out=reads.fq len=100 reads=10000</pre></div><p><span>You can specify paired reads, an insert size distribution, read lengths (or length ranges), and so forth. But because I developed it to benchmark mapping algorithms, it is specifically designed to give excellent control over mutations. You can specify the number of snps, insertions, deletions, and Ns per read, either exactly or probabilistically; the lengths of these events is individually customizable, the quality values can alternately be set to allow errors to be generated on the basis of quality; there's a PacBio error model; and all of the reads are annotated with their genomic origin, so you will know the correct answer when mapping.</span><br /><br /><span>Bear in mind that 50% of the reads are going to be generated from the plus strand and 50% from the minus strand. So, either a read will match the reference perfectly, OR its reverse-complement will match perfectly.</span><br /><br /><span>You can generate the same set of reads with and without SNPs by fixing the seed to a positive number, like this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ randomreads.sh maxsnps=0 adderrors=false out=perfect.fastq reads=1000 minlength=18 maxlength=55 seed=5

$ randomreads.sh maxsnps=2 snprate=1 adderrors=false out=2snps.fastq reads=1000 minlength=18 maxlength=55 seed=5</pre></div><p><span>[As of BBmap v. 36.59] rendomreads.sh gains the ability to simulate metagenomes.&nbsp;</span><br /><br /><span>coverage=X will automatically set "reads" to a level that will give X average coverage (decimal point is allowed).</span><br /><br /><span>metagenome will assign each scaffold a random exponential variable, which decides the probability that a read be generated from that scaffold. So, if you concatenate together 20 bacterial genomes, you can run randomreads and get a metagenomic-like distribution. It could also be used for RNA-seq when using a transcriptome reference.</span><br /><br /><span>The coverage is decided on a per-reference-sequence level, so if a bacterial assembly has more than one contig, you may want to glue them together first with fuse.sh before concatenating them with the other references.&nbsp;</span><br /><br /></p><ul>
<li><strong>Simulate a jump library</strong></li>
</ul><p><br /><span>You can simulate a 4000bp jump library from your existing data like this.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ cat assembly1.fa assembly2.fa &gt; combined.fa
$ bbmap.sh ref=combined.fa
$ randomreads.sh reads=1000000 length=100 paired interleaved mininsert=3500 maxinsert=4500 bell perfect=1 q=35 out=jump.fq.gz</pre></div><p><span>--------------------------------------------------------------</span><br /><strong>Shred.sh</strong><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ shred.sh in=ref.fasta out=reads.fastq length=200</pre></div><p><span>The difference is that RandomReads will make reads in a random order from random locations, ensuring flat coverage on average, but it won't ensure 100% coverage unless you generate many fold depth. Shred, on the other hand, gives you exactly 1x depth and exactly 100% coverage (and is not capable of modelling errors). So, the use-cases are different.&nbsp;</span><br /><span>---------------------------------------------------------------</span><br /><strong>Demuxbyname.sh</strong></p><ul>
<li><strong>Demultiplex fastq files when the tag is present in the fastq read header (illumina)</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ demuxbyname.sh in=r#.fq out=out_%_#.fq prefixmode=f names=GGACTCCT+GCGATCTA,TAAGGCGA+TCTACTCT,...
outu=filename</pre></div><p><span>"Names" can also be a text file with one barcode per line (in exactly the format found in the read header). You do have to include all of the expected barcodes, though.</span><br /><br /><span>In the output filename, the "%" symbol gets replaced by the barcode; in both the input and output names, the "#" symbol gets replaced by 1 or 2 for read 1 or read 2. It's optional, though; you can leave it out for interleaved input/output, or specify in1=/in2=/out1=/out2= if you want custom naming.</span><br /><br /><span>----------------------------------------------------------------</span><br /><br /><strong>Readlength.sh</strong></p><ul>
<li><strong>Plotting the length distribution of reads</strong></li>
</ul><div><div>Code:</div><pre dir="ltr">$ readlength.sh in=file out=histogram.txt bin=10 max=80000</pre></div><p><span>That will plot the result in bins of size 10, with everything above 80k placed in the same bin. The defaults are set for relatively short sequences so if they are many megabases long you may need to add the flag "-Xmx8g" and increase "max=" to something much higher.</span><br /><br /><span>Alternatively, if these are assemblies and you're interested in continuity information (L50, N50, etc), you can run stats on each or statswrapper on all of them:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">stats.sh in=file</pre></div><p><span>or</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">statswrapper.sh in=file,file,file,file&hellip;</pre></div><p><span>----------------------------------------------------------------</span><br /><strong>Filterbyname.sh</strong><br /><br /><span>By default, "filterbyname" discards reads with names in your name list, and keeps the rest. To include them and discard the others, do this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ filterbyname.sh in=003.fastq out=filter003.fq names=names003.txt include=t</pre></div><p><span>----------------------------------------------------------------</span><br /><strong>getreads.sh</strong><br /><br /><span>If you only know the number(s) of the fasta/fastq record(s) in a file (records start at 0) then you can use the following command to extract those reads in a new file.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ getreads.sh in= id=&lt;number,number,number...&gt; out=</pre></div><p><span>The first read (or pair) has ID 0, the second read (or pair) has ID 1, etc.</span><br /><br /><span>Parameters:</span><br /><span>in= Specify the input file, or stdin.</span><br /><span>out= Specify the output file, or stdout.</span><br /><span>id= Comma delimited list of numbers or ranges, in any order.</span><br /><span>For example: id=5,93,17-31,8,0,12-13&nbsp;</span><br /><span>----------------------------------------------------------------</span><br /><strong>Splitsam.sh</strong></p><ul>
<li><strong>Splits a sam file into forward and reverse reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">splitsam.sh mapped.sam plus.sam minus.sam unmapped.sam
reformat.sh in=plus.sam out=plus.fq
reformat.sh in=minus.sam out=minus.fq rcomp</pre></div><p><span>----------------------------------------------------------------</span><br /><strong>BBSplit.sh</strong><br /><br /><span>BBSplit now has the ability to output paired reads in dual files using the # symbol. For example:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ bbsplit.sh ref=x.fa,y.fa in1=read1.fq in2=read2.fq basename=o%_#.fq</pre></div><p><span>will produce ox_1.fq, ox_2.fq, oy_1.fq, and oy_2.fq</span><br /><br /><span>You can use the # symbol for input also, like "in=read#.fq", and it will get expanded into 1 and 2.&nbsp;&nbsp;</span><br /><br /><strong>Added feature:&nbsp;</strong><span>One can specify a directory for the "ref=" argument. If anything in the list is a directory, it will use all fasta files in that directory. They need a fasta extension, like .fa or .fasta, but can be compressed with an additional .gz after that. Reason this is useful is to use BBSplit is to have it split input into one output file per reference file.</span><br /><br /><br /><strong>NOTE: 1</strong><span>&nbsp;By default BBSplit uses fairly strict mapping parameters; you can get the same sensitivity as BBMap by adding the flags "minid=0.76 maxindel=16k minhits=1". With those parameters it is extremely sensitive.</span><br /><br /><strong>NOTE: 2</strong><span>&nbsp;BBSplit has different ambiguity settings for dealing with reads that map to multiple genomes. In any case, if the alignment score is higher to one genome than another, it will be associated with that genome only (this considers the combined scores of read pairs - pairs are always kept together). But when a read or pair has two identically-scoring mapping locations, on different genomes, the behavior is controlled by the "ambig2" flag - "ambig2=toss" will discard the read, "all" will send it to all output files, and "split" will send it to a separate file for ambiguously-mapped reads (one per genome to which it maps).</span><br /><br /><strong>NOTE: 3</strong><span>&nbsp;Zero-count lines are suppressed by default, but they should be printed if you include the flag "nzo=f" (nonzeroonly=false).&nbsp;</span><br /><br /><strong>NOTE: 4</strong><span>&nbsp;BBSplit needs multiple reference files as input; one per organism, or one for target and another for everything else. It only outputs one file per reference file.</span><br /><br /><span>Seal.sh, on the other hand, which is similar, can use a single concatenated file, as it (by default) will output one file per reference sequence within a concatenated set of references.&nbsp;</span><br /><span>--------------------------------------------------------------</span><br /><strong>Pileup.sh</strong></p><ul>
<li><strong>To generate transcript coverage stats</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ pileup.sh in=mapped.sam normcov=normcoverage.txt normb=20 stats=stats.txt</pre></div><p><span>That will generate coverage per transcript, with 20 lines per transcript, each line showing the coverage for that fraction of the transcript. "stats" will contain other information like the fraction of bases in each transcript that was covered.&nbsp;</span></p><ul>
<li><strong>To calculate physical coverage stats (region covered by paired-end reads)&nbsp;</strong></li>
</ul><p><span>BBMap has a "physcov" flag that allows it to report physical rather than sequenced coverage. It can be used directly in BBMap, or with pileup, if you already have a sam file. For example:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ pileup.sh in=mapped.sam covstats=coverage.txt</pre></div><ul>
<li><strong>Calculating coverage of the genome</strong></li>
</ul><p><br /><span>Program will take sam or bam, sorted or unsorted.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ pileup.sh in=mapped.sam out=stats.txt hist=histogram.txt</pre></div><p><span>stats.txt will contain the average depth and percent covered of each reference sequence; the histogram will contain the exact number of bases with a each coverage level. You can also get per-base coverage or binned coverage if you want to plot the coverage. It also generates median and standard deviation, and so forth.</span><br /><br /><span>It's also possible to generate coverage directly from BBMap, without an intermediate sam file, like this:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ bbmap.sh in=reads.fq ref=reference.fasta nodisk covstats=stats.txt covhist=histogram.txt</pre></div><p><span>We use this a lot in situations where all you care about is coverage distributions, which is somewhat common in metagenome assemblies. It also supports most of the flags that pileup.sh supports, though the syntax is slightly different to prevent collisions. In each case you can see all the possible flags by running the shellscript with no arguments.</span></p><ul>
<li><strong>To bin aligned reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ pileup.sh in=mapped.sam out=stats.txt bincov=coverage.txt binsize=1000</pre></div><p><span>That will give coverage within each bin. For read density regardless of read length, add the "startcov=t" flag.&nbsp;&nbsp;</span><br /><br /><span>--------------------------------------------------------------</span><br /><strong>Dedupe.sh</strong><br /><br /><span>Dedupe ensures that there is at most one copy of any input sequence, optionally allowing contaminants (substrings) to be removed, and a variable hamming or edit distance to be specified. Usage:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ dedupe.sh in=assembly1.fa,assembly2.fa out=merged.fa</pre></div><p><span>That will absorb exact duplicates and containments. You can use "hdist" and "edist" flags to allow mismatches, or get a complete list of flags by running the shellscript with no arguments.&nbsp;&nbsp;</span><br /><br /><span>Dedupe&nbsp;</span><span style="text-decoration: underline;">will merge assemblies</span><span>, but it&nbsp;</span><span style="text-decoration: underline;">will not produce consensus sequences or join overlapping reads</span><span>; it only removes sequences that are fully contained within other sequences (allowing the specified number of mismatches or edits).</span><br /><br /><span>Dedupe can remove duplicate reads from multiple files simultaneously, if they are comma-delimited (e.g. in=file1.fastq,file2.fastq,file3.fastq). And if you set the flag "uniqueonly=t" then ALL copies of duplicate reads will be removed, as opposed to the default behavior of leaving one copy of duplicate reads.</span><br /><br /><span>However, it does not care which file a read came from; in other words, it can't remove only reads that are duplicates across multiple files but leave the ones that are duplicates within a file. That can still be accomplished, though, like this:</span><br /><br /><span>1) Run dedupe on each sample individually, so now there are at most 1 copy of a read per sample.</span><br /><span>2) Run dedupe again on all of the samples together, with "uniqueonly=t". The only remaining duplicate reads will be the ones duplicated between samples, so that's all that will be removed.&nbsp;&nbsp;</span><br /><br /><span>--------------------------------------------------------------</span></p><ul>
<li><strong>Generate ROC curves from any aligner</strong></li>
</ul><p><br /><strong>[*]index the reference<br /><br /></strong></p><div><div>Code:</div><pre dir="ltr">$ bbmap.sh ref=reference.fasta</pre></div><p><br /><strong>[*]Generate random reads</strong><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ randomreads.sh reads=100000 length=100 out=synth.fastq maxq=35 midq=25 minq=15</pre></div><p><strong>[*]Map to produce a sam file</strong><br /><br /><span>...substitute this command with the appropriate one from your aligner of choice</span></p><div><div>Code:</div><pre dir="ltr">$ bbmap.sh in=synth.fq out=mapped.sam</pre></div><p><strong>[*]Generate ROC curve</strong><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ samtoroc.sh in=mapped.sam reads=100000</pre></div><p><span>--------------------------------------------------------------</span></p><ul>
<li><strong>Calculate heterozygous rate for sequence data</strong></li>
</ul><p><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ kmercountexact.sh in=reads.fq khist=histogram.txt peaks=peaks.txt</pre></div><p><span>You can examine the histogram manually, or use the "peaks" file which tells you the number of unique kmers in each peak on the histogram. For a diploid, the first peak will be the het peak, the second will be the homozygous peak, and the rest will be repeat peaks. The peak caller is not perfect, though, so particularly with noisy data I would only rely on it for the first two peaks, and try to quantify the higher-order peaks manually if you need to (which you generally don't).</span><br /><br /><span>-----------------------------------------------------------------</span></p><ul>
<li><strong>Compare mapped reads between two files</strong></li>
</ul><p><br /><span>To see how many mapped reads (can be mapped concordant or discordant, doesn't matter) are shared between the two alignment files and how many mapped reads are unique to one file or the other.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ reformat.sh in=file1.sam out=mapped1.sam mappedonly
$ reformat.sh in=file2.sam out=mapped2.sam mappedonly</pre></div><p><span>That gets you the mapped reads only. Then:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ filterbyname.sh in=mapped1.sam names=mapped2.sam out=shared.sam include=t</pre></div><p><span>...which gets you the set intersection;</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ filterbyname.sh in=mapped1.sam names=mapped2.sam out=only1.sam include=f
$ filterbyname.sh in=mapped2.sam names=mapped1.sam out=only2.sam include=f</pre></div><p><span>...which get you the set subtractions.&nbsp;&nbsp;</span><br /><br /><span>--------------------------------------------------------------</span><br /><br /><strong>BBrename.sh</strong></p><div><div>Code:</div><pre dir="ltr">$ bbrename.sh in=old.fasta out=new.fasta</pre></div><p><span>That will rename the reads as 1, 2, 3, 4, ... 222.</span><br /><br /><span>You can also give a custom prefix if you want. The input has to be text format, not .doc.&nbsp;&nbsp;</span><br /><br /><span>---------------------------------------------------------------------</span><br /><br /><strong>BBfakereads.sh</strong></p><ul>
<li><strong>Generating &ldquo;fake&rdquo; paired end reads from a single end read file</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ bfakereads.sh in=reads.fastq out1=r1.fastq out2=r2.fastq length=100</pre></div><p><span>That will generate fake pairs from the input file, with whatever length you want (maximum of input read length). We use it in some cases for generating a fake LMP library for scaffolding from a set of contigs. Read 1 will be from the left end, and read 2 will be reverse-complemented and from the right end; both will retain the correct original qualities. And " /1" " /2" will be suffixed after the read name.&nbsp;&nbsp;</span><br /><br /><span>------------------------------------------------------------------</span><br /><strong>Randomreads.sh</strong></p><ul>
<li><strong>Generate random reads</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ randomreads.sh ref=genome.fasta out=reads.fq len=100 reads=10000</pre></div><p><span>"seed=-1" will use a random seed; any other value will use that specific number as the seed</span><br /><br /><span>You can specify paired reads, an insert size distribution, read lengths (or length ranges), and so forth. But because I developed it to benchmark mapping algorithms, it is specifically designed to give excellent control over mutations. You can specify the number of snps, insertions, deletions, and Ns per read, either exactly or probabilistically; the lengths of these events is individually customizable, the quality values can alternately be set to allow errors to be generated on the basis of quality; there's a PacBio error model; and all of the reads are annotated with their genomic origin, so you will know the correct answer when mapping.</span><br /><br /><span>--------------------------------------------------------------------</span></p><ul>
<li><strong>Generate saturation curves to assess sequencing depth</strong></li>
</ul><p>&nbsp;</p><div><div>Code:</div><pre dir="ltr">$ bbcountunique.sh in=reads.fq out=histogram.txt</pre></div><p><span>It works by pulling kmers from each input read, and testing whether it has been seen before, then storing it in a table.</span><br /><br /><span>The bottom line, "first", tracks whether the first kmer of the read has been seen before (independent of whether it is read 1 or read 2).</span><br /><br /><span>The top line, "pair", indicates whether a combined kmer from both read 1 and read 2 has been seen before. The other lines are generally safe to ignore but they track other things, like read1- or read2-specific data, and random kmers versus the first kmer.</span><br /><br /><span>It plots a point every X reads (configurable, default 25000).</span><br /><br /><span>In noncumulative mode (default), a point indicates "for the last X reads, this percentage had never been seen before". In this mode, once the line hits zero, sequencing more is not useful.</span><br /><br /><span>In cumulative mode, a point indicates "for all reads, this percentage had never been seen before", but still only one point is plotted per X reads.</span><br /><br /><span>-----------------------------------------------------------------</span><br /><strong>CalcTrueQuality.sh</strong><br /><br /><a href="http://seqanswers.com/forums/showthread.php?p=170904" target="_blank">http://seqanswers.com/forums/showthread.php?p=170904</a><br /><br /><span>In light of the quality-score issues with the NextSeq platform, and the possibility of future Illumina platforms (HiSeq 3000 and 4000) also using quantized quality scores, I developed it for recalibrating the scores to ensure accuracy and restore the full range of values.</span><br /><br /><span>-----------------------------------------------------------------</span><br /><br /><strong>BBMapskimmer.sh</strong><br /><br /><span>BBMap is designed to find the best mapping, and heuristics will cause it to ignore mappings that are valid but substantially worse. Therefore, I made a different version of it, BBMapSkimmer, which is designed to find all of the mappings above a certain threshold. The shellscript is bbmapskimmer.sh and the usage is similar to bbmap.sh or mapPacBio.sh. For primers, which I assume will be short, you may wish to use a lower than default K of, say, 10 or 11, and add the "slow" flag.</span><br /><br /><span>--------------------------------------------------------------</span><br /><br /><strong>msa.sh and curprimers.sh</strong><br /><br /><span>Quoted from Brian's response directly.</span><br /><br /><span>I also wrote another pair of programs specifically for working with primer pairs, msa.sh and cutprimers.sh. msa.sh will forcibly align a primer sequence (or a set of primer sequences) against a set of reference sequences to find the single best matching location per reference sequence - in other words, if you have 3 primers and 100 ref sequences, it will output a sam file with exactly 100 alignments - one per ref sequence, using the primer sequence that matched best. Of course you can also just run it with 1 primer sequence.</span><br /><br /><span>So you run msa twice - once for the left primer, and once for the right primer - and generate 2 sam files. Then you feed those into cutprimers.sh, which will create a new fasta file containing the sequence between the primers, for each reference sequence. We used these programs to synthetically cut V4 out of full-length 16S sequences.</span><br /><br /><span>I should say, though, that the primer sites identified are based on the normal BBMap scoring, which is not necessarily the same as where the primers would bind naturally, though with highly conserved regions there should be no difference.</span><br /><br /><span>------------------------------------------------------</span><br /><strong>testformat.sh</strong><br /><br /><strong>Identify type of Q-score encoding in sequence files</strong><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ testformat.sh in=seq.fq.gz
sanger    fastq    gz    interleaved    150bp</pre></div><p><span>--------------------------------------------------</span><br /><strong>kcompress.sh</strong><br /><br /><span>Newest member of BBTools. Identify constituent k-mers.&nbsp;</span><br /><a href="http://seqanswers.com/forums/showthread.php?t=63258" target="_blank">http://seqanswers.com/forums/showthread.php?t=63258</a><br /><br /><span>----------------------------------------------------</span><br /><strong>commonkmers.sh</strong><br /><br /><span>Find all k-mers for a given sequence.</span></p><div><div>Code:</div><pre dir="ltr">$ commonkmers.sh in=reads.fq out=kmers.txt k=4 count=t display=999</pre></div><p><span>Will produce output that looks like</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">MISEQ05:239:000000000-A74HF:1:2110:14788:23085	ATGA=8	ATGC=6	GTCA=6	AAAT=5	AAGC=5	AATG=5	AGCA=5	ATAA=5	ATTA=5	CAAA=5	CATA=5	CATC=5	CTGC=5	AACC=4	AACG=4	AAGA=4	ACAT=4	ACCA=4	AGAA=4	ATCA=4	ATGG=4	CAAG=4	CCAA=4	CCTC=4	CTCA=4	CTGA=4	CTTC=4	GAGC=4	GGTA=4	GTAA=4	GTTA=4	AAAA=3	AAAC=3	AAGT=3	ACCG=3	ACGG=3	ACTG=3	AGAT=3	AGCT=3	AGGA=3	AGTA=3	AGTC=3	CAGC=3	CATG=3	CGAG=3	CGGA=3	CGTC=3	CTAA=3	CTCC=3	CTTA=3	GAAA=3	GACA=3	GACC=3	GAGA=3	GCAA=3	GGAC=3	TCAA=3	TGCA=3	AAAG=2	AACA=2	AATA=2	AATC=2	ACAA=2	ACCC=2	ACCT=2	ACGA=2	ACGC=2	AGAC=2	AGCG=2	AGGC=2	CAAC=2	CAGG=2	CCGC=2	GCCA=2	GCTA=2	GGAA=2	GGCA=2	TAAA=2	TAGA=2	TCCA=2	TGAA=2	AAGG=1	AATT=1	ACGT=1	AGAG=1	AGCC=1	AGGG=1	ATAC=1	ATAG=1	ATTG=1	CACA=1	CACG=1	CAGA=1	CCAC=1	CCCA=1	CCGA=1	CCTA=1	CGAC=1	CGCA=1	CGCC=1	CGCG=1	CGTA=1	CTAC=1	GAAC=1	GCGA=1	GCGC=1	GTAC=1	GTGA=1	TTAA=1</pre></div><p><span>-----------------------------------------------------</span><br /><strong>Mutate.sh</strong><br /><br /><span>Simulate multiple mutants from a known reference (e.g.&nbsp;</span><em>E. coli</em><span>).</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">$ mutate.sh in=e_coli.fasta out=mutant.fasta id=99 
$ randomreads.sh ref=mutant.fasta out=reads.fq.gz reads=5m length=150 paired adderrors</pre></div><p><span>That will create a mutant version of E.coli with 99% identity to the original, and then generate 5 million simulated read pairs from the new genome. You can repeat this multiple times; each mutant will be different.</span><br /><br /><span>------------------------------------</span><br /><br /><strong>Partition.sh</strong><br /><br /><span>One can partition a large dataset with partition.sh into smaller subsets (example below splits data into 8 chunks).</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">partition.sh in=r1.fq in2=r2.fq out=r1_part%.fq out2=r2_part%.fq ways=8</pre></div><p><span>-----------------------------------</span><br /><strong>clumpify.sh</strong><br /><br /><span>If you are concerned about file size and want the files to be as small as possible, give Clumpify a try. It can reduce filesize by around 30% losslessly by reordering the reads. I've found that this also typically accelerates subsequent analysis pipelines by a similar factor (up to 30%). Usage:</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">clumpify.sh in=reads.fastq.gz out=clumped.fastq.gz</pre></div><div><div>Code:</div><pre dir="ltr">clumpify.sh in1=reads_R1.fastq.gz in2=reads_R2.fastq.gz out1=clumped_R1.fastq.gz out2=clumped_R2.fastq.gz</pre></div><ul>
<li><strong>Clumpify.sh can now mark/remove sequence duplicates (optical/PCR/otherwise) from NGS data</strong></li>
</ul><p><br /><span>This does NOT require alignments so it should prove more useful compared to Picard MarkDuplicates. Relevant options for clumpify.sh command are listed below.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">dedupe=f optical=f (default)
Nothing happens with regards to duplicates.

dedupe=t optical=f
All duplicates are detected, whether optical or not.  All copies except one are removed for each duplicate.

dedupe=f optical=t
Nothing happens.

dedupe=t optical=t

Only optical duplicates (those with an X or Y coordinate within dist) are detected.  All copies except one are removed for each duplicate.
The allduplicates flag makes all copies of duplicates removed, rather than leaving a single copy.  But like optical, it has no effect unless dedupe=t.

Note: If you set "dupedist" to anything greater than 0, "optical" gets enabled automatically.</pre></div><p><span>-------------------------------------</span><br /><strong>fuse.sh</strong><br /><br /><span>Fuse will automatically reverse-complement read 2. Pad (N) amount can be adjusted as necessary. This will for example create a full size amplicon that can be used for alignments.</span><br /><br /></p><div><div>Code:</div><pre dir="ltr">fuse.sh in1=r1.fq in2=r2.fq pad=130 out=fused.fq fusepairs</pre></div>]]></description>
	<dc:creator>Surabhi Chaudhary</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/44593/bear-better-emulation-for-artificial-reads</guid>
	<pubDate>Sat, 06 Jul 2024 04:27:53 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/44593/bear-better-emulation-for-artificial-reads</link>
	<title><![CDATA[BEAR: Better Emulation for Artificial Reads]]></title>
	<description><![CDATA[<p dir="auto">Created by Stephen Johnson, Brett Trost, Dr. Jeffrey R. Long, Dr. Anthony Kusalik University of Saskatchewan, Department of Computer Science</p>
<p dir="auto">BEAR is intended to be an easy-to-use collection of scripts for generating simulated WGS metagenomic reads with read lengths, quality scores, error profiles, and species abundances derived from real user-supplied WGS data.</p><p>Address of the bookmark: <a href="https://github.com/sej917/BEAR" rel="nofollow">https://github.com/sej917/BEAR</a></p>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/926/list-of-popular-bioinformatics-softwaretools</guid>
	<pubDate>Tue, 16 Jul 2013 14:30:30 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/926/list-of-popular-bioinformatics-softwaretools</link>
	<title><![CDATA[List of popular bioinformatics software/tools]]></title>
	<description><![CDATA[<p><a href="http://samtools.sourceforge.net/swlist.shtml">I</a>n current genome era, our day to day work is to handle the huge geneome sequences, expression data, several other datasets. This link provide a comprehensive list of commonly used sofware/tools.</p><p>Address of the bookmark: <a href="http://samtools.sourceforge.net/swlist.shtml" rel="nofollow">http://samtools.sourceforge.net/swlist.shtml</a></p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/8265/list-of-generic-simulation-softwaretoolsresource-with-brief-description-and-homepage</guid>
	<pubDate>Mon, 10 Feb 2014 05:57:29 -0600</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/8265/list-of-generic-simulation-softwaretoolsresource-with-brief-description-and-homepage</link>
	<title><![CDATA[List of generic simulation software/tools/resource with brief description and homepage !!!]]></title>
	<description><![CDATA[<p>List of generic simulation software/tools/resource with brief description and homepage</p><p><img src="http://www.evolution-of-life.com/fileadmin/images/carousel/genetic.PNG" alt="image" style="border: 0px;"></p><p>ALF <br />A Simulation Framework for Genome Evolution <br />http://www.cbrg.ethz.ch/alf<br /><br />Bayesian Serial SimCoal <br />Bayesian Serial SimCoal, (BayeSSC) is a modification of SIMCOAL 1.0, a program written by Laurent Excoffier, John Novembre, and Stefan Schneider. <br />http://www.stanford.edu/group/hadlylab/ssc/index.html<br /><br />BEERS <br />BEERS was designed to benchmark RNA-Seq alignment algorithms and also algorithms that aim to reconstruct different isoforms and alternate splicing from RNA-Seq data <br />http://cbil.upenn.edu/beers/<br /><br />BOTTLENECK <br />Bottleneck is a program for detecting recent effective population size reductions from allele data frequencies <br />http://www.ensam.inra.fr/urlb/bottleneck/bottleneck.html<br /><br />BottleSim <br />BottleSim is a computer simulation program for simulating the process of population bottlenecks <br />http://chkuo.name/software/bottlesim.html<br /><br />CASS <br />Protein Sequence Simulation <br />http://www.wyomingbioinformatics.org/liberlesgroup/cass/<br /><br />CDPOP <br />CDPOP is a landscape genetics tool for simulating the emergence of spatial genetic structure in populations resulting from specified landscape processes governing organism movement behavior. <br />http://cel.dbs.umt.edu/cdpop<br /><br />CoalFace <br />CoalFace is a simulation of the coalescent process with the visual display of gene genealogies. <br />http://web.up.ac.za/default.asp?ipkcategoryid=3283<br /><br />CoaSim <br />CoaSim is a tool for simulating the coalescent process with recombination and geneconversion under various demographic models. <br />http://users-birc.au.dk/mailund/coasim/index.html<br /><br />cosi <br />The cosi package is written in C and is available as a tar file. <br />http://www.broadinstitute.org/~sfs/cosi/<br /><br />CS-PSeq-Gen <br />A program to simulate the evolution of protein sequences under the constraints of the information of a particular reconstructed phylogeny <br />http://bioserv.rpbs.univ-paris-diderot.fr/software/cs-pseq-gen.html<br /><br />DAWG <br />An application designed to simulate the evolution of recombinant DNA sequences in continuous time <br />http://scit.us/projects/dawg<br /><br />Easypop <br />EASYPOP is an individual based model intended to simulate datasets under a very broad range of conditions <br />http://www.unil.ch/dee/page36926_fr.html<br /><br />EggLib <br />EggLib is a C++/Python library and program package for evolutionary genetics and genomics. <br />http://egglib.sourceforge.net/<br /><br />EvolSimulator <br />A simulation test bed for hypotheses of genome evolution <br />http://acb.qfab.org/acb/evolsim/<br /><br />EvolveAGene <br />A realistic coding sequence simulation program that separates mutation from selection and allows the user to set selection conditions <br />http://bellinghamresearchinstitute.com/software/index.html<br /><br />fastsimcoal <br />A continuous-&not;‐time coalescent simulator of genomic diversity under arbitrarily complex evolutionary scenarios <br />http://cmpg.unibe.ch/software/fastsimcoal/<br /><br />FastSLINK <br />Simulation of Marker and Phenotype Data in Pedigrees <br />http://watson.hgen.pitt.edu/<br /><br />FFPopSim <br />C++/Python library for population genetics. <br />http://webdav.tuebingen.mpg.de/ffpopsim/<br /><br />FLUX SIMULATOR <br />The Flux Simulator aims at providing a deterministic in silico reproduction of the experimental pipelines for RNA-Seq, employing a minimal set of parameters. <br />http://flux.sammeth.net/simulator.html<br /><br />ForSim <br />ForSim: A Forward Evolutionary Computer Simulation <br />http://www.anthro.psu.edu/weiss_lab/research.shtml<br /><br />ForwSim <br />The program given below is based on the algorithm described in Padhukasahasram et al. 2008 to simulate genetic drift in a standard Wright-Fisher process. <br />http://badri-populationgeneticsimulators.blogspot.com/<br /><br />FPG <br />Forward Population Genetic simulation <br />http://genfaculty.rutgers.edu/hey/software#fpg<br /><br />FREGENE <br />FREGENE is a C++ program that simulates sequence-like data over large genomic regions in large diploid populations. <br />http://www.ebi.ac.uk/projects/bargen/download/fregen/documentation_html.html<br /><br />GAMETES <br />Genetic Architecture Model Emulator for Testing and Evaluating Software: Simulates complex SNP models with pure, strict epistatic interactions with n-loci. <br />http://sourceforge.net/projects/gametes/?source=navbar<br /><br />GASP <br />Genometric Analysis Simulation Program. A software tool for testing and investigating methods in statistical genetics by generating samples of family data based on user specified models. <br />http://research.nhgri.nih.gov/gasp/<br /><br />GemSIM <br />Next generation sequencing read simulator <br />http://sourceforge.net/projects/gemsim/<br /><br />GeneArtisan <br />Simulation of Markers in Case-Control Study Designs <br />http://www.rannala.org/?page_id=241<br /><br />GENOME <br />A rapid coalescent-based whole genome simulator <br />http://www.sph.umich.edu/csg/liang/genome/<br /><br />GenomePop2 <br />GenomePop2 is a specialization of the program GenomePop just to manage SNPs under more flexible and useful settings. If you need models with more than 2 alleles please use the GenomePop program version. <br />http://webs.uvigo.es/acraaj/genomepop2.htm<br /><br />GenomeSimla <br />GenomeSIMLA is currently under development- however, we have a beta release that we are asking to be tested <br />http://chgr.mc.vanderbilt.edu/genomesimla/<br /><br />GENS2 <br />Simulates interactions among two genetic and one environmental factor and also allows for epistatic interactions. <br />https://sourceforge.net/projects/gensim/<br /><br />GWAsimulator <br />A rapid whole genome simulation program <br />http://biostat.mc.vanderbilt.edu/wiki/main/gwasimulator<br /><br />HAP-SAMPLE <br />An association simulator for candidate regions or genome scans <br />http://www.hapsample.org/<br /><br />HAPGEN <br />A simulator for the simulation of case control datasets at SNP markers <br />https://mathgen.stats.ox.ac.uk/genetics_software/hapgen/hapgen2.html<br /><br />HapSim <br />A simulation tool for generating haplotype data with pre-specified allele frequencies and LD coefficients <br />http://cran.r-project.org/web/packages/hapsim/index.html<br /><br />HAPSIMU <br />A program that simulates heterogeneous populations with various known and controllable structures under the continuous migration model or the discrete model <br />http://l.web.umkc.edu/liujian/<br /><br />IBDsim <br />IBDSim is a computer package for the simulation of genotypic data under general isolation by distance models. <br />http://raphael.leblois.free.fr/<br /><br />indel-Seq-Gen <br />A biological sequence simulation program that simulates highly divergent DNA sequences and protein superfamilies <br />http://bioinfolab.unl.edu/~cstrope/isg/<br /><br />Indelible <br />A powerful and flexible simulator of biological evolution <br />http://abacus.gene.ucl.ac.uk/software/indelible/<br /><br />invertFREGENE <br />InvertFREGENE is a forward-in-time simulator of inversions in population genetic data <br />http://www.ebi.ac.uk/projects/bargen/<br /><br />kernalPop <br />A spatially explicit population genetic simulation engine <br />http://cran.r-project.org/src/contrib/archive/kernelpop/<br /><br />MaCS <br />Markovian Coalescent Simulator <br />http://www-hsc.usc.edu/~garykche/<br /><br />Mason <br />A package for the simulation of nucleotide data. <br />http://www.seqan.de/projects/mason/<br /><br />mbs <br />modifying Hudson's ms software to generate samples of DNA sequences with a biallelic site under selection <br />http://www.sendou.soken.ac.jp/esb/innan/innanlab/software.html<br /><br />Mendel's Accountant <br />Mendel's Accountant (MENDEL) is an advanced numerical simulation program for modeling genetic change over time and was developed collaboratively by Sanford, Baumgardner, Brewer, Gibson and ReMine <br />http://mendelsaccount.sourceforge.net/<br /><br />MetaSim <br />A tool to generate collections of synthetic reads that reflect the diverse taxonomical composition of typical metagenome data sets <br />http://ab.inf.uni-tuebingen.de/software/metasim/<br /><br />mlcoalsim <br />Multilocus Coalescent Simulations <br />http://code.google.com/p/mlcoalsim-v1/<br /><br />ms <br />The purpose of this program is to allow one to investigate the statistical properties of such samples, to evaluate estimators or statistical tests, and generally to aid in the interpretation of polymorphism data sets. <br />http://home.uchicago.edu/~rhudson1/source/mksamples.html<br /><br />msHOT <br />The purpose of this program is to allow one to investigate the statistical properties of such samples, to evaluate estimators or statistical tests, and generally to aid in the interpretation of polymorphism data sets. <br />http://home.uchicago.edu/~rhudson1/<br /><br />msms <br />A coalescent Simlation tool with selection. <br />http://www.mabs.at/ewing/msms/index.shtml<br /><br />MySSP <br />A program for the simulation of DNA sequence evolution across a phylogenetic tree <br />http://www.rosenberglab.net/software.php<br /><br />Nemo <br />A forward-time, individual-based, genetically explicit, and stochastic simulation program designed to study the evolution of genetic markers, life history traits, and phenotypic traits in a flexible (meta-)population framework. <br />http://nemo2.sourceforge.net/<br /><br />NetRecodon <br />Coalescent simulation of coding DNA sequences with recombination (inter and intracodon), migration and demography <br />http://code.google.com/p/netrecodon/<br /><br />PEDAGOG <br />Software for simulating eco-evolutionary population dynamics <br />https://bcrc.bio.umass.edu/pedigreesoftware/node/5<br /><br />phenosim <br />A tool to add phenotypes to simulated genotypes <br />http://evoplant.uni-hohenheim.de/doku.php?id=software:software<br /><br />PhyloSim <br />An R package for the Monte Carlo simulation of sequence evolution <br />http://bit.ly/rlsim-git<br /><br />pIRS <br />Profile-based Illumina pair-end reads simulator <br />https://code.google.com/p/pirs/<br /><br />ProteinEvolver <br />Simulation of protein evolution along phylogenies under structure-based substitution models <br />http://code.google.com/p/proteinevolver/<br /><br />QMSim <br />QTL and Marker Simulator <br />http://www.aps.uoguelph.ca/~msargol/qmsim/<br /><br />quantiNEMO <br />An individual-based program for the analysis of quantitative traits with explicit genetic architecture potentially under selection in a structured population <br />http://www2.unil.ch/popgen/softwares/quantinemo/<br /><br />RECOAL <br />Simulates new haplotype data from a reference population of haplotypes. <br />ftp://popgen.usc.edu/<br /><br />Recodon <br />Coalescent simulation of coding DNA sequences with recombination, migration and demography <br />http://code.google.com/p/recodon/<br /><br />rlsim <br />A package for simulating RNA-seq library preparation with parameter estimation <br />http://bit.ly/rlsim-git<br /><br />Rmetasim <br />Rmetasim is a front-end for the metasim engine that is implemented as a package that runs in the statistical computing environment R <br />http://linum.cofc.edu/software.html#metasim<br /><br />RNA Seq Simulator <br />RSS takes SAM alignment files from RNA-Seq data and simulates over dispersed, multiple replica, differential, non-stranded RNA-Seq datasets. <br />http://useq.sourceforge.net/cmdlnmenus.html#rnaseqsimulator<br /><br />Rose <br />Random model of sequence evolution <br />http://bibiserv.techfak.uni-bielefeld.de/rose/<br /><br />SelSim <br />SelSim is a program for Monte Carlo simulation of DNA polymorphism data for a recom- bining region within which a single bi-allelic site has experienced natural selection <br />http://www.well.ox.ac.uk/~spencer/selsim/<br /><br />Seq-Gen <br />An application for the Monte Carlo simulation of molecular sequence evolution along phylogenetic trees. <br />http://tree.bio.ed.ac.uk/software/seqgen/<br /><br />SEQPower <br />Statistical power analysis for sequence-based association studies <br />http://bioinformatics.org/spower/<br /><br />SeqSIMLA <br />SeqSIMLA can simulate sequence data with user-specified disease and quantitative trait models. Family or unrelated case-control data can be simulated. <br />http://seqsimla.sourceforge.net/<br /><br />Serial NetEvolve <br />A flexible utility for generating serially-sampled sequences along a tree or recombinant network <br />http://biorg.cis.fiu.edu/sne/<br /><br />SFS_CODE <br />SFS_CODE can perform forward population genetic simulations under a general Wright-Fisher model with arbitrary migration, demographic, selective, and mutational effects. <br />http://sfscode.sourceforge.net/sfs_code/index/index.html<br /><br />SIBSIM <br />Quantitative phenotype simulation in extended pedigrees <br />http://sourceforge.net/projects/sibsim/<br /><br />SIMCOAL2 <br />A coalescent program for the simulation of complex recombination patterns over large genomic regions under various demographic models <br />http://cmpg.unibe.ch/software/simcoal2/<br /><br />SimCopy <br />An R package simulating the evolution of copy number profiles along a tree. <br />http://bit.ly/simcopy<br /><br />SIMLA <br />SIMLA is a SIMuLAtion program that generates data sets of families for use in Linkage and Association studies. <br />http://www.chg.duke.edu/research/simla.html<br /><br />SimPed <br />A Simulation Program to Generate Haplotype and Genotype Data for Pedigree Structures <br />http://www.hgsc.bcm.tmc.edu/content/simped<br /><br />Simprot <br />A program to simulate protein evolution by substitution, insertion and deletion <br />http://www.uhnresearch.ca/labs/tillier/software.htm#3<br /><br />SimRare <br />Rare variant simulation and analysis tool <br />http://code.google.com/p/simrare/<br /><br />simuGWAS <br />A forward-time simulator that simulates realistic samples for genome-wide association studies. <br />http://simupop.sourceforge.net/cookbook/simucomplexdisease<br /><br />simuPOP <br />simuPOP is a general-purpose individual-based forward-time population genetics simulation environment. <br />http://simupop.sourceforge.net/<br /><br />SISSI <br />A software tool to generate data of related sequences along a given phylogeny, taking into account user defined system of neighbourhoods and instantaneous rate matrices. <br />http://www.cibiv.at/software/sissi/<br /><br />SNPsim <br />Coalescent simulation of hotspot recombination <br />http://code.google.com/p/phylosoftware/<br /><br />SPIP <br />SPIP simulates the transmission of genes from parents to offspring in a population having demographic structure defined by the user <br />http://swfsc.noaa.gov/textblock.aspx?division=fed&amp;id=3434<br /><br />Splatche <br />Spatial and Temporal Coalescences in Heterogeneous Environment <br />http://www.splatche.com/<br /><br />srv <br />Simulator of Rare Varaints (srv) is a simulator for the simulation of the introduction and evolution of (rare) genetic variants. <br />http://simupop.sourceforge.net/cookbook/simurarevariants<br /><br />SUP <br />SLINK/FastSLINK utility program <br />http://mlemire.freeshell.org/software.html<br /><br />TreesimJ <br />A flexible, forward-time population genetic simulator <br />http://code.google.com/p/treesimj/<br /><br />Vortex <br />VORTEX is an individual-based simulation model for population viability analysis (PVA). <br />http://www.vortex9.org/vortex.html<br /><br />References:</p><p>Image www.evolution-of-life.com</p><p>www.cancer.gov</p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/26617/list-of-bioinformatics-software-tools-for-next-generation-sequencing</guid>
	<pubDate>Fri, 11 Mar 2016 20:22:14 -0600</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/26617/list-of-bioinformatics-software-tools-for-next-generation-sequencing</link>
	<title><![CDATA[List of Bioinformatics Software Tools for Next Generation Sequencing]]></title>
	<description><![CDATA[<p><strong>Commercial tools</strong></p><ol>
<li><strong><a href="http://www.strand-ngs.com/">Strand NGS</a></strong>
<ul>
<li>offers many different tools including alignment, RNA-Seq, DNA-Seq, ChIP-Seq, Small RNA-Seq, Genome Browser, visualizations, Biological Interpretation, etc. Supports workflows &ldquo;one can import the sample data in FASTA, FASTQ or tag-count format. In addition, prealigned data in SAM, BAM or Illumina-specific ELAND format can be directly imported for analysis.&rdquo;</li>
<li>Alignment feature: Supports alignment from Illumina, Ion Torrent, 454 (Roche), and Pac Bio</li>
<li>DNA-Seq Feature, can annotate with dbSNP</li>
</ul>
</li>
<li><strong><a href="http://www.clcbio.com/desktop-applications/top-features/">CLC Genomics Workbench</a></strong><br />
<ul>
<li>(QIAGEN). Features include: resequencing, workflow, read mapping, de novo assembly, variant detection, RNA-Seq, ChIP-Seq, Genome Browser, etc (entire list on website); Main Workbench offers database search (Genbank, Blast, Pubmed); 2000 organizations have invested in CLC</li>
<li>Accepts VCF files from 1000 Genomes Project</li>
<li>Accepts downloaded tracks from dbSNP</li>
<li>Also accepts: FASTA, GFF/GTF/GVF, BED, Wiggle, Cosmic, UCSC variant database, complete genomics master var file</li>
<li>Read mapping: &ldquo;In addition to Sanger sequence data, reads from these high-throughput sequencing machines are supported: The 454 FLX System and the 454 GS Junior System from Roche, Illumina Genome Analyzer, Illumina HiSeq, Illumina HiScan, and Illumina MiSeq sequencing systems, SOLiD system from Life Technologies, Ion Torrent system from Life Technologies, Helicos from Helicos BioSciences&rdquo;</li>
<li>De novo assembly: &ldquo;In addition to Sanger sequence data, reads from these high-throughput sequencing machines are supported The 454 FLX System and the 454 GS Junior System from Roche, Illumina Genome Analyzer, Illumina HiSeq, Illumina HiScan, and Illumina MiSeq sequencing systems, SOLiD system from Life Technologies, Ion Torrent system from Life Technologies&rdquo;</li>
<li>Annotation tracks from Ensembl</li>
</ul>
</li>
<li><strong><a href="https://www.dnanexus.com/product-overview">DNAnexus</a></strong>
<ul>
<li>Private cloud repository -- formerly a redistributor of SRA and other NCBI resources; command-line or via web, can fetch data from a URL, build custom pipeline/ workflow has sra.dnanexus.com site: data downloads come directly from NCBI</li>
</ul>
</li>
<li><strong><a href="http://www.ingenuity.com/products/variant-analysis">Ingenuity Variant Analysis</a></strong>
<ul>
<li>(QIAGEN) allows for variant identification and analysis, uses NCI-60 data set for cancer, Supported third part informatin: Entrez Gene, RefSeq, ClinVar; gives contextual details of results instead of just A to B relationship</li>
<li>Has own database-- &ldquo;knowledge base&rdquo; based on COSMIC, OMIM, and TCGA databases</li>
</ul>
</li>
<li><strong><a href="http://www.dnastar.com/t-products-dnastar-lasergene-genomics.aspx">Lasergene Genomics Suite</a></strong>
<ul>
<li>Comprehensive NGS software pipeline for assembly, alignment, variant calling and analysis of NGS data</li>
<li>Supported workflows include: reference-guided and de novo genome and transcriptome assembly and analysis, metagenomics sample assembly, targeted resequencing, exome alignment, gene panels with validation control, variant analysis, and RNA-Seq, ChIP-Seq and miRNA alignment and analysis.</li>
<li>#1 in accuracy: fewer false negatives and better sensitivity compared to results obtained from other aligners</li>
<li>Aligns exome data and performs variant calling an average of 3 times faster than alternative pipelines</li>
<li>Annotates genomic data with allele and genotype frequency, functional impact predictions, evolutionary conservation scores and pathogenicity</li>
<li>Supports all major NGS technologies (Illumina, Ion Torrent, Pac Bio and Roche 454) and project types</li>
<li>Available on Windows, Mac OS X, Linux, and the Amazon Cloud</li>
</ul>
</li>
<li><strong><a href="http://www.softgenetics.com/NextGENe.html">NextGENe</a></strong>
<ul>
<li>&ldquo;perfect analytical partner for the analysis of desktop sequencing data produced by the ION PGM&trade;, Roche Junior, Illumina MiSeq as well as high throughput systems as the Ion Torrent Proton, Roche FLX, Applied BioSystems SOLiD&trade; and Illumina&reg; platforms.&rdquo; runs on Windows, free-standing multi-application package-- SNP/Indel analysis, CNV prediction and disease discovery, whole genome alignment, etc.</li>
<li>Data can be imported from Clinvar, dbSNP, Genbank:<a href="http://www.softgenetics.com/PDF/NextGene_UsersManual_web.pdf">http://www.softgenetics.com/PDF/NextGene_UsersManual_web.pdf</a></li>
</ul>
</li>
<li><strong><a href="http://www.partek.com/pgs">Partek Genomics Suite</a></strong>
<ul>
<li>Cited in over 3,500 peer-reviewed scientific publications</li>
<li>Workflows for microarray and PCR data include: Gene expression including alternative splicing, miRNA expression, Genome Wide Association Studies, Mother-Father-Child Trio analysis, DNA Copy number including allele specific copy number and Loss of Heterozygosity (LOH), and ChIP, and methylation. Next Generation Sequencing (NGS) workflows include: RNA-Seq, miRNA-Seq, ChIP-Seq, DNA-Seq, and Methylation</li>
<li>Powerful statistics and interactive, publication ready visualizations</li>
<li>Supports all commercial next generation sequencing and microarray file format as well as text files</li>
<li>Can input GEO SOFT files</li>
</ul>
</li>
<li><strong><a href="http://www.partek.com/partekflow">Partek Flow</a></strong>
<ul>
<li>Installation can be cloud-based or on a local cluster or Linux server</li>
<li>Easy to use point-and-click interface</li>
<li>Takes NGS data (.fastq, BAM, SAM), microarrays (Affymetrix, Illumina) and text files</li>
<li>Supports custom genome builds and annotation databases</li>
<li>Performs base trimming, alignment, quantification, quality analysis, statistics, and visualization</li>
<li>Includes ten fully customizable aligners (Bowtie, Bowtie 2, BWA, GSNAP, Isaac 2, SHRiMP 2, STAR, TMAP, TopHat and TopHat 2)</li>
<li>Applications for RNA-Seq, Small RNA-Seq, WGS/WES, Pathway enrichment, Fusion detection and Variant calling</li>
<li>Allows users to create, save, share, or download analysis pipelines for automated and repeatable analysis</li>
<li>Collaborate with others without transferring data</li>
<li>Integrates microarray and next generation sequencing data</li>
</ul>
</li>
<li><strong><a href="http://goldenhelix.com/SNP_Variation/">Golden Helix: SNP and Variation Suite</a></strong>
<ul>
<li>used for managing, analyzing and visualizing genotypic and phenotypic data; Features: Genome-wide association studies, genomic prediction, copy number analysis, small sample DNA-Seq workflows, large sample DNA-seq analysis, RNA-seq analysis. Supported files: .txt, excel XLS &amp; XLSX, CEL, CHP, CNT, Illumina, Plink PED, TPED, BED, Agilent files, NimbleGen data summary files, VCF files, Impute2 GWAS files, HapMap format, MACH output, + 50 other formats consumes NCBI data directly</li>
</ul>
</li>
<li><strong><a href="https://www.genomatix.de/">Genomatix</a></strong>
<ul>
<li>Applications: ChIP-Seq, DNA-Seq, RNA-Seq, DNA methylation; enable personalized medicine,</li>
<li>Mining Stations: Supports all established NGS sequencing platforms- SOLiD, 454 Life Sciences, Genome Analyzer, HiSeq, MiSeq, IonTorrent</li>
<li>Software Suite: can upload sequence of BED files</li>
<li>Genome browser: BED and BAM files, Public data- 1500 BED files available for every user</li>
</ul>
</li>
<li><strong><a href="http://www.biodatomics.com/">Biodatomics</a></strong>
<ul>
<li>Open source platform (SaaS), analysis and genome sequencing tools, integrates over 400 genomic analysis open source tools and pipelines, have a private and public cloud version. Features: genomic data visualization, drag and drop interface, accelerated analysis, real-time collaboration</li>
<li>They have a couple modules to do so, and have enabled parts of the sra toolkit</li>
</ul>
</li>
<li><strong><a href="https://www.solvebio.com/">SolveBio</a></strong>
<ul>
<li>Software product, for clinical genomics professionals, manage, curate, report genomic variation</li>
<li>Has own data library -- data from NCBI</li>
</ul>
</li>
<li><strong><a href="http://www.basepairtech.com">Basepair</a></strong>
<ul>
<li>Offers high quality workflows for all common NGS applications (RNA-Seq, ChIP-Seq, DNA-Seq, etc.)</li>
<li>Very fast - get all results in a 1-2 hours. Cloud-based, no storage or computing limits.</li>
<li>Easy to use - less than a minute to run an analysis</li>
<li>REST and Python API to mange large projects.</li>
</ul>
<div>&nbsp;</div>
</li>
</ol><h2><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#variant-identification"></a>Variant Identification</h2><h3><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#germline-callers"></a>Germline Callers</h3><ol>
<li><strong><a href="http://mathgen.stats.ox.ac.uk/impute/impute_v2.html">IMPUTE2</a></strong>
<ul>
<li>Description: phasing observed genotypes and imputing missing genotypes uses reference panels to provide all available halotypes, does not use population labels or genome-wide measures; designed to represent variation in one population; Fairly popular</li>
<li>Input:</li>
<li>Reference Haplotypes: Links to 1000 Genomes and HapMap downloads</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/ekg/freebayes">FreeBayes</a></strong>
<ul>
<li>Description: finds SNPs, Indels, MNPs; reports variants based on alignment; haplotype based</li>
<li>Input: BAM- uses BAMtools API to parse</li>
<li>Reference genome: FASTA</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://soap.genomics.org.cn/soapindel.html">SOAPindel</a></strong>
<ul>
<li>Description: detects indels from NGS paired-end sequencing</li>
<li>Input: files with read alignment can be SOAP or SAM formats, users must also give raw reads in Fasta or Fastq</li>
<li>Reference Sequence used to align reads: FASTA</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/danmaclean/2kplus2">2Kplus2</a></strong>
<ul>
<li>Description: algorithm searches graphs produced by de novo assembler Cortex; c++ source code for SNP detection &ldquo;2kplus2.cpp is a c++ source code for the detection and the classification of single nucleotide polymorphisms in transformed De Bruijn graphs using Cortex assembler.&rdquo;</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://www.hgsc.bcm.edu/software/atlas-2">Atlas 2</a></strong>
<ul>
<li>Description: specializes in separation of true SNPs and indels from sequencing and mapping errors, last update January 2013</li>
<li>Input: takes BAM file,</li>
<li>Reference Genome: FASTA</li>
<li>Output: produces VCF</li>
</ul>
</li>
<li><strong><a href="https://sites.google.com/site/vibansal/software/crisp">CRISP</a></strong>
<ul>
<li>Description: identifies SNPs and INDELs from pooled high-throughput NGS, not used for analysis of single samples; implemented in C and uses SAMtools API; latest version should work with diploid genomes</li>
<li>Input: requires BAM files (aligned with GATK)</li>
<li>Reference Genome: indexed FASTA file</li>
<li>Output: VCF files</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/dindel/">Dindel</a></strong>
<ul>
<li>Description: (Wellcome Trust Sanger) calls small indels from short-read sequences, only can handle Illumina data; cannot test candidate indels; written in C++, used on Linux based and Mac computers (not tested in windows)</li>
<li>Input: BAM files</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://colibread.inria.fr/software/discosnp/">discoSnp++</a></strong>
<ul>
<li>Description: detects homozygous and heterozygous SNPs and Indels; software composed of 2 modules (kissnp2 and kissreads)</li>
<li>Input: raw NGS datasets; fasta, fastq, gzipped or not;</li>
<li>no reference genome required; read pairs can be given</li>
<li>Output: FASTA</li>
</ul>
</li>
<li><strong><a href="http://odin.mdacc.tmc.edu/~wwang7/FamSeqIndex.html">FamSeq</a></strong>
<ul>
<li>Description: family-based sequencing studies- provides probability of an individual carrying variant based on family&rsquo;s raw measurements; accommodates de novo mutations, can perform variant calling at chrX;</li>
<li>Input: VCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/p/%20geneticthesaurus/wiki/Example/">GeneticThesaurus</a></strong>
<ul>
<li>Description: &ldquo;Annotation of genetic variants in repetitive regions&rdquo;</li>
<li>Input: Initial variant calling from bam &rarr; vcf output</li>
<li>Reference Genome: need to provide own fasta file for hg19 genome,</li>
<li>Output: vcf.gz, vtf.gz, and baf.tsv.gz output</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/GlfMultiples">glfMultiples</a></strong>
<ul>
<li>Description: command-line, variant caller</li>
<li>Input: GLF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/GlfSingle">glfSingle</a></strong>
<ul>
<li>Description: uses likelihood-based model for variant calling, starts from genotype likelihoods that have been computed from other tools (ex. Samtools BAQ), the likelihoods combine with individual-based prior p(genotype) to generate posterior probabilities</li>
<li>Input: GLF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://github.com/ddcap/halvade">Halvade</a></strong>
<ul>
<li>Description: command-line; written in Java, &ldquo;to run halvade a reference is needed for both GATK and BWA and a SNP (dbSNP!) database is required</li>
<li>Input: FASTQ</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://github.com/aakrosh/indelMINER">indelMINER</a></strong>
<ul>
<li>Description: identifies indels from paired-end reads</li>
<li>Input: BAM (aligned in SAMtools API)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://www.broadinstitute.org/cancer/cga/indelocator">Indelocator</a></strong>
<ul>
<li>Description: (Broad Institute): does not perform realignment, relies on alignments in BAM files (BAM files need aligned before put into indelocator); recommended to use GATK prior;</li>
<li>Input: 2 BAM files(tumor &amp; normal), annotated as germline or somatic; also has single sample mode</li>
<li>Output: &ldquo;Output of Indelocator is a high-sensitivity list of putative indel events containing large numbers of false positives. The statistics reported for each event have to be used to custom-filter the list in order to lower false positive rate&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/sequencing/isaac_variant_caller">Isaac Variant Caller</a></strong>
<ul>
<li>Description: detects SNPs and small indels from diploid sample; designed to run on &ldquo;nux-like platforms&rdquo;</li>
<li>Input: BAM</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.swisstph.ch/kvarq">KvarQ</a></strong>
<ul>
<li>Description: in silico genotyping for selected loci in bacterial genome, written in Python and C</li>
<li>Input: FASTQ</li>
<li>reference genome or de novo assembly not needed</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/lofreq/files/">LoFreq</a></strong>
<ul>
<li>Description: SNV caller, Python language, standalone program, uncovers cell-population heterogeneity from high-throughput sequencing datasets; calls variants found in &lt;.05% of the population</li>
<li>Input: BAM file input&rarr; suggest running through GATK</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/Illumina/manta">Manta</a></strong>
<ul>
<li>Description: Calls indels and SVs from paired end reads; standalone, command line program; Written in C++ and Python</li>
<li>Input: BAM (can tolerate non-paired-end reads); a matched tumor sample may be provided as well</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://github.com/benedictpaten/marginAlign">MarginAlign</a></strong>
<ul>
<li>Description: SNV caller, specifically tailored to Oxford Nanopore Reads, written in Python; Package comes with 3 programs, marginAlign, marginCaller (calls SNVs), marginStats (computes qc stats on sam files)</li>
<li>Input: SAM</li>
<li>Output: SAM</li>
</ul>
</li>
<li><strong><a href="http://gmt.genome.wustl.edu/packages/mendelscan/">MendelScan</a></strong>
<ul>
<li>Description: Last release March 2014; for analyzing sequencing data in family studies of inherited diseases; variant calls for a family in VCF file; still in alpha-testing on github, example data uses 1000 genomes dataset</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/mitenjain/nanopore">nanopore</a></strong>
<ul>
<li>Description: UCSC Nanopore group (group at UCSC studying using ion channels for analysis of single RNA/DNA structures) software pipeline; tailored to Oxford Nanopore Reads; command line program</li>
<li>Input: FASTQ</li>
<li>Reference files: FASTA</li>
<li>Output: &ldquo;For each possible pair of read file, reference genome and mapping algorithm an experiment directory will be created in the nanopore/output directory.&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://omictools.com/platypus-s1989.html">Platypus</a></strong>
<ul>
<li>Description: Package program, written in C, Python, Cython; Can identify SNPs, MNPs, short indels, and larger variants; has been tested on very large datasets (1000 genomes)</li>
<li>Input: BAM</li>
<li>Reference Genome: FASTA (files must be indexed using Samtools or similar program</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.bioinformatics.nl/QualitySNPng/">QualitySNPng</a></strong>
<ul>
<li>Description: detection of SNPs; &ldquo;can be used as a standalone application with graphical user interface as part of pipeline system&rdquo;; does not require fully sequenced reference genome; haplotype strategy</li>
<li>Input:SAM, ACE</li>
<li>Output: GUI</li>
</ul>
</li>
<li><strong><a href="http://revister.sourceforge.net/">ReviSTER</a></strong>
<ul>
<li>Description: command line program; automated pipeline; utilizes BWA, BLAT, and SAMTools; utilizes BWA mapping program;</li>
<li>Input: FASTQ,</li>
<li>Reference sequence file and list file containing STR locations as inputs</li>
<li>Output: SAM</li>
</ul>
</li>
<li><strong><a href="http://dna-discovery.stanford.edu/software/rvd/">RVD</a></strong>
<ul>
<li>Description: command-line program, detection of rare SNVs, relies upon Samtools, can be run in MATLAB</li>
<li>Input: BAM</li>
<li>Reference Genome: FASTA</li>
<li>Output: &ldquo;The algorithm output is a call table -- a comma-separated file with one line for each base position and each line in the following format:</li>
<li>AlginmentReferencePosition, AlignmentBase, Call ,SecondBase, CenteredErrorPrc, ReferenceErrorPrc, SecondBasePrc&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://snver.sourceforge.net/">SNVer</a></strong>
<ul>
<li>Description: calls common and rare variants in pool or individual NGS data, reports overall p-value, operating system independent statistical tool, identifies SNPs and INDELs, written in Java, no dependencies, straightforward command-line</li>
<li>(SNVerGUI=GUI version) --SNVerGUI: desktop tool for variant detection</li>
<li>Input: chrX annotation, sam.zip, bam.zip</li>
<li>reference file must be aligned to the data file</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://compbio.bccrc.ca/software/snvmix/">SNVMix</a></strong>
<ul>
<li>Description: detects SNVs from NGS, post-alignment tool</li>
<li>Input: pileupformat (Maq or Samtools)</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.bsse.ethz.ch/mlcb/research/bioinformatics-and-computational-biology/structural-variant-machine--sv-m-.html">SV-M</a></strong>
<ul>
<li>Description: Structural Variant Machine - predicts indels, uses split read alignment profiles, validated by Sanger Sequencng</li>
<li>Input:paired-end Illumina reads from 1001 genomes project (uses ref plant- 1001genomes.org)</li>
<li>Ouptut:</li>
</ul>
</li>
<li><strong><a href="https://github.com/slindgreen/SNPest">SNPest</a></strong>
<ul>
<li>Description: Standalone program, language C++, Perl</li>
<li>Input: mpileup (SAMtools)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/TrioCaller">TrioCaller</a></strong>
<ul>
<li>Description:Command line program, relies on BWA and samtools; genotype calling for unrelated individuals and parent-offspring trios</li>
<li>Input: BAM (that has been aligned in BWA and Samtools</li>
<li>Output: BCF that can be formatted to VCF using bcftools</li>
</ul>
</li>
<li><strong><a href="http://www.vicbioinformatics.com/software.snippy.shtml">Snippy</a></strong>
<ul>
<li>Description: finds indels between haploid reference genome and NGS sequence reads</li>
<li>Input:read files- FASTQ or FASTA (can be .gz compressed), output- .aln, .tab, .txt</li>
<li>Reference genome in FASTA or GENBANK</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://orca.bu.edu/vntrseek/">VntrSeek</a></strong>
<ul>
<li>Description: pipeline for discovering microsatellite tandem repeats with high-throughput sequencing data</li>
<li>Input: gzip-compressed FASTA or FASTQ</li>
<li>Output: VCF files; one for TRs and observed alleles, another file contains link to viewer</li>
</ul>
</li>
</ol><h3><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#somatic-callers"></a>Somatic Callers</h3><ol>
<li><strong><a href="http://cakesomatic.sourceforge.net/">Cake</a></strong>
<ul>
<li>Description: standalone program, &ldquo;pipeline for the integrated analysis of somatic variants in cancer genomes&rdquo;; integrates four algorithms; written in Perl; required tools: samtools, tabix, vcftools, VarScan2, bambino, cmake, somaticsniper (User guide; workflow page)</li>
<li>Input: tumor and normal reads in BAM files, run through variant calling programs to generate intermediate VCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/cancer/cga/mutect">MuTect</a></strong>
<ul>
<li>Description: Broad Institute, identification of somatic point mutations in cancer genomes; requires preprocessing of reads (GATK)</li>
<li>Input: same as GATK (FASTA reference genome, SAM read files)</li>
<li>Output: call-stats, VCF, wiggle files</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/Polymutt">Polymutt</a></strong>
<ul>
<li>Description: calls SNVs and detects de novo point mutations in families</li>
<li>Input: GLF or BAM or VCF (must have identical chromosome orders)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://tvap.genome.wustl.edu/tools/bassovac/">Bassovac</a></strong>
<ul>
<li>Description: Improved Bayesian inversion somatic caller; unlike other software packages, treats effects fully probabilisticallys instead of using ad-hoc modeling; effects are integrated at the atomic level and standard probability theory integrates read tallies to the sample level and to the tumor-normal pair level; "pending public release"</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://bioinformatics.ustc.edu.cn/CLImAT/">CLImAT</a></strong>
<ul>
<li>Description: standalone program; &ldquo;accurate detection of copy number alteration and loss of heterozygosity in impure and aneuploid tumor samples using whole genome sequencing data&rdquo;</li>
<li>Input: depth file generated by DFExtract and a config file</li>
<li>Output: .results file, .Gtype, LOG.txt, also generates visualization</li>
</ul>
</li>
<li><strong><a href="http://denovogear.sourceforge.net/">DeNovoGear</a></strong>
<ul>
<li>Description: de-novo variant calling and interpretation; standalone program; dependencies C++ compiler, CMake, HTSlib, Eigen, Boost</li>
<li>Input: PED and BCF</li>
<li>Output: &ldquo;The output format is a single row for each putative de novo mutation (DNM), with the following fields&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/friend1ws/EBCall">EBCall</a></strong>
<ul>
<li>Description: Empirical Baysian Mutation Calling; standalone program; uses tumor/normal paired reads and non-paired normal reference samples; dependent on samtools, R and VGAM pack for R</li>
<li>Input: BAM</li>
<li>Output: not sure what exact type of file- &ldquo;The format of the result is suitable for adding annotation by annovar.&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/usuyama/hapmuc">HapMuc</a></strong>
<ul>
<li>Description: standalone program; &ldquo;utilizes the information of heterozygous germline variants near candidate mutations&rdquo;; Dependent upon- Boost, SAMtools, BEDtools; 3 step workflow</li>
<li>Input: BAM</li>
<li>Output: BED</li>
</ul>
</li>
<li><strong><a href="https://github.com/cui-lab/multigems">MultiGeMS</a></strong>
<ul>
<li>Description: Multi-sample Genotype Model Selection</li>
<li>Input: .txt, pileup (SAM/BAM converted to pileup format)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://bitbucket.org/joseph07/multisnv/wiki/Home">MultiSNV</a></strong>
<ul>
<li>Description: command-line program; calls SNVs from NGS data from multiple samples from the same patient; dependent on R, Git, cmake, Boost and compile libraries</li>
<li>Input: BAM or pileup</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://compbio.bccrc.ca/software/mutationseq/">MutationSeq</a></strong>
<ul>
<li>Description: standalone program, somatic SNV detection in tumor/normal samples; dependent on python, bamtools, boost, and LAPACK</li>
<li>Input: BAM</li>
<li>Output: VCF4.1 consisting of two parts (meta information &amp; data lines)</li>
</ul>
</li>
<li><strong><a href="http://www.qcmg.org/bioinformatics/tiki-index.php">qSNP</a></strong>
<ul>
<li>Description: standalone program; SNV caller for somatic variants in &ldquo;low cellularity cancer samples&rdquo;</li>
<li>Input: BAM, dbSNP data, Illumina data, chrConv</li>
<li>Output: &ldquo;qSNP output files are named using a 4-element pattern: ...&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/aradenbaugh/radia/">RADIA</a></strong>
<ul>
<li>Description: RNA and DNA Integrated Analysis for Somatic Mutation Detection; DNA only Method(tumor/normal pair, ignores RNA) or Triple BAM Method (uses all three datasets from same patient); dependent upon python, samtoools, pysam API, BLAT, SnpEff</li>
<li>Input: BAM</li>
<li>Reference Genome: FASTA indexed with SAMtools faidx</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genomics.wpi.edu/rvd2/">RVD2</a></strong>
<ul>
<li>Description: sensitive, variant detection for low-depth targeted NGS data; python module or command- line program;</li>
<li>Input: tab- deliminted depth chart format (converted from pileup files)</li>
<li>Output: three hdf5 files and a vcf file</li>
</ul>
</li>
<li><strong><a href="https://github.com/nhansen/Shimmer">Shimmer</a></strong>
<ul>
<li>Description: standalone program; detects somatic SNVs with multiple testing correction, uses Fisher&rsquo;s exact test; dependent on git, samtools, R, R statmod package; for tumor/normal matched samples</li>
<li>Input: BAM</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.cs.helsinki.fi/en/gsa/snv-ppilp/">SNV-PPILP</a></strong>
<ul>
<li>Description: Refines GATK&rsquo;s Unified Genotyper SNV calls for &ldquo;multiple samples assumed to form a phylogeny&rdquo;</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://gmt.genome.wustl.edu/packages/somatic-sniper/">SomaticSniper</a></strong>
<ul>
<li>Description: command-line application to identify SNPs between tumor/normal pairs- predicts probability of difference between two</li>
<li>Input: BAM</li>
<li>Reference Genome in FASTA</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://sites.google.com/site/strelkasomaticvariantcaller/">Strelka</a></strong>
<ul>
<li>Description: somatic variant calling workflow for matched tumor-normal samples; detects indels; runs on *nux-like platform</li>
<li>Input: BAM (must be sorted and indexed)- Strelka does own realignment around indels-- don&rsquo;t need to do this type of pre-processing</li>
<li>Output: pair of VCF files</li>
</ul>
</li>
<li><strong><a href="http://www.pitt.edu/~wec47/triodenovo.html">Triodenovo</a></strong>
<ul>
<li>Description: Bayesian framework for calling de novo mutations in trios</li>
<li>Input: VCF file with PL or GL fields (recommend using GATK or samtools to generate)</li>
<li>Output: out_vcf</li>
</ul>
</li>
<li><strong><a href="http://lbg.med.unc.edu/~mwilkers/unceqr_dist/">UNCeqr</a></strong>
<ul>
<li>Description: finds somatic mutations using integration of DNA and RNA seq data-- boosts sensitivity for low purity tumors and rare mutations;</li>
<li>Input:&rdquo;can accept a variety of sequencing inputs and configurations&rdquo;</li>
<li>Output: &ldquo;table of somatically mutated sites and associated information. These somatic mutations can be annotated with predicted transcript and protein effects using third party tools, such as Annovar&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/virmid/">Virmid</a></strong>
<ul>
<li>Description: Virtual Microdissection for SNP calling; Java based; for disease-control matched samples; uncovers SNPs with low allele frequency by considering alpha contamination</li>
<li>Input: BAM (must be sorted and indexed- samtools sort)</li>
<li>Output: VCF and report file</li>
</ul>
</li>
</ol><h3><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#germline--somatic--callers"></a>Germline + Somatic Callers</h3><ol>
<li><strong><a href="http://massgenomics.org/varscan">VarScan 2</a></strong>
<ul>
<li>Description: identify germline variants, private and shared variants, somatic mutations, and somatic CNVs; detects indels</li>
<li>Input: SAMtools pileup</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genformatic.com/baysic/">BAYSIC</a></strong>
<ul>
<li>Description: Bayesian method; combines variant calls from different methods (GATK, FreeBayes, Atlas, Samtools, etc)</li>
<li>Input: VCF format from one or more variant calling programs</li>
<li>Output: VCF file containing integrated set of variant calls</li>
</ul>
</li>
<li><strong><a href="https://github.com/ding-lab/msisensor">MSIsensor</a></strong>
<ul>
<li>Description: Microsatellite instability detection; C++ program, detects somatic and germline variants in tumor-normal paired data</li>
<li>Input: BAM index files (normal and tumor)</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://faculty.washington.edu/browning/beagle/beagle.html">Beagle version 4</a></strong>
<ul>
<li>Description: software package: genotype calling, phasing, imputation of ungenotyped markers, and identity-by-descent segment detection:unsure if this one is in the right category; genotype calling, phasing, imputation of ungenotyped markers, and identity-by-descent segment detection;</li>
<li>Input: VCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.iro.umontreal.ca/~csuros/quadgt/">QuadGT</a></strong>
<ul>
<li>Description: software package, SNV calling from normal-tumor pair and two parent genomes; quantifies descent-by-modification relationships; Written in Java</li>
<li>Input: BAM files (parsed by Picard/Samtools API)</li>
<li>Reference Genome; FASTA</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/rarevator/">RAREVATOR</a></strong>
<ul>
<li>Description: RAre REference VAriant annotaTOR; command line; &ldquo;identification and annotation of germline and somatic variants in rare reference allele loci from second generation sequencing data&rdquo;; Bayesian genotype likelihood model</li>
<li>Input: BED or VCF files from GATK</li>
<li>Output: two VCF files (one for SNVs, one for Indels)</li>
</ul>
</li>
<li><strong><a href="http://scalpel.sourceforge.net/">Scalpel</a></strong>
<ul>
<li>Description: Used for detecting indels in a reference genome; performs localized micro-assembly of specific regions of interest; can do single, de novo, somatic reads; requires that raw reads are aligned with BWA</li>
<li>Input: BAM</li>
<li>Output: either VCF or ANNOVAR</li>
</ul>
</li>
<li><strong><a href="http://soap.genomics.org.cn/soapsnp.html">SOAPsnp</a></strong>
<ul>
<li>Description: based on Baye&rsquo;s theorem; calls consensus genotype</li>
<li>Input:SOAP short read alignment results</li>
<li>Output: GLF, option of flat tabular format</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/variantmaster/">VariantMaster</a></strong>
<ul>
<li>Description: &ldquo;extract causative variants for monogenic and sporadic genetic diseases&rdquo;; uses ANNOVAR;</li>
<li>Input: BAM or VCF files (from SAMtools, GATK)</li>
<li>Output:</li>
</ul>
</li>
</ol><h2><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#downstream-analysis-of-variants"></a>Downstream Analysis of Variants</h2><ol>
<li><strong><a href="https://github.com/hakyimlab/PrediXcan%20https://github.com/hriordan/PrediXcan/">PrediXcan</a></strong>
<ul>
<li>Description: command-line, standalone package program; available in Perl, Python, and R versions; predicts liklihood of a gene being related to a certain phenotype- &ldquo;that directly tests the molecular mechanisms through which genetic variation affects phenotype.&rdquo;; no actual expression data used, only in silico expression; &ldquo;PrediXcan can detect known and novel genes associated with disease traits and provide insights into the mechanism of these associations.&rdquo;</li>
<li>Input: genotype and phenotype file (doesn&rsquo;t specify file type)</li>
<li>Output:default values: genelist, dosages (file format: snpid rsid) , dosage_prefix, weights, output</li>
</ul>
</li>
<li><strong><a href="http://ritchielab.psu.edu/software/athena-downloads">ATHENA</a></strong>
<ul>
<li>Description: Analysis Tool for Heritable and Environmental Network Associations; software package, combines machine learning model with biology and statistics to predict non-linear interactions</li>
<li>Input: Configuration file, Data file, Map file (includes rsID)</li>
<li>Output: Summary file, Best model file, dot file, individual score file, cross-validation file</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/rarevariant/#t_2">CCRaVAT and QuTie</a></strong>
<ul>
<li>Description: (Wellcome Trust Sanger) Case-Control Rare Variant Analysis Tool and Quantitative Trait; software packages for large-scale analysis of rare variants</li>
<li>Input: PED file and MAP file</li>
<li>Output: Five tab-delimited txt files</li>
</ul>
</li>
<li><strong><a href="http://cnsgenomics.com/software/gcta/">GCTA</a></strong>
<ul>
<li>Description: Genome Wide Complex Trait Analysis; package program, command line interface; estimates variance by all SNPs; 5 main functions: &ldquo;data management, estimation of the genetic relationships from SNPs, mixed linear model analysis of variance explained by the SNPs, estimation of the linkage disequilibrium structure, and GWAS simulation&rdquo;</li>
<li>Input: PLINK binary PED files, MACH output format</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://genomecomb.sourceforge.net/">GenomeComb</a></strong>
<ul>
<li>Description: package for analysis of complete genome data; annotation using public data or custom tracks, automated primer desing for Sanger or Sequenom validation; &ldquo;The cg process_illumina command can be used to generate annotated multisample data starting from fastq files, using tools such as bwa for alignment and GATK and samtools for variant calling. Sequencing data can also be imported from Complete Genomics (cg_process_sample command), Real Time Genomics (cg_process_rtgsample command) and VariantCallFormat (VCF) variant files (vcf2sft command).&rdquo;</li>
<li>Input: Sequencing data from Complete Genomics, Illumina, SOLiD and VCF;</li>
<li>Output: standard file format used is a simple tab delimited file (.sft, .tsv)</li>
</ul>
</li>
<li><strong><a href="http://ancorr.eimb.ru/">Genome Track Analyzer</a></strong>
<ul>
<li>Description: compares genome tracks; allows user to compare DNA expression/binding;</li>
<li>Input: multiple: SGR/TXT, BED, BED6, GFF; if using prealigned sequence data- use MACS peak caller: BAM, BED, SAM, ELAND</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://animalgene.umn.edu/gvcblub">GVCBLUP</a></strong>
<ul>
<li>Description: animal gene mapping; &ldquo;genomic prediction and variance component estimation of additive and dominance effects&rdquo;; standalone program, command line interface, writting in C++ and Java</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.jurgott.org/linkage/homog.htm">HOMOG</a></strong>
<ul>
<li>Description: Analyzes heterogeneity with respect to single marker loci or known maps of markers; Carries out homogeneity test for alternative hypothesis &ldquo;Two family types, one with linkage betweeen a trait to a marker or map of markers, the other without linkage&rdquo;</li>
<li>Input: HOMOG.DAT - described on website</li>
<li>Output: HOMOG.OUT</li>
</ul>
</li>
<li><strong><a href="http://intersnp.meb.uni-bonn.de/">INTERSNP</a></strong>
<ul>
<li>Description: GWIA for case-control SNP and quantitative traits; selected for joint analysis using priori information; Provides linear regression framework, Pathway Association Analysis, Genome-wide Haplotype Analysis,</li>
<li>Input: PLINK input formats (ped/map, tped/tfam, bed/bim/fam) Compatible with SetID files</li>
<li>Gene reference file: Ensembl Release 75</li>
<li>Output: covariance matrix for regression models</li>
</ul>
</li>
<li><strong><a href="https://github.com/PMBio/mtSet">mtSet</a></strong>
<ul>
<li>Description: Currently only the standalone version available, but moving to LIMIX software suite; offers set tests- allows for testing between variants and traits; accounts for confounding factors ex. relatedness</li>
<li>Input: sample-to-sample genetic covariance matrix needs to be computed; multiple types of input; simulator requires input genotype and relatedness component;</li>
<li>Output: resdir (result file of analysis), outfile (test statistics and p-values), manhattan_plot (flag)</li>
</ul>
</li>
<li><strong><a href="http://dougspeed.com/multiblup/">MultiBLUP</a></strong>
<ul>
<li>Description: Package program, command line interface; constructs linear prediction models; Best Linear Unbiased Prediction; improves upon BLUP involving kinship matrices; options: pre-specified kinships, regional kinships, adaptive multiblups, LD weightings</li>
<li>Input: PLINK format</li>
<li>Output:.reml, .indi.blp</li>
</ul>
</li>
</ol><h2><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#variant-annotation"></a>Variant Annotation</h2><ol>
<li><strong><a href="http://annovar.openbioinformatics.org/en/latest/">ANNOVAR</a></strong>
<ul>
<li>Description: command-line tool, supports SNPs, INDELs, CNVs and block substitutions, provides wide variety of annotation techniques, depends upon multiple databases (each needing to be downloaded); annotates genetic variants; utilizes RefSeq, UCSC Genes, and the Ensembl gene annotation systems; can compare mutations detected in dpSNP or 1000 Genomes Project; Very popular *&ldquo;The final command run TABLE_ANNOVAR, using dbSNP version 138, 1000 Genomes Project 2014 Oct version, NIH-NHLBI 6500 exome database version 2 (referred to as esp6400siv2), dbNFSP version 2.6 (referred to as ljb26), dbSNP version 138 (referred to as snp138) databases and remove all temporary files, and generates the output file called myanno.hg19_multianno.txt&rdquo;</li>
<li>Input: VCF, ANNOVAR input format (simple text-based format); can convert other formats into ANNOVAR input format</li>
<li>Output: VCF (if input VCF), output file with multiple columns, tab-delimited output file</li>
</ul>
</li>
<li><strong><a href="http://wannovar.usc.edu/">wANNOVAR</a></strong>
<ul>
<li>provides web-based access to ANNOVAR software</li>
</ul>
</li>
<li><strong><a href="http://genetics.bwh.harvard.edu/pph2/">PolyPhen-2</a></strong>
<ul>
<li>Description: Very popular; Polymorphism Phenotyping; Web application; predicts impact of amino acid substitution on protein; Calculates Bayes posterior probability (Last update July 2015)</li>
<li>Input: FASTA</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://sift.jcvi.org/">SIFT</a></strong>
<ul>
<li>Description: predicts how an amino acid substitution will affect protein function; Based on degree of conservation of amino acid residues- collected though PSI-BLAST; can be applied to nonsynonymous polymorphisms or laboratory-induced missense mutations; links to dbSNP 132, GRCh37; Standalone or web app program; Very popular</li>
<li>Input: Uniprot ID or Accession, Go term ID, Function name, Species Name or ID, etc</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://snpeff.sourceforge.net/">snpEff</a></strong>
<ul>
<li>Description: Genetic variant annotation and effect prediction toolbox; integrated with Galaxy, GATK, and GNKO; can annotate SNPs, INDELs, and multiple-nucleotide polymorphisms; categorizes effects into classes by functionality; Very popular; Standalone or Web app; Claims to calculate all SNPs in 1000 genomes (EMBI) in less than 15 minutes; can annotate SNPs, MNPs, and insertions and deletions; Provides assessment of impact of the variant ( low, medium or high)</li>
<li>Input: VCF, BED</li>
<li>Output: VCF (with new ANN field, also used in ANNOVAR and VEP), HTML summary files</li>
</ul>
</li>
<li><strong><a href="http://snpeff.sourceforge.net/SnpSift.html">SnpSIFT</a></strong>
<ul>
<li>Description: Filter and manipulate annotated files; Part of SnpEff main distribution; one variants have been annotated, this can be used to filter your data to find relevant variants</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.yandell-lab.org/software/vaast.html">VAAST 2</a></strong>
<ul>
<li>Description: Variant Annotation, Analysis, and Search Tool; probabilistic search tool for identifying damage genes and the disease causing variants; can score both coding and non-coding variants; Four tools: VAT (Variant annotation tool), VST (Variant Selection Tool), VAAST, pVAAST (for pedigree data); updated April 2015</li>
<li>Input: FASTA, GFF3, GVF</li>
<li>Output: CDR (condenser file), VAAST file (both unique to VAAST)</li>
</ul>
</li>
<li><strong><a href="http://useast.ensembl.org/info/docs/tools/vep/index.html?redirect=no">VEP</a></strong>
<ul>
<li>Description: (Ensembl) Variant Effect Predictor; determines effect of variants on genes, transcripts, and protein sequence; uses SIFT and PolyPhen</li>
<li>Input: Coordinates of variants and nucleotide changes; whitespace- separated format, VCF, pileup, HGVS</li>
<li>Output: VCF, JSON, Statistics</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/cancer/cga/absolute">ABSOLUTE</a></strong>
<ul>
<li>Description: (Broad Institute); can estimate purity and ploidy to compute absolute copy number and mutation multiplicitie; reextracts data from the mixed DNA population</li>
<li>Input: HAPSEQ segdat or segmentation file</li>
<li>Output: per-sample output directory and subdirectory providing per-sample text files containing standard out being emitted from R</li>
</ul>
</li>
<li><strong><a href="http://www.interactive-biosoftware.com/alamut-batch/">Alamut Batch</a></strong>
<ul>
<li>Description: high-throughput annotation software for NGS analysis; for &ldquo;intensive variant analysis workflows&rdquo;; &ldquo;enriches raw NGS variants with dozens of attributes&rdquo;; based on clinically oriented Alamut database; Supports human genes; easy to integrate into pipeline (Latest Release- July 2015)</li>
<li>Input:VCF, tab-delimted file</li>
<li>Output: tab-separated file of annotations</li>
</ul>
</li>
<li><strong><a href="http://avia.abcc.ncifcrf.gov/apps/site/index">AVIA</a></strong>
<ul>
<li>Description: Annotation, Visualization, and Impact Analysis; &ldquo;The tool is based on coupling a comprehensive annotation pipeline with a flexible visualization method. We leveraged the ANNOVAR (Wang et. al, 2010) framework for assigning functional impact to genomic variations by extending its list of reference annotation databases (RefSeq, UCSC, SIFT, Polyphen etc.) with additional in-house developed sources (Non-B DB, PolyBrowse).&rdquo;</li>
<li>Input: BED</li>
<li>Output: Table of annotations with gene annotation features</li>
</ul>
</li>
<li><strong><a href="http://bioinformaticstools.mayo.edu/research/bior/">BioR</a></strong>
<ul>
<li>Description: (Mayo Clinic) (Page last updated June 2015) Biological Reference Repository; &ldquo;data integration tool that enables coordinate based searches and joins based on strings&rdquo;; &ldquo;BioR consists of two parts 1) the BioR toolkit which depends on Java&hellip;. 2) the BioR catalogs which are the data files used by the system&rdquo;</li>
<li>Input: VCF</li>
<li>BioR-Supported Catalogs (tar-gzip files): dbSNP, 1000 genomes, HapMap, OMIM, NCBIGene</li>
<li>Output: VCF + JSON</li>
</ul>
</li>
<li><strong><a href="http://cadd.gs.washington.edu/">CADD</a></strong>
<ul>
<li>Description: Combined Annotation Dependent Depletion; tool for scoring SNV deletions/insertions; &ldquo;integrates multiple annotations into one metric&rdquo;; Score strongly correlates with allelic diversity and pathogenicity; links to 1000 Genome variants; uses Ensembl Variant Effect Predictor</li>
<li>Input: VCF</li>
<li>Output: CADD score</li>
</ul>
</li>
<li><strong><a href="http://www2.hu-berlin.de/wikizbnutztier/software/CandiSNPer/">CandiSNPer</a></strong>
<ul>
<li>Description: web application, characterizes SNPs located in vicinity of SNP of interest;</li>
<li>Input: enter SNP ID (rsID), choose population, region, measure for LD, threshold plot format, color of SNPs, and chose to show genes</li>
<li>Output: Imagefile</li>
</ul>
</li>
<li><strong><a href="https://github.com/UppsalaGenomeCenter/CanvasDB">CanvasDB</a></strong>
<ul>
<li>Description: &ldquo;local database infrastructure for analysis of targeted- and whole genome re-sequencing projects&rdquo;; dependent on MySQL, R, and ANNOVAR</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/carol/">CAROL</a></strong>
<ul>
<li>Description: (Wellcome Trust Sanger); Combined Annotation scoRing toOL; Combined functional annotation score of nonsynonymous coding variants; Combines information from PolyPhen-2 and SIFT</li>
<li>Input: tab-delimited with columns obtained from PolyPhen-2 and SIFT output</li>
<li>Output: tab-delimited file</li>
</ul>
</li>
<li><strong><a href="http://wiki.chasmsoftware.org/index.php/Main_Page">CHASM</a></strong>
<ul>
<li>Description: Cancer-specific High-throughput Annotation of Somatic Mutations; Last updated May 2014; uses Random Forest Method to &ldquo;distinguish between driver and passenger somatic mutations&rdquo;; Positive driver class curated from COSMIC database; packed together with SNVBox (database)</li>
<li>Input:Passenger mutation rates, Transcript and amino acid change, Genomic coordinates</li>
<li>Output: CHASM score, p-value, FDR</li>
</ul>
</li>
<li><strong><a href="http://www.cravat.us/">CRAVAT</a></strong>
<ul>
<li>Description: Cancer-Related Analysis of Variants Toolkit; Web application; Uses CHASM, VEST, SNVGet; &ldquo;CRAVAT provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases&rdquo; Latest Release May 2015;</li>
<li>Input: VCF, CRAVAT format</li>
<li>Output: CRAVAT report- MS Excel spreadsheet or tab-separated file (emailed)</li>
</ul>
</li>
<li><strong><a href="http://cupsat.tu-bs.de/">CUPSAT</a></strong>
<ul>
<li>Description: Cologne University Protein Stability Analysis Tool; &ldquo;tool to predict changes in protein stability upon point mutations&rdquo;; web service program; Can predict mutant stability from existing PDB structures or custom protein structures</li>
<li>Input:for PDB- provide PDB ID and Amino Acid Residue Number; for custom- PDB file format</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://cbcl.ics.uci.edu/public_data/DANN/">DANN</a></strong>
<ul>
<li>Description: Deleterious Annotation of genetic variants; standalone program, uses &ldquo;the same feature set and training data as CADD to train a deep neural network&rdquo;; can catch nonlinear relationships; &ldquo;There are four different datasets: training, validation, testing, and ClinVar_ESP...The ClinVar_ESP dataset is also a testing set containing a set of &ldquo;gold standard&rdquo; pathogenic and benign variants&rdquo;</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://rulai.cshl.edu/cgi-bin/tools/ESE3/esefinder.cgi?process=matrices">ESEfinder</a></strong>
<ul>
<li>Description: Exonic Splicing Enhancer; useful for interpretation of point mutations/polymorphisms that are disease-associated; GUI interface; web app program</li>
<li>Input: FASTA</li>
<li>Output: html or plain text format, graphical display of results</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/exomiser/">Exomiser</a></strong>
<ul>
<li>Description: Wellcome Trust Sanger; functionally annotates variants from whole-exome sequencing data; Based on Jannovar and uses UCSC KnownGene; Java program; web app program (Page last modified Feb 2015)</li>
<li>Input: VCF</li>
<li>Output: TSV, VCF</li>
</ul>
</li>
<li><strong><a href="https://sites.google.com/site/famannotation/home">FamAn</a></strong>
<ul>
<li>Description: Automated variant annotation pipeline for family-based sequencing studies; Annotaties SNVs and INDELs; 4 models- autosomal dominant, autosomal recessive, de novo mutations and a general model; &ldquo;A variety of annotations are provided for each segregating variant: number of family (and family ID) each variant hits, variant genomic location and coding effect (based on snpEff), loss-of-function mutation annotation, selected ENCODE annotation, allele frequency in the 1000 Genomes Project, allele frequency in the Exome Variant Server (ESP6500), segmental duplication annotation, SIFT, PolyPhen2, LRT, MutationTaster, GERP++, PhyloP, SiPhy, etc.&rdquo; (Last updated May 2014)</li>
<li>Input: VCF</li>
<li>Output: two excel compatible outputs</li>
</ul>
</li>
<li><strong><a href="http://www.gene-talk.de/">GeneTalk</a></strong>
<ul>
<li>Description: Combines tool for filtering and data analysis with an online network for genetic professionals; Different degrees- basic license, premium license, in-house solution (the last ones are paid for- Commercial tool?)</li>
<li>Input: VCF</li>
<li>Output: GeneTalk Annotation- includes clinical data, medical relevance, scientific relevance (<a href="http://www.gene-talk.de/public/GeneTalk_Whitepaper_Annotations.pdf">http://www.gene-talk.de/public/GeneTalk_Whitepaper_Annotations.pdf</a>)</li>
</ul>
</li>
<li><strong><a href="http://genevetter.kidneyomics.org/">GeneVetter</a></strong>
<ul>
<li>Description: &ldquo;GeneVetter is a tool designed for investigation of the background prevalence of exonic variation in the Phase 3 1000 Genomes data under user defined filtering criteria&rdquo;; web app program; GeneVetter uses GRch37p4 (hs37d5.fa.gz), dbSNP build 138, 1000G Phase 3, clinvar_2014072</li>
<li>Input: VCF</li>
<li>Output: TIMS score, summary table, PCA plot</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/software/cprg/?q=node/31">GSITIC</a></strong>
<ul>
<li>Description: (Broad Institute) Last update- July 2014; Identifies genomic regions that are significantly &ldquo;amplified or deleted&rdquo;; Each is given a G score; gives genomic locations and q-values from aberrant regions</li>
<li>Input: segmentation file -seg, markers file -mk (required); -array file list -alf, CNV file -cnv</li>
<li>Reference genome: -refgene (created in MATLAB, GISITIC provides four reference genomes: hg16.mat, hg17.mat, hg18.mat, hg19.mat</li>
<li>Output: All lesions file (text file), amplifications file (text file), deletion genes file (text file), Gistic Scores file, Segmented copy number (pdf file), amplification score GISTIC plot (pdf file), Deletion score/q-vale GISTIC plot (pdf file)</li>
</ul>
</li>
<li><strong><a href="http://www.cmbi.ru.nl/hope/about">HOPE</a></strong>
<ul>
<li>Description: Have yOur Protein Explained; Web app program; Automatic mutant analysis server that provides structural effects of a mutation; Uses BLAST against UniProt and PDB along with homology modeling</li>
<li>Input: FASTA protein sequence, or accession code of protein of interest</li>
<li>Output: a report containing information from a &ldquo;decision tree&rdquo; and illustrated figures and animations</li>
</ul>
</li>
<li><strong><a href="http://umd.be/HSF/">Human Splicing Finder</a></strong>
<ul>
<li>Description: Last update: May 2013; aimed to help study pre-mRNA splicing; combines 12 algorithms to identify mutations&rsquo; effect on splicing motifs; uses ensembl database 70</li>
<li>Input: Gene Name, Ensembl transcript ID, Ensembl Gene ID, Consensus CDS, RefSeq Peptide ID, or own sequence (looks like you can enter FASTA)</li>
<li>Output: Chart with columns for predicted signal, predicted algorithm, cDNA position and interpretation</li>
</ul>
</li>
<li><strong><a href="http://larva.gersteinlab.org/">LARVA</a></strong>
<ul>
<li>Description: Large-scale Analysis of Variants in noncoding Annotations; New version released July 2015; Command-line program; used for studying noncoding variants; integrates comprehensive set of noncoding elements, modeling their mutation count; Dependent on C++ and BEDtools</li>
<li>Input: multiple</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.jurgott.org/linkage/LinkagePC.html">LINKAGE</a></strong>
<ul>
<li>Description:three main programs: mlink (calculates lod scores at fixed values for the recombination fraction in one interval of a genetic map), linkmap (calculates location scores for positions of a disease locus along a marker), and ilink (estimates parameters including recombination fractions, allele frequencies, penetrances, etc)</li>
<li>Input: pedfile (processed by MAKEPED) and datafile (reflects loci for each individual; set in PREPLINK)</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/mnvannotationcorrector/">MAC</a></strong>
<ul>
<li>Description: MNV Annotation Corrector; Ad hoc software, fixes incorrect amino acid predictions that are caused by multiple nucleotide variations; Uses existing annotators ANNOVAR, SnpEff, VEP (last update April 2015) (only 1 download this week &rarr; not popular)</li>
<li>Input: List of called SNVs and corresponding BAM</li>
<li>Output: Report identifying block of mutation within codon (BMCs)</li>
</ul>
</li>
<li><strong><a href="http://genome.igib.res.in/mitomatic/">mit-o-matic</a></strong>
<ul>
<li>Description: focuses on mtDNA, provides clinically relevant information from different resources; two component pipeline: command link for alignment of NGS reads and online version that provides genetic report on mitocondrial variants</li>
<li>Input:FASTQ, pileup</li>
<li>Reference sequence: rCRSm</li>
<li>Output: Online version gives comprehensive genetic report</li>
</ul>
</li>
<li><strong><a href="http://krauthammerlab.med.yale.edu/mutadelic/index.html">Mutadelic</a></strong>
<ul>
<li>Description: Web App program; &ldquo;This application generates reports on inherited mutations in five genes (ANK1, SLC4A1, SPTA1, SPTB and EPB42) associated with the following rare Mendelian blood disorders: Hereditary Spherocytosis (HS), Hereditary Elliptocytosis (HE) and Hereditary Pyropoikilocytosis&rdquo;; Newer program- recently validated on omictools</li>
<li>Input: Can upload coordinates of DNA variants or VEP</li>
<li>Output: Displayed on web or can be downloaded in Excel or RDF format</li>
</ul>
</li>
<li><strong><a href="http://www.mutationtaster.org/">MutationTaster</a></strong>
<ul>
<li>Description: (Last post on site 2014) Web app program; Rapid evaluation of disease causing alterations; uses NCBI 37 and Ensembl 69</li>
<li>Input: HGNC symbol, NCBI GeneID, or Ensembl ID,</li>
<li>Output: Report containing prediction, summary, name of alteration, etc</li>
</ul>
</li>
<li><strong><a href="http://mutpred.mutdb.org/">MutPred</a></strong>
<ul>
<li>Description: web app tool; Classifies amino acids substituation as disease associated or neutral in humans; Last modified Feb. 2014; Based on SIFT, trained using Human Gene Mutation Database</li>
<li>Input:</li>
<li>Output: &ldquo;The output of MutPred contains a general score (g), i.e., the probability that the amino acid substitution is deleterious/disease-associated, and top 5 property scores (p), where p is the P-value that certain structural and functional properties are impacted.&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/cancer/cga/mutsig">MutSigCV</a></strong>
<ul>
<li>Description: (Broad Institute) Mutation Significance (CV= covariates); Analyzes mutations discovered in DNA sequencing to identify genes that were mutated more often than expected</li>
<li>Input: mutations.maf, coverage.txt, covariates.txt</li>
<li>Output: output.txt</li>
</ul>
</li>
<li><strong><a href="http://stothard.afns.ualberta.ca/downloads/NGS-SNP/">NGS-SNP</a></strong>
<ul>
<li>Description: Collection of command-line scripts for providing rich SNP annotations; &ldquo;NCBI, Ensembl, and Uniprot IDs are provided for genes, transcripts and proteins when applicable&rdquo;;</li>
<li>Input: Samtools consensus pileup, Maq, diBayes, Genetic format, VCF</li>
<li>Output: File containing annotated SNPs is copied from SNP list and some classes are added</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/oncotator">Oncotator</a></strong>
<ul>
<li>Description: (Broad Institute) &ldquo;Tool for annotating human genomic point mutations and data relevant to cancer researchers&rdquo;; Web app; Supports annotation of data from ClinVar, dbSNP, 1000 genomes (plus many other external sites); Only GRCh27 coordinates supported; Last update: April 2015</li>
<li>Input: tal-delimited file</li>
<li>Output: tab-delimited MAF</li>
</ul>
</li>
<li><strong><a href="http://omictools.com/panther-s649.html">PANTHER</a></strong>
<ul>
<li>Description: Protein ANalysis THrough Evolutionary Relationships; Web app program, also has its own database; Classification system used to classify proteins and their genes; Also, &ldquo;Estimates the likelihood of a particular nonsynonymous (amino-acid changing) coding SNP to cause a functional impact on the protein&rdquo;; Updated in 2015</li>
<li>Input: Data from PANTHER, IDs from Ensembl, EntrezGene, NCBI GI numbers, NCBI UniGene IDs HUGO, UniProt; if ID type is not one of the above, can input txt file or excel format</li>
<li>Output: Analysis results displayed online</li>
</ul>
</li>
<li><strong><a href="http://cubio.biology.columbia.edu/pesx/pesx/">PESX</a></strong>
<ul>
<li>Description: Putative Exonic Splicing Enhancers/Silencers; (Can&rsquo;t tell if this is outdated or not)</li>
<li>Input: FASTA or plain text</li>
<li>Output: Excel spread sheet</li>
</ul>
</li>
<li><strong><a href="http://phen-gen.org/index.html">Phen-Gen</a></strong>
<ul>
<li>Description: Combines patient's&rsquo; disease symptoms with sequencing data; Standalone or Web app version; Only excepts 1 family per run, in order to evaluate unrelated individuals, each sample needs to be run individually</li>
<li>Input: Variant- VCF; Pheotype- HPO; Pedigree- PED</li>
<li>Output: Combined scores file, variants for top genes file</li>
</ul>
</li>
<li><strong><a href="http://mmb.pcb.ub.es/PMut/">PMUT</a></strong>
<ul>
<li>Description: Aimed at annotation and prediction of pathological mutations; based on different kinds of sequence info and neural networks to process information</li>
<li>Input: FASTA</li>
<li>Output; Simple yes/no and reliability index</li>
</ul>
</li>
<li><strong><a href="http://provean.jcvi.org/index.php">PROVEAN</a></strong>
<ul>
<li>Description: Protein Variation Effect Analyzer; predicts whether an amino acid substitution or indel has impact on biological function of the protein; &ldquo;comparable to SIFT or Polyphen-2&rdquo;; Standalone, Web app, Command line or GUI; Last update May 2014</li>
<li>Input: FASTA, list of variants;</li>
<li>Output: tab-separated columns including Variant, Provean Score and prediciton</li>
</ul>
</li>
<li><strong><a href="http://genes.mit.edu/burgelab/rescue-ese/">Rescue-ESE</a></strong>
<ul>
<li>Description: &ldquo;An online tool for identifying candidate ESEs in vertebrate exons&rdquo;; Web application; For human, mouse, zebrafish, pufferfish</li>
<li>Input: multi-FASTA or plain text</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://scandb.org/newinterface/index_v1.html">SCAN</a></strong>
<ul>
<li>Description: Web application program, includes a database as well; Database contains physical-based SNP annotations and functional annotations; &ldquo;Information on physical, functional, and LD annotation served on the SCAN database comes directly from public resources, including the HapMap (release 23a), NCBI (dbSNP 129), or is information created by us using data downloaded from these public resources&rdquo;; &ldquo;SCAN can be utilized in several ways including: (i) queries of the SNP and gene databases; (ii) analysis using the attached tools and algorithms; (iii) downloading files with SNP annotation for various GWA platforms&rdquo;</li>
<li>Input:</li>
<li>Output: HTML, comma-delimited, tab-delimited</li>
</ul>
</li>
<li><strong><a href="http://snp.gs.washington.edu/SeattleSeqAnnotation137/">SeattleSeq Annotation</a></strong>
<ul>
<li>Description: &ldquo;SeattleSeqAnnotation137 was most recently updated October 13, 2013. The current version is 8.08. The most recent site, based on dbSNP build 141, and hg38/NCBI 38&rdquo;; Provides annotations for SNVs and Indels- includes dbSNP rsID, gene names and accession numbers, variation functions, protein positions and amino acid changes, conservation scores, HapMap frequencies, PolyPhen predictions and clinical association.</li>
<li>Input: Maq, gff, CASAVA, VCF, GATK bed, custom</li>
<li>Output: &ldquo;default output file format is a header line (starting with "#") followed by tab-separated annotations&rdquo;; VCF</li>
</ul>
</li>
<li><strong><a href="https://cran.r-project.org/web/packages/seqminer/">seqminer 3.7</a></strong>
<ul>
<li>Description: &ldquo;Efficiently Read Sequence Data (VCF Format, BCF Format and METAL Format) into R&rdquo;; Command line package program; Published August 2015</li>
<li>Input: VCF, BCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://genomics.scripps.edu/ADVISER/Home.jsp">SG Adviser</a></strong>
<ul>
<li>Description: Scripps Genome Annotation and Distributed Variant Interpretation Server, web developed applications for variant annotation, &ldquo;Downstream applications of variant annotation include: Clinical sequencing applications including: carrier testing, or identification of causal variants in molecular diagnosis, tumor sequencing, or diagnostic odyssey. Prioritization of variants prior to statistical analysis of sequence based disease association studies, especially for automated set-generation and enrichment of likely functional variants within sets. Identification of causal variants in post-GWAS/linkage sequencing studies. Identification of causal variants in forward genetic screens (stay tuned for non-human annotation)&rdquo;</li>
<li>Input: SNV- VCF, BED, and a few others; CNV- BED, CNVator, plus others</li>
<li>Output: tab-delimited file</li>
</ul>
</li>
<li><strong><a href="https://rostlab.org/services/snap/">SNAP-2</a></strong>
<ul>
<li>Descriptio</li></ul></li></ol>]]></description>
	<dc:creator>Jitendra Prajapati</dc:creator>
</item>

</channel>
</rss>