<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/12206?offset=70</link>
	<atom:link href="https://bioinformaticsonline.com/related/12206?offset=70" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/26617/list-of-bioinformatics-software-tools-for-next-generation-sequencing</guid>
	<pubDate>Fri, 11 Mar 2016 20:22:14 -0600</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/26617/list-of-bioinformatics-software-tools-for-next-generation-sequencing</link>
	<title><![CDATA[List of Bioinformatics Software Tools for Next Generation Sequencing]]></title>
	<description><![CDATA[<p><strong>Commercial tools</strong></p><ol>
<li><strong><a href="http://www.strand-ngs.com/">Strand NGS</a></strong>
<ul>
<li>offers many different tools including alignment, RNA-Seq, DNA-Seq, ChIP-Seq, Small RNA-Seq, Genome Browser, visualizations, Biological Interpretation, etc. Supports workflows &ldquo;one can import the sample data in FASTA, FASTQ or tag-count format. In addition, prealigned data in SAM, BAM or Illumina-specific ELAND format can be directly imported for analysis.&rdquo;</li>
<li>Alignment feature: Supports alignment from Illumina, Ion Torrent, 454 (Roche), and Pac Bio</li>
<li>DNA-Seq Feature, can annotate with dbSNP</li>
</ul>
</li>
<li><strong><a href="http://www.clcbio.com/desktop-applications/top-features/">CLC Genomics Workbench</a></strong><br />
<ul>
<li>(QIAGEN). Features include: resequencing, workflow, read mapping, de novo assembly, variant detection, RNA-Seq, ChIP-Seq, Genome Browser, etc (entire list on website); Main Workbench offers database search (Genbank, Blast, Pubmed); 2000 organizations have invested in CLC</li>
<li>Accepts VCF files from 1000 Genomes Project</li>
<li>Accepts downloaded tracks from dbSNP</li>
<li>Also accepts: FASTA, GFF/GTF/GVF, BED, Wiggle, Cosmic, UCSC variant database, complete genomics master var file</li>
<li>Read mapping: &ldquo;In addition to Sanger sequence data, reads from these high-throughput sequencing machines are supported: The 454 FLX System and the 454 GS Junior System from Roche, Illumina Genome Analyzer, Illumina HiSeq, Illumina HiScan, and Illumina MiSeq sequencing systems, SOLiD system from Life Technologies, Ion Torrent system from Life Technologies, Helicos from Helicos BioSciences&rdquo;</li>
<li>De novo assembly: &ldquo;In addition to Sanger sequence data, reads from these high-throughput sequencing machines are supported The 454 FLX System and the 454 GS Junior System from Roche, Illumina Genome Analyzer, Illumina HiSeq, Illumina HiScan, and Illumina MiSeq sequencing systems, SOLiD system from Life Technologies, Ion Torrent system from Life Technologies&rdquo;</li>
<li>Annotation tracks from Ensembl</li>
</ul>
</li>
<li><strong><a href="https://www.dnanexus.com/product-overview">DNAnexus</a></strong>
<ul>
<li>Private cloud repository -- formerly a redistributor of SRA and other NCBI resources; command-line or via web, can fetch data from a URL, build custom pipeline/ workflow has sra.dnanexus.com site: data downloads come directly from NCBI</li>
</ul>
</li>
<li><strong><a href="http://www.ingenuity.com/products/variant-analysis">Ingenuity Variant Analysis</a></strong>
<ul>
<li>(QIAGEN) allows for variant identification and analysis, uses NCI-60 data set for cancer, Supported third part informatin: Entrez Gene, RefSeq, ClinVar; gives contextual details of results instead of just A to B relationship</li>
<li>Has own database-- &ldquo;knowledge base&rdquo; based on COSMIC, OMIM, and TCGA databases</li>
</ul>
</li>
<li><strong><a href="http://www.dnastar.com/t-products-dnastar-lasergene-genomics.aspx">Lasergene Genomics Suite</a></strong>
<ul>
<li>Comprehensive NGS software pipeline for assembly, alignment, variant calling and analysis of NGS data</li>
<li>Supported workflows include: reference-guided and de novo genome and transcriptome assembly and analysis, metagenomics sample assembly, targeted resequencing, exome alignment, gene panels with validation control, variant analysis, and RNA-Seq, ChIP-Seq and miRNA alignment and analysis.</li>
<li>#1 in accuracy: fewer false negatives and better sensitivity compared to results obtained from other aligners</li>
<li>Aligns exome data and performs variant calling an average of 3 times faster than alternative pipelines</li>
<li>Annotates genomic data with allele and genotype frequency, functional impact predictions, evolutionary conservation scores and pathogenicity</li>
<li>Supports all major NGS technologies (Illumina, Ion Torrent, Pac Bio and Roche 454) and project types</li>
<li>Available on Windows, Mac OS X, Linux, and the Amazon Cloud</li>
</ul>
</li>
<li><strong><a href="http://www.softgenetics.com/NextGENe.html">NextGENe</a></strong>
<ul>
<li>&ldquo;perfect analytical partner for the analysis of desktop sequencing data produced by the ION PGM&trade;, Roche Junior, Illumina MiSeq as well as high throughput systems as the Ion Torrent Proton, Roche FLX, Applied BioSystems SOLiD&trade; and Illumina&reg; platforms.&rdquo; runs on Windows, free-standing multi-application package-- SNP/Indel analysis, CNV prediction and disease discovery, whole genome alignment, etc.</li>
<li>Data can be imported from Clinvar, dbSNP, Genbank:<a href="http://www.softgenetics.com/PDF/NextGene_UsersManual_web.pdf">http://www.softgenetics.com/PDF/NextGene_UsersManual_web.pdf</a></li>
</ul>
</li>
<li><strong><a href="http://www.partek.com/pgs">Partek Genomics Suite</a></strong>
<ul>
<li>Cited in over 3,500 peer-reviewed scientific publications</li>
<li>Workflows for microarray and PCR data include: Gene expression including alternative splicing, miRNA expression, Genome Wide Association Studies, Mother-Father-Child Trio analysis, DNA Copy number including allele specific copy number and Loss of Heterozygosity (LOH), and ChIP, and methylation. Next Generation Sequencing (NGS) workflows include: RNA-Seq, miRNA-Seq, ChIP-Seq, DNA-Seq, and Methylation</li>
<li>Powerful statistics and interactive, publication ready visualizations</li>
<li>Supports all commercial next generation sequencing and microarray file format as well as text files</li>
<li>Can input GEO SOFT files</li>
</ul>
</li>
<li><strong><a href="http://www.partek.com/partekflow">Partek Flow</a></strong>
<ul>
<li>Installation can be cloud-based or on a local cluster or Linux server</li>
<li>Easy to use point-and-click interface</li>
<li>Takes NGS data (.fastq, BAM, SAM), microarrays (Affymetrix, Illumina) and text files</li>
<li>Supports custom genome builds and annotation databases</li>
<li>Performs base trimming, alignment, quantification, quality analysis, statistics, and visualization</li>
<li>Includes ten fully customizable aligners (Bowtie, Bowtie 2, BWA, GSNAP, Isaac 2, SHRiMP 2, STAR, TMAP, TopHat and TopHat 2)</li>
<li>Applications for RNA-Seq, Small RNA-Seq, WGS/WES, Pathway enrichment, Fusion detection and Variant calling</li>
<li>Allows users to create, save, share, or download analysis pipelines for automated and repeatable analysis</li>
<li>Collaborate with others without transferring data</li>
<li>Integrates microarray and next generation sequencing data</li>
</ul>
</li>
<li><strong><a href="http://goldenhelix.com/SNP_Variation/">Golden Helix: SNP and Variation Suite</a></strong>
<ul>
<li>used for managing, analyzing and visualizing genotypic and phenotypic data; Features: Genome-wide association studies, genomic prediction, copy number analysis, small sample DNA-Seq workflows, large sample DNA-seq analysis, RNA-seq analysis. Supported files: .txt, excel XLS &amp; XLSX, CEL, CHP, CNT, Illumina, Plink PED, TPED, BED, Agilent files, NimbleGen data summary files, VCF files, Impute2 GWAS files, HapMap format, MACH output, + 50 other formats consumes NCBI data directly</li>
</ul>
</li>
<li><strong><a href="https://www.genomatix.de/">Genomatix</a></strong>
<ul>
<li>Applications: ChIP-Seq, DNA-Seq, RNA-Seq, DNA methylation; enable personalized medicine,</li>
<li>Mining Stations: Supports all established NGS sequencing platforms- SOLiD, 454 Life Sciences, Genome Analyzer, HiSeq, MiSeq, IonTorrent</li>
<li>Software Suite: can upload sequence of BED files</li>
<li>Genome browser: BED and BAM files, Public data- 1500 BED files available for every user</li>
</ul>
</li>
<li><strong><a href="http://www.biodatomics.com/">Biodatomics</a></strong>
<ul>
<li>Open source platform (SaaS), analysis and genome sequencing tools, integrates over 400 genomic analysis open source tools and pipelines, have a private and public cloud version. Features: genomic data visualization, drag and drop interface, accelerated analysis, real-time collaboration</li>
<li>They have a couple modules to do so, and have enabled parts of the sra toolkit</li>
</ul>
</li>
<li><strong><a href="https://www.solvebio.com/">SolveBio</a></strong>
<ul>
<li>Software product, for clinical genomics professionals, manage, curate, report genomic variation</li>
<li>Has own data library -- data from NCBI</li>
</ul>
</li>
<li><strong><a href="http://www.basepairtech.com">Basepair</a></strong>
<ul>
<li>Offers high quality workflows for all common NGS applications (RNA-Seq, ChIP-Seq, DNA-Seq, etc.)</li>
<li>Very fast - get all results in a 1-2 hours. Cloud-based, no storage or computing limits.</li>
<li>Easy to use - less than a minute to run an analysis</li>
<li>REST and Python API to mange large projects.</li>
</ul>
<div>&nbsp;</div>
</li>
</ol><h2><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#variant-identification"></a>Variant Identification</h2><h3><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#germline-callers"></a>Germline Callers</h3><ol>
<li><strong><a href="http://mathgen.stats.ox.ac.uk/impute/impute_v2.html">IMPUTE2</a></strong>
<ul>
<li>Description: phasing observed genotypes and imputing missing genotypes uses reference panels to provide all available halotypes, does not use population labels or genome-wide measures; designed to represent variation in one population; Fairly popular</li>
<li>Input:</li>
<li>Reference Haplotypes: Links to 1000 Genomes and HapMap downloads</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/ekg/freebayes">FreeBayes</a></strong>
<ul>
<li>Description: finds SNPs, Indels, MNPs; reports variants based on alignment; haplotype based</li>
<li>Input: BAM- uses BAMtools API to parse</li>
<li>Reference genome: FASTA</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://soap.genomics.org.cn/soapindel.html">SOAPindel</a></strong>
<ul>
<li>Description: detects indels from NGS paired-end sequencing</li>
<li>Input: files with read alignment can be SOAP or SAM formats, users must also give raw reads in Fasta or Fastq</li>
<li>Reference Sequence used to align reads: FASTA</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/danmaclean/2kplus2">2Kplus2</a></strong>
<ul>
<li>Description: algorithm searches graphs produced by de novo assembler Cortex; c++ source code for SNP detection &ldquo;2kplus2.cpp is a c++ source code for the detection and the classification of single nucleotide polymorphisms in transformed De Bruijn graphs using Cortex assembler.&rdquo;</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://www.hgsc.bcm.edu/software/atlas-2">Atlas 2</a></strong>
<ul>
<li>Description: specializes in separation of true SNPs and indels from sequencing and mapping errors, last update January 2013</li>
<li>Input: takes BAM file,</li>
<li>Reference Genome: FASTA</li>
<li>Output: produces VCF</li>
</ul>
</li>
<li><strong><a href="https://sites.google.com/site/vibansal/software/crisp">CRISP</a></strong>
<ul>
<li>Description: identifies SNPs and INDELs from pooled high-throughput NGS, not used for analysis of single samples; implemented in C and uses SAMtools API; latest version should work with diploid genomes</li>
<li>Input: requires BAM files (aligned with GATK)</li>
<li>Reference Genome: indexed FASTA file</li>
<li>Output: VCF files</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/dindel/">Dindel</a></strong>
<ul>
<li>Description: (Wellcome Trust Sanger) calls small indels from short-read sequences, only can handle Illumina data; cannot test candidate indels; written in C++, used on Linux based and Mac computers (not tested in windows)</li>
<li>Input: BAM files</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://colibread.inria.fr/software/discosnp/">discoSnp++</a></strong>
<ul>
<li>Description: detects homozygous and heterozygous SNPs and Indels; software composed of 2 modules (kissnp2 and kissreads)</li>
<li>Input: raw NGS datasets; fasta, fastq, gzipped or not;</li>
<li>no reference genome required; read pairs can be given</li>
<li>Output: FASTA</li>
</ul>
</li>
<li><strong><a href="http://odin.mdacc.tmc.edu/~wwang7/FamSeqIndex.html">FamSeq</a></strong>
<ul>
<li>Description: family-based sequencing studies- provides probability of an individual carrying variant based on family&rsquo;s raw measurements; accommodates de novo mutations, can perform variant calling at chrX;</li>
<li>Input: VCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/p/%20geneticthesaurus/wiki/Example/">GeneticThesaurus</a></strong>
<ul>
<li>Description: &ldquo;Annotation of genetic variants in repetitive regions&rdquo;</li>
<li>Input: Initial variant calling from bam &rarr; vcf output</li>
<li>Reference Genome: need to provide own fasta file for hg19 genome,</li>
<li>Output: vcf.gz, vtf.gz, and baf.tsv.gz output</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/GlfMultiples">glfMultiples</a></strong>
<ul>
<li>Description: command-line, variant caller</li>
<li>Input: GLF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/GlfSingle">glfSingle</a></strong>
<ul>
<li>Description: uses likelihood-based model for variant calling, starts from genotype likelihoods that have been computed from other tools (ex. Samtools BAQ), the likelihoods combine with individual-based prior p(genotype) to generate posterior probabilities</li>
<li>Input: GLF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://github.com/ddcap/halvade">Halvade</a></strong>
<ul>
<li>Description: command-line; written in Java, &ldquo;to run halvade a reference is needed for both GATK and BWA and a SNP (dbSNP!) database is required</li>
<li>Input: FASTQ</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://github.com/aakrosh/indelMINER">indelMINER</a></strong>
<ul>
<li>Description: identifies indels from paired-end reads</li>
<li>Input: BAM (aligned in SAMtools API)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://www.broadinstitute.org/cancer/cga/indelocator">Indelocator</a></strong>
<ul>
<li>Description: (Broad Institute): does not perform realignment, relies on alignments in BAM files (BAM files need aligned before put into indelocator); recommended to use GATK prior;</li>
<li>Input: 2 BAM files(tumor &amp; normal), annotated as germline or somatic; also has single sample mode</li>
<li>Output: &ldquo;Output of Indelocator is a high-sensitivity list of putative indel events containing large numbers of false positives. The statistics reported for each event have to be used to custom-filter the list in order to lower false positive rate&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/sequencing/isaac_variant_caller">Isaac Variant Caller</a></strong>
<ul>
<li>Description: detects SNPs and small indels from diploid sample; designed to run on &ldquo;nux-like platforms&rdquo;</li>
<li>Input: BAM</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.swisstph.ch/kvarq">KvarQ</a></strong>
<ul>
<li>Description: in silico genotyping for selected loci in bacterial genome, written in Python and C</li>
<li>Input: FASTQ</li>
<li>reference genome or de novo assembly not needed</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/lofreq/files/">LoFreq</a></strong>
<ul>
<li>Description: SNV caller, Python language, standalone program, uncovers cell-population heterogeneity from high-throughput sequencing datasets; calls variants found in &lt;.05% of the population</li>
<li>Input: BAM file input&rarr; suggest running through GATK</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/Illumina/manta">Manta</a></strong>
<ul>
<li>Description: Calls indels and SVs from paired end reads; standalone, command line program; Written in C++ and Python</li>
<li>Input: BAM (can tolerate non-paired-end reads); a matched tumor sample may be provided as well</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://github.com/benedictpaten/marginAlign">MarginAlign</a></strong>
<ul>
<li>Description: SNV caller, specifically tailored to Oxford Nanopore Reads, written in Python; Package comes with 3 programs, marginAlign, marginCaller (calls SNVs), marginStats (computes qc stats on sam files)</li>
<li>Input: SAM</li>
<li>Output: SAM</li>
</ul>
</li>
<li><strong><a href="http://gmt.genome.wustl.edu/packages/mendelscan/">MendelScan</a></strong>
<ul>
<li>Description: Last release March 2014; for analyzing sequencing data in family studies of inherited diseases; variant calls for a family in VCF file; still in alpha-testing on github, example data uses 1000 genomes dataset</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://github.com/mitenjain/nanopore">nanopore</a></strong>
<ul>
<li>Description: UCSC Nanopore group (group at UCSC studying using ion channels for analysis of single RNA/DNA structures) software pipeline; tailored to Oxford Nanopore Reads; command line program</li>
<li>Input: FASTQ</li>
<li>Reference files: FASTA</li>
<li>Output: &ldquo;For each possible pair of read file, reference genome and mapping algorithm an experiment directory will be created in the nanopore/output directory.&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://omictools.com/platypus-s1989.html">Platypus</a></strong>
<ul>
<li>Description: Package program, written in C, Python, Cython; Can identify SNPs, MNPs, short indels, and larger variants; has been tested on very large datasets (1000 genomes)</li>
<li>Input: BAM</li>
<li>Reference Genome: FASTA (files must be indexed using Samtools or similar program</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.bioinformatics.nl/QualitySNPng/">QualitySNPng</a></strong>
<ul>
<li>Description: detection of SNPs; &ldquo;can be used as a standalone application with graphical user interface as part of pipeline system&rdquo;; does not require fully sequenced reference genome; haplotype strategy</li>
<li>Input:SAM, ACE</li>
<li>Output: GUI</li>
</ul>
</li>
<li><strong><a href="http://revister.sourceforge.net/">ReviSTER</a></strong>
<ul>
<li>Description: command line program; automated pipeline; utilizes BWA, BLAT, and SAMTools; utilizes BWA mapping program;</li>
<li>Input: FASTQ,</li>
<li>Reference sequence file and list file containing STR locations as inputs</li>
<li>Output: SAM</li>
</ul>
</li>
<li><strong><a href="http://dna-discovery.stanford.edu/software/rvd/">RVD</a></strong>
<ul>
<li>Description: command-line program, detection of rare SNVs, relies upon Samtools, can be run in MATLAB</li>
<li>Input: BAM</li>
<li>Reference Genome: FASTA</li>
<li>Output: &ldquo;The algorithm output is a call table -- a comma-separated file with one line for each base position and each line in the following format:</li>
<li>AlginmentReferencePosition, AlignmentBase, Call ,SecondBase, CenteredErrorPrc, ReferenceErrorPrc, SecondBasePrc&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://snver.sourceforge.net/">SNVer</a></strong>
<ul>
<li>Description: calls common and rare variants in pool or individual NGS data, reports overall p-value, operating system independent statistical tool, identifies SNPs and INDELs, written in Java, no dependencies, straightforward command-line</li>
<li>(SNVerGUI=GUI version) --SNVerGUI: desktop tool for variant detection</li>
<li>Input: chrX annotation, sam.zip, bam.zip</li>
<li>reference file must be aligned to the data file</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://compbio.bccrc.ca/software/snvmix/">SNVMix</a></strong>
<ul>
<li>Description: detects SNVs from NGS, post-alignment tool</li>
<li>Input: pileupformat (Maq or Samtools)</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.bsse.ethz.ch/mlcb/research/bioinformatics-and-computational-biology/structural-variant-machine--sv-m-.html">SV-M</a></strong>
<ul>
<li>Description: Structural Variant Machine - predicts indels, uses split read alignment profiles, validated by Sanger Sequencng</li>
<li>Input:paired-end Illumina reads from 1001 genomes project (uses ref plant- 1001genomes.org)</li>
<li>Ouptut:</li>
</ul>
</li>
<li><strong><a href="https://github.com/slindgreen/SNPest">SNPest</a></strong>
<ul>
<li>Description: Standalone program, language C++, Perl</li>
<li>Input: mpileup (SAMtools)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/TrioCaller">TrioCaller</a></strong>
<ul>
<li>Description:Command line program, relies on BWA and samtools; genotype calling for unrelated individuals and parent-offspring trios</li>
<li>Input: BAM (that has been aligned in BWA and Samtools</li>
<li>Output: BCF that can be formatted to VCF using bcftools</li>
</ul>
</li>
<li><strong><a href="http://www.vicbioinformatics.com/software.snippy.shtml">Snippy</a></strong>
<ul>
<li>Description: finds indels between haploid reference genome and NGS sequence reads</li>
<li>Input:read files- FASTQ or FASTA (can be .gz compressed), output- .aln, .tab, .txt</li>
<li>Reference genome in FASTA or GENBANK</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://orca.bu.edu/vntrseek/">VntrSeek</a></strong>
<ul>
<li>Description: pipeline for discovering microsatellite tandem repeats with high-throughput sequencing data</li>
<li>Input: gzip-compressed FASTA or FASTQ</li>
<li>Output: VCF files; one for TRs and observed alleles, another file contains link to viewer</li>
</ul>
</li>
</ol><h3><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#somatic-callers"></a>Somatic Callers</h3><ol>
<li><strong><a href="http://cakesomatic.sourceforge.net/">Cake</a></strong>
<ul>
<li>Description: standalone program, &ldquo;pipeline for the integrated analysis of somatic variants in cancer genomes&rdquo;; integrates four algorithms; written in Perl; required tools: samtools, tabix, vcftools, VarScan2, bambino, cmake, somaticsniper (User guide; workflow page)</li>
<li>Input: tumor and normal reads in BAM files, run through variant calling programs to generate intermediate VCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/cancer/cga/mutect">MuTect</a></strong>
<ul>
<li>Description: Broad Institute, identification of somatic point mutations in cancer genomes; requires preprocessing of reads (GATK)</li>
<li>Input: same as GATK (FASTA reference genome, SAM read files)</li>
<li>Output: call-stats, VCF, wiggle files</li>
</ul>
</li>
<li><strong><a href="http://genome.sph.umich.edu/wiki/Polymutt">Polymutt</a></strong>
<ul>
<li>Description: calls SNVs and detects de novo point mutations in families</li>
<li>Input: GLF or BAM or VCF (must have identical chromosome orders)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://tvap.genome.wustl.edu/tools/bassovac/">Bassovac</a></strong>
<ul>
<li>Description: Improved Bayesian inversion somatic caller; unlike other software packages, treats effects fully probabilisticallys instead of using ad-hoc modeling; effects are integrated at the atomic level and standard probability theory integrates read tallies to the sample level and to the tumor-normal pair level; "pending public release"</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://bioinformatics.ustc.edu.cn/CLImAT/">CLImAT</a></strong>
<ul>
<li>Description: standalone program; &ldquo;accurate detection of copy number alteration and loss of heterozygosity in impure and aneuploid tumor samples using whole genome sequencing data&rdquo;</li>
<li>Input: depth file generated by DFExtract and a config file</li>
<li>Output: .results file, .Gtype, LOG.txt, also generates visualization</li>
</ul>
</li>
<li><strong><a href="http://denovogear.sourceforge.net/">DeNovoGear</a></strong>
<ul>
<li>Description: de-novo variant calling and interpretation; standalone program; dependencies C++ compiler, CMake, HTSlib, Eigen, Boost</li>
<li>Input: PED and BCF</li>
<li>Output: &ldquo;The output format is a single row for each putative de novo mutation (DNM), with the following fields&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/friend1ws/EBCall">EBCall</a></strong>
<ul>
<li>Description: Empirical Baysian Mutation Calling; standalone program; uses tumor/normal paired reads and non-paired normal reference samples; dependent on samtools, R and VGAM pack for R</li>
<li>Input: BAM</li>
<li>Output: not sure what exact type of file- &ldquo;The format of the result is suitable for adding annotation by annovar.&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/usuyama/hapmuc">HapMuc</a></strong>
<ul>
<li>Description: standalone program; &ldquo;utilizes the information of heterozygous germline variants near candidate mutations&rdquo;; Dependent upon- Boost, SAMtools, BEDtools; 3 step workflow</li>
<li>Input: BAM</li>
<li>Output: BED</li>
</ul>
</li>
<li><strong><a href="https://github.com/cui-lab/multigems">MultiGeMS</a></strong>
<ul>
<li>Description: Multi-sample Genotype Model Selection</li>
<li>Input: .txt, pileup (SAM/BAM converted to pileup format)</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://bitbucket.org/joseph07/multisnv/wiki/Home">MultiSNV</a></strong>
<ul>
<li>Description: command-line program; calls SNVs from NGS data from multiple samples from the same patient; dependent on R, Git, cmake, Boost and compile libraries</li>
<li>Input: BAM or pileup</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://compbio.bccrc.ca/software/mutationseq/">MutationSeq</a></strong>
<ul>
<li>Description: standalone program, somatic SNV detection in tumor/normal samples; dependent on python, bamtools, boost, and LAPACK</li>
<li>Input: BAM</li>
<li>Output: VCF4.1 consisting of two parts (meta information &amp; data lines)</li>
</ul>
</li>
<li><strong><a href="http://www.qcmg.org/bioinformatics/tiki-index.php">qSNP</a></strong>
<ul>
<li>Description: standalone program; SNV caller for somatic variants in &ldquo;low cellularity cancer samples&rdquo;</li>
<li>Input: BAM, dbSNP data, Illumina data, chrConv</li>
<li>Output: &ldquo;qSNP output files are named using a 4-element pattern: ...&rdquo;</li>
</ul>
</li>
<li><strong><a href="https://github.com/aradenbaugh/radia/">RADIA</a></strong>
<ul>
<li>Description: RNA and DNA Integrated Analysis for Somatic Mutation Detection; DNA only Method(tumor/normal pair, ignores RNA) or Triple BAM Method (uses all three datasets from same patient); dependent upon python, samtoools, pysam API, BLAT, SnpEff</li>
<li>Input: BAM</li>
<li>Reference Genome: FASTA indexed with SAMtools faidx</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genomics.wpi.edu/rvd2/">RVD2</a></strong>
<ul>
<li>Description: sensitive, variant detection for low-depth targeted NGS data; python module or command- line program;</li>
<li>Input: tab- deliminted depth chart format (converted from pileup files)</li>
<li>Output: three hdf5 files and a vcf file</li>
</ul>
</li>
<li><strong><a href="https://github.com/nhansen/Shimmer">Shimmer</a></strong>
<ul>
<li>Description: standalone program; detects somatic SNVs with multiple testing correction, uses Fisher&rsquo;s exact test; dependent on git, samtools, R, R statmod package; for tumor/normal matched samples</li>
<li>Input: BAM</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.cs.helsinki.fi/en/gsa/snv-ppilp/">SNV-PPILP</a></strong>
<ul>
<li>Description: Refines GATK&rsquo;s Unified Genotyper SNV calls for &ldquo;multiple samples assumed to form a phylogeny&rdquo;</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://gmt.genome.wustl.edu/packages/somatic-sniper/">SomaticSniper</a></strong>
<ul>
<li>Description: command-line application to identify SNPs between tumor/normal pairs- predicts probability of difference between two</li>
<li>Input: BAM</li>
<li>Reference Genome in FASTA</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://sites.google.com/site/strelkasomaticvariantcaller/">Strelka</a></strong>
<ul>
<li>Description: somatic variant calling workflow for matched tumor-normal samples; detects indels; runs on *nux-like platform</li>
<li>Input: BAM (must be sorted and indexed)- Strelka does own realignment around indels-- don&rsquo;t need to do this type of pre-processing</li>
<li>Output: pair of VCF files</li>
</ul>
</li>
<li><strong><a href="http://www.pitt.edu/~wec47/triodenovo.html">Triodenovo</a></strong>
<ul>
<li>Description: Bayesian framework for calling de novo mutations in trios</li>
<li>Input: VCF file with PL or GL fields (recommend using GATK or samtools to generate)</li>
<li>Output: out_vcf</li>
</ul>
</li>
<li><strong><a href="http://lbg.med.unc.edu/~mwilkers/unceqr_dist/">UNCeqr</a></strong>
<ul>
<li>Description: finds somatic mutations using integration of DNA and RNA seq data-- boosts sensitivity for low purity tumors and rare mutations;</li>
<li>Input:&rdquo;can accept a variety of sequencing inputs and configurations&rdquo;</li>
<li>Output: &ldquo;table of somatically mutated sites and associated information. These somatic mutations can be annotated with predicted transcript and protein effects using third party tools, such as Annovar&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/virmid/">Virmid</a></strong>
<ul>
<li>Description: Virtual Microdissection for SNP calling; Java based; for disease-control matched samples; uncovers SNPs with low allele frequency by considering alpha contamination</li>
<li>Input: BAM (must be sorted and indexed- samtools sort)</li>
<li>Output: VCF and report file</li>
</ul>
</li>
</ol><h3><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#germline--somatic--callers"></a>Germline + Somatic Callers</h3><ol>
<li><strong><a href="http://massgenomics.org/varscan">VarScan 2</a></strong>
<ul>
<li>Description: identify germline variants, private and shared variants, somatic mutations, and somatic CNVs; detects indels</li>
<li>Input: SAMtools pileup</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://genformatic.com/baysic/">BAYSIC</a></strong>
<ul>
<li>Description: Bayesian method; combines variant calls from different methods (GATK, FreeBayes, Atlas, Samtools, etc)</li>
<li>Input: VCF format from one or more variant calling programs</li>
<li>Output: VCF file containing integrated set of variant calls</li>
</ul>
</li>
<li><strong><a href="https://github.com/ding-lab/msisensor">MSIsensor</a></strong>
<ul>
<li>Description: Microsatellite instability detection; C++ program, detects somatic and germline variants in tumor-normal paired data</li>
<li>Input: BAM index files (normal and tumor)</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://faculty.washington.edu/browning/beagle/beagle.html">Beagle version 4</a></strong>
<ul>
<li>Description: software package: genotype calling, phasing, imputation of ungenotyped markers, and identity-by-descent segment detection:unsure if this one is in the right category; genotype calling, phasing, imputation of ungenotyped markers, and identity-by-descent segment detection;</li>
<li>Input: VCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://www.iro.umontreal.ca/~csuros/quadgt/">QuadGT</a></strong>
<ul>
<li>Description: software package, SNV calling from normal-tumor pair and two parent genomes; quantifies descent-by-modification relationships; Written in Java</li>
<li>Input: BAM files (parsed by Picard/Samtools API)</li>
<li>Reference Genome; FASTA</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/rarevator/">RAREVATOR</a></strong>
<ul>
<li>Description: RAre REference VAriant annotaTOR; command line; &ldquo;identification and annotation of germline and somatic variants in rare reference allele loci from second generation sequencing data&rdquo;; Bayesian genotype likelihood model</li>
<li>Input: BED or VCF files from GATK</li>
<li>Output: two VCF files (one for SNVs, one for Indels)</li>
</ul>
</li>
<li><strong><a href="http://scalpel.sourceforge.net/">Scalpel</a></strong>
<ul>
<li>Description: Used for detecting indels in a reference genome; performs localized micro-assembly of specific regions of interest; can do single, de novo, somatic reads; requires that raw reads are aligned with BWA</li>
<li>Input: BAM</li>
<li>Output: either VCF or ANNOVAR</li>
</ul>
</li>
<li><strong><a href="http://soap.genomics.org.cn/soapsnp.html">SOAPsnp</a></strong>
<ul>
<li>Description: based on Baye&rsquo;s theorem; calls consensus genotype</li>
<li>Input:SOAP short read alignment results</li>
<li>Output: GLF, option of flat tabular format</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/variantmaster/">VariantMaster</a></strong>
<ul>
<li>Description: &ldquo;extract causative variants for monogenic and sporadic genetic diseases&rdquo;; uses ANNOVAR;</li>
<li>Input: BAM or VCF files (from SAMtools, GATK)</li>
<li>Output:</li>
</ul>
</li>
</ol><h2><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#downstream-analysis-of-variants"></a>Downstream Analysis of Variants</h2><ol>
<li><strong><a href="https://github.com/hakyimlab/PrediXcan%20https://github.com/hriordan/PrediXcan/">PrediXcan</a></strong>
<ul>
<li>Description: command-line, standalone package program; available in Perl, Python, and R versions; predicts liklihood of a gene being related to a certain phenotype- &ldquo;that directly tests the molecular mechanisms through which genetic variation affects phenotype.&rdquo;; no actual expression data used, only in silico expression; &ldquo;PrediXcan can detect known and novel genes associated with disease traits and provide insights into the mechanism of these associations.&rdquo;</li>
<li>Input: genotype and phenotype file (doesn&rsquo;t specify file type)</li>
<li>Output:default values: genelist, dosages (file format: snpid rsid) , dosage_prefix, weights, output</li>
</ul>
</li>
<li><strong><a href="http://ritchielab.psu.edu/software/athena-downloads">ATHENA</a></strong>
<ul>
<li>Description: Analysis Tool for Heritable and Environmental Network Associations; software package, combines machine learning model with biology and statistics to predict non-linear interactions</li>
<li>Input: Configuration file, Data file, Map file (includes rsID)</li>
<li>Output: Summary file, Best model file, dot file, individual score file, cross-validation file</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/rarevariant/#t_2">CCRaVAT and QuTie</a></strong>
<ul>
<li>Description: (Wellcome Trust Sanger) Case-Control Rare Variant Analysis Tool and Quantitative Trait; software packages for large-scale analysis of rare variants</li>
<li>Input: PED file and MAP file</li>
<li>Output: Five tab-delimited txt files</li>
</ul>
</li>
<li><strong><a href="http://cnsgenomics.com/software/gcta/">GCTA</a></strong>
<ul>
<li>Description: Genome Wide Complex Trait Analysis; package program, command line interface; estimates variance by all SNPs; 5 main functions: &ldquo;data management, estimation of the genetic relationships from SNPs, mixed linear model analysis of variance explained by the SNPs, estimation of the linkage disequilibrium structure, and GWAS simulation&rdquo;</li>
<li>Input: PLINK binary PED files, MACH output format</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://genomecomb.sourceforge.net/">GenomeComb</a></strong>
<ul>
<li>Description: package for analysis of complete genome data; annotation using public data or custom tracks, automated primer desing for Sanger or Sequenom validation; &ldquo;The cg process_illumina command can be used to generate annotated multisample data starting from fastq files, using tools such as bwa for alignment and GATK and samtools for variant calling. Sequencing data can also be imported from Complete Genomics (cg_process_sample command), Real Time Genomics (cg_process_rtgsample command) and VariantCallFormat (VCF) variant files (vcf2sft command).&rdquo;</li>
<li>Input: Sequencing data from Complete Genomics, Illumina, SOLiD and VCF;</li>
<li>Output: standard file format used is a simple tab delimited file (.sft, .tsv)</li>
</ul>
</li>
<li><strong><a href="http://ancorr.eimb.ru/">Genome Track Analyzer</a></strong>
<ul>
<li>Description: compares genome tracks; allows user to compare DNA expression/binding;</li>
<li>Input: multiple: SGR/TXT, BED, BED6, GFF; if using prealigned sequence data- use MACS peak caller: BAM, BED, SAM, ELAND</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://animalgene.umn.edu/gvcblub">GVCBLUP</a></strong>
<ul>
<li>Description: animal gene mapping; &ldquo;genomic prediction and variance component estimation of additive and dominance effects&rdquo;; standalone program, command line interface, writting in C++ and Java</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.jurgott.org/linkage/homog.htm">HOMOG</a></strong>
<ul>
<li>Description: Analyzes heterogeneity with respect to single marker loci or known maps of markers; Carries out homogeneity test for alternative hypothesis &ldquo;Two family types, one with linkage betweeen a trait to a marker or map of markers, the other without linkage&rdquo;</li>
<li>Input: HOMOG.DAT - described on website</li>
<li>Output: HOMOG.OUT</li>
</ul>
</li>
<li><strong><a href="http://intersnp.meb.uni-bonn.de/">INTERSNP</a></strong>
<ul>
<li>Description: GWIA for case-control SNP and quantitative traits; selected for joint analysis using priori information; Provides linear regression framework, Pathway Association Analysis, Genome-wide Haplotype Analysis,</li>
<li>Input: PLINK input formats (ped/map, tped/tfam, bed/bim/fam) Compatible with SetID files</li>
<li>Gene reference file: Ensembl Release 75</li>
<li>Output: covariance matrix for regression models</li>
</ul>
</li>
<li><strong><a href="https://github.com/PMBio/mtSet">mtSet</a></strong>
<ul>
<li>Description: Currently only the standalone version available, but moving to LIMIX software suite; offers set tests- allows for testing between variants and traits; accounts for confounding factors ex. relatedness</li>
<li>Input: sample-to-sample genetic covariance matrix needs to be computed; multiple types of input; simulator requires input genotype and relatedness component;</li>
<li>Output: resdir (result file of analysis), outfile (test statistics and p-values), manhattan_plot (flag)</li>
</ul>
</li>
<li><strong><a href="http://dougspeed.com/multiblup/">MultiBLUP</a></strong>
<ul>
<li>Description: Package program, command line interface; constructs linear prediction models; Best Linear Unbiased Prediction; improves upon BLUP involving kinship matrices; options: pre-specified kinships, regional kinships, adaptive multiblups, LD weightings</li>
<li>Input: PLINK format</li>
<li>Output:.reml, .indi.blp</li>
</ul>
</li>
</ol><h2><a href="https://github.com/NCBI-Hackathons/Community_Software_Tools_for_NGS/blob/master/NGS_Tools_List.md#variant-annotation"></a>Variant Annotation</h2><ol>
<li><strong><a href="http://annovar.openbioinformatics.org/en/latest/">ANNOVAR</a></strong>
<ul>
<li>Description: command-line tool, supports SNPs, INDELs, CNVs and block substitutions, provides wide variety of annotation techniques, depends upon multiple databases (each needing to be downloaded); annotates genetic variants; utilizes RefSeq, UCSC Genes, and the Ensembl gene annotation systems; can compare mutations detected in dpSNP or 1000 Genomes Project; Very popular *&ldquo;The final command run TABLE_ANNOVAR, using dbSNP version 138, 1000 Genomes Project 2014 Oct version, NIH-NHLBI 6500 exome database version 2 (referred to as esp6400siv2), dbNFSP version 2.6 (referred to as ljb26), dbSNP version 138 (referred to as snp138) databases and remove all temporary files, and generates the output file called myanno.hg19_multianno.txt&rdquo;</li>
<li>Input: VCF, ANNOVAR input format (simple text-based format); can convert other formats into ANNOVAR input format</li>
<li>Output: VCF (if input VCF), output file with multiple columns, tab-delimited output file</li>
</ul>
</li>
<li><strong><a href="http://wannovar.usc.edu/">wANNOVAR</a></strong>
<ul>
<li>provides web-based access to ANNOVAR software</li>
</ul>
</li>
<li><strong><a href="http://genetics.bwh.harvard.edu/pph2/">PolyPhen-2</a></strong>
<ul>
<li>Description: Very popular; Polymorphism Phenotyping; Web application; predicts impact of amino acid substitution on protein; Calculates Bayes posterior probability (Last update July 2015)</li>
<li>Input: FASTA</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://sift.jcvi.org/">SIFT</a></strong>
<ul>
<li>Description: predicts how an amino acid substitution will affect protein function; Based on degree of conservation of amino acid residues- collected though PSI-BLAST; can be applied to nonsynonymous polymorphisms or laboratory-induced missense mutations; links to dbSNP 132, GRCh37; Standalone or web app program; Very popular</li>
<li>Input: Uniprot ID or Accession, Go term ID, Function name, Species Name or ID, etc</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://snpeff.sourceforge.net/">snpEff</a></strong>
<ul>
<li>Description: Genetic variant annotation and effect prediction toolbox; integrated with Galaxy, GATK, and GNKO; can annotate SNPs, INDELs, and multiple-nucleotide polymorphisms; categorizes effects into classes by functionality; Very popular; Standalone or Web app; Claims to calculate all SNPs in 1000 genomes (EMBI) in less than 15 minutes; can annotate SNPs, MNPs, and insertions and deletions; Provides assessment of impact of the variant ( low, medium or high)</li>
<li>Input: VCF, BED</li>
<li>Output: VCF (with new ANN field, also used in ANNOVAR and VEP), HTML summary files</li>
</ul>
</li>
<li><strong><a href="http://snpeff.sourceforge.net/SnpSift.html">SnpSIFT</a></strong>
<ul>
<li>Description: Filter and manipulate annotated files; Part of SnpEff main distribution; one variants have been annotated, this can be used to filter your data to find relevant variants</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.yandell-lab.org/software/vaast.html">VAAST 2</a></strong>
<ul>
<li>Description: Variant Annotation, Analysis, and Search Tool; probabilistic search tool for identifying damage genes and the disease causing variants; can score both coding and non-coding variants; Four tools: VAT (Variant annotation tool), VST (Variant Selection Tool), VAAST, pVAAST (for pedigree data); updated April 2015</li>
<li>Input: FASTA, GFF3, GVF</li>
<li>Output: CDR (condenser file), VAAST file (both unique to VAAST)</li>
</ul>
</li>
<li><strong><a href="http://useast.ensembl.org/info/docs/tools/vep/index.html?redirect=no">VEP</a></strong>
<ul>
<li>Description: (Ensembl) Variant Effect Predictor; determines effect of variants on genes, transcripts, and protein sequence; uses SIFT and PolyPhen</li>
<li>Input: Coordinates of variants and nucleotide changes; whitespace- separated format, VCF, pileup, HGVS</li>
<li>Output: VCF, JSON, Statistics</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/cancer/cga/absolute">ABSOLUTE</a></strong>
<ul>
<li>Description: (Broad Institute); can estimate purity and ploidy to compute absolute copy number and mutation multiplicitie; reextracts data from the mixed DNA population</li>
<li>Input: HAPSEQ segdat or segmentation file</li>
<li>Output: per-sample output directory and subdirectory providing per-sample text files containing standard out being emitted from R</li>
</ul>
</li>
<li><strong><a href="http://www.interactive-biosoftware.com/alamut-batch/">Alamut Batch</a></strong>
<ul>
<li>Description: high-throughput annotation software for NGS analysis; for &ldquo;intensive variant analysis workflows&rdquo;; &ldquo;enriches raw NGS variants with dozens of attributes&rdquo;; based on clinically oriented Alamut database; Supports human genes; easy to integrate into pipeline (Latest Release- July 2015)</li>
<li>Input:VCF, tab-delimted file</li>
<li>Output: tab-separated file of annotations</li>
</ul>
</li>
<li><strong><a href="http://avia.abcc.ncifcrf.gov/apps/site/index">AVIA</a></strong>
<ul>
<li>Description: Annotation, Visualization, and Impact Analysis; &ldquo;The tool is based on coupling a comprehensive annotation pipeline with a flexible visualization method. We leveraged the ANNOVAR (Wang et. al, 2010) framework for assigning functional impact to genomic variations by extending its list of reference annotation databases (RefSeq, UCSC, SIFT, Polyphen etc.) with additional in-house developed sources (Non-B DB, PolyBrowse).&rdquo;</li>
<li>Input: BED</li>
<li>Output: Table of annotations with gene annotation features</li>
</ul>
</li>
<li><strong><a href="http://bioinformaticstools.mayo.edu/research/bior/">BioR</a></strong>
<ul>
<li>Description: (Mayo Clinic) (Page last updated June 2015) Biological Reference Repository; &ldquo;data integration tool that enables coordinate based searches and joins based on strings&rdquo;; &ldquo;BioR consists of two parts 1) the BioR toolkit which depends on Java&hellip;. 2) the BioR catalogs which are the data files used by the system&rdquo;</li>
<li>Input: VCF</li>
<li>BioR-Supported Catalogs (tar-gzip files): dbSNP, 1000 genomes, HapMap, OMIM, NCBIGene</li>
<li>Output: VCF + JSON</li>
</ul>
</li>
<li><strong><a href="http://cadd.gs.washington.edu/">CADD</a></strong>
<ul>
<li>Description: Combined Annotation Dependent Depletion; tool for scoring SNV deletions/insertions; &ldquo;integrates multiple annotations into one metric&rdquo;; Score strongly correlates with allelic diversity and pathogenicity; links to 1000 Genome variants; uses Ensembl Variant Effect Predictor</li>
<li>Input: VCF</li>
<li>Output: CADD score</li>
</ul>
</li>
<li><strong><a href="http://www2.hu-berlin.de/wikizbnutztier/software/CandiSNPer/">CandiSNPer</a></strong>
<ul>
<li>Description: web application, characterizes SNPs located in vicinity of SNP of interest;</li>
<li>Input: enter SNP ID (rsID), choose population, region, measure for LD, threshold plot format, color of SNPs, and chose to show genes</li>
<li>Output: Imagefile</li>
</ul>
</li>
<li><strong><a href="https://github.com/UppsalaGenomeCenter/CanvasDB">CanvasDB</a></strong>
<ul>
<li>Description: &ldquo;local database infrastructure for analysis of targeted- and whole genome re-sequencing projects&rdquo;; dependent on MySQL, R, and ANNOVAR</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/carol/">CAROL</a></strong>
<ul>
<li>Description: (Wellcome Trust Sanger); Combined Annotation scoRing toOL; Combined functional annotation score of nonsynonymous coding variants; Combines information from PolyPhen-2 and SIFT</li>
<li>Input: tab-delimited with columns obtained from PolyPhen-2 and SIFT output</li>
<li>Output: tab-delimited file</li>
</ul>
</li>
<li><strong><a href="http://wiki.chasmsoftware.org/index.php/Main_Page">CHASM</a></strong>
<ul>
<li>Description: Cancer-specific High-throughput Annotation of Somatic Mutations; Last updated May 2014; uses Random Forest Method to &ldquo;distinguish between driver and passenger somatic mutations&rdquo;; Positive driver class curated from COSMIC database; packed together with SNVBox (database)</li>
<li>Input:Passenger mutation rates, Transcript and amino acid change, Genomic coordinates</li>
<li>Output: CHASM score, p-value, FDR</li>
</ul>
</li>
<li><strong><a href="http://www.cravat.us/">CRAVAT</a></strong>
<ul>
<li>Description: Cancer-Related Analysis of Variants Toolkit; Web application; Uses CHASM, VEST, SNVGet; &ldquo;CRAVAT provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases&rdquo; Latest Release May 2015;</li>
<li>Input: VCF, CRAVAT format</li>
<li>Output: CRAVAT report- MS Excel spreadsheet or tab-separated file (emailed)</li>
</ul>
</li>
<li><strong><a href="http://cupsat.tu-bs.de/">CUPSAT</a></strong>
<ul>
<li>Description: Cologne University Protein Stability Analysis Tool; &ldquo;tool to predict changes in protein stability upon point mutations&rdquo;; web service program; Can predict mutant stability from existing PDB structures or custom protein structures</li>
<li>Input:for PDB- provide PDB ID and Amino Acid Residue Number; for custom- PDB file format</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="https://cbcl.ics.uci.edu/public_data/DANN/">DANN</a></strong>
<ul>
<li>Description: Deleterious Annotation of genetic variants; standalone program, uses &ldquo;the same feature set and training data as CADD to train a deep neural network&rdquo;; can catch nonlinear relationships; &ldquo;There are four different datasets: training, validation, testing, and ClinVar_ESP...The ClinVar_ESP dataset is also a testing set containing a set of &ldquo;gold standard&rdquo; pathogenic and benign variants&rdquo;</li>
<li>Input:</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://rulai.cshl.edu/cgi-bin/tools/ESE3/esefinder.cgi?process=matrices">ESEfinder</a></strong>
<ul>
<li>Description: Exonic Splicing Enhancer; useful for interpretation of point mutations/polymorphisms that are disease-associated; GUI interface; web app program</li>
<li>Input: FASTA</li>
<li>Output: html or plain text format, graphical display of results</li>
</ul>
</li>
<li><strong><a href="http://www.sanger.ac.uk/resources/software/exomiser/">Exomiser</a></strong>
<ul>
<li>Description: Wellcome Trust Sanger; functionally annotates variants from whole-exome sequencing data; Based on Jannovar and uses UCSC KnownGene; Java program; web app program (Page last modified Feb 2015)</li>
<li>Input: VCF</li>
<li>Output: TSV, VCF</li>
</ul>
</li>
<li><strong><a href="https://sites.google.com/site/famannotation/home">FamAn</a></strong>
<ul>
<li>Description: Automated variant annotation pipeline for family-based sequencing studies; Annotaties SNVs and INDELs; 4 models- autosomal dominant, autosomal recessive, de novo mutations and a general model; &ldquo;A variety of annotations are provided for each segregating variant: number of family (and family ID) each variant hits, variant genomic location and coding effect (based on snpEff), loss-of-function mutation annotation, selected ENCODE annotation, allele frequency in the 1000 Genomes Project, allele frequency in the Exome Variant Server (ESP6500), segmental duplication annotation, SIFT, PolyPhen2, LRT, MutationTaster, GERP++, PhyloP, SiPhy, etc.&rdquo; (Last updated May 2014)</li>
<li>Input: VCF</li>
<li>Output: two excel compatible outputs</li>
</ul>
</li>
<li><strong><a href="http://www.gene-talk.de/">GeneTalk</a></strong>
<ul>
<li>Description: Combines tool for filtering and data analysis with an online network for genetic professionals; Different degrees- basic license, premium license, in-house solution (the last ones are paid for- Commercial tool?)</li>
<li>Input: VCF</li>
<li>Output: GeneTalk Annotation- includes clinical data, medical relevance, scientific relevance (<a href="http://www.gene-talk.de/public/GeneTalk_Whitepaper_Annotations.pdf">http://www.gene-talk.de/public/GeneTalk_Whitepaper_Annotations.pdf</a>)</li>
</ul>
</li>
<li><strong><a href="http://genevetter.kidneyomics.org/">GeneVetter</a></strong>
<ul>
<li>Description: &ldquo;GeneVetter is a tool designed for investigation of the background prevalence of exonic variation in the Phase 3 1000 Genomes data under user defined filtering criteria&rdquo;; web app program; GeneVetter uses GRch37p4 (hs37d5.fa.gz), dbSNP build 138, 1000G Phase 3, clinvar_2014072</li>
<li>Input: VCF</li>
<li>Output: TIMS score, summary table, PCA plot</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/software/cprg/?q=node/31">GSITIC</a></strong>
<ul>
<li>Description: (Broad Institute) Last update- July 2014; Identifies genomic regions that are significantly &ldquo;amplified or deleted&rdquo;; Each is given a G score; gives genomic locations and q-values from aberrant regions</li>
<li>Input: segmentation file -seg, markers file -mk (required); -array file list -alf, CNV file -cnv</li>
<li>Reference genome: -refgene (created in MATLAB, GISITIC provides four reference genomes: hg16.mat, hg17.mat, hg18.mat, hg19.mat</li>
<li>Output: All lesions file (text file), amplifications file (text file), deletion genes file (text file), Gistic Scores file, Segmented copy number (pdf file), amplification score GISTIC plot (pdf file), Deletion score/q-vale GISTIC plot (pdf file)</li>
</ul>
</li>
<li><strong><a href="http://www.cmbi.ru.nl/hope/about">HOPE</a></strong>
<ul>
<li>Description: Have yOur Protein Explained; Web app program; Automatic mutant analysis server that provides structural effects of a mutation; Uses BLAST against UniProt and PDB along with homology modeling</li>
<li>Input: FASTA protein sequence, or accession code of protein of interest</li>
<li>Output: a report containing information from a &ldquo;decision tree&rdquo; and illustrated figures and animations</li>
</ul>
</li>
<li><strong><a href="http://umd.be/HSF/">Human Splicing Finder</a></strong>
<ul>
<li>Description: Last update: May 2013; aimed to help study pre-mRNA splicing; combines 12 algorithms to identify mutations&rsquo; effect on splicing motifs; uses ensembl database 70</li>
<li>Input: Gene Name, Ensembl transcript ID, Ensembl Gene ID, Consensus CDS, RefSeq Peptide ID, or own sequence (looks like you can enter FASTA)</li>
<li>Output: Chart with columns for predicted signal, predicted algorithm, cDNA position and interpretation</li>
</ul>
</li>
<li><strong><a href="http://larva.gersteinlab.org/">LARVA</a></strong>
<ul>
<li>Description: Large-scale Analysis of Variants in noncoding Annotations; New version released July 2015; Command-line program; used for studying noncoding variants; integrates comprehensive set of noncoding elements, modeling their mutation count; Dependent on C++ and BEDtools</li>
<li>Input: multiple</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://www.jurgott.org/linkage/LinkagePC.html">LINKAGE</a></strong>
<ul>
<li>Description:three main programs: mlink (calculates lod scores at fixed values for the recombination fraction in one interval of a genetic map), linkmap (calculates location scores for positions of a disease locus along a marker), and ilink (estimates parameters including recombination fractions, allele frequencies, penetrances, etc)</li>
<li>Input: pedfile (processed by MAKEPED) and datafile (reflects loci for each individual; set in PREPLINK)</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://sourceforge.net/projects/mnvannotationcorrector/">MAC</a></strong>
<ul>
<li>Description: MNV Annotation Corrector; Ad hoc software, fixes incorrect amino acid predictions that are caused by multiple nucleotide variations; Uses existing annotators ANNOVAR, SnpEff, VEP (last update April 2015) (only 1 download this week &rarr; not popular)</li>
<li>Input: List of called SNVs and corresponding BAM</li>
<li>Output: Report identifying block of mutation within codon (BMCs)</li>
</ul>
</li>
<li><strong><a href="http://genome.igib.res.in/mitomatic/">mit-o-matic</a></strong>
<ul>
<li>Description: focuses on mtDNA, provides clinically relevant information from different resources; two component pipeline: command link for alignment of NGS reads and online version that provides genetic report on mitocondrial variants</li>
<li>Input:FASTQ, pileup</li>
<li>Reference sequence: rCRSm</li>
<li>Output: Online version gives comprehensive genetic report</li>
</ul>
</li>
<li><strong><a href="http://krauthammerlab.med.yale.edu/mutadelic/index.html">Mutadelic</a></strong>
<ul>
<li>Description: Web App program; &ldquo;This application generates reports on inherited mutations in five genes (ANK1, SLC4A1, SPTA1, SPTB and EPB42) associated with the following rare Mendelian blood disorders: Hereditary Spherocytosis (HS), Hereditary Elliptocytosis (HE) and Hereditary Pyropoikilocytosis&rdquo;; Newer program- recently validated on omictools</li>
<li>Input: Can upload coordinates of DNA variants or VEP</li>
<li>Output: Displayed on web or can be downloaded in Excel or RDF format</li>
</ul>
</li>
<li><strong><a href="http://www.mutationtaster.org/">MutationTaster</a></strong>
<ul>
<li>Description: (Last post on site 2014) Web app program; Rapid evaluation of disease causing alterations; uses NCBI 37 and Ensembl 69</li>
<li>Input: HGNC symbol, NCBI GeneID, or Ensembl ID,</li>
<li>Output: Report containing prediction, summary, name of alteration, etc</li>
</ul>
</li>
<li><strong><a href="http://mutpred.mutdb.org/">MutPred</a></strong>
<ul>
<li>Description: web app tool; Classifies amino acids substituation as disease associated or neutral in humans; Last modified Feb. 2014; Based on SIFT, trained using Human Gene Mutation Database</li>
<li>Input:</li>
<li>Output: &ldquo;The output of MutPred contains a general score (g), i.e., the probability that the amino acid substitution is deleterious/disease-associated, and top 5 property scores (p), where p is the P-value that certain structural and functional properties are impacted.&rdquo;</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/cancer/cga/mutsig">MutSigCV</a></strong>
<ul>
<li>Description: (Broad Institute) Mutation Significance (CV= covariates); Analyzes mutations discovered in DNA sequencing to identify genes that were mutated more often than expected</li>
<li>Input: mutations.maf, coverage.txt, covariates.txt</li>
<li>Output: output.txt</li>
</ul>
</li>
<li><strong><a href="http://stothard.afns.ualberta.ca/downloads/NGS-SNP/">NGS-SNP</a></strong>
<ul>
<li>Description: Collection of command-line scripts for providing rich SNP annotations; &ldquo;NCBI, Ensembl, and Uniprot IDs are provided for genes, transcripts and proteins when applicable&rdquo;;</li>
<li>Input: Samtools consensus pileup, Maq, diBayes, Genetic format, VCF</li>
<li>Output: File containing annotated SNPs is copied from SNP list and some classes are added</li>
</ul>
</li>
<li><strong><a href="http://www.broadinstitute.org/oncotator">Oncotator</a></strong>
<ul>
<li>Description: (Broad Institute) &ldquo;Tool for annotating human genomic point mutations and data relevant to cancer researchers&rdquo;; Web app; Supports annotation of data from ClinVar, dbSNP, 1000 genomes (plus many other external sites); Only GRCh27 coordinates supported; Last update: April 2015</li>
<li>Input: tal-delimited file</li>
<li>Output: tab-delimited MAF</li>
</ul>
</li>
<li><strong><a href="http://omictools.com/panther-s649.html">PANTHER</a></strong>
<ul>
<li>Description: Protein ANalysis THrough Evolutionary Relationships; Web app program, also has its own database; Classification system used to classify proteins and their genes; Also, &ldquo;Estimates the likelihood of a particular nonsynonymous (amino-acid changing) coding SNP to cause a functional impact on the protein&rdquo;; Updated in 2015</li>
<li>Input: Data from PANTHER, IDs from Ensembl, EntrezGene, NCBI GI numbers, NCBI UniGene IDs HUGO, UniProt; if ID type is not one of the above, can input txt file or excel format</li>
<li>Output: Analysis results displayed online</li>
</ul>
</li>
<li><strong><a href="http://cubio.biology.columbia.edu/pesx/pesx/">PESX</a></strong>
<ul>
<li>Description: Putative Exonic Splicing Enhancers/Silencers; (Can&rsquo;t tell if this is outdated or not)</li>
<li>Input: FASTA or plain text</li>
<li>Output: Excel spread sheet</li>
</ul>
</li>
<li><strong><a href="http://phen-gen.org/index.html">Phen-Gen</a></strong>
<ul>
<li>Description: Combines patient's&rsquo; disease symptoms with sequencing data; Standalone or Web app version; Only excepts 1 family per run, in order to evaluate unrelated individuals, each sample needs to be run individually</li>
<li>Input: Variant- VCF; Pheotype- HPO; Pedigree- PED</li>
<li>Output: Combined scores file, variants for top genes file</li>
</ul>
</li>
<li><strong><a href="http://mmb.pcb.ub.es/PMut/">PMUT</a></strong>
<ul>
<li>Description: Aimed at annotation and prediction of pathological mutations; based on different kinds of sequence info and neural networks to process information</li>
<li>Input: FASTA</li>
<li>Output; Simple yes/no and reliability index</li>
</ul>
</li>
<li><strong><a href="http://provean.jcvi.org/index.php">PROVEAN</a></strong>
<ul>
<li>Description: Protein Variation Effect Analyzer; predicts whether an amino acid substitution or indel has impact on biological function of the protein; &ldquo;comparable to SIFT or Polyphen-2&rdquo;; Standalone, Web app, Command line or GUI; Last update May 2014</li>
<li>Input: FASTA, list of variants;</li>
<li>Output: tab-separated columns including Variant, Provean Score and prediciton</li>
</ul>
</li>
<li><strong><a href="http://genes.mit.edu/burgelab/rescue-ese/">Rescue-ESE</a></strong>
<ul>
<li>Description: &ldquo;An online tool for identifying candidate ESEs in vertebrate exons&rdquo;; Web application; For human, mouse, zebrafish, pufferfish</li>
<li>Input: multi-FASTA or plain text</li>
<li>Output:</li>
</ul>
</li>
<li><strong><a href="http://scandb.org/newinterface/index_v1.html">SCAN</a></strong>
<ul>
<li>Description: Web application program, includes a database as well; Database contains physical-based SNP annotations and functional annotations; &ldquo;Information on physical, functional, and LD annotation served on the SCAN database comes directly from public resources, including the HapMap (release 23a), NCBI (dbSNP 129), or is information created by us using data downloaded from these public resources&rdquo;; &ldquo;SCAN can be utilized in several ways including: (i) queries of the SNP and gene databases; (ii) analysis using the attached tools and algorithms; (iii) downloading files with SNP annotation for various GWA platforms&rdquo;</li>
<li>Input:</li>
<li>Output: HTML, comma-delimited, tab-delimited</li>
</ul>
</li>
<li><strong><a href="http://snp.gs.washington.edu/SeattleSeqAnnotation137/">SeattleSeq Annotation</a></strong>
<ul>
<li>Description: &ldquo;SeattleSeqAnnotation137 was most recently updated October 13, 2013. The current version is 8.08. The most recent site, based on dbSNP build 141, and hg38/NCBI 38&rdquo;; Provides annotations for SNVs and Indels- includes dbSNP rsID, gene names and accession numbers, variation functions, protein positions and amino acid changes, conservation scores, HapMap frequencies, PolyPhen predictions and clinical association.</li>
<li>Input: Maq, gff, CASAVA, VCF, GATK bed, custom</li>
<li>Output: &ldquo;default output file format is a header line (starting with "#") followed by tab-separated annotations&rdquo;; VCF</li>
</ul>
</li>
<li><strong><a href="https://cran.r-project.org/web/packages/seqminer/">seqminer 3.7</a></strong>
<ul>
<li>Description: &ldquo;Efficiently Read Sequence Data (VCF Format, BCF Format and METAL Format) into R&rdquo;; Command line package program; Published August 2015</li>
<li>Input: VCF, BCF</li>
<li>Output: VCF</li>
</ul>
</li>
<li><strong><a href="https://genomics.scripps.edu/ADVISER/Home.jsp">SG Adviser</a></strong>
<ul>
<li>Description: Scripps Genome Annotation and Distributed Variant Interpretation Server, web developed applications for variant annotation, &ldquo;Downstream applications of variant annotation include: Clinical sequencing applications including: carrier testing, or identification of causal variants in molecular diagnosis, tumor sequencing, or diagnostic odyssey. Prioritization of variants prior to statistical analysis of sequence based disease association studies, especially for automated set-generation and enrichment of likely functional variants within sets. Identification of causal variants in post-GWAS/linkage sequencing studies. Identification of causal variants in forward genetic screens (stay tuned for non-human annotation)&rdquo;</li>
<li>Input: SNV- VCF, BED, and a few others; CNV- BED, CNVator, plus others</li>
<li>Output: tab-delimited file</li>
</ul>
</li>
<li><strong><a href="https://rostlab.org/services/snap/">SNAP-2</a></strong>
<ul>
<li>Descriptio</li></ul></li></ol>]]></description>
	<dc:creator>Jitendra Prajapati</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/31566/software-and-tools-to-detect-structure-variation-with-long-reads</guid>
	<pubDate>Wed, 15 Mar 2017 14:31:09 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/31566/software-and-tools-to-detect-structure-variation-with-long-reads</link>
	<title><![CDATA[Software and Tools to detect structure variation with long reads !!]]></title>
	<description><![CDATA[<p>Uncovering the connection between genetics and heritable diseases requires an approach that looks at all the variant bases and types in a genome. While a PacBio&nbsp;<em>de novo</em>&nbsp;assembly resolves the most novel SV variants. 8-10X PacBio coverage of single genomes or trios reveals triple the SVs detectable by short-read data.</p><p>With&nbsp;<span style="text-decoration: underline;"><a href="http://www.pacb.com/smrt-science/">Single Molecule, Real-Time (SMRT) Sequencing</a></span>, you can access structural variations having a broad range of sizes, types, and GC content with the ability to:</p><ul>
<li>Uncover missing heritability linked to structural variation</li>
<li>Unambiguously identify genomic context and variant breakpoints at the sequence level to unravel the genetic etiology of disease</li>
<li>Resolve structural variation across the complete size spectrum with basepair resolution</li>
</ul><p>Following are the SV tools, which can assist you to achieve your goal.</p><p><strong>Sniffles:</strong>&nbsp;Structural variation caller using third generation sequencing</p><p>Sniffles is a structural variation caller using third generation sequencing (PacBio or Oxford Nanopore). It detects all types of SVs using evidence from split-read alignments, high-mismatch regions, and coverage analysis. Please note the current version of Sniffles requires sorted output from BWA-MEM (use -M and -x parameter) or NGM-LR with the optional SAM attributes enabled!&nbsp;</p><p>More at&nbsp;https://github.com/fritzsedlazeck/Sniffles</p><p><strong style="font-size: 12.8px;"><br />MultiBreak-SV:</strong> It identifies structural variants from next-generation paired end data, third-generation long read data, or data from a combination of sequencing platforms.</p><p>There are two pieces of software in this release: (1) a pre-processor that takes machineformat (.m5) BLASR files, and (2) MultiBreak-SV. For installation and usage instructions, see doc/MultiBreakSV-Manual.txt.</p><p>More at&nbsp;https://github.com/raphael-group/multibreak-sv</p><p><strong style="font-size: 12.8px;"><br />Parliament:</strong>&nbsp;A Structural Variation Tool. Why ask a single sv-detection approach to find every variant when you can have a parliament of tools deciding?</p><p>Publication about the algorithm and &ldquo;&hellip;the first long-read characterization of structural variation in a diploid human personal genome&hellip;&rdquo; (HS1011) -&nbsp;<a href="http://www.biomedcentral.com/1471-2164/16/286">&ldquo;Assessing structural variation in a personal genome&mdash;towards a human reference diploid genome&rdquo;</a></p><p>More at&nbsp;https://sourceforge.net/projects/parliamentsv/</p><p>https://www.dnanexus.com/papers/Parliament_Info_Sheet.pdf</p><p><br /><strong>PBHoney:</strong>&nbsp;the structural variation discovery tool&nbsp;<br /><br />PBHoney is an implementation of two variant-identification approaches designed to exploit the high mappability of long reads (i.e., greater than 10,000 bp). PBHoney considers both intra-read discordance and soft-clipped tails of long reads to identify structural variants.</p><p>Read The Paper&nbsp;<a href="http://www.biomedcentral.com/1471-2105/15/180/abstract" target="_blank">http://www.biomedcentral.com/1471-2105/15/180/abstract</a></p><p>More at&nbsp;https://sourceforge.net/projects/pb-jelly/</p><p><strong><br />SMRT-SV:</strong> Structural variant and indel caller for PacBio reads</p><p>Structural variant (SV) and indel caller for PacBio reads based on methods from&nbsp;<a href="http://www.nature.com/nature/journal/vaop/ncurrent/full/nature13907.html">Chaisson et al. 2014</a>.</p><p>SMRT-SV provides an official software package for tools described in&nbsp;<a href="http://www.nature.com/nature/journal/vaop/ncurrent/full/nature13907.html">Chaisson et al. 2014</a>&nbsp;and adds several key features including the following.</p><ul>
<li>Unified variant calling user interface with built-in cluster compute support</li>
<li>Small indel calling (2-49 bp)</li>
<li>Improved inversion calling (<code>screenInversions</code>)</li>
<li>Quality metric for SV calls based on number of local assemblies supporting each call</li>
<li>Higher sensitivity for SV calls using tiled local assemblies across the entire genome instead of "signature" regions</li>
<li>Genotyping of SVs with Illumina paired-end reads from WGS samples</li>
</ul><p>More at&nbsp;https://github.com/EichlerLab/pacbio_variant_caller</p>]]></description>
	<dc:creator>Archana Malhotra</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/researchlabs/view/25993/hoffman-lab</guid>
  <pubDate>Tue, 12 Jan 2016 02:47:41 -0600</pubDate>
  <link></link>
  <title><![CDATA[Hoffman Lab]]></title>
  <description><![CDATA[
<p>They develop machine learning techniques to better understand chromatin biology. These models and algorithms transform high-dimensional functional genomics data into interpretable patterns and lead to new biological insight.</p>

<p>https://www.pmgenomics.ca/hoffmanlab/</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/1161/genomics-for-bioinformatician</guid>
	<pubDate>Sat, 20 Jul 2013 07:03:00 -0500</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/1161/genomics-for-bioinformatician</link>
	<title><![CDATA[Genomics for Bioinformatician]]></title>
	<description><![CDATA[<p>Genomics is the study of the genomes of organisms. The field includes intensive efforts to determine the entire DNA sequence of organisms and fine-scale genetic mapping efforts. The field also includes studies of intragenomic phenomena such as heterosis, epistasis, pleiotropy and other interactions between loci and alleles within the genome. In contrast, the investigation of the roles and functions of single genes is a primary focus of molecular biology or genetics and is a common topic of modern medical and biological research. Research of single genes does not fall into the definition of genomics unless the aim of this genetic, pathway, and functional information analysis is to elucidate its effect on, place in, and response to the entire genome's networks.<br /><br />Genomics was established by Fred Sanger when he first sequenced the complete genomes of a virus and a mitochondrion. His group established techniques of sequencing, genome mapping, data storage, and bioinformatic analyses in the 1970-1980s. A major branch of genomics is still concerned with sequencing the genomes of various organisms, but the knowledge of full genomes has created the possibility for the field of functional genomics, mainly concerned with patterns of gene expression during various conditions. The most important tools here are microarrays and bioinformatics. Study of the full set of proteins in a cell type or tissue, and the changes during various conditions, is called proteomics. A related concept is materiomics, which is defined as the study of the material properties of biological materials (e.g. hierarchical protein structures and materials, mineralized biological tissues, etc.) and their effect on the macroscopic function and failure in their biological context, linking processes, structure and properties at multiple scales through a materials science approach. The actual term 'genomics' is thought to have been coined by Dr. Tom Roderick, a geneticist at the Jackson Laboratory (Bar Harbor, ME) over beer at a meeting held in Maryland on the mapping of the human genome in 1986.<br /><br />The outcome of almost two years of intense discussions with literally hundreds of scientists and members of the public, has three major areas of focus: Genomics to Biology, Genomics to Health, and Genomics to Society.<br /><br /><strong><em>Genomics to Biology:</em></strong>&nbsp;<br />The human genome sequence provides foundational information that now will allow development of a comprehensive catalog of all of the genome's components, determination of the function of all human genes, and deciphering of how genes and proteins work together in pathways and networks.<br /><br /><strong><em>Genomics to Health:<br /></em></strong>Completion of the human genome sequence offers a unique opportunity to understand the role of genetic factors in health and disease, and to apply that understanding rapidly to prevention, diagnosis, and treatment. This opportunity will be realized through such genomics-based approaches as identification of genes and pathways and determining how they interact with environmental factors in health and disease, more precise prediction of disease susceptibility and drug response, early detection of illness, and development of entirely new therapeutic approaches.<br /><br /><strong><em>Genomics to Society:</em>&nbsp;<br /></strong>Just as the HGP has spawned new areas of research in basic biology and in health, it has created new opportunities in exploring the ethical, legal, and social implications (ELSI) of such work. These include defining policy options regarding the use of genomic information in both medical and non-medical settings and analysis of the impact of genomics on such concepts as race, ethnicity, kinship, individual and group identity, health, disease, and "normality" for traits and behaviors.<br /><br />This vision for the future of genomics is not just about the NHGRI. It encompasses the whole field of genomics, including the work of all the other Institutes and Centers at the NIH and of a number of other federal agencies. All of the NIH Institutes are already taking full advantage of the sequence and will apply its data to the better understanding of both rare and common diseases, almost all of which have a genetic component. A recent example of the way that the HGP and the knowledge and new technologies it has spawned are already facilitating science is the extremely rapid sequencing by groups in Canada and at the Centers for Disease Control and Prevention (CDC) in Atlanta of the genome of the virus that causes Severe Acute Respiratory Syndrome (SARS). The sequencing of the SARS virus genome provides insight into this new and deadly disease at a speed never before possible in science. In turn, this should lead to the rapid development of diagnostic tests and, in time, vaccines and effective treatments.<br /><br /><strong>Links for the addition material available on Net</strong></p><p><a href="http://pevsnerlab.kennedykrieger.org/bioinformatics/bioinf10_genomes.htm">Genomes and genomics:</a></p><p><a href="http://www.123genomics.com/learning.html">Bioinformatics and Genomics:</a></p><p><a href="http://www.ebi.ac.uk/pdbe/docs/roadshow_tutorial/strgenomics/tutorial.html">Structural genomics tutorial:</a></p><p><a href="http://www.hgu.mrc.ac.uk/Users/Philippe.Gautier/tutorial/index.html">Comparative Genomics Tutorial:</a></p><p><a href="http://www.scfbio-iitd.res.in/tutorial/genomics.html">GENOME TUTORIAL:</a></p><p><a href="http://genomebiology.com/content/pdf/gb-2001-3-1-reviews2001.pdf">Tools and resources for identifying protein families, domains and motifs</a></p><p><a href="http://www.ornl.gov/sci/techresources/Human_Genome/posters/chromosome/tools.shtml">Bioinformatics Tools</a><a href="http://www.ornl.gov/sci/techresources/Human_Genome/posters/chromosome/tools.shtml">&nbsp;<br />Tips, Tutorials, and Terminology for Using Selected Resources in Genome Database Guide:</a></p><p><a href="http://www.doe-mbi.ucla.edu/Reprints/R31%20Strong%20A%20Web-based%20Comparative%20Genomics%20tutorial%20Microbiology%20Eduction%202004.pdf">A Web-Based Comparative Genomics Tutorial for Investigating Microbial Genomes:</a></p><p><a href="http://www.genome.gov/27530225">Free Online Tutorials Teach Anyone How to Use Genome Databases:</a></p><p><a href="http://mkweb.bcgsc.ca/circos/?tutorials">Circos to create concise, explanatory, unique and print-ready visualizations of your data:</a></p><p><a href="http://www.igd.cornell.edu/Comparative%20Genomics/Comparative%20Genomics%20Proj.html">Genomics and Comparative Genomics</a><a href="http://www.igd.cornell.edu/Comparative%20Genomics/Comparative%20Genomics%20Proj.html">&nbsp;Learning Module:</a></p><p><a href="http://psb.stanford.edu/psb10/conference-materials/tutorials/compgen-notes.pdf">Computational Challenges in Comparative Genomics</a></p><p><a href="http://psb.stanford.edu/psb10/conference-materials/tutorials/compgen-notes.pdf">A Tutorial:</a></p><p><a href="http://gramene.agrinome.org/tutorials/modules_tutorial.pdf">A Comparative Genomics Resource for Grains</a>:</p><p><a href="http://www.plantcell.org/cgi/content/full/21/12/3718">PLAZA: A Comparative Genomics Resource to Study Gene and Genome Evolution in Plants:</a></p><p><a href="http://en.wikipedia.org/wiki/VISTA_(comparative_genomics)">VISTA</a><a href="http://en.wikipedia.org/wiki/VISTA_(comparative_genomics)">:</a></p><p>Software for Genomics</p><ol>
<li><strong>Artemis</strong>&nbsp;Artemis is a free genome viewer and annotation tool that allows visualization of sequence features and the results of analyses within the context of the sequence, and its six-frame translation.</li>
<li><strong>Chromas&nbsp;</strong>It will display and prints chromatogram files from ABI automated DNA sequencers, and Staden SCF files which the analysis programs for ALF, Li-Cor and Visible Genetics OpenGene sequencers can create.</li>
<li><strong>Glimmer</strong>&nbsp;A system for finding genes in microbial DNA, especially the genomes of bacteria and archaea.Glimmer (Gene Locator and Interpolated Markov Modeler) uses interpolated Markov models (IMMs) to identify the coding regions and distinguish them from noncoding DN</li>
<li><strong>Glimmer</strong>&nbsp;HMM&nbsp;A fast and accurate gene finder based on a GHMM architecture, developed specifically for eukaryotes. It incorporates splice site models adapted from the GeneSplicer program and uses interpolated Markov models for evaluating the coding regions.</li>
<li><strong>Glimmer</strong>&nbsp;M&nbsp;A gene finder derived from Glimmer, but developed specifically for eukaryotes. It is based on a dynamic programming algorithm that considers all combinations of possible exons for inclusion in a gene model and chooses the best of these combinations. The d</li>
<li><strong>MUMmer</strong>&nbsp;MUMmer is a system for rapidly aligning entire genomes, whether in complete or draft form.</li>
<li><strong>pDRAW</strong>&nbsp;pDRAW32 is being developed as a free time hobby project. It is far from finished, but as it has reached a point where it could be helpful for many labs, it is now available to the scientific community.</li>
<li><strong>Sequin</strong>&nbsp;Sequin is a stand-alone software tool developed by the NCBI for submitting and updating entries to the GenBank, EMBL, or DDBJ sequence databases. It is capable of handling simple submissions that contain a single short mRNA sequence, and complex submissio</li>
<li><strong>Staden&nbsp;</strong>The Staden Package consists of a series of tools for DNA sequence preparation (pregap4), assembly (gap4), editing (gap4) and DNA/protein sequence analysis (spin).</li>
</ol><p>For more software @&nbsp;<a href="http://bioinformaticsonline.com/bookmarks/view/926/list-of-popular-bioinformatics-softwaretools">http://bioinformaticsonline.com/bookmarks/view/926/list-of-popular-bioinformatics-softwaretools</a></p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/1295/five-points-for-bioinformatics-softwaretools</guid>
	<pubDate>Mon, 05 Aug 2013 04:12:32 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/1295/five-points-for-bioinformatics-softwaretools</link>
	<title><![CDATA[Five points for bioinformatics software/tools]]></title>
	<description><![CDATA[<p><span>In the bioinformatics sector we mostly spend time on computational analysis of huge amounts of data and try to make sense of it, biologically. But, most of the newbie bioinformaticians are faced with dilemma when they receive biological sequence data for the first time. They mostly found confusing over open source, user friendly GUI, and commercial bioinformatics software. Don&rsquo;t be surprise this is true and also not an easy task to decide, because analytical step is the most crucial part and believe to be the biggest bottleneck in publishing paper in high impact journals. Through this blog I would like to address the pros and cons of both kind of software/tools and try to assist (Hmmm not really, It looks convince) you to make decision on your software selections.</span></p><p><span><img src="http://bioinformaticsonline.com/mod/photo/five.jpg" alt="image" style="border: 0px;"></span></p><p><span>The most common newbie questions are:</span><span></span></p><p><span>Should I try to use these free open source programs? &nbsp;Why are we not trying GUI software for computational analysis? Should I use commercial bioinformatics programs/software?&rdquo;</span><span><br /></span><span><br />1. Let&rsquo;s be open</span><span></span></p><p><span>We generally think free and cheap are useless. But this concept is not applicable when we discuss open source software. Mostly, the bioinformatics software is developed by highly competitive biological programmers who believe in open sharing of knowledge. They come under Open Bioinformatics Foundation or O|B|F which is a non-profit, volunteer run organization focused on supporting open source programming in bioinformatics. The best part about open source tools/software is that they&rsquo;re free to download the source code and read exactly what the program does. If you are so inclined, you can view all of the parts of the program and see the logical flow of the pipeline. In addition, open source makes an excellent learning tool for any beginning bioinformatician. Moreover, you can modify existing open source programs to deal with cutting-edge problems or to customize your pipeline.</span><span>&nbsp;</span><span>Apart from your computational and analysis work, most of the reviewer also prefers the open source based results so that they can validate the results if validation required.</span></p><p><span>2. Code headache</span><span></span></p><p><span>As a bioinformatician you are supposed to know the basics of programming languages, and if you are not good at it, then please learn it as soon as possible because you are not a bio-analyst but biological programmers. The<span>&nbsp;</span>open source programs usually lack dedicated service and support teams (often because they were the product of an overworked doc/postdoc!) so you are responsible for troubleshooting your own errors most of the time.<span>&nbsp;</span>We commonly receive the HELP email to support and assist to setup the pipeline; you can also find this kind of request on any QA forum. I personally believe this coding horror brings the biggest downside of open-source programs; where you need some programming skills in order to implement the program in your pipeline. But, if you are not able to fix the pipeline and modify the open source code according to your requirements them you should re-think on your bioinformatician name tag!!!</span><span></span></p><p><span>3. Dive into the codes</span><span></span></p><p><span>Some of the biologist turn bioinformatician says &ldquo;if you can do the same thing with commercial software then why to get migraine with weird codes&rdquo;, well this statement looks to me that guys are keen to learn swimming but still don&rsquo;t like to get wet. If you are still using paid software and doing your work by customer support and clicking some of the well-designed GUI button then perhaps you are not interested in learning and trying new and challenging bioinformatics works. You are missing the basic flavour of bioinformatics. Let&rsquo;s dive into the coding world, I am sure your will enjoy it. I recommend your to swim freely in code&rsquo;s sea, and enjoy the journey; do not merely watch it from the outside. &nbsp;</span></p><p><span>4. Paid does not mean better</span><span></span></p><p><span>The bioinformatics company which are specializes in bioinformatics solutions develop well designed/packed, user friendly software by using a large number of specialised scientist, programmers and support staff. They also provide good services to accomplice your biological analysis work. This means that if you hit a &lsquo;snag&rsquo; with your data, help is likely only a phone call away! These companies price their products competitively against the cost of a dedicated bioinformatician. You may be able to afford the program, but not the additional staff! Additionally, most of the functionality that you need in your analysis is already coded into the program. Need to plot a graph? Just click this button right here. It is that easy.</span><span>&nbsp;</span><span>But, as a bioinformatician this is not generally well encouraged approach in biological analysis work, because the software is not available to everyone and your data can&rsquo;t be validated. Moreover, there is very less chances that anyone will repeat your work or love to do similar kind of research (because not all the labs in the world are rich like yours).</span></p><p><span>5. Take a caution<br /><br />In biological analysis work, in which you deal GB/TB of data are having maximum chances of getting errors, so please be careful and always cross check your data before coming to any conclusion. Even an error in two line code can alter your entire analysis and display weird results. Some of the scientist blindly believes on commercial software, which is entirely wrong. Using proprietary tools does not absolve you of the need to actually read and research the type of analysis that you are doing. This is particularly true in the case of genome assembly and annotation.</span></p><p><span><br />At the end, I would like to tell only one think that open source solutions allows you to do more cutting edge analysis than the commercial tools. So let&rsquo;s go for it.</span></p><p>Disclaimer:</p><p>This is my personal view. I have nothing to do with any company or open source community.&nbsp;The views expressed on these pages are mine alone and not those of my current/past employers. I do reserve the right to remove comments left by spammers or off-topic comments.</p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/1469/prime-minister%E2%80%99s-100k-genome-project</guid>
	<pubDate>Thu, 08 Aug 2013 09:40:39 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/1469/prime-minister%E2%80%99s-100k-genome-project</link>
	<title><![CDATA[Prime Minister’s 100k Genome Project]]></title>
	<description><![CDATA[<p>Genomics Ebgland is destined to sequence 100,000 patients over the next five year in England.&nbsp; A landmark project by british government.</p><p>Genomics England will play a key role in building on the UK&rsquo;s long track record as leader in medical science advances to push the boundaries by unlocking the power of DNA data. The UK will become the first ever country to introduce this technology in its mainstream health system &ndash; leading the global race for better tests, better drugs and above all better, more personalised care.</p><p>http://www.genomicsengland.co.uk/100k-genome-project/</p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/view/2021</guid>
	<pubDate>Mon, 12 Aug 2013 09:27:57 -0500</pubDate>
	<link>https://bioinformaticsonline.com/view/2021</link>
	<title><![CDATA[What are the difference between BioRuby and BioGem?]]></title>
	<description><![CDATA[<p>I came across two diferent but matching term BioRuby and BioGem. What are the difference between these two term? If both are using same Ruby language for development then why did they develope two different biological packages.</p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/videolist/watch/4090/computational-biology-in-the-21st-century-making-sense-out-of-massive-data</guid>
	<pubDate>Thu, 29 Aug 2013 08:32:26 -0500</pubDate>
	<link>https://bioinformaticsonline.com/videolist/watch/4090/computational-biology-in-the-21st-century-making-sense-out-of-massive-data</link>
	<title><![CDATA[Computational Biology in the 21st Century: Making Sense out of Massive Data]]></title>
	<description><![CDATA[<iframe width="" height="" src="https://www.youtube-nocookie.com/embed/I99UiA_vaJQ" frameborder="0" allowfullscreen></iframe>Computational Biology in the 21st Century: Making Sense out of Massive Data    
    
Air date:  Wednesday, February 01, 2012, 3:00:00 PM
Category:  Wednesday Afternoon Lectures  
 
Description:  The last two decades have seen an exponential increase in genomic and biomedical data, which will soon outstrip advances in computing power to perform current methods of analysis. Extracting new science from these massive datasets will require not only faster computers; it will require smarter algorithms. We show how ideas from cutting-edge algorithms, including spectral graph theory and modern data structures, can be used to attack challenges in sequencing, medical genomics and biological networks. 

The NIH Wednesday Afternoon Lecture Series includes weekly scientific talks by some of the top researchers in the biomedical sciences worldwide. 

Author:  Dr. Bonnie Berger  
Runtime:  00:58:06  
Permanent link:  http://videocast.nih.gov/launch.asp?17563]]></description>
	
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/2334/binc-bioinformatics-national-certification-website-address</guid>
	<pubDate>Wed, 14 Aug 2013 09:40:22 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/2334/binc-bioinformatics-national-certification-website-address</link>
	<title><![CDATA[BINC (BioInformatics National Certification) Website address]]></title>
	<description><![CDATA[<p><span>BINC (BioInformatics National Certification) is an initiative of Department of Biotechnology(DBT), Government Of India in coordination with Bioinformatics Center, University of Pune. The objective of the examination is to recognize trained manpower in the area of Bioinformatics. Currently, various Indian universities, Government and private institutions are involved in imparting courses in Bioinformatics in India.</span></p>
<p>Foreign nationals intending to have certification are eligible to appear for BINC examination.<br>Minimum qualification includes a degree from a recognized university/institute in the areas listed in FAQ.<br>Formal training in the area of Bioinformatics is not a prerequisite.<br>Note that the foreign students will only be certified by DBT and are not eligible for the cash award as well as junior research fellowship.</p><p>Address of the bookmark: <a href="http://binc.scisjnu.ernet.in/" rel="nofollow">http://binc.scisjnu.ernet.in/</a></p>]]></description>
	<dc:creator>Kamalakshi Mukherjee</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/7674/useful-publications-and-websites-for-deep-sequencing-data-analysis</guid>
	<pubDate>Sun, 29 Dec 2013 22:30:45 -0600</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/7674/useful-publications-and-websites-for-deep-sequencing-data-analysis</link>
	<title><![CDATA[Useful Publications and Websites for Deep Sequencing Data Analysis]]></title>
	<description><![CDATA[<h3>Global overview papers</h3><p>Next generation quantitative genetics in plants. Jim&eacute;nez-G&oacute;mez, Frontiers in Plant Science 2:77, 2011 <span style="text-decoration: underline;"><a href="http://www.frontiersin.org/Plant_Physiology/10.3389/fpls.2011.00077/full">Full Text</a> </span><em>[equally relevant to animal and microbial systems]</em></p><p>Sense from sequence reads: methods for alignment and assembly. Flicek &amp; Birney, Nat Methods 6(11 Suppl):S6-S12, 2009. <a href="http://www.nature.com/nmeth/journal/v6/n11s/full/nmeth.1376.html"><span style="text-decoration: underline;">Full Text</span></a></p><h3>Library construction and experimental design</h3><p>Statistical design and analysis of RNA sequencing data. Auer &amp; Doerge, Genetics 185(2):405-16, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2881125"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Biases in Illumina transcriptome sequencing caused by random hexamer priming. Hansen et al., Nucleic Acids Res. 38(12): e131, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2896536"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries. Aird et al, Genome Biology 12:R18, 2011 <a href="http://genomebiology.com/2011/12/2/R18"><span style="text-decoration: underline;">Full Text</span></a></p><p>Amplification-free Illumina sequencing-library preparation facilitates improved mapping and assembly of GC-biased genomes. Kozarewa et al, Nature Methods 6(4):291-5, 2009 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2664327/"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Cost-effective, high-throughput DNA sequencing libraries for multiplexed target capture. Rohland &amp; Reich, Genome Research 22(5): 939&ndash;946. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3337438/"><span style="text-decoration: underline;">PubMedCentral</span></a></p><h3>Data formats, data management, and alignment software tools<span style="text-decoration: underline;"> </span></h3><p>The Sequence Alignment/Map format and SAMtools. Li et al, Bioinformatics 25(16):2078-9, 2009 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2723002"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>SAM format specification <a href="http://samtools.sourceforge.net/SAM1.pdf"><span style="text-decoration: underline;">file</span></a></p><p>Efficient storage of high throughput sequencing data using reference-based compression. Fritz et al, Genome Res 21(5):734-40, 2011. <a href="http://genome.cshlp.org/content/21/5/734.long"><span style="text-decoration: underline;">Full Text</span></a></p><p>Compression of DNA sequence reads in FASTQ format. Deorowicz &amp; Grabowski, Bioinformatics 27(6):860-2, 2011. <a href="http://www.ncbi.nlm.nih.gov/pubmed/21252073"><span style="text-decoration: underline;">PubMed</span></a></p><p>Fast and accurate short read alignment with Burrows-Wheeler transform. Li &amp; Durbin, Bioinformatics 25(14):1754-60, 2009. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2705234"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Improving SNP discovery by base alignment quality. Li H, Bioinformatics 27(8):1157-8, 2011. <a href="http://www.ncbi.nlm.nih.gov/pubmed/21320865"><span style="text-decoration: underline;">PubMed</span></a></p><p>BEDTools: a flexible suite of utilities for comparing genomic features. Quinlan and Hall, Bioinformatics 26:841-842, 2010. <a href="http://bioinformatics.oxfordjournals.org/content/26/6/841.full.pdf+html"><span style="text-decoration: underline;">Publisher Website</span></a></p><h3>Data quality assessment, filtering, and correction</h3><p>SolexaQA: At-a-glance quality assessment of Illumina second-generation sequencing data. Cox et al, BMC Bioinformatics 11:485, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2956736"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>TileQC: a system for tile-based quality control of Solexa data. Dolan &amp; Denver, BMC Bioinformatics 9:250, 2008 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2443380"><span style="text-decoration: underline;">PubMedCentral</span></a> <em>[requires a reference sequence]</em></p><p>Quake: quality-aware detection and correction of sequencing errors. Kelley et al, Genome Biol 11(11):R116, 2010. <a href="http://www.ncbi.nlm.nih.gov/pubmed/21114842"> <span style="text-decoration: underline;">PubMed</span></a></p><p>FastQC: a quality control tool for high-throughput sequence data. <a href="http://www.bioinformatics.bbsrc.ac.uk/projects/fastqc/"><span style="text-decoration: underline;">Home Page</span></a></p><p>FASTX-toolkit: FASTQ/A short-reads pre-processing tools <a href="http://hannonlab.cshl.edu/fastx_toolkit/"><span style="text-decoration: underline;">Home Page</span></a></p><p>Reference-free validation of short read data. Schr&ouml;der et al, PLoS One 5(9):e12681, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2943903"> <span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Correction of sequencing errors in a mixed set of reads. Salmela, Bioinformatics 26(10):1284, 2010. <a href="http://bioinformatics.oxfordjournals.org/content/26/10/1284.long"><span style="text-decoration: underline;">Full Text</span></a> <em>[includes error correction of SOLiD reads in colorspace]</em></p><p>Repeat-aware modeling and correction of short read errors. Yang et al, BMC Bioinformatics 12(Supp1):S52, 2011 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3044310"> <span style="text-decoration: underline;">PubMedCentral</span></a> <em>[requires a reference sequence]</em></p><p>HiTEC: accurate error correction in high-throughput sequencing data. Ilie et al, Bioinformatics 27(3):295, 2011 <a href="http://bioinformatics.oxfordjournals.org/content/27/3/295.long"><span style="text-decoration: underline;">Full Text</span></a></p><p>Error correction of high-throughput sequencing datasets with non-uniform coverage. Medvedev et al., Bioinformatics 27(13):i137-41, 2011. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3117386"><span style="text-decoration: underline;">PubMedCentral</span></a></p><h3>De novo assembly<span style="text-decoration: underline;"> </span></h3><p>Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Zerbino &amp; Birney, Genome Res 18(5):821-9, 2008. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2336801">u&gt;PubMedCentral</a></p><p>Assembly of large genomes using second-generation sequencing. Schatz et al, Genome Res 20(9):1165-73, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2928494"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>High-quality draft assemblies of mammalian genomes from massively parallel sequence data. Gnerre et al, PNAS 108(4): 1513-18, 2011 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3029755"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Genome assembly has a major impact on gene content: a comparison of annotation in two <em>Bos taurus </em> assemblies. Florea&nbsp; et al., PLoS One 6(6):e21400, 2011. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3120881/"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Artemis: an integrated platform for visualization and analysis of high-throughput sequence-based experimental data. Carver et al, Bioinformatics 28(4):464 - 469, 2012 <span style="text-decoration: underline;"><a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3278759/">PubMedCentral</a></span></p><p>Efficient de novo assembly of large genomes using compressed data structures. Simpson &amp; Durbin, Genome Research 22:549-556, 2012 <span style="text-decoration: underline;"><a href="http://genome.cshlp.org/content/22/3/549.full">Full Text</a></span> <em>[Describes the String Graph Assembler (SGA), which assembled a human genome in less than 6 days using 54 Gb of RAM and a 123-processor compute cluster for calculation of an FM-index of the 1.2 billion reads]</em></p><p>Readjoiner: a fast and memory efficient string graph-based sequence assembler. Gonnella &amp; Kurtz, BMC Bioinformatics 13: 82, 2012 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3507659"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Assemblathon 1: A competitive assessment of de novo short read assembly methods. Earl et al, Genome Research 21:2224-2241, 2011 <span style="text-decoration: underline;"><a href="http://genome.cshlp.org/content/early/2011/09/16/gr.126599.111.full.pdf+html">Full Text</a></span></p><h3>Chromatin immunoprecipation analysis: ChIP-seq</h3><p>ChIP-seq: advantages and challenges of a maturing technology. Park, Nat Rev Genet. 10:669-80, 2009 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3191340/"><span style="text-decoration: underline;">PubMed</span></a></p><p>ChIP-seq and Beyond: new and improved methodologies to detect and characterize protein-DNA interactions. Furey, Nat Rev Genet 13: 840&ndash;852, 2012 <a href="http://www.nature.com/nrg/journal/v13/n12/full/nrg3306.html"> <span style="text-decoration: underline;">Publisher Web Site</span></a></p><p>MuMoD: a Bayesian approach to detect multiple modes of protein&ndash;DNA binding from genome-wide ChIP data. Narlikar, Nucleic Acids Res 41:21&ndash;32, 2013 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3592440/"><span style="text-decoration: underline;">PubMed</span></a></p><h3>Transcriptome analysis</h3><h3>Assembly and comparison to genome</h3><p>Full-length transcriptome assembly from RNA-Seq data without a reference genome. Grabherr et al, Nature Biotechnology 29:644 - 652, 2011. <a href="http://www.ncbi.nlm.nih.gov/pubmed/21572440"><span style="text-decoration: underline;">PubMed</span></a> <em>[The software is called <a href="http://trinityrnaseq.sourceforge.net/"><span style="text-decoration: underline;">Trinity</span></a>, and is available on Sourceforge.]</em></p><p>Comprehensive analysis of RNA-Seq data reveals extensive RNA editing in a human transcriptome. Peng et al, Nature Biotechnology 30:253 - 260, 2012. <span style="text-decoration: underline;"><a href="http://www.ncbi.nlm.nih.gov/pubmed/22327324">PubMed</a></span> <em>[Several comments on this paper question whether the reported differences are in fact evidence of editing or are simply sequencing errors - the authors stand by their conclusions, but the controversy demonstrates the importance of robust data analysis methods.] </em></p><p>Optimization of de novo transcriptome assembly from next-generation sequencing data. Surget-Groba &amp; Montoya-Burgos, Genome Res 20(10):1432-40, 2010. <a href="http://genome.cshlp.org/content/20/10/1432.long"><span style="text-decoration: underline;">Full Text</span></a></p><p>Rnnotator: an automated <em>de novo</em> transcriptome assembly pipeline from stranded RNA-Seq reads. Martin et al, BMC Genomics 11:663, 2010 <a href="http://www.biomedcentral.com/1471-2164/11/663"><span style="text-decoration: underline;">Full Text</span></a></p><p><em>De novo</em> assembly and analysis of RNA-seq data. Robertson et al, Nature Methods 7:909-912, 2010 <a href="http://www.nature.com/nmeth/journal/v7/n11/full/nmeth.1517.html"><span style="text-decoration: underline;">Full Text</span></a> <em>[describes Trans-ABySS, a pipeline to use the ABySS parallel assembler for de novo transcriptome analysis]</em></p><h3>Differential expression analysis</h3><p>R-SAP: a multi-threading computational pipeline for the characterization of high-throughput RNA-sequencing data. Mittal &amp; McDonald, Nucleic Acids Res, 2012 <span style="text-decoration: underline;"><a href="http://nar.oxfordjournals.org/content/early/2012/01/28/nar.gks047.long">Full Text</a></span></p><p>Targeted RNA sequencing reveals the deep complexity of the human transcriptome. Mercer et al, Nature Biotechnology 30:99 - 104, 2012 <span style="text-decoration: underline;"><a href="http://www.nature.com/nbt/journal/v30/n1/full/nbt.2024.html"> Publisher Website</a></span></p><p>Differential gene and transcript expression analysis of RNA-Seq experiments with TopHat and Cufflinks. Trapnell et al, Nature Protocols 7:562 - 578, 2012 <span style="text-decoration: underline;"><a href="http://www.nature.com/nprot/journal/v7/n3/full/nprot.2012.016.html"> Publisher Website</a></span></p><p>Characterization and improvement of RNA-Seq precision in quantitative transcript expression profiling. Łabaj et al, Bioinformatics 27:i383 - i391, 2011 <span style="text-decoration: underline;"><a href="http://bioinformatics.oxfordjournals.org/content/27/13/i383.full.pdf+html"> Full Text</a></span></p><p>Improving RNA-Seq expression estimates by correcting for fragment bias. Roberts et al, Genome Biol 12:R22, 2011 <span style="text-decoration: underline;"><a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3129672/">PubMed Central</a></span></p><p>Cloud-scale RNA-sequencing differential expression analysis with Myrna. Langmead et al, Genome Biol 11:R83, 2010 <a href="http://genomebiology.com/2010/11/8/R83"><span style="text-decoration: underline;">Full Text</span></a></p><p>From RNA-seq reads to differential expression results. Oshlack et al, Genome Biol 11(12):220, 2010 <a href="http://genomebiology.com/content/11/12/220"><span style="text-decoration: underline;">Full Text</span></a></p><p>DEGseq: an R package for identifying differentially expressed genes from RNA-seq data. Wang et al., Bioinformatics. 26(1):136-8. 2010 <a href="http://www.ncbi.nlm.nih.gov/pubmed/19855105"><span style="text-decoration: underline;"> PubMed</span></a></p><p>DEseq: Differential expression analysis for sequence count data. Anders and Huber, Genome Biology 11:R106, 2010 <a href="http://genomebiology.com/2010/11/10/R106"><span style="text-decoration: underline;">Full Text</span></a></p><p>edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Robinson et al., Bioinformatics 26(1):139-40 2010 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2796818"> <span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Two-stage Poisson model for testing RNA-seq data. Auer and Doerge, SAGMB 10(1), article 26 <a href="http://www.bepress.com/sagmb/vol10/iss1/art26/"><span style="text-decoration: underline;">Full Text</span></a></p><p>Experimental design, preprocessing, normalization and differential expression analysis of small RNA sequencing experiments. McCormick et al., Silence2(1):2, 2011 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3055805"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>RNA-Seq gene expression estimation with read mapping uncertainty. Li et al, Bioinformatics 26:493-500, 2010 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2820677">PubMedCentral</a> <em>[describes the RSEM software package]</em></p><h3>Comparing genomes and assemblies; variant detection<span style="text-decoration: underline;"> </span></h3><p>Versatile and open software for comparing large genomes. Kurtz et al, Genome Biol (5(2):R12, 2004. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC395750"><span style="text-decoration: underline;">PubMedCentral</span></a> <em>[describes the MUMmer software for full-genome alignment &amp; comparisons]</em></p><p>Searching for SNPs with cloud computing. Langmead et al, Genome Biol 10(11):R134, 2009 <a href="http://genomebiology.com/content/10/11/R134"><span style="text-decoration: underline;">Full Text</span></a></p><p>Calling SNPs without a reference sequence. Ratan et al, BMC Bioinformatics 11:130, 2010 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2851604"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Microindel detection in short-read sequence data. Krawitz et al, Bioinformatics 26(6):722-9, 2010. <a href="http://bioinformatics.oxfordjournals.org/content/26/6/722.long"><span style="text-decoration: underline;">Full Text</span></a></p><p>vipR: variant identification in pooled DNA using R. Altmann et al., Bioinformatics 27: i77-i84, 2011. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3117388"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Geoseq: a tool for dissecting deep-sequencing datasets. Gurtowski et al, BMC Bioinformatics 11:506, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2972303/"><span style="text-decoration: underline;">PubMedCentral</span></a> <em>[Geoseq is a web service that allows searching deep sequencing datasets with a reference sequence of a gene of interest]</em></p><p>Detecting and annotating genetic variations using the HugeSeq pipeline. Lam et al, Nature Biotechnology 30:226 - 229, 2012 <span style="text-decoration: underline;"><a href="http://www.nature.com/nbt/journal/v30/n3/full/nbt.2134.html">Publisher Website</a></span>, <span style="text-decoration: underline;"><a href="http://hugeseq.snyderlab.org/">Home Page</a></span></p><p>Genome-wide LORE1 retrotransposon mutagenesis and high-throughput insertion detection in <em>Lotus japonicus</em>. Urbański et al, Plant J 64:731-741, 2012. <span style="text-decoration: underline;"><a href="http://onlinelibrary.wiley.com/doi/10.1111/j.1365-313X.2011.04827.x/abstract">Publisher Website</a></span> <em>[This paper describes a 2-dimensional pooling strategy with barcoding to allow use of Illumina sequencing to screen for retrotransposon insertion mutations, and includes a software package called FSTpoolit for analysis of the resulting sequence reads.]</em></p><h3>Genotyping by sequencing</h3><p>Genome-wide genetic marker discovery and genotyping using next-generation sequencing. Davey et al., Nat Rev Genet 12(7):499-510, 2011 <a href="http://www.ncbi.nlm.nih.gov/pubmed/21681211"><span style="text-decoration: underline;">PubMed</span></a> <em>[A review of methods available at the time]</em></p><p>A robust, simple genotyping-by-sequencing (GBS) approach for high diversity species. Elshire et al., PLoS One 6(5):e19379, 2011. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3087801"><span style="text-decoration: underline;">Full Text</span></a></p><p>Development of high-density genetic maps for barley and wheat using a novel two-enzyme genotyping-by-sequencing approach. Poland et al., PLoS One 7(2): e32253, 2012. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3289635/"><span style="text-decoration: underline;">Full Text</span></a></p><p>Double digest RADseq: an inexpensive method for de novo SNP discovery and genotyping in model and non-model species. Peterson et al, PLoS One 7(5):e37135, . 2012. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3365034/"><span style="text-decoration: underline;">Full Text</span></a></p><p>Imputation of unordered markers and the impact on genomic selection accuracy. Rutkowski et al, G3 3(3):427-39, 2013. <a href="http://www.g3journal.org/content/3/3/427.long"><span style="text-decoration: underline;">Full Text</span></a></p><p>Diversity Arrays Technology (DArT) and next-generation sequencing combined: genome-wide, high-throughput, highly informative genotyping for molecular breeding of <em>Eucalyptus</em>. Sansaloni et al., BMC Proceedings 5(Suppl 7):P54, 2011 <span style="text-decoration: underline;"><a href="http://www.biomedcentral.com/1753-6561/5/S7/P54">Full Text</a></span></p><p>High-throughput genotyping by whole-genome resequencing. Huang et al., Genome Res 19(6):1068-76, 2009. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2694477"><span style="text-decoration: underline;">Full Text</span></a></p><p>Multiplexed shotgun genotyping for rapid and efficient genetic mapping. Andolfatto et al. Genome Res 21(4):610-7, 2011. <a href="http://genome.cshlp.org/content/21/4/610.long"><span style="text-decoration: underline;">Full Text</span></a></p><h3>Restriction-site Associated DNA (RAD) markers</h3><p>Rapid SNP discovery and genetic mapping using sequenced RAD markers. Baird et al, PLoS One 3(10):e3376, 2008 <span style="text-decoration: underline;"><a href="http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0003376">Full Text</a></span></p><p>Linkage mapping and comparative genomics using next-generation RAD sequencing of a non-model organism. Baxter et al., PLoS One 6(4):e19315, 2011. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3082572"><span style="text-decoration: underline;">Full Text</span></a></p><p>Genome evolution and meiotic maps by massively parallel DNA sequencing: spotted gar, an outgroup for the teleost genome duplication. Amores et al, Genetics 188(4):799-808, 2011. <a href="http://www.ncbi.nlm.nih.gov/pubmed/21828280"><span style="text-decoration: underline;"> PubMed</span></a></p><p>Construction and application for QTL analysis of a Restriction-site Associated DNA (RAD) linkage map in barley. Chutimanitsakun et al, BMC Genomics 4; 12:4, 2011. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3023751"><span style="text-decoration: underline;">Full Text</span></a></p><p>RAD tag sequencing as a source of SNP markers in <em>Cynara cardunculus </em>L. Scaglione et al., BMC Genomics 13:3, 2012. <span style="text-decoration: underline;"><a href="http://www.biomedcentral.com/1471-2164/13/3">Full Text</a></span></p><p>Paired-end RAD-seq for de novo assembly and marker design without available reference. Willing et al., Bioinformatics 27(16):2187-93, 2011. <a href="http://bioinformatics.oxfordjournals.org/content/27/16/2187.long"><span style="text-decoration: underline;">Publisher Website</span></a></p><p>Local de novo assembly of RAD paired-end contigs using short sequencing reads. Etter et al., PLOS ONE 6(4): e18561, 2011. <a href="http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0018561"><span style="text-decoration: underline;">Full Text</span></a></p><p>Stacks: building and genotyping loci de novo from short-read sequences. Catchen et al., G3: Genes, Genomes, Genetics, 1:171-182, 2011. <span style="text-decoration: underline;"> Full Text</span>, <a href="http://creskolab.uoregon.edu/stacks/"><span style="text-decoration: underline;">Home Page</span></a></p><p>Rainbow: an integrated tool for efficient clustering and assembling RAD-seq reads. Chong et al, Bioinformatics 28(21):2732-7, 2012. <a href="http://bioinformatics.oxfordjournals.org/content/28/21/2732.long"> <span style="text-decoration: underline;">Publisher Website</span></a></p><p>UK RAD Sequencing Wiki page, with bibliography and RADTools software download <a href="https://www.wiki.ed.ac.uk/display/RADSequencing/Home"><span style="text-decoration: underline;">Home Page</span></a></p><h3>Workspace environments</h3><p><span style="text-decoration: underline;">Papers</span></p><p>Galaxy: a comprehensive approach for supporting accessible, reproducible, and transparent computational research in the life sciences. Goecks et al, Genome Biol 11(8):R86, 2010 <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2945788"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>Galaxy Cloudman: Delivering compute clusters. BMC Bioinformatics 11(Suppl. 12):S4, 2010 <a href="http://www.biomedcentral.com/content/pdf/1471-2105-11-S12-S4.pdf"><span style="text-decoration: underline;">Full Text</span></a></p><p><a href="http://www.broadinstitute.org/gsa/wiki/index.php/The_Genome_Analysis_Toolkit"><span style="text-decoration: underline;">The Genome Analysis Toolkit</span></a>: a MapReduce framework for analyzing next-generation DNA sequencing data. McKenna et al, Genome Res 20(9):1297-303, 2010. <a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2928508"><span style="text-decoration: underline;">PubMedCentral</span></a></p><p>A framework for variation discovery and genotyping using next-generation DNA sequencing data. DePristo et al., Nat Genet 43(5):491-8, 2011. <a href="http://www.ncbi.nlm.nih.gov/pubmed/21478889"><span style="text-decoration: underline;"> PubMed</span></a></p><p><span style="text-decoration: underline;">Online resources</span></p><p>The <a href="http://cran.r-project.org/"><span style="text-decoration: underline;">R statistical computing</span></a> environment includes<a href="http://www.bioconductor.org/"><span style="text-decoration: underline;"> Bioconductor</span></a>, a specialized set of tools for analysis of microarray and high-throughput sequencing data. Introductory materials from on-line or short workshops are widely available online; examples are <span style="text-decoration: underline;"><a href="http://bioconductor.org/help/course-materials/2012/Evomics2012/Bioconductor-tutorial.pdf">Evomics2012 Bioconductor-tutorial.pdf</a></span>, and <a href="http://bcb.dfci.harvard.edu/%7Eaedin/courses/Bioconductor/"><span style="text-decoration: underline;">Intro to Bioconductor</span></a>. Materials from an advanced course on high-throughput genetic data analysis are at <span style="text-decoration: underline;"><a href="http://bioconductor.org/help/course-materials/2012/SeattleFeb2012/">Seattle 2012 materials</a></span>. Thomas Girke of UC-Riverside has written a very complete set of manuals describing the use of R and Bioconductor for analysis of genomic datasets, available at <a href="http://manuals.bioinformatics.ucr.edu/home/R_BioCondManual">R and Bioconductor Manuals</a>. <br /> <a href="http://cran.r-project.org/manuals.html"><span style="text-decoration: underline;">Manuals</span></a> and contributed <a href="http://cran.r-project.org/other-docs.html"><span style="text-decoration: underline;">documentation</span></a> for R are available at the R-project.org website, and video tutorials are also available on Youtube; those posted by Tutorlol are brief, clear, and to the point. <br /> Materials from a series of mini-courses in R taught in 2010 at UCLA are available:</p><ul>
<li><a href="http://scc.stat.ucla.edu/page_attachments/0000/0141/10S-basicR.pdf">Intro to programming and graphics</a></li>
<li><a href="http://scc.stat.ucla.edu/page_attachments/0000/0143/S10_RProgII.pdf">Data manipulation and functions</a></li>
<li><a href="http://scc.stat.ucla.edu/page_attachments/0000/0185/Graphics_course.pdf">Graphics for exploratory data analysis</a></li>
<li><a href="http://scc.stat.ucla.edu/page_attachments/0000/0147/20100503_IntroStats.pdf">Introductory statistics</a></li>
<li><a href="http://scc.stat.ucla.edu/page_attachments/0000/0188/reg_R_1_09S_slides.pdf">Linear regression</a></li>
</ul><p><a href="http://a-little-book-of-r-for-bioinformatics.readthedocs.org/en/latest/"> <span style="text-decoration: underline;">A Little Book of R for Bioinformatics</span></a> is an on-line resource with information and exercises to provide practice in bioinformatics analysis of DNA sequences and other biological data in R. <br /> Many books on specific topics in R programming are also available through Amazon or other vendors.</p><h3>Cloud computing resources</h3><p>The case for cloud computing in genome informatics. Lincoln Stein, Genome Biol. 11(5):207, 2010 <a href="http://www.ncbi.nlm.nih.gov/pubmed/20441614"><span style="text-decoration: underline;">Pubmed</span></a></p><p>Galaxy Cloudman: delivering cloud compute clusters. Afgan et al, BMC Bioinformatics <span style="text-decoration: underline;">11</span>(Suppl 12):S4, 2010 <a href="http://www.biomedcentral.com/1471-2105/11/S12/S4"><span style="text-decoration: underline;">Full Text</span></a></p><p><a href="http://cloudbiolinux.com/">CloudBioLinux</a> is an open-source project that provides a bioinformatics Linux system for cloud computing, pre-configured with a variety of software tools installed and ready to use.</p><p>A <a href="https://github.com/chapmanb/cloudbiolinux/blob/master/doc/intro/gettingStarted_CloudBioLinux.pdf?raw=true"><span style="text-decoration: underline;">tutorial</span></a> on getting started with CloudBioLinux on the Amazon Web Services Elastic Compute Cloud (EC2)</p><p><a href="http://userwww.service.emory.edu/%7Eeafgan/content/ppt/EnisAfgan_BOSC_2010.pdf"><span style="text-decoration: underline;">Deploying Galaxy on the Cloud</span></a>  slides from a presentation by Enis Afgan (Emory University) at the <br /> &nbsp;Bioinformatics Open Source Conference in Boston, July 2010</p><p>A <a href="http://screencast.g2.bx.psu.edu/cloud/"><span style="text-decoration: underline;"> screencast</span></a> that provides a step-by-step guide to starting a Galaxy cluster in the EC2 environment</p><p>A <a href="https://bitbucket.org/galaxy/galaxy-central/wiki/cloud"><span style="text-decoration: underline;">webpage</span></a> that has the same information in text form, and is the basis for the screencast</p><p>The iPlant Collaborative, an NSF-funded project to create computational resources for plant biology research, provides access to cloud computing resources through <span style="text-decoration: underline;"><a href="http://www.iplantcollaborative.org/discover/atmosphere">Atmosphere</a></span></p><p>SeqWare Query Engine: storing and searching sequence data in the cloud. OConnor et al, BMC Bioinformatics <strong>11</strong>(Suppl 12)<strong>:</strong>S2, 2010 <a href="http://www.biomedcentral.com/1471-2105/11/S12/S2"><span style="text-decoration: underline;">Full Text</span></a></p><p>An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics. Taylor, BMC Bioinformatics <strong>11</strong>(Suppl 12)<strong>:</strong>S1, 2010 <a href="http://www.biomedcentral.com/1471-2105/11/S12/S1"><span style="text-decoration: underline;">Full Text</span></a></p><h3>Links to Linux command-line tutorials and resources</h3><p>Tutorials for AWK, a powerful tool for handling data tables</p><ul>
<li>A set of <a href="http://people.bu.edu/scottm/AWK.NOTES"><span style="text-decoration: underline;">awk notes</span></a> from Boston University</li>
<li>Bruce Barnett's <a href="http://www.grymoire.com/Unix/Awk.html"><span style="text-decoration: underline;">awk tutorial</span></a></li>
<li>Greg Goebel's <a href="http://www.vectorsite.net/tsawk.html"><span style="text-decoration: underline;">awk tutorial</span></a></li>
<li><a href="http://teaching.software-carpentry.org/2013/01/16/1433/"><span style="text-decoration: underline;">Executing an awk command from R</span></a> to simplify data exploratory analysis, from Lex Nederbragt</li>
</ul><p>Tutorials for bash shell scripting</p><ul>
<li>A <a href="http://www.linuxconfig.org/bash-scripting-tutorial"><span style="text-decoration: underline;">tutorial</span></a> at linuxconfig.org</li>
<li>A <a href="http://www.hypexr.org/bash_tutorial.php"><span style="text-decoration: underline;">Getting Started With Bash</span></a> tutorial at hypexr.org</li>
<li>Mendel Cooper's <a href="http://tldp.org/LDP/abs/html/"><span style="text-decoration: underline;">Advanced Bash Shell-Scripting Guide</span></a></li>
</ul><p>Tutorials for sed, the command-line stream editor</p><ul>
<li>A <a href="http://www.panix.com/%7Eelflord/unix/sed.html"><span style="text-decoration: underline;">tutorial</span></a> at Rutgers</li>
<li>Peteris Krumins claims to have the <a href="http://www.catonmat.net/blog/worlds-best-introduction-to-sed/"><span style="text-decoration: underline;"> World's Best Introduction to Sed</span></a>; take a look and judge for yourself.</li>
<li>Bruce Barnett's <a href="http://www.grymoire.com/Unix/Sed.html"><span style="text-decoration: underline;">sed tutorial</span></a>.</li>
</ul><h3>Links to other useful sites</h3><p>The<a href="http://seqanswers.com/"><span style="text-decoration: underline;"> SEQanswers</span></a> online community has forums on several topics related to sequencing; the bioinformatics forum is the most active.</p><p>The SEQanswers <span style="text-decoration: underline;"><a href="http://seqanswers.com/wiki/Software">Software Wiki</a></span> is a list of software for analysis of sequencing data</p><p><a href="http://biostar.stackexchange.com/">Biostar</a> is another online community for questions and answers on bioinformatics and computational genomics.</p><p>Information on file formats used by the University of California - Santa Cruz Genome Browser is on the <a href="http://genome.ucsc.edu/FAQ/FAQformat"><span style="text-decoration: underline;"> FAQ list</span></a></p><p>A manual for the Integrated Genome Browser visualization tool is <a href="http://wiki.transvar.org/confluence/display/igbman/Home"><span style="text-decoration: underline;">here</span></a></p><p>Course materials for a short course entitled <a href="http://bioconductor.org/help/course-materials/2010/SeattleIntro/"><span style="text-decoration: underline;">Introduction to R and Bioconductor</span></a>, held in Seattle in Dec 2010</p><p><a href="http://great.stanford.edu/"><span style="text-decoration: underline;">Genomic Regions Enrichment of Annotations Tool</span></a> - A web service to test for over-representation of specific ontology categories among genes near ChIP-seq peaks</p><p><a href="http://www.animalgenome.org/bioinfo/resources/nextgensoft.html"><span style="text-decoration: underline;">Next-gen-seq software</span></a> - a list of software packages, both commercial and open-source, related to analysis of deep sequencing datasets</p><p><a href="http://www.cbcb.umd.edu/software/"><span style="text-decoration: underline;">Software</span></a> from the Center for Bioinformatics and Computational Biology, University of Maryland - many useful programs, all open-source</p><p><a href="http://bioinformatics.psb.ugent.be/plaza/"><span style="text-decoration: underline;"> PLAZA</span></a>: a comparative genomics resource to study gene and genome evolution in plants; described by Proost et al, Plant Cell 21:3718, 2010 <a href="http://www.plantcell.org/content/21/12/3718.full"><span style="text-decoration: underline;">Full Text</span></a></p><p>The European Bioinformatics Institute provides tools <a href="http://www.ebi.ac.uk/Tools/rcloud/"><span style="text-decoration: underline;">ArrayExpressHTS</span><span style="text-decoration: underline;"> and R-Cloud</span></a> for analysis of transcriptome data</p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>

</channel>
</rss>