BOL: Related items

List of generic simulation software/tools/resource with brief description and homepage !!!

Jit — Mon, 10 Feb 2014 05:57:29 -0600

List of generic simulation software/tools/resource with brief description and homepage

ALF
A Simulation Framework for Genome Evolution
http://www.cbrg.ethz.ch/alf

Bayesian Serial SimCoal
Bayesian Serial SimCoal, (BayeSSC) is a modification of SIMCOAL 1.0, a program written by Laurent Excoffier, John Novembre, and Stefan Schneider.
http://www.stanford.edu/group/hadlylab/ssc/index.html

BEERS
BEERS was designed to benchmark RNA-Seq alignment algorithms and also algorithms that aim to reconstruct different isoforms and alternate splicing from RNA-Seq data
http://cbil.upenn.edu/beers/

BOTTLENECK
Bottleneck is a program for detecting recent effective population size reductions from allele data frequencies
http://www.ensam.inra.fr/urlb/bottleneck/bottleneck.html

BottleSim
BottleSim is a computer simulation program for simulating the process of population bottlenecks
http://chkuo.name/software/bottlesim.html

CASS
Protein Sequence Simulation
http://www.wyomingbioinformatics.org/liberlesgroup/cass/

CDPOP
CDPOP is a landscape genetics tool for simulating the emergence of spatial genetic structure in populations resulting from specified landscape processes governing organism movement behavior.
http://cel.dbs.umt.edu/cdpop

CoalFace
CoalFace is a simulation of the coalescent process with the visual display of gene genealogies.
http://web.up.ac.za/default.asp?ipkcategoryid=3283

CoaSim
CoaSim is a tool for simulating the coalescent process with recombination and geneconversion under various demographic models.
http://users-birc.au.dk/mailund/coasim/index.html

cosi
The cosi package is written in C and is available as a tar file.
http://www.broadinstitute.org/~sfs/cosi/

CS-PSeq-Gen
A program to simulate the evolution of protein sequences under the constraints of the information of a particular reconstructed phylogeny
http://bioserv.rpbs.univ-paris-diderot.fr/software/cs-pseq-gen.html

DAWG
An application designed to simulate the evolution of recombinant DNA sequences in continuous time
http://scit.us/projects/dawg

Easypop
EASYPOP is an individual based model intended to simulate datasets under a very broad range of conditions
http://www.unil.ch/dee/page36926_fr.html

EggLib
EggLib is a C++/Python library and program package for evolutionary genetics and genomics.
http://egglib.sourceforge.net/

EvolSimulator
A simulation test bed for hypotheses of genome evolution
http://acb.qfab.org/acb/evolsim/

EvolveAGene
A realistic coding sequence simulation program that separates mutation from selection and allows the user to set selection conditions
http://bellinghamresearchinstitute.com/software/index.html

fastsimcoal
A continuous-¬‐time coalescent simulator of genomic diversity under arbitrarily complex evolutionary scenarios
http://cmpg.unibe.ch/software/fastsimcoal/

FastSLINK
Simulation of Marker and Phenotype Data in Pedigrees
http://watson.hgen.pitt.edu/

FFPopSim
C++/Python library for population genetics.
http://webdav.tuebingen.mpg.de/ffpopsim/

FLUX SIMULATOR
The Flux Simulator aims at providing a deterministic in silico reproduction of the experimental pipelines for RNA-Seq, employing a minimal set of parameters.
http://flux.sammeth.net/simulator.html

ForSim
ForSim: A Forward Evolutionary Computer Simulation
http://www.anthro.psu.edu/weiss_lab/research.shtml

ForwSim
The program given below is based on the algorithm described in Padhukasahasram et al. 2008 to simulate genetic drift in a standard Wright-Fisher process.
http://badri-populationgeneticsimulators.blogspot.com/

FPG
Forward Population Genetic simulation
http://genfaculty.rutgers.edu/hey/software#fpg

FREGENE
FREGENE is a C++ program that simulates sequence-like data over large genomic regions in large diploid populations.
http://www.ebi.ac.uk/projects/bargen/download/fregen/documentation_html.html

GAMETES
Genetic Architecture Model Emulator for Testing and Evaluating Software: Simulates complex SNP models with pure, strict epistatic interactions with n-loci.
http://sourceforge.net/projects/gametes/?source=navbar

GASP
Genometric Analysis Simulation Program. A software tool for testing and investigating methods in statistical genetics by generating samples of family data based on user specified models.
http://research.nhgri.nih.gov/gasp/

GemSIM
Next generation sequencing read simulator
http://sourceforge.net/projects/gemsim/

GeneArtisan
Simulation of Markers in Case-Control Study Designs
http://www.rannala.org/?page_id=241

GENOME
A rapid coalescent-based whole genome simulator
http://www.sph.umich.edu/csg/liang/genome/

GenomePop2
GenomePop2 is a specialization of the program GenomePop just to manage SNPs under more flexible and useful settings. If you need models with more than 2 alleles please use the GenomePop program version.
http://webs.uvigo.es/acraaj/genomepop2.htm

GenomeSimla
GenomeSIMLA is currently under development- however, we have a beta release that we are asking to be tested
http://chgr.mc.vanderbilt.edu/genomesimla/

GENS2
Simulates interactions among two genetic and one environmental factor and also allows for epistatic interactions.
https://sourceforge.net/projects/gensim/

GWAsimulator
A rapid whole genome simulation program
http://biostat.mc.vanderbilt.edu/wiki/main/gwasimulator

HAP-SAMPLE
An association simulator for candidate regions or genome scans
http://www.hapsample.org/

HAPGEN
A simulator for the simulation of case control datasets at SNP markers
https://mathgen.stats.ox.ac.uk/genetics_software/hapgen/hapgen2.html

HapSim
A simulation tool for generating haplotype data with pre-specified allele frequencies and LD coefficients
http://cran.r-project.org/web/packages/hapsim/index.html

HAPSIMU
A program that simulates heterogeneous populations with various known and controllable structures under the continuous migration model or the discrete model
http://l.web.umkc.edu/liujian/

IBDsim
IBDSim is a computer package for the simulation of genotypic data under general isolation by distance models.
http://raphael.leblois.free.fr/

indel-Seq-Gen
A biological sequence simulation program that simulates highly divergent DNA sequences and protein superfamilies
http://bioinfolab.unl.edu/~cstrope/isg/

Indelible
A powerful and flexible simulator of biological evolution
http://abacus.gene.ucl.ac.uk/software/indelible/

invertFREGENE
InvertFREGENE is a forward-in-time simulator of inversions in population genetic data
http://www.ebi.ac.uk/projects/bargen/

kernalPop
A spatially explicit population genetic simulation engine
http://cran.r-project.org/src/contrib/archive/kernelpop/

MaCS
Markovian Coalescent Simulator
http://www-hsc.usc.edu/~garykche/

Mason
A package for the simulation of nucleotide data.
http://www.seqan.de/projects/mason/

mbs
modifying Hudson's ms software to generate samples of DNA sequences with a biallelic site under selection
http://www.sendou.soken.ac.jp/esb/innan/innanlab/software.html

Mendel's Accountant
Mendel's Accountant (MENDEL) is an advanced numerical simulation program for modeling genetic change over time and was developed collaboratively by Sanford, Baumgardner, Brewer, Gibson and ReMine
http://mendelsaccount.sourceforge.net/

MetaSim
A tool to generate collections of synthetic reads that reflect the diverse taxonomical composition of typical metagenome data sets
http://ab.inf.uni-tuebingen.de/software/metasim/

mlcoalsim
Multilocus Coalescent Simulations
http://code.google.com/p/mlcoalsim-v1/

ms
The purpose of this program is to allow one to investigate the statistical properties of such samples, to evaluate estimators or statistical tests, and generally to aid in the interpretation of polymorphism data sets.
http://home.uchicago.edu/~rhudson1/source/mksamples.html

msHOT
The purpose of this program is to allow one to investigate the statistical properties of such samples, to evaluate estimators or statistical tests, and generally to aid in the interpretation of polymorphism data sets.
http://home.uchicago.edu/~rhudson1/

msms
A coalescent Simlation tool with selection.
http://www.mabs.at/ewing/msms/index.shtml

MySSP
A program for the simulation of DNA sequence evolution across a phylogenetic tree
http://www.rosenberglab.net/software.php

Nemo
A forward-time, individual-based, genetically explicit, and stochastic simulation program designed to study the evolution of genetic markers, life history traits, and phenotypic traits in a flexible (meta-)population framework.
http://nemo2.sourceforge.net/

NetRecodon
Coalescent simulation of coding DNA sequences with recombination (inter and intracodon), migration and demography
http://code.google.com/p/netrecodon/

PEDAGOG
Software for simulating eco-evolutionary population dynamics
https://bcrc.bio.umass.edu/pedigreesoftware/node/5

phenosim
A tool to add phenotypes to simulated genotypes
http://evoplant.uni-hohenheim.de/doku.php?id=software:software

PhyloSim
An R package for the Monte Carlo simulation of sequence evolution
http://bit.ly/rlsim-git

pIRS
Profile-based Illumina pair-end reads simulator
https://code.google.com/p/pirs/

ProteinEvolver
Simulation of protein evolution along phylogenies under structure-based substitution models
http://code.google.com/p/proteinevolver/

QMSim
QTL and Marker Simulator
http://www.aps.uoguelph.ca/~msargol/qmsim/

quantiNEMO
An individual-based program for the analysis of quantitative traits with explicit genetic architecture potentially under selection in a structured population
http://www2.unil.ch/popgen/softwares/quantinemo/

RECOAL
Simulates new haplotype data from a reference population of haplotypes.
ftp://popgen.usc.edu/

Recodon
Coalescent simulation of coding DNA sequences with recombination, migration and demography
http://code.google.com/p/recodon/

rlsim
A package for simulating RNA-seq library preparation with parameter estimation
http://bit.ly/rlsim-git

Rmetasim
Rmetasim is a front-end for the metasim engine that is implemented as a package that runs in the statistical computing environment R
http://linum.cofc.edu/software.html#metasim

RNA Seq Simulator
RSS takes SAM alignment files from RNA-Seq data and simulates over dispersed, multiple replica, differential, non-stranded RNA-Seq datasets.
http://useq.sourceforge.net/cmdlnmenus.html#rnaseqsimulator

Rose
Random model of sequence evolution
http://bibiserv.techfak.uni-bielefeld.de/rose/

SelSim
SelSim is a program for Monte Carlo simulation of DNA polymorphism data for a recom- bining region within which a single bi-allelic site has experienced natural selection
http://www.well.ox.ac.uk/~spencer/selsim/

Seq-Gen
An application for the Monte Carlo simulation of molecular sequence evolution along phylogenetic trees.
http://tree.bio.ed.ac.uk/software/seqgen/

SEQPower
Statistical power analysis for sequence-based association studies
http://bioinformatics.org/spower/

SeqSIMLA
SeqSIMLA can simulate sequence data with user-specified disease and quantitative trait models. Family or unrelated case-control data can be simulated.
http://seqsimla.sourceforge.net/

Serial NetEvolve
A flexible utility for generating serially-sampled sequences along a tree or recombinant network
http://biorg.cis.fiu.edu/sne/

SFS_CODE
SFS_CODE can perform forward population genetic simulations under a general Wright-Fisher model with arbitrary migration, demographic, selective, and mutational effects.
http://sfscode.sourceforge.net/sfs_code/index/index.html

SIBSIM
Quantitative phenotype simulation in extended pedigrees
http://sourceforge.net/projects/sibsim/

SIMCOAL2
A coalescent program for the simulation of complex recombination patterns over large genomic regions under various demographic models
http://cmpg.unibe.ch/software/simcoal2/

SimCopy
An R package simulating the evolution of copy number profiles along a tree.
http://bit.ly/simcopy

SIMLA
SIMLA is a SIMuLAtion program that generates data sets of families for use in Linkage and Association studies.
http://www.chg.duke.edu/research/simla.html

SimPed
A Simulation Program to Generate Haplotype and Genotype Data for Pedigree Structures
http://www.hgsc.bcm.tmc.edu/content/simped

Simprot
A program to simulate protein evolution by substitution, insertion and deletion
http://www.uhnresearch.ca/labs/tillier/software.htm#3

SimRare
Rare variant simulation and analysis tool
http://code.google.com/p/simrare/

simuGWAS
A forward-time simulator that simulates realistic samples for genome-wide association studies.
http://simupop.sourceforge.net/cookbook/simucomplexdisease

simuPOP
simuPOP is a general-purpose individual-based forward-time population genetics simulation environment.
http://simupop.sourceforge.net/

SISSI
A software tool to generate data of related sequences along a given phylogeny, taking into account user defined system of neighbourhoods and instantaneous rate matrices.
http://www.cibiv.at/software/sissi/

SNPsim
Coalescent simulation of hotspot recombination
http://code.google.com/p/phylosoftware/

SPIP
SPIP simulates the transmission of genes from parents to offspring in a population having demographic structure defined by the user
http://swfsc.noaa.gov/textblock.aspx?division=fed&id=3434

Splatche
Spatial and Temporal Coalescences in Heterogeneous Environment
http://www.splatche.com/

srv
Simulator of Rare Varaints (srv) is a simulator for the simulation of the introduction and evolution of (rare) genetic variants.
http://simupop.sourceforge.net/cookbook/simurarevariants

SUP
SLINK/FastSLINK utility program
http://mlemire.freeshell.org/software.html

TreesimJ
A flexible, forward-time population genetic simulator
http://code.google.com/p/treesimj/

Vortex
VORTEX is an individual-based simulation model for population viability analysis (PVA).
http://www.vortex9.org/vortex.html

References:

Image www.evolution-of-life.com

www.cancer.gov

Bioinformatics Scientist, Production Bioinformatics @ South San Francisco, CA

Thu, 19 Aug 2021 08:45:24 -0500

wist is looking for a Bioinformatics Scientist to join our Production Bioinformatics Team. You will work alongside research scientists, software engineers and data scientists to further deliver on our mission to expand access to best-in-class synthetic biology and next-generation sequencing applications. You will be developing and engineering tools to better evaluate and build hardened, production quality pipelines, optimize data quality, and automate lab and bioinformatics processes. Our ideal candidate is an organized problem solver with a background in developing and building novel production-quality bioinformatics tools and packages. Equally excellent communication skills and a proven ability to work independently are required.

More at https://boards.greenhouse.io/twistbioscience/jobs/3135495?gh_src=9ecc0b941us

Virus Bioinformatics Tools

LEGE — Wed, 24 Apr 2024 06:19:55 -0500

Bioinformatics tools play a crucial role in studying viruses, enabling researchers to analyze their genetic makeup, structure, function, and evolution. Here are some commonly used bioinformatics tools for virus research

https://evirusbioinfc.notion.site/18e21bc49827484b8a2f84463cb40b8d?v=92e7eb6703be4720abf17a901bc9a947

Address of the bookmark: https://evirusbioinfc.notion.site/18e21bc49827484b8a2f84463cb40b8d?v=92e7eb6703be4720abf17a901bc9a947

Exploring Bacterial Comparative Genomics: A Bioinformatics Approach

LEGE — Sat, 14 Dec 2024 12:31:14 -0600

In the world of microbiology, bacteria have long fascinated scientists for their diversity, adaptability, and crucial roles in ecosystems and human health. Comparative genomics—a field that involves analyzing and comparing the genomes of different organisms—has revolutionized our understanding of bacterial evolution, adaptation, and pathogenicity. By leveraging bioinformatics tools and techniques, researchers can uncover genomic insights that were once hidden. This blog delves into the principles, methodologies, and applications of bacterial comparative genomics from a bioinformatics perspective.

What is Bacterial Comparative Genomics?

Comparative genomics involves the systematic comparison of genomes across different bacterial species or strains. This approach allows scientists to:

Identify conserved and unique genes.
Explore genetic determinants of pathogenicity.
Understand bacterial evolution and phylogenetics.
Investigate horizontal gene transfer and its role in antibiotic resistance.

Bioinformatics is central to these analyses, enabling the processing and interpretation of large-scale genomic data.

Key Steps in Bacterial Comparative Genomics

Genome Sequencing and Assembly: The process begins with obtaining high-quality bacterial genome sequences. Advances in next-generation sequencing (NGS) technologies have made it faster and more affordable to sequence bacterial genomes. Tools such as SPAdes and Velvet are commonly used for genome assembly.
Genome Annotation: Annotating a genome involves identifying genes, regulatory elements, and other genomic features. Automated tools like Prokka and RAST provide functional annotations, allowing researchers to predict the roles of genes and proteins.
Genome Alignment: Aligning genomes is crucial for identifying conserved regions, single-nucleotide polymorphisms (SNPs), and structural variations. Tools like Mauve and progressiveMauve are commonly employed for whole-genome alignments.
Comparative Analyses:
- Core and Pan-genome Analysis: The core genome consists of genes shared across all strains of a species, while the pan-genome includes all genes found in any strain. Software like Roary and BPGA can perform core and pan-genome analyses.
- Phylogenetic Analysis: Comparative genomics often involves reconstructing evolutionary relationships. Tools such as MEGA and IQ-TREE facilitate phylogenetic tree construction based on genomic data.
- Functional Enrichment Analysis: To understand the biological significance of unique or shared genes, functional enrichment analysis using databases like GO (Gene Ontology) and KEGG is essential.

Recommended Bioinformatics Tools for Comparative Genomics

Here are some additional bioinformatics tools that can aid bacterial comparative genomics:

OrthoFinder: For accurate ortholog identification across multiple genomes.
PanOCT: Specifically designed for pan-genome clustering and annotation.
FASTANI: A tool for calculating Average Nucleotide Identity (ANI) for microbial genome comparisons.
CIRCOS: For visually comparing genomic data through circular genome plots.
Galaxy Platform: A user-friendly web-based platform offering numerous genomic analysis tools.
BLAST: Essential for sequence alignment and similarity searches.
PhyloSift: Focused on phylogenetic analysis of microbial genomes using marker genes.

These tools, in combination with the methods discussed, provide a robust framework for conducting comprehensive comparative genomic studies.

Applications of Bacterial Comparative Genomics

Understanding Pathogenicity: Comparative genomics helps identify virulence factors that distinguish pathogenic strains from non-pathogenic relatives. For instance, comparing genomes of Escherichia coli strains has revealed key genetic determinants of pathogenicity in enterohemorrhagic strains.
Antibiotic Resistance Research: The spread of antibiotic resistance genes through horizontal gene transfer is a major global concern. Comparative analyses can trace the origins and dissemination of resistance genes, aiding in the development of countermeasures.
Microbial Ecology and Evolution: By studying genomic variations, researchers can understand how bacteria adapt to different environments. This is particularly relevant for extremophiles and symbiotic bacteria.
Vaccine Development: Identifying conserved antigens across pathogenic strains is critical for vaccine design. Comparative genomics has been instrumental in developing vaccines against pathogens like Neisseria meningitidis.
Biotechnology Applications: Comparative studies can uncover unique metabolic pathways in bacteria, paving the way for applications in bioremediation, synthetic biology, and industrial microbiology.

Challenges in Bacterial Comparative Genomics

While the field has made significant strides, several challenges remain:

Data Overload: The rapid growth of sequencing data requires robust computational infrastructure and efficient algorithms.
Genome Plasticity: High rates of horizontal gene transfer and genome rearrangements in bacteria complicate comparative analyses.
Annotation Accuracy: Automated annotation tools are not infallible, and manual curation is often needed for high-confidence results.
Interpreting Non-Coding Regions: Understanding the functional significance of non-coding genomic regions remains a challenge.

Future Directions

The integration of bacterial comparative genomics with other ‘omics’ approaches—such as transcriptomics, proteomics, and metabolomics—promises a more comprehensive understanding of bacterial biology. Additionally, advancements in machine learning and artificial intelligence are likely to further enhance bioinformatics analyses, enabling the prediction of complex phenotypes from genomic data.

Conclusion

Bacterial comparative genomics, driven by bioinformatics, continues to unravel the complexities of bacterial life. From combating antibiotic resistance to uncovering the secrets of microbial evolution, this interdisciplinary field holds immense potential for addressing pressing challenges in microbiology and beyond. As technology advances, so too will our ability to harness the power of comparative genomics for scientific and societal benefit.

Predicting Pathogen Virulence Using Bioinformatics Tools

BioStar — Tue, 04 Nov 2025 07:55:53 -0600

In the genomic era, the ability to predict the virulence potential of pathogens has become an indispensable part of infectious disease research. With the exponential growth of microbial genome data, bioinformatics tools now enable scientists to identify virulence factors, model pathogen behavior, and even forecast outbreak risks — all from sequence data.

In an age where pathogens continue to evolve and cross boundaries, understanding what makes them virulent—that is, capable of causing disease—has become a critical focus in modern microbiology and genomics. Virulence prediction bridges computational biology, genomics, and machine learning to forecast the pathogenic potential of microbes before they strike.

What Is Virulence?

Virulence refers to the degree of damage a pathogen can inflict on its host. It is determined by a combination of genetic factors—called virulence factors (VFs)—that allow the organism to attach, invade, evade, and harm the host. These include genes coding for toxins, secretion systems, adhesins, and enzymes that disrupt host defenses.

Understanding virulence factors not only helps in deciphering the mechanisms of infection but also provides early warning signs for emerging threats.

Why Predict Virulence?

Traditional virulence studies relied heavily on experimental infection models, which, although accurate, are time-consuming, expensive, and ethically constrained.
Today, the availability of whole-genome sequences and large-scale pathogen databases has paved the way for in silico virulence prediction—a computational approach that can screen thousands of genomes within hours.

This approach enables researchers to:

Rapidly identify potential high-risk strains.
Prioritize pathogens for containment, surveillance, or further study.
Guide vaccine development and drug target discovery.
Support One Health frameworks, linking animal, human, and environmental health data.

How Is Virulence Predicted?

Virulence prediction combines bioinformatics pipelines with machine learning and comparative genomics. The process generally involves:

Genome Annotation: Identifying genes and coding sequences in microbial genomes.
Feature Extraction: Comparing sequences with curated databases like VFDB (Virulence Factor Database), PATRIC, or Victors.
Pattern Recognition: Using algorithms (e.g., Random Forest, SVM, or deep learning models) to classify genes or strains as virulent or non-virulent based on sequence patterns, motifs, and protein domains.
Scoring and Visualization: Assigning a virulence score or confidence level and visualizing it through heatmaps or genome maps.

Tools and Resources for Virulence Prediction

A number of tools and databases make virulence prediction accessible to the scientific community:

VFanalyzer – For identifying virulence genes based on VFDB.
PathoFact – Predicts virulence, antimicrobial resistance (AMR), and toxin genes from metagenomic data.
Pangenome-based models – Identify virulence-associated gene clusters across strains.
Machine learning models – Use features like GC content, codon usage bias, or protein domains to predict pathogenicity.

Emerging tools now integrate multi-omic data—including transcriptomics, proteomics, and metabolomics—to understand virulence in a systems biology framework.

Applications in the Real World

Virulence prediction has major implications across public health and research sectors:

Epidemic preparedness: Early identification of virulent strains in outbreak samples.
AMR surveillance: Linking virulence profiles with antibiotic resistance determinants.
Environmental monitoring: Predicting pathogenic potential of soil or waterborne microbes.
Clinical diagnostics: Supporting personalized treatment through pathogen profiling.

For instance, integrating virulence prediction pipelines into national surveillance networks could enable faster risk assessment and response to infectious outbreaks.

The Road Ahead

As machine learning and genomics advance, virulence prediction will evolve from simple gene-based detection to dynamic, context-aware models that account for host–pathogen interactions, environmental signals, and evolutionary adaptation.

Future tools may predict not just if a strain is virulent, but under what conditions it expresses that virulence—bridging the gap between genotype and phenotype.

In Summary

Virulence prediction is redefining how we understand and anticipate infectious diseases. By coupling genomic insights with computational intelligence, researchers can identify potential threats earlier, design smarter interventions, and ultimately, strengthen our preparedness against emerging pathogens.

Software for genome assembly !

LEGE — Sun, 30 Aug 2020 09:51:38 -0500

List of bioinformatics tools/Software Website References for genome assembly:

1 Falcon https://github.com/PacificBiosciences/pb-assembly

2 Canu assembler http://canu.readthedocs.io/en/latest/index.html

3 Miniasm assembler https://github.com/lh3/miniasm

4 PBJelly scaffolding tool https://sourceforge.net/projects/pb-jelly/

5 ARCS scaffolding tool https://github.com/bcgsc/arcs

6 Redundans reduction and scaffolding tool https://github.com/Gabaldonlab/redundans

7 Arrow error correction https://github.com/PacificBiosciences/ GenomicConsensus

8 PILON error correction https://github.com/broadinstitute/pilon/wiki

9 BUSCO single copy gene markers http://busco.ezlab.org/

10 Bandage graph assembly viewer https://rrwick.github.io/Bandage/

11 Gepard dotter http://cube.univie.ac.at/gepard

12 MUMmer aligner and plotter http://mummer.sourceforge.net/

Frequently used bioinformatics tools for viral genome analysis !

Neel — Wed, 23 Jun 2021 07:40:41 -0500

IVA: accurate de novo assembly of RNA virus genomes.
Hunt M, Gall A, Ong SH, Brener J, Ferns B, Goulder P, Nastouli E, Keane JA, Kellam P, Otto TD.
Bioinformatics. 2015 Jul 15;31(14):2374-6. doi: 10.1093/bioinformatics/btv120. Epub 2015 Feb 28.

Adapter sequences:
Optimal enzymes for amplifying sequencing libraries.
Quail, M. a et al. Nat. Methods 9, 10-1 (2012).

GAGE:
GAGE: A critical evaluation of genome assemblies and assembly algorithms.
Salzberg, S. L. et al. Genome Res. 22, 557-67 (2012).

KMC:
Disk-based k-mer counting on a PC.
Deorowicz, S., Debudaj-Grabysz, A. & Grabowski, S. BMC Bioinformatics 14, 160 (2013).

Kraken:
Kraken: ultrafast metagenomic sequence classification using exact alignments.
Wood, D. E. & Salzberg, S. L. Genome Biol. 15, R46 (2014).

MUMmer:
Versatile and open software for comparing large genomes.
Kurtz, S. et al. Genome Biol. 5, R12 (2004).

R:
R: A language and environment for statistical computing.
R Core Team (2013). R Foundation for Statistical Computing, Vienna, Austria. URL http://www.R-project.org/.

RATT:
RATT: Rapid Annotation Transfer Tool.
Otto, T. D., Dillon, G. P., Degrave, W. S. & Berriman, M. Nucleic Acids Res. 39, e57 (2011).

SAMtools:
The Sequence Alignment/Map format and SAMtools.
Li, H. et al. Bioinformatics 25, 2078-9 (2009).

Trimmomatic:
Trimmomatic: A flexible trimmer for Illumina Sequence Data.
Bolger, A. M., Lohse, M. & Usadel, B. Bioinformatics 1-7 (2014).

Commercial and public next-gen-seq (NGS) software

Surabhi Chaudhary — Tue, 03 Jun 2014 20:45:11 -0500

Integrated solutions
CLCbio Genomics Workbench - de novo and reference assembly of Sanger, Roche FLX, Illumina, Helicos, and SOLiD data. Commercial next-gen-seq software that extends the CLCbio Main Workbench software. Includes SNP detection, CHiP-seq, browser and other features. Commercial. Windows, Mac OS X and Linux.
Galaxy - Galaxy = interactive and reproducible genomics. A job webportal.
Genomatix - Integrated Solutions for Next Generation Sequencing data analysis.
JMP Genomics - Next gen visualization and statistics tool from SAS. They are working with NCGR to refine this tool and produce others.
NextGENe - de novo and reference assembly of Illumina, SOLiD and Roche FLX data. Uses a novel Condensation Assembly Tool approach where reads are joined via "anchors" into mini-contigs before assembly. Includes SNP detection, CHiP-seq, browser and other features. Commercial. Win or MacOS.
Partek - Commercial software for NGS, microarray, and qPCR data analysis. Streamlined analysis workflows for: ChIP-Seq, RNA-Seq, DNA-Seq, DNA Methylation, Gene Expression, Exon, miRNA Expression, Copy Number, Allele-Specific Copy Number, LOH, Association, Trio Analysis, and Tiling. Supports all commercial sequencing and microarray technologies.
SeqMan Genome Analyser - Software for Next Generation sequence assembly of Illumina, Roche FLX and Sanger data integrating with Lasergene Sequence Analysis software for additional analysis and visualization capabilities. Can use a hybrid templated/de novo approach. Commercial. Win or Mac OS X.
SHORE - SHORE, for Short Read, is a mapping and analysis pipeline for short DNA sequences produced on a Illumina Genome Analyzer. A suite created by the 1001 Genomes project. Source for POSIX.
SlimSearch - Fledgling commercial product.
Synamatix has SXOligoSearch (http://synasite.mgrc.com.my:8080/sxo...ligoSearch.php)
The SWIFT suit is a software collection for fast index-based sequence comparison. It contains the following programs: SWIFT — fast local alignment search, guaranteeing to find epsilon-matches between two sequences; SWIFT BALSAM — a very fast program to find semiglobal non-gapped alignments based on k-mer seeds. http://bibiserv.techfak.uni-bielefeld.de/swift/
biolib.is library and a set of script targeted to NGS. There are modules to: clean sequences (sanger, 454, ilumina), parse caf, ace and bowtie map files, clean and filter contigs, look for snps and indels., filter snps, do statistics for: reads, contigs and snps.

Align/Assemble to a reference
BFAST - Blat-like Fast Accurate Search Tool. Written by Nils Homer, Stanley F. Nelson and Barry Merriman at UCLA.
Bowtie - Ultrafast, memory-efficient short read aligner. It aligns short DNA sequences (reads) to the human genome at a rate of 25 million reads per hour on a typical workstation with 2 gigabytes of memory. Uses a Burrows-Wheeler-Transformed (BWT) index. Link to discussion thread here. Written by Ben Langmead and Cole Trapnell. Linux, Windows, and Mac OS X.
BWA - Heng Lee's BWT Alignment program - a progression from Maq. BWA is a fast light-weighted tool that aligns short sequences to a sequence database, such as the human reference genome. By default, BWA finds an alignment within edit distance 2 to the query sequence. C++ source.
ELAND - Efficient Large-Scale Alignment of Nucleotide Databases. Whole genome alignments to a reference genome. Written by Illumina author Anthony J. Cox for the Solexa 1G machine.
Exonerate - Various forms of pairwise alignment (including Smith-Waterman-Gotoh) of DNA/protein against a reference. Authors are Guy St C Slater and Ewan Birney from EMBL. C for POSIX.
GenomeMapper - GenomeMapper is a short read mapping tool designed for accurate read alignments. It quickly aligns millions of reads either with ungapped or gapped alignments. A tool created by the 1001 Genomes project. Source for POSIX.
GMAP - GMAP (Genomic Mapping and Alignment Program) for mRNA and EST Sequences. Developed by Thomas Wu and Colin Watanabe at Genentec. C/Perl for Unix.
gnumap - The Genomic Next-generation Universal MAPper (gnumap) is a program designed to accurately map sequence data obtained from next-generation sequencing machines (specifically that of Solexa/Illumina) back to a genome of any size. It seeks to align reads from nonunique repeats using statistics. From authors at Brigham Young University. C source/Unix.
MAQ - Mapping and Assembly with Qualities (renamed from MAPASS2). Particularly designed for Illumina with preliminary functions to handle ABI SOLiD data. Written by Heng Li from the Sanger Centre. Features extensive supporting tools for DIP/SNP detection, etc. C++ source
MOSAIK - MOSAIK produces gapped alignments using the Smith-Waterman algorithm. Features a number of support tools. Support for Roche FLX, Illumina, SOLiD, and Helicos. Written by Michael Strömberg at Boston College. Win/Linux/MacOSX
MrFAST and MrsFAST - mrFAST & mrsFAST are designed to map short reads generated with the Illumina platform to reference genome assemblies; in a fast and memory-efficient manner. Robust to INDELs and MrsFAST has a bisulphite mode. Authors are from the University of Washington. C as source.
MUMmer - MUMmer is a modular system for the rapid whole genome alignment of finished or draft sequence. Released as a package providing an efficient suffix tree library, seed-and-extend alignment, SNP detection, repeat detection, and visualization tools. Version 3.0 was developed by Stefan Kurtz, Adam Phillippy, Arthur L Delcher, Michael Smoot, Martin Shumway, Corina Antonescu and Steven L Salzberg - most of whom are at The Institute for Genomic Research in Maryland, USA. POSIX OS required.
Novocraft - Tools for reference alignment of paired-end and single-end Illumina reads. Uses a Needleman-Wunsch algorithm. Can support Bis-Seq. Commercial. Available free for evaluation, educational use and for use on open not-for-profit projects. Requires Linux or Mac OS X.
PASS - It supports Illumina, SOLiD and Roche-FLX data formats and allows the user to modulate very finely the sensitivity of the alignments. Spaced seed intial filter, then NW dynamic algorithm to a SW(like) local alignment. Authors are from CRIBI in Italy. Win/Linux.
RMAP - Assembles 20 - 64 bp Illumina reads to a FASTA reference genome. By Andrew D. Smith and Zhenyu Xuan at CSHL. (published in BMC Bioinformatics). POSIX OS required.
SeqMap - Supports up to 5 or more bp mismatches/INDELs. Highly tunable. Written by Hui Jiang from the Wong lab at Stanford. Builds available for most OS's.
SHRiMP - Assembles to a reference sequence. Developed with Applied Biosystem's colourspace genomic representation in mind. Authors are Michael Brudno and Stephen Rumble at the University of Toronto. POSIX.
Slider- An application for the Illumina Sequence Analyzer output that uses the probability files instead of the sequence files as an input for alignment to a reference sequence or a set of reference sequences. Authors are from BCGSC. Paper is here.
SOAP - SOAP (Short Oligonucleotide Alignment Program). A program for efficient gapped and ungapped alignment of short oligonucleotides onto reference sequences. The updated version uses a BWT. Can call SNPs and INDELs. Author is Ruiqiang Li at the Beijing Genomics Institute. C++, POSIX.
SSAHA - SSAHA (Sequence Search and Alignment by Hashing Algorithm) is a tool for rapidly finding near exact matches in DNA or protein databases using a hash table. Developed at the Sanger Centre by Zemin Ning, Anthony Cox and James Mullikin. C++ for Linux/Alpha.
SOCS - Aligns SOLiD data. SOCS is built on an iterative variation of the Rabin-Karp string search algorithm, which uses hashing to reduce the set of possible matches, drastically increasing search speed. Authors are Ondov B, Varadarajan A, Passalacqua KD and Bergman NH.
SWIFT - The SWIFT suit is a software collection for fast index-based sequence comparison. It contains: SWIFT — fast local alignment search, guaranteeing to find epsilon-matches between two sequences. SWIFT BALSAM — a very fast program to find semiglobal non-gapped alignments based on k-mer seeds. Authors are Kim Rasmussen (SWIFT) and Wolfgang Gerlach (SWIFT BALSAM)
SXOligoSearch - SXOligoSearch is a commercial platform offered by the Malaysian based Synamatix. Will align Illumina reads against a range of Refseq RNA or NCBI genome builds for a number of organisms. Web Portal. OS independent.
Vmatch - A versatile software tool for efficiently solving large scale sequence matching tasks. Vmatch subsumes the software tool REPuter, but is much more general, with a very flexible user interface, and improved space and time requirements. Essentially a large string matching toolbox. POSIX.
Zoom - ZOOM (Zillions Of Oligos Mapped) is designed to map millions of short reads, emerged by next-generation sequencing technology, back to the reference genomes, and carry out post-analysis. ZOOM is developed to be highly accurate, flexible, and user-friendly with speed being a critical priority. Commercial. Supports Illumina and SOLiD data.
NCGR uses GMAP (http://www.gene.com/share/gmap/) to alignment Solexa reads. GMAP is free, though.
Exonerate (http://www.ebi.ac.uk/~guy/exonerate/)
MUMmer (http://mummer.sourceforge.net/)
The mapping short reads called gnumap (http://dna.cs.byu.edu/gnumap/) made to increase the accuracy with duplicate matches. Open source, creates viewable output (with Affy's Integrated Genome Browser), and produces results very similar to novocraft's.
SOCS (short oligonucleotides in color space)
BFAST https://secure.genome.ucla.edu/index.php/BFAST

De novo Align/Assemble
ABySS - Assembly By Short Sequences. ABySS is a de novo sequence assembler that is designed for very short reads. The single-processor version is useful for assembling genomes up to 40-50 Mbases in size. The parallel version is implemented using MPI and is capable of assembling larger genomes. By Simpson JT and others at the Canada's Michael Smith Genome Sciences Centre. C++ as source.
ALLPATHS - ALLPATHS: De novo assembly of whole-genome shotgun microreads. ALLPATHS is a whole genome shotgun assembler that can generate high quality assemblies from short reads. Assemblies are presented in a graph form that retains ambiguities, such as those arising from polymorphism, thereby providing information that has been absent from previous genome assemblies. Broad Institute.
Edena - Edena (Exact DE Novo Assembler) is an assembler dedicated to process the millions of very short reads produced by the Illumina Genome Analyzer. Edena is based on the traditional overlap layout paradigm. By D. Hernandez, P. François, L. Farinelli, M. Osteras, and J. Schrenzel. Linux/Win.
EULER-SR - Short read de novo assembly. By Mark J. Chaisson and Pavel A. Pevzner from UCSD (published in Genome Research). Uses a de Bruijn graph approach.
MIRA2 - MIRA (Mimicking Intelligent Read Assembly) is able to perform true hybrid de-novo assemblies using reads gathered through 454 sequencing technology (GS20 or GS FLX). Compatible with 454, Solexa and Sanger data. Linux OS required.
SEQAN - A Consistency-based Consensus Algorithm for De Novo and Reference-guided Sequence Assembly of Short Reads. By Tobias Rausch and others. C++, Linux/Win.
SHARCGS - De novo assembly of short reads. Authors are Dohm JC, Lottaz C, Borodina T and Himmelbauer H. from the Max-Planck-Institute for Molecular Genetics.
SSAKE - The Short Sequence Assembly by K-mer search and 3' read Extension (SSAKE) is a genomics application for aggressively assembling millions of short nucleotide sequences by progressively searching for perfect 3'-most k-mers using a DNA prefix tree. Authors are René Warren, Granger Sutton, Steven Jones and Robert Holt from the Canada's Michael Smith Genome Sciences Centre. Perl/Linux.
SOAPdenovo - Part of the SOAP suite. See above.
VCAKE - De novo assembly of short reads with robust error correction. An improvement on early versions of SSAKE.
Velvet - Velvet is a de novo genomic assembler specially designed for short read sequencing technologies, such as Solexa or 454. Need about 20-25X coverage and paired reads. Developed by Daniel Zerbino and Ewan Birney at the European Bioinformatics Institute (EMBL-EBI).
SOAP (http://soap.genomics.org.cn) by Ruiqiang Li, as has been pointed by ECO.
Euler-SR (Euler-Short Reads Assembly, http://euler-assembler.ucsd.edu/portal/) by Mark J. Chaisson and Pavel A. Pevzner from UCSD. (published in Genome Research)
RMAP (A program for mapping Solexa reads, http://rulai.cshl.edu/rmap/) by Andrew D. Smith and Zhenyu Xuan at CSHL. (published in BMC Bioinformatics)
Short read aligner called Bowtie (http://bowtie-bio.sourceforge.net/) designed for fast mapping of Illumina reads

SNP/Indel Discovery
ssahaSNP - ssahaSNP is a polymorphism detection tool. It detects homozygous SNPs and indels by aligning shotgun reads to the finished genome sequence. Highly repetitive elements are filtered out by ignoring those kmer words with high occurrence numbers. More tuned for ABI Sanger reads. Developers are Adam Spargo and Zemin Ning from the Sanger Centre. Compaq Alpha, Linux-64, Linux-32, Solaris and Mac
PolyBayesShort - A re-incarnation of the PolyBayes SNP discovery tool developed by Gabor Marth at Washington University. This version is specifically optimized for the analysis of large numbers (millions) of high-throughput next-generation sequencer reads, aligned to whole chromosomes of model organism or mammalian genomes. Developers at Boston College. Linux-64 and Linux-32.
PyroBayes - PyroBayes is a novel base caller for pyrosequences from the 454 Life Sciences sequencing machines. It was designed to assign more accurate base quality estimates to the 454 pyrosequences. Developers at Boston College.
Maq is also able to find SNPs with its own alignment. It has a graphical viewer, but again for its own alignment format.
SSAHA has been optimized for short-reads, too. But yes, SSAHASNP appears in your "SNP/INDEL discovery" category.

Genome Annotation/Genome Browser/Alignment Viewer/Assembly Database
EagleView - An information-rich genome assembler viewer. EagleView can display a dozen different types of information including base quality and flowgram signal. Developers at Boston College.
LookSeq - LookSeq is a web-based application for alignment visualization, browsing and analysis of genome sequence data. LookSeq supports multiple sequencing technologies, alignment sources, and viewing modes; low or high-depth read pileups; and easy visualization of putative single nucleotide and structural variation. From the Sanger Centre.
MapView - MapView: visualization of short reads alignment on desktop computer. From the Evolutionary Genomics Lab at Sun-Yat Sen University, China. Linux.
SAM - Sequence Assembly Manager. Whole Genome Assembly (WGA) Management and Visualization Tool. It provides a generic platform for manipulating, analyzing and viewing WGA data, regardless of input type. Developers are Rene Warren, Yaron Butterfield, Asim Siddiqui and Steven Jones at Canada's Michael Smith Genome Sciences Centre. MySQL backend and Perl-CGI web-based frontend/Linux.
STADEN - Includes GAP4. GAP5 once completed will handle next-gen sequencing data. A partially implemented test version is available here
XMatchView - A visual tool for analyzing cross_match alignments. Developed by Rene Warren and Steven Jones at Canada's Michael Smith Genome Sciences Centre. Python/Win or Linux.

Counting e.g. CHiP-Seq, Bis-Seq, CNV-Seq
BS-Seq - The source code and data for the "Shotgun Bisulphite Sequencing of the Arabidopsis Genome Reveals DNA Methylation Patterning" Nature paper by Cokus et al. (Steve Jacobsen's lab at UCLA). POSIX.
CHiPSeq - Program used by Johnson et al. (2007) in their Science publication
CNV-Seq - CNV-seq, a new method to detect copy number variation using high-throughput sequencing. Chao Xie and Martti T Tammi at the National University of Singapore. Perl/R.
FindPeaks - perform analysis of ChIP-Seq experiments. It uses a naive algorithm for identifying regions of high coverage, which represent Chromatin Immunoprecipitation enrichment of sequence fragments, indicating the location of a bound protein of interest. Original algorithm by Matthew Bainbridge, in collaboration with Gordon Robertson. Current code and implementation by Anthony Fejes. Authors are from the Canada's Michael Smith Genome Sciences Centre. JAVA/OS independent. Latest versions available as part of the Vancouver Short Read Analysis Package
MACS - Model-based Analysis for ChIP-Seq. MACS empirically models the length of the sequenced ChIP fragments, which tends to be shorter than sonication or library construction size estimates, and uses it to improve the spatial resolution of predicted binding sites. MACS also uses a dynamic Poisson distribution to effectively capture local biases in the genome sequence, allowing for more sensitive and robust prediction. Written by Yong Zhang and Tao Liu from Xiaole Shirley Liu's Lab.
PeakSeq - PeakSeq: Systematic Scoring of ChIP-Seq Experiments Relative to Controls. a two-pass approach for scoring ChIP-Seq data relative to controls. The first pass identifies putative binding sites and compensates for variation in the mappability of sequences across the genome. The second pass filters out sites that are not significantly enriched compared to the normalized input DNA and computes a precise enrichment and significance. By Rozowsky J et al. C/Perl.
QuEST - Quantitative Enrichment of Sequence Tags. Sidow and Myers Labs at Stanford. From the 2008 publication Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data. (C++)
SISSRs - Site Identification from Short Sequence Reads. BED file input. Raja Jothi @ NIH. Perl.
SeqMap (http://biogibbs.stanford.edu/~jiangh/SeqMap/) - work like ELand, can do 3 or more bp mismatches and also insdel
ChIPSeq analysis is: http://dir.nhlbi.nih.gov/papers/lmi/epigenomes/sissrs/

See also this thread for ChIP-Seq, until I get time to update this list.

Alternate Base Calling
Rolexa - R-based framework for base calling of Solexa data. Project publication
Alta-cyclic - "a novel Illumina Genome-Analyzer (Solexa) base caller"

Transcriptomics
ERANGE - Mapping and Quantifying Mammalian Transcriptomes by RNA-Seq. Supports Bowtie, BLAT and ELAND. From the Wold lab.
G-Mo.R-Se - G-Mo.R-Se is a method aimed at using RNA-Seq short reads to build de novo gene models. First, candidate exons are built directly from the positions of the reads mapped on the genome (without any ab initio assembly of the reads), and all the possible splice junctions between those exons are tested against unmapped reads. From CNS in France.
MapNext - MapNext: A software tool for spliced and unspliced alignments and SNP detection of short sequence reads. From the Evolutionary Genomics Lab at Sun-Yat Sen University, China.
QPalma - Optimal Spliced Alignments of Short Sequence Reads. Authors are Fabio De Bona, Stephan Ossowski, Korbinian Schneeberger, and Gunnar Rätsch. A paper is available.
RSAT - RSAT: RNA-Seq Analysis Tools. RNASAT is developed and maintained by Hui Jiang at Stanford University.
TopHat - TopHat is a fast splice junction mapper for RNA-Seq reads. It aligns RNA-Seq reads to mammalian-sized genomes using the ultra high-throughput short read aligner Bowtie, and then analyzes the mapping results to identify splice junctions between exons. TopHat is a collaborative effort between the University of Maryland and the University of California, Berkeley
NGS-Trex: Next Generation Sequencing Transcriptome profile explorer http://www.biomedcentral.com/1471-2105/14/S7/S10

Reference

Illumina has a software list: http://www.illumina.com/pagesnrn.ilmn?ID=245.

Some softwares in his blog (http://www.fejes.ca/labels/DNA.html)

http://seqanswers.com/wiki/Software

BioinfoLab

Fri, 25 Mar 2016 11:05:35 -0500

Laboratory of Statistics and Computational tools for Bioinformatics

The Laboratory of Statistics and Computational tools for Bioinformatics (BioinfoLab) is hosted at the Istituto per le Applicazioni del Calcolo "Mauro Picone" - CNR . The laboratory has been officially opened in 2012 with the support of Programma Operativo Nazionale "Ricerca e Competitività" 2007-2013 (PON "R&C"), and it incorporates several expertise and research activities started since 2007, and supported by several CNR projects. Main interest of BioinfoLab is to develop novel statistical methods and computational tools for the analysis of high dimensional data arising from "Multi-omics" applications. In particular, current activities involve the analysis of ChIP-seq and RNA-seq experiments.

More at http://bioinfo.na.iac.cnr.it/BioinfoLab/index.html

List of Bioinformatics Software Tools for Next Generation Sequencing

Jitendra Prajapati — Fri, 11 Mar 2016 20:22:14 -0600

Commercial tools

Strand NGS
- offers many different tools including alignment, RNA-Seq, DNA-Seq, ChIP-Seq, Small RNA-Seq, Genome Browser, visualizations, Biological Interpretation, etc. Supports workflows “one can import the sample data in FASTA, FASTQ or tag-count format. In addition, prealigned data in SAM, BAM or Illumina-specific ELAND format can be directly imported for analysis.”
- Alignment feature: Supports alignment from Illumina, Ion Torrent, 454 (Roche), and Pac Bio
- DNA-Seq Feature, can annotate with dbSNP
CLC Genomics Workbench
- (QIAGEN). Features include: resequencing, workflow, read mapping, de novo assembly, variant detection, RNA-Seq, ChIP-Seq, Genome Browser, etc (entire list on website); Main Workbench offers database search (Genbank, Blast, Pubmed); 2000 organizations have invested in CLC
- Accepts VCF files from 1000 Genomes Project
- Accepts downloaded tracks from dbSNP
- Also accepts: FASTA, GFF/GTF/GVF, BED, Wiggle, Cosmic, UCSC variant database, complete genomics master var file
- Read mapping: “In addition to Sanger sequence data, reads from these high-throughput sequencing machines are supported: The 454 FLX System and the 454 GS Junior System from Roche, Illumina Genome Analyzer, Illumina HiSeq, Illumina HiScan, and Illumina MiSeq sequencing systems, SOLiD system from Life Technologies, Ion Torrent system from Life Technologies, Helicos from Helicos BioSciences”
- De novo assembly: “In addition to Sanger sequence data, reads from these high-throughput sequencing machines are supported The 454 FLX System and the 454 GS Junior System from Roche, Illumina Genome Analyzer, Illumina HiSeq, Illumina HiScan, and Illumina MiSeq sequencing systems, SOLiD system from Life Technologies, Ion Torrent system from Life Technologies”
- Annotation tracks from Ensembl
DNAnexus
- Private cloud repository -- formerly a redistributor of SRA and other NCBI resources; command-line or via web, can fetch data from a URL, build custom pipeline/ workflow has sra.dnanexus.com site: data downloads come directly from NCBI
Ingenuity Variant Analysis
- (QIAGEN) allows for variant identification and analysis, uses NCI-60 data set for cancer, Supported third part informatin: Entrez Gene, RefSeq, ClinVar; gives contextual details of results instead of just A to B relationship
- Has own database-- “knowledge base” based on COSMIC, OMIM, and TCGA databases
Lasergene Genomics Suite
- Comprehensive NGS software pipeline for assembly, alignment, variant calling and analysis of NGS data
- Supported workflows include: reference-guided and de novo genome and transcriptome assembly and analysis, metagenomics sample assembly, targeted resequencing, exome alignment, gene panels with validation control, variant analysis, and RNA-Seq, ChIP-Seq and miRNA alignment and analysis.
- #1 in accuracy: fewer false negatives and better sensitivity compared to results obtained from other aligners
- Aligns exome data and performs variant calling an average of 3 times faster than alternative pipelines
- Annotates genomic data with allele and genotype frequency, functional impact predictions, evolutionary conservation scores and pathogenicity
- Supports all major NGS technologies (Illumina, Ion Torrent, Pac Bio and Roche 454) and project types
- Available on Windows, Mac OS X, Linux, and the Amazon Cloud
NextGENe
- “perfect analytical partner for the analysis of desktop sequencing data produced by the ION PGM™, Roche Junior, Illumina MiSeq as well as high throughput systems as the Ion Torrent Proton, Roche FLX, Applied BioSystems SOLiD™ and Illumina® platforms.” runs on Windows, free-standing multi-application package-- SNP/Indel analysis, CNV prediction and disease discovery, whole genome alignment, etc.
- Data can be imported from Clinvar, dbSNP, Genbank:http://www.softgenetics.com/PDF/NextGene_UsersManual_web.pdf
Partek Genomics Suite
- Cited in over 3,500 peer-reviewed scientific publications
- Workflows for microarray and PCR data include: Gene expression including alternative splicing, miRNA expression, Genome Wide Association Studies, Mother-Father-Child Trio analysis, DNA Copy number including allele specific copy number and Loss of Heterozygosity (LOH), and ChIP, and methylation. Next Generation Sequencing (NGS) workflows include: RNA-Seq, miRNA-Seq, ChIP-Seq, DNA-Seq, and Methylation
- Powerful statistics and interactive, publication ready visualizations
- Supports all commercial next generation sequencing and microarray file format as well as text files
- Can input GEO SOFT files
Partek Flow
- Installation can be cloud-based or on a local cluster or Linux server
- Easy to use point-and-click interface
- Takes NGS data (.fastq, BAM, SAM), microarrays (Affymetrix, Illumina) and text files
- Supports custom genome builds and annotation databases
- Performs base trimming, alignment, quantification, quality analysis, statistics, and visualization
- Includes ten fully customizable aligners (Bowtie, Bowtie 2, BWA, GSNAP, Isaac 2, SHRiMP 2, STAR, TMAP, TopHat and TopHat 2)
- Applications for RNA-Seq, Small RNA-Seq, WGS/WES, Pathway enrichment, Fusion detection and Variant calling
- Allows users to create, save, share, or download analysis pipelines for automated and repeatable analysis
- Collaborate with others without transferring data
- Integrates microarray and next generation sequencing data
Golden Helix: SNP and Variation Suite
- used for managing, analyzing and visualizing genotypic and phenotypic data; Features: Genome-wide association studies, genomic prediction, copy number analysis, small sample DNA-Seq workflows, large sample DNA-seq analysis, RNA-seq analysis. Supported files: .txt, excel XLS & XLSX, CEL, CHP, CNT, Illumina, Plink PED, TPED, BED, Agilent files, NimbleGen data summary files, VCF files, Impute2 GWAS files, HapMap format, MACH output, + 50 other formats consumes NCBI data directly
Genomatix
- Applications: ChIP-Seq, DNA-Seq, RNA-Seq, DNA methylation; enable personalized medicine,
- Mining Stations: Supports all established NGS sequencing platforms- SOLiD, 454 Life Sciences, Genome Analyzer, HiSeq, MiSeq, IonTorrent
- Software Suite: can upload sequence of BED files
- Genome browser: BED and BAM files, Public data- 1500 BED files available for every user
Biodatomics
- Open source platform (SaaS), analysis and genome sequencing tools, integrates over 400 genomic analysis open source tools and pipelines, have a private and public cloud version. Features: genomic data visualization, drag and drop interface, accelerated analysis, real-time collaboration
- They have a couple modules to do so, and have enabled parts of the sra toolkit
SolveBio
- Software product, for clinical genomics professionals, manage, curate, report genomic variation
- Has own data library -- data from NCBI
Basepair
- Offers high quality workflows for all common NGS applications (RNA-Seq, ChIP-Seq, DNA-Seq, etc.)
- Very fast - get all results in a 1-2 hours. Cloud-based, no storage or computing limits.
- Easy to use - less than a minute to run an analysis
- REST and Python API to mange large projects.

Variant Identification

Germline Callers

IMPUTE2
- Description: phasing observed genotypes and imputing missing genotypes uses reference panels to provide all available halotypes, does not use population labels or genome-wide measures; designed to represent variation in one population; Fairly popular
- Input:
- Reference Haplotypes: Links to 1000 Genomes and HapMap downloads
- Output:
FreeBayes
- Description: finds SNPs, Indels, MNPs; reports variants based on alignment; haplotype based
- Input: BAM- uses BAMtools API to parse
- Reference genome: FASTA
- Output: VCF
SOAPindel
- Description: detects indels from NGS paired-end sequencing
- Input: files with read alignment can be SOAP or SAM formats, users must also give raw reads in Fasta or Fastq
- Reference Sequence used to align reads: FASTA
- Output:
2Kplus2
- Description: algorithm searches graphs produced by de novo assembler Cortex; c++ source code for SNP detection “2kplus2.cpp is a c++ source code for the detection and the classification of single nucleotide polymorphisms in transformed De Bruijn graphs using Cortex assembler.”
- Input:
- Output:
Atlas 2
- Description: specializes in separation of true SNPs and indels from sequencing and mapping errors, last update January 2013
- Input: takes BAM file,
- Reference Genome: FASTA
- Output: produces VCF
CRISP
- Description: identifies SNPs and INDELs from pooled high-throughput NGS, not used for analysis of single samples; implemented in C and uses SAMtools API; latest version should work with diploid genomes
- Input: requires BAM files (aligned with GATK)
- Reference Genome: indexed FASTA file
- Output: VCF files
Dindel
- Description: (Wellcome Trust Sanger) calls small indels from short-read sequences, only can handle Illumina data; cannot test candidate indels; written in C++, used on Linux based and Mac computers (not tested in windows)
- Input: BAM files
- Output: VCF
discoSnp++
- Description: detects homozygous and heterozygous SNPs and Indels; software composed of 2 modules (kissnp2 and kissreads)
- Input: raw NGS datasets; fasta, fastq, gzipped or not;
- no reference genome required; read pairs can be given
- Output: FASTA
FamSeq
- Description: family-based sequencing studies- provides probability of an individual carrying variant based on family’s raw measurements; accommodates de novo mutations, can perform variant calling at chrX;
- Input: VCF
- Output: VCF
GeneticThesaurus
- Description: “Annotation of genetic variants in repetitive regions”
- Input: Initial variant calling from bam → vcf output
- Reference Genome: need to provide own fasta file for hg19 genome,
- Output: vcf.gz, vtf.gz, and baf.tsv.gz output
glfMultiples
- Description: command-line, variant caller
- Input: GLF
- Output: VCF
glfSingle
- Description: uses likelihood-based model for variant calling, starts from genotype likelihoods that have been computed from other tools (ex. Samtools BAQ), the likelihoods combine with individual-based prior p(genotype) to generate posterior probabilities
- Input: GLF
- Output: VCF
Halvade
- Description: command-line; written in Java, “to run halvade a reference is needed for both GATK and BWA and a SNP (dbSNP!) database is required
- Input: FASTQ
- Output: VCF
indelMINER
- Description: identifies indels from paired-end reads
- Input: BAM (aligned in SAMtools API)
- Output: VCF
Indelocator
- Description: (Broad Institute): does not perform realignment, relies on alignments in BAM files (BAM files need aligned before put into indelocator); recommended to use GATK prior;
- Input: 2 BAM files(tumor & normal), annotated as germline or somatic; also has single sample mode
- Output: “Output of Indelocator is a high-sensitivity list of putative indel events containing large numbers of false positives. The statistics reported for each event have to be used to custom-filter the list in order to lower false positive rate”
Isaac Variant Caller
- Description: detects SNPs and small indels from diploid sample; designed to run on “nux-like platforms”
- Input: BAM
- Output: VCF
KvarQ
- Description: in silico genotyping for selected loci in bacterial genome, written in Python and C
- Input: FASTQ
- reference genome or de novo assembly not needed
- Output:
LoFreq
- Description: SNV caller, Python language, standalone program, uncovers cell-population heterogeneity from high-throughput sequencing datasets; calls variants found in <.05% of the population
- Input: BAM file input→ suggest running through GATK
- Output:
Manta
- Description: Calls indels and SVs from paired end reads; standalone, command line program; Written in C++ and Python
- Input: BAM (can tolerate non-paired-end reads); a matched tumor sample may be provided as well
- Output: VCF
MarginAlign
- Description: SNV caller, specifically tailored to Oxford Nanopore Reads, written in Python; Package comes with 3 programs, marginAlign, marginCaller (calls SNVs), marginStats (computes qc stats on sam files)
- Input: SAM
- Output: SAM
MendelScan
- Description: Last release March 2014; for analyzing sequencing data in family studies of inherited diseases; variant calls for a family in VCF file; still in alpha-testing on github, example data uses 1000 genomes dataset
- Input:
- Output:
nanopore
- Description: UCSC Nanopore group (group at UCSC studying using ion channels for analysis of single RNA/DNA structures) software pipeline; tailored to Oxford Nanopore Reads; command line program
- Input: FASTQ
- Reference files: FASTA
- Output: “For each possible pair of read file, reference genome and mapping algorithm an experiment directory will be created in the nanopore/output directory.”
Platypus
- Description: Package program, written in C, Python, Cython; Can identify SNPs, MNPs, short indels, and larger variants; has been tested on very large datasets (1000 genomes)
- Input: BAM
- Reference Genome: FASTA (files must be indexed using Samtools or similar program
- Output: VCF
QualitySNPng
- Description: detection of SNPs; “can be used as a standalone application with graphical user interface as part of pipeline system”; does not require fully sequenced reference genome; haplotype strategy
- Input:SAM, ACE
- Output: GUI
ReviSTER
- Description: command line program; automated pipeline; utilizes BWA, BLAT, and SAMTools; utilizes BWA mapping program;
- Input: FASTQ,
- Reference sequence file and list file containing STR locations as inputs
- Output: SAM
RVD
- Description: command-line program, detection of rare SNVs, relies upon Samtools, can be run in MATLAB
- Input: BAM
- Reference Genome: FASTA
- Output: “The algorithm output is a call table -- a comma-separated file with one line for each base position and each line in the following format:
- AlginmentReferencePosition, AlignmentBase, Call ,SecondBase, CenteredErrorPrc, ReferenceErrorPrc, SecondBasePrc”
SNVer
- Description: calls common and rare variants in pool or individual NGS data, reports overall p-value, operating system independent statistical tool, identifies SNPs and INDELs, written in Java, no dependencies, straightforward command-line
- (SNVerGUI=GUI version) --SNVerGUI: desktop tool for variant detection
- Input: chrX annotation, sam.zip, bam.zip
- reference file must be aligned to the data file
- Output:
SNVMix
- Description: detects SNVs from NGS, post-alignment tool
- Input: pileupformat (Maq or Samtools)
- Output:
SV-M
- Description: Structural Variant Machine - predicts indels, uses split read alignment profiles, validated by Sanger Sequencng
- Input:paired-end Illumina reads from 1001 genomes project (uses ref plant- 1001genomes.org)
- Ouptut:
SNPest
- Description: Standalone program, language C++, Perl
- Input: mpileup (SAMtools)
- Output: VCF
TrioCaller
- Description:Command line program, relies on BWA and samtools; genotype calling for unrelated individuals and parent-offspring trios
- Input: BAM (that has been aligned in BWA and Samtools
- Output: BCF that can be formatted to VCF using bcftools
Snippy
- Description: finds indels between haploid reference genome and NGS sequence reads
- Input:read files- FASTQ or FASTA (can be .gz compressed), output- .aln, .tab, .txt
- Reference genome in FASTA or GENBANK
- Output:
VntrSeek
- Description: pipeline for discovering microsatellite tandem repeats with high-throughput sequencing data
- Input: gzip-compressed FASTA or FASTQ
- Output: VCF files; one for TRs and observed alleles, another file contains link to viewer

Somatic Callers

Cake
- Description: standalone program, “pipeline for the integrated analysis of somatic variants in cancer genomes”; integrates four algorithms; written in Perl; required tools: samtools, tabix, vcftools, VarScan2, bambino, cmake, somaticsniper (User guide; workflow page)
- Input: tumor and normal reads in BAM files, run through variant calling programs to generate intermediate VCF
- Output: VCF
MuTect
- Description: Broad Institute, identification of somatic point mutations in cancer genomes; requires preprocessing of reads (GATK)
- Input: same as GATK (FASTA reference genome, SAM read files)
- Output: call-stats, VCF, wiggle files
Polymutt
- Description: calls SNVs and detects de novo point mutations in families
- Input: GLF or BAM or VCF (must have identical chromosome orders)
- Output: VCF
Bassovac
- Description: Improved Bayesian inversion somatic caller; unlike other software packages, treats effects fully probabilisticallys instead of using ad-hoc modeling; effects are integrated at the atomic level and standard probability theory integrates read tallies to the sample level and to the tumor-normal pair level; "pending public release"
- Input:
- Output:
CLImAT
- Description: standalone program; “accurate detection of copy number alteration and loss of heterozygosity in impure and aneuploid tumor samples using whole genome sequencing data”
- Input: depth file generated by DFExtract and a config file
- Output: .results file, .Gtype, LOG.txt, also generates visualization
DeNovoGear
- Description: de-novo variant calling and interpretation; standalone program; dependencies C++ compiler, CMake, HTSlib, Eigen, Boost
- Input: PED and BCF
- Output: “The output format is a single row for each putative de novo mutation (DNM), with the following fields”
EBCall
- Description: Empirical Baysian Mutation Calling; standalone program; uses tumor/normal paired reads and non-paired normal reference samples; dependent on samtools, R and VGAM pack for R
- Input: BAM
- Output: not sure what exact type of file- “The format of the result is suitable for adding annotation by annovar.”
HapMuc
- Description: standalone program; “utilizes the information of heterozygous germline variants near candidate mutations”; Dependent upon- Boost, SAMtools, BEDtools; 3 step workflow
- Input: BAM
- Output: BED
MultiGeMS
- Description: Multi-sample Genotype Model Selection
- Input: .txt, pileup (SAM/BAM converted to pileup format)
- Output: VCF
MultiSNV
- Description: command-line program; calls SNVs from NGS data from multiple samples from the same patient; dependent on R, Git, cmake, Boost and compile libraries
- Input: BAM or pileup
- Output: VCF
MutationSeq
- Description: standalone program, somatic SNV detection in tumor/normal samples; dependent on python, bamtools, boost, and LAPACK
- Input: BAM
- Output: VCF4.1 consisting of two parts (meta information & data lines)
qSNP
- Description: standalone program; SNV caller for somatic variants in “low cellularity cancer samples”
- Input: BAM, dbSNP data, Illumina data, chrConv
- Output: “qSNP output files are named using a 4-element pattern: ...”
RADIA
- Description: RNA and DNA Integrated Analysis for Somatic Mutation Detection; DNA only Method(tumor/normal pair, ignores RNA) or Triple BAM Method (uses all three datasets from same patient); dependent upon python, samtoools, pysam API, BLAT, SnpEff
- Input: BAM
- Reference Genome: FASTA indexed with SAMtools faidx
- Output: VCF
RVD2
- Description: sensitive, variant detection for low-depth targeted NGS data; python module or command- line program;
- Input: tab- deliminted depth chart format (converted from pileup files)
- Output: three hdf5 files and a vcf file
Shimmer
- Description: standalone program; detects somatic SNVs with multiple testing correction, uses Fisher’s exact test; dependent on git, samtools, R, R statmod package; for tumor/normal matched samples
- Input: BAM
- Output: VCF
SNV-PPILP
- Description: Refines GATK’s Unified Genotyper SNV calls for “multiple samples assumed to form a phylogeny”
- Input:
- Output:
SomaticSniper
- Description: command-line application to identify SNPs between tumor/normal pairs- predicts probability of difference between two
- Input: BAM
- Reference Genome in FASTA
- Output: VCF
Strelka
- Description: somatic variant calling workflow for matched tumor-normal samples; detects indels; runs on *nux-like platform
- Input: BAM (must be sorted and indexed)- Strelka does own realignment around indels-- don’t need to do this type of pre-processing
- Output: pair of VCF files
Triodenovo
- Description: Bayesian framework for calling de novo mutations in trios
- Input: VCF file with PL or GL fields (recommend using GATK or samtools to generate)
- Output: out_vcf
UNCeqr
- Description: finds somatic mutations using integration of DNA and RNA seq data-- boosts sensitivity for low purity tumors and rare mutations;
- Input:”can accept a variety of sequencing inputs and configurations”
- Output: “table of somatically mutated sites and associated information. These somatic mutations can be annotated with predicted transcript and protein effects using third party tools, such as Annovar”
Virmid
- Description: Virtual Microdissection for SNP calling; Java based; for disease-control matched samples; uncovers SNPs with low allele frequency by considering alpha contamination
- Input: BAM (must be sorted and indexed- samtools sort)
- Output: VCF and report file

Germline + Somatic Callers

VarScan 2
- Description: identify germline variants, private and shared variants, somatic mutations, and somatic CNVs; detects indels
- Input: SAMtools pileup
- Output: VCF
BAYSIC
- Description: Bayesian method; combines variant calls from different methods (GATK, FreeBayes, Atlas, Samtools, etc)
- Input: VCF format from one or more variant calling programs
- Output: VCF file containing integrated set of variant calls
MSIsensor
- Description: Microsatellite instability detection; C++ program, detects somatic and germline variants in tumor-normal paired data
- Input: BAM index files (normal and tumor)
- Output:
Beagle version 4
- Description: software package: genotype calling, phasing, imputation of ungenotyped markers, and identity-by-descent segment detection:unsure if this one is in the right category; genotype calling, phasing, imputation of ungenotyped markers, and identity-by-descent segment detection;
- Input: VCF
- Output: VCF
QuadGT
- Description: software package, SNV calling from normal-tumor pair and two parent genomes; quantifies descent-by-modification relationships; Written in Java
- Input: BAM files (parsed by Picard/Samtools API)
- Reference Genome; FASTA
- Output: VCF
RAREVATOR
- Description: RAre REference VAriant annotaTOR; command line; “identification and annotation of germline and somatic variants in rare reference allele loci from second generation sequencing data”; Bayesian genotype likelihood model
- Input: BED or VCF files from GATK
- Output: two VCF files (one for SNVs, one for Indels)
Scalpel
- Description: Used for detecting indels in a reference genome; performs localized micro-assembly of specific regions of interest; can do single, de novo, somatic reads; requires that raw reads are aligned with BWA
- Input: BAM
- Output: either VCF or ANNOVAR
SOAPsnp
- Description: based on Baye’s theorem; calls consensus genotype
- Input:SOAP short read alignment results
- Output: GLF, option of flat tabular format
VariantMaster
- Description: “extract causative variants for monogenic and sporadic genetic diseases”; uses ANNOVAR;
- Input: BAM or VCF files (from SAMtools, GATK)
- Output:

Downstream Analysis of Variants

PrediXcan
- Description: command-line, standalone package program; available in Perl, Python, and R versions; predicts liklihood of a gene being related to a certain phenotype- “that directly tests the molecular mechanisms through which genetic variation affects phenotype.”; no actual expression data used, only in silico expression; “PrediXcan can detect known and novel genes associated with disease traits and provide insights into the mechanism of these associations.”
- Input: genotype and phenotype file (doesn’t specify file type)
- Output:default values: genelist, dosages (file format: snpid rsid) , dosage_prefix, weights, output
ATHENA
- Description: Analysis Tool for Heritable and Environmental Network Associations; software package, combines machine learning model with biology and statistics to predict non-linear interactions
- Input: Configuration file, Data file, Map file (includes rsID)
- Output: Summary file, Best model file, dot file, individual score file, cross-validation file
CCRaVAT and QuTie
- Description: (Wellcome Trust Sanger) Case-Control Rare Variant Analysis Tool and Quantitative Trait; software packages for large-scale analysis of rare variants
- Input: PED file and MAP file
- Output: Five tab-delimited txt files
GCTA
- Description: Genome Wide Complex Trait Analysis; package program, command line interface; estimates variance by all SNPs; 5 main functions: “data management, estimation of the genetic relationships from SNPs, mixed linear model analysis of variance explained by the SNPs, estimation of the linkage disequilibrium structure, and GWAS simulation”
- Input: PLINK binary PED files, MACH output format
- Output:
GenomeComb
- Description: package for analysis of complete genome data; annotation using public data or custom tracks, automated primer desing for Sanger or Sequenom validation; “The cg process_illumina command can be used to generate annotated multisample data starting from fastq files, using tools such as bwa for alignment and GATK and samtools for variant calling. Sequencing data can also be imported from Complete Genomics (cg_process_sample command), Real Time Genomics (cg_process_rtgsample command) and VariantCallFormat (VCF) variant files (vcf2sft command).”
- Input: Sequencing data from Complete Genomics, Illumina, SOLiD and VCF;
- Output: standard file format used is a simple tab delimited file (.sft, .tsv)
Genome Track Analyzer
- Description: compares genome tracks; allows user to compare DNA expression/binding;
- Input: multiple: SGR/TXT, BED, BED6, GFF; if using prealigned sequence data- use MACS peak caller: BAM, BED, SAM, ELAND
- Output:
GVCBLUP
- Description: animal gene mapping; “genomic prediction and variance component estimation of additive and dominance effects”; standalone program, command line interface, writting in C++ and Java
- Input:
- Output:
HOMOG
- Description: Analyzes heterogeneity with respect to single marker loci or known maps of markers; Carries out homogeneity test for alternative hypothesis “Two family types, one with linkage betweeen a trait to a marker or map of markers, the other without linkage”
- Input: HOMOG.DAT - described on website
- Output: HOMOG.OUT
INTERSNP
- Description: GWIA for case-control SNP and quantitative traits; selected for joint analysis using priori information; Provides linear regression framework, Pathway Association Analysis, Genome-wide Haplotype Analysis,
- Input: PLINK input formats (ped/map, tped/tfam, bed/bim/fam) Compatible with SetID files
- Gene reference file: Ensembl Release 75
- Output: covariance matrix for regression models
mtSet
- Description: Currently only the standalone version available, but moving to LIMIX software suite; offers set tests- allows for testing between variants and traits; accounts for confounding factors ex. relatedness
- Input: sample-to-sample genetic covariance matrix needs to be computed; multiple types of input; simulator requires input genotype and relatedness component;
- Output: resdir (result file of analysis), outfile (test statistics and p-values), manhattan_plot (flag)
MultiBLUP
- Description: Package program, command line interface; constructs linear prediction models; Best Linear Unbiased Prediction; improves upon BLUP involving kinship matrices; options: pre-specified kinships, regional kinships, adaptive multiblups, LD weightings
- Input: PLINK format
- Output:.reml, .indi.blp

Variant Annotation

ANNOVAR
- Description: command-line tool, supports SNPs, INDELs, CNVs and block substitutions, provides wide variety of annotation techniques, depends upon multiple databases (each needing to be downloaded); annotates genetic variants; utilizes RefSeq, UCSC Genes, and the Ensembl gene annotation systems; can compare mutations detected in dpSNP or 1000 Genomes Project; Very popular *“The final command run TABLE_ANNOVAR, using dbSNP version 138, 1000 Genomes Project 2014 Oct version, NIH-NHLBI 6500 exome database version 2 (referred to as esp6400siv2), dbNFSP version 2.6 (referred to as ljb26), dbSNP version 138 (referred to as snp138) databases and remove all temporary files, and generates the output file called myanno.hg19_multianno.txt”
- Input: VCF, ANNOVAR input format (simple text-based format); can convert other formats into ANNOVAR input format
- Output: VCF (if input VCF), output file with multiple columns, tab-delimited output file
wANNOVAR
- provides web-based access to ANNOVAR software
PolyPhen-2
- Description: Very popular; Polymorphism Phenotyping; Web application; predicts impact of amino acid substitution on protein; Calculates Bayes posterior probability (Last update July 2015)
- Input: FASTA
- Output:
SIFT
- Description: predicts how an amino acid substitution will affect protein function; Based on degree of conservation of amino acid residues- collected though PSI-BLAST; can be applied to nonsynonymous polymorphisms or laboratory-induced missense mutations; links to dbSNP 132, GRCh37; Standalone or web app program; Very popular
- Input: Uniprot ID or Accession, Go term ID, Function name, Species Name or ID, etc
- Output:
snpEff
- Description: Genetic variant annotation and effect prediction toolbox; integrated with Galaxy, GATK, and GNKO; can annotate SNPs, INDELs, and multiple-nucleotide polymorphisms; categorizes effects into classes by functionality; Very popular; Standalone or Web app; Claims to calculate all SNPs in 1000 genomes (EMBI) in less than 15 minutes; can annotate SNPs, MNPs, and insertions and deletions; Provides assessment of impact of the variant ( low, medium or high)
- Input: VCF, BED
- Output: VCF (with new ANN field, also used in ANNOVAR and VEP), HTML summary files
SnpSIFT
- Description: Filter and manipulate annotated files; Part of SnpEff main distribution; one variants have been annotated, this can be used to filter your data to find relevant variants
- Input:
- Output:
VAAST 2
- Description: Variant Annotation, Analysis, and Search Tool; probabilistic search tool for identifying damage genes and the disease causing variants; can score both coding and non-coding variants; Four tools: VAT (Variant annotation tool), VST (Variant Selection Tool), VAAST, pVAAST (for pedigree data); updated April 2015
- Input: FASTA, GFF3, GVF
- Output: CDR (condenser file), VAAST file (both unique to VAAST)
VEP
- Description: (Ensembl) Variant Effect Predictor; determines effect of variants on genes, transcripts, and protein sequence; uses SIFT and PolyPhen
- Input: Coordinates of variants and nucleotide changes; whitespace- separated format, VCF, pileup, HGVS
- Output: VCF, JSON, Statistics
ABSOLUTE
- Description: (Broad Institute); can estimate purity and ploidy to compute absolute copy number and mutation multiplicitie; reextracts data from the mixed DNA population
- Input: HAPSEQ segdat or segmentation file
- Output: per-sample output directory and subdirectory providing per-sample text files containing standard out being emitted from R
Alamut Batch
- Description: high-throughput annotation software for NGS analysis; for “intensive variant analysis workflows”; “enriches raw NGS variants with dozens of attributes”; based on clinically oriented Alamut database; Supports human genes; easy to integrate into pipeline (Latest Release- July 2015)
- Input:VCF, tab-delimted file
- Output: tab-separated file of annotations
AVIA
- Description: Annotation, Visualization, and Impact Analysis; “The tool is based on coupling a comprehensive annotation pipeline with a flexible visualization method. We leveraged the ANNOVAR (Wang et. al, 2010) framework for assigning functional impact to genomic variations by extending its list of reference annotation databases (RefSeq, UCSC, SIFT, Polyphen etc.) with additional in-house developed sources (Non-B DB, PolyBrowse).”
- Input: BED
- Output: Table of annotations with gene annotation features
BioR
- Description: (Mayo Clinic) (Page last updated June 2015) Biological Reference Repository; “data integration tool that enables coordinate based searches and joins based on strings”; “BioR consists of two parts 1) the BioR toolkit which depends on Java…. 2) the BioR catalogs which are the data files used by the system”
- Input: VCF
- BioR-Supported Catalogs (tar-gzip files): dbSNP, 1000 genomes, HapMap, OMIM, NCBIGene
- Output: VCF + JSON
CADD
- Description: Combined Annotation Dependent Depletion; tool for scoring SNV deletions/insertions; “integrates multiple annotations into one metric”; Score strongly correlates with allelic diversity and pathogenicity; links to 1000 Genome variants; uses Ensembl Variant Effect Predictor
- Input: VCF
- Output: CADD score
CandiSNPer
- Description: web application, characterizes SNPs located in vicinity of SNP of interest;
- Input: enter SNP ID (rsID), choose population, region, measure for LD, threshold plot format, color of SNPs, and chose to show genes
- Output: Imagefile
CanvasDB
- Description: “local database infrastructure for analysis of targeted- and whole genome re-sequencing projects”; dependent on MySQL, R, and ANNOVAR
- Input:
- Output:
CAROL
- Description: (Wellcome Trust Sanger); Combined Annotation scoRing toOL; Combined functional annotation score of nonsynonymous coding variants; Combines information from PolyPhen-2 and SIFT
- Input: tab-delimited with columns obtained from PolyPhen-2 and SIFT output
- Output: tab-delimited file
CHASM
- Description: Cancer-specific High-throughput Annotation of Somatic Mutations; Last updated May 2014; uses Random Forest Method to “distinguish between driver and passenger somatic mutations”; Positive driver class curated from COSMIC database; packed together with SNVBox (database)
- Input:Passenger mutation rates, Transcript and amino acid change, Genomic coordinates
- Output: CHASM score, p-value, FDR
CRAVAT
- Description: Cancer-Related Analysis of Variants Toolkit; Web application; Uses CHASM, VEST, SNVGet; “CRAVAT provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases” Latest Release May 2015;
- Input: VCF, CRAVAT format
- Output: CRAVAT report- MS Excel spreadsheet or tab-separated file (emailed)
CUPSAT
- Description: Cologne University Protein Stability Analysis Tool; “tool to predict changes in protein stability upon point mutations”; web service program; Can predict mutant stability from existing PDB structures or custom protein structures
- Input:for PDB- provide PDB ID and Amino Acid Residue Number; for custom- PDB file format
- Output:
DANN
- Description: Deleterious Annotation of genetic variants; standalone program, uses “the same feature set and training data as CADD to train a deep neural network”; can catch nonlinear relationships; “There are four different datasets: training, validation, testing, and ClinVar_ESP...The ClinVar_ESP dataset is also a testing set containing a set of “gold standard” pathogenic and benign variants”
- Input:
- Output:
ESEfinder
- Description: Exonic Splicing Enhancer; useful for interpretation of point mutations/polymorphisms that are disease-associated; GUI interface; web app program
- Input: FASTA
- Output: html or plain text format, graphical display of results
Exomiser
- Description: Wellcome Trust Sanger; functionally annotates variants from whole-exome sequencing data; Based on Jannovar and uses UCSC KnownGene; Java program; web app program (Page last modified Feb 2015)
- Input: VCF
- Output: TSV, VCF
FamAn
- Description: Automated variant annotation pipeline for family-based sequencing studies; Annotaties SNVs and INDELs; 4 models- autosomal dominant, autosomal recessive, de novo mutations and a general model; “A variety of annotations are provided for each segregating variant: number of family (and family ID) each variant hits, variant genomic location and coding effect (based on snpEff), loss-of-function mutation annotation, selected ENCODE annotation, allele frequency in the 1000 Genomes Project, allele frequency in the Exome Variant Server (ESP6500), segmental duplication annotation, SIFT, PolyPhen2, LRT, MutationTaster, GERP++, PhyloP, SiPhy, etc.” (Last updated May 2014)
- Input: VCF
- Output: two excel compatible outputs
GeneTalk
- Description: Combines tool for filtering and data analysis with an online network for genetic professionals; Different degrees- basic license, premium license, in-house solution (the last ones are paid for- Commercial tool?)
- Input: VCF
- Output: GeneTalk Annotation- includes clinical data, medical relevance, scientific relevance (http://www.gene-talk.de/public/GeneTalk_Whitepaper_Annotations.pdf)
GeneVetter
- Description: “GeneVetter is a tool designed for investigation of the background prevalence of exonic variation in the Phase 3 1000 Genomes data under user defined filtering criteria”; web app program; GeneVetter uses GRch37p4 (hs37d5.fa.gz), dbSNP build 138, 1000G Phase 3, clinvar_2014072
- Input: VCF
- Output: TIMS score, summary table, PCA plot
GSITIC
- Description: (Broad Institute) Last update- July 2014; Identifies genomic regions that are significantly “amplified or deleted”; Each is given a G score; gives genomic locations and q-values from aberrant regions
- Input: segmentation file -seg, markers file -mk (required); -array file list -alf, CNV file -cnv
- Reference genome: -refgene (created in MATLAB, GISITIC provides four reference genomes: hg16.mat, hg17.mat, hg18.mat, hg19.mat
- Output: All lesions file (text file), amplifications file (text file), deletion genes file (text file), Gistic Scores file, Segmented copy number (pdf file), amplification score GISTIC plot (pdf file), Deletion score/q-vale GISTIC plot (pdf file)
HOPE
- Description: Have yOur Protein Explained; Web app program; Automatic mutant analysis server that provides structural effects of a mutation; Uses BLAST against UniProt and PDB along with homology modeling
- Input: FASTA protein sequence, or accession code of protein of interest
- Output: a report containing information from a “decision tree” and illustrated figures and animations
Human Splicing Finder
- Description: Last update: May 2013; aimed to help study pre-mRNA splicing; combines 12 algorithms to identify mutations’ effect on splicing motifs; uses ensembl database 70
- Input: Gene Name, Ensembl transcript ID, Ensembl Gene ID, Consensus CDS, RefSeq Peptide ID, or own sequence (looks like you can enter FASTA)
- Output: Chart with columns for predicted signal, predicted algorithm, cDNA position and interpretation
LARVA
- Description: Large-scale Analysis of Variants in noncoding Annotations; New version released July 2015; Command-line program; used for studying noncoding variants; integrates comprehensive set of noncoding elements, modeling their mutation count; Dependent on C++ and BEDtools
- Input: multiple
- Output:
LINKAGE
- Description:three main programs: mlink (calculates lod scores at fixed values for the recombination fraction in one interval of a genetic map), linkmap (calculates location scores for positions of a disease locus along a marker), and ilink (estimates parameters including recombination fractions, allele frequencies, penetrances, etc)
- Input: pedfile (processed by MAKEPED) and datafile (reflects loci for each individual; set in PREPLINK)
- Output:
MAC
- Description: MNV Annotation Corrector; Ad hoc software, fixes incorrect amino acid predictions that are caused by multiple nucleotide variations; Uses existing annotators ANNOVAR, SnpEff, VEP (last update April 2015) (only 1 download this week → not popular)
- Input: List of called SNVs and corresponding BAM
- Output: Report identifying block of mutation within codon (BMCs)
mit-o-matic
- Description: focuses on mtDNA, provides clinically relevant information from different resources; two component pipeline: command link for alignment of NGS reads and online version that provides genetic report on mitocondrial variants
- Input:FASTQ, pileup
- Reference sequence: rCRSm
- Output: Online version gives comprehensive genetic report
Mutadelic
- Description: Web App program; “This application generates reports on inherited mutations in five genes (ANK1, SLC4A1, SPTA1, SPTB and EPB42) associated with the following rare Mendelian blood disorders: Hereditary Spherocytosis (HS), Hereditary Elliptocytosis (HE) and Hereditary Pyropoikilocytosis”; Newer program- recently validated on omictools
- Input: Can upload coordinates of DNA variants or VEP
- Output: Displayed on web or can be downloaded in Excel or RDF format
MutationTaster
- Description: (Last post on site 2014) Web app program; Rapid evaluation of disease causing alterations; uses NCBI 37 and Ensembl 69
- Input: HGNC symbol, NCBI GeneID, or Ensembl ID,
- Output: Report containing prediction, summary, name of alteration, etc
MutPred
- Description: web app tool; Classifies amino acids substituation as disease associated or neutral in humans; Last modified Feb. 2014; Based on SIFT, trained using Human Gene Mutation Database
- Input:
- Output: “The output of MutPred contains a general score (g), i.e., the probability that the amino acid substitution is deleterious/disease-associated, and top 5 property scores (p), where p is the P-value that certain structural and functional properties are impacted.”
MutSigCV
- Description: (Broad Institute) Mutation Significance (CV= covariates); Analyzes mutations discovered in DNA sequencing to identify genes that were mutated more often than expected
- Input: mutations.maf, coverage.txt, covariates.txt
- Output: output.txt
NGS-SNP
- Description: Collection of command-line scripts for providing rich SNP annotations; “NCBI, Ensembl, and Uniprot IDs are provided for genes, transcripts and proteins when applicable”;
- Input: Samtools consensus pileup, Maq, diBayes, Genetic format, VCF
- Output: File containing annotated SNPs is copied from SNP list and some classes are added
Oncotator
- Description: (Broad Institute) “Tool for annotating human genomic point mutations and data relevant to cancer researchers”; Web app; Supports annotation of data from ClinVar, dbSNP, 1000 genomes (plus many other external sites); Only GRCh27 coordinates supported; Last update: April 2015
- Input: tal-delimited file
- Output: tab-delimited MAF
PANTHER
- Description: Protein ANalysis THrough Evolutionary Relationships; Web app program, also has its own database; Classification system used to classify proteins and their genes; Also, “Estimates the likelihood of a particular nonsynonymous (amino-acid changing) coding SNP to cause a functional impact on the protein”; Updated in 2015
- Input: Data from PANTHER, IDs from Ensembl, EntrezGene, NCBI GI numbers, NCBI UniGene IDs HUGO, UniProt; if ID type is not one of the above, can input txt file or excel format
- Output: Analysis results displayed online
PESX
- Description: Putative Exonic Splicing Enhancers/Silencers; (Can’t tell if this is outdated or not)
- Input: FASTA or plain text
- Output: Excel spread sheet
Phen-Gen
- Description: Combines patient's’ disease symptoms with sequencing data; Standalone or Web app version; Only excepts 1 family per run, in order to evaluate unrelated individuals, each sample needs to be run individually
- Input: Variant- VCF; Pheotype- HPO; Pedigree- PED
- Output: Combined scores file, variants for top genes file
PMUT
- Description: Aimed at annotation and prediction of pathological mutations; based on different kinds of sequence info and neural networks to process information
- Input: FASTA
- Output; Simple yes/no and reliability index
PROVEAN
- Description: Protein Variation Effect Analyzer; predicts whether an amino acid substitution or indel has impact on biological function of the protein; “comparable to SIFT or Polyphen-2”; Standalone, Web app, Command line or GUI; Last update May 2014
- Input: FASTA, list of variants;
- Output: tab-separated columns including Variant, Provean Score and prediciton
Rescue-ESE
- Description: “An online tool for identifying candidate ESEs in vertebrate exons”; Web application; For human, mouse, zebrafish, pufferfish
- Input: multi-FASTA or plain text
- Output:
SCAN
- Description: Web application program, includes a database as well; Database contains physical-based SNP annotations and functional annotations; “Information on physical, functional, and LD annotation served on the SCAN database comes directly from public resources, including the HapMap (release 23a), NCBI (dbSNP 129), or is information created by us using data downloaded from these public resources”; “SCAN can be utilized in several ways including: (i) queries of the SNP and gene databases; (ii) analysis using the attached tools and algorithms; (iii) downloading files with SNP annotation for various GWA platforms”
- Input:
- Output: HTML, comma-delimited, tab-delimited
SeattleSeq Annotation
- Description: “SeattleSeqAnnotation137 was most recently updated October 13, 2013. The current version is 8.08. The most recent site, based on dbSNP build 141, and hg38/NCBI 38”; Provides annotations for SNVs and Indels- includes dbSNP rsID, gene names and accession numbers, variation functions, protein positions and amino acid changes, conservation scores, HapMap frequencies, PolyPhen predictions and clinical association.
- Input: Maq, gff, CASAVA, VCF, GATK bed, custom
- Output: “default output file format is a header line (starting with "#") followed by tab-separated annotations”; VCF
seqminer 3.7
- Description: “Efficiently Read Sequence Data (VCF Format, BCF Format and METAL Format) into R”; Command line package program; Published August 2015
- Input: VCF, BCF
- Output: VCF
SG Adviser
- Description: Scripps Genome Annotation and Distributed Variant Interpretation Server, web developed applications for variant annotation, “Downstream applications of variant annotation include: Clinical sequencing applications including: carrier testing, or identification of causal variants in molecular diagnosis, tumor sequencing, or diagnostic odyssey. Prioritization of variants prior to statistical analysis of sequence based disease association studies, especially for automated set-generation and enrichment of likely functional variants within sets. Identification of causal variants in post-GWAS/linkage sequencing studies. Identification of causal variants in forward genetic screens (stay tuned for non-human annotation)”
- Input: SNV- VCF, BED, and a few others; CNV- BED, CNVator, plus others
- Output: tab-delimited file
SNAP-2
- Descriptio