BOL: Related items

R programming and Jobs website

Pragati Singh — Sun, 25 May 2014 14:43:57 -0500

Welcome to the R Jobs section of ProgrammingR.com. If your organization has an R employment opportunity that you would like to have posted here, submit it via the contact page. Prospective employees: use the contact information provided in the position listing to apply or contact the hiring organization.

Address of the bookmark: http://www.programmingr.com/category/stype/r-job-listings/

HNADOCK: a nucleic acid docking server for modeling RNA/DNA–RNA/DNA 3D complex structures

Poonam Mahapatra — Thu, 04 Jun 2020 23:19:07 -0500

The HNADOCK server is to predict the binding complex structure between two nucleic acid molecules through a hierarchical docking algorihtm of an FFT-based global search strategy and an intrinsic scoring function for nucleic acid interactions. Users are required to provide the three-dimensional (3D) structures of the two molecules to be docked.

Address of the bookmark: http://huanglab.phys.hust.edu.cn/hnadock/

Perl one-liner for bioinformatician !!!

Abhimanyu Singh — Fri, 30 May 2014 05:49:07 -0500

With the emergence of NGS technologies, and sequencing data most of the bioinformaticians mung and wrangle around massive amounts of genomics text. There are several "standardized" file formats (FASTQ, SAM, VCF, etc.) and some tools for manipulating them (fastx toolkit, samtools, vcftools, etc.), there are still times where knowing a little bit of Perl onliner is extremely helpful.

Perl one-liners are small and awesome Perl programs that fit in a single line of code and they do one thing really well. These things include changing line spacing, numbering lines, doing calculations, converting and substituting text, deleting and printing certain lines, parsing logs, editing files in-place, doing statistics, carrying out system administration tasks, updating a bunch of files at once, and many more. Perl one-liners will make you the shell warrior. Anything that took you minutes to solve, will now take you seconds!

perl -pe '$\="\n"'
#double space a file

perl -pe '$_ .= "\n" unless /^$/'
#double space a file except blank lines

perl -pe '$_.="\n"x7'
#7 space in a line.

perl -ne 'print unless /^$/'
#remove all blank lines

perl -lne 'print if length($_) < 20'
#print all lines with length less than 20.

perl -00 -pe ''
#If there are multiple spaces, delete all leaving one(make the file a single spaced file).

perl -00 -pe '$_.="\n"x4'
#Expand single blank lines into 4 consecutive blank lines

perl -pe '$_ = "$. $_"'
#Number all lines in a file

perl -pe '$_ = ++$a." $_" if /./'
#Number only non-empty lines in a file

perl -ne 'print ++$a." $_" if /./'
#Number and print only non-empty lines in a file

perl -pe '$_ = ++$a." $_" if /regex/'
#Number only lines that match a pattern

perl -ne 'print ++$a." $_" if /regex/'
#Number and print only lines that match a pattern

perl -ne 'printf "%-5d %s", $., $_ if /regex/'
#Left align lines with 5 white spaces if matches a pattern (perl -ne 'printf "%-5d %s", $., $_' : for all the lines)

perl -le 'print scalar(grep{/./}<>)'
#prints the total number of non-empty lines in a file

perl -lne '$a++ if /regex/; END {print $a+0}'
#print the total number of lines that matches the pattern

perl -alne 'print scalar @F'
#print the total number fields(words) in each line.

perl -alne '$t += @F; END { print $t}'
#Find total number of words in the file

perl -alne 'map { /regex/ && $t++ } @F; END { print $t }'
#find total number of fields that match the pattern

perl -lne '/regex/ && $t++; END { print $t }'
#Find total number of lines that match a pattern

perl -le '$n = 20; $m = 35; ($m,$n) = ($n,$m%$n) while $n; print $m'
#will calculate the GCD of two numbers.

perl -le '$a = $n = 20; $b = $m = 35; ($m,$n) = ($n,$m%$n) while $n; print $a*$b/$m'
#will calculate lcd of 20 and 35.

perl -le '$n=10; $min=5; $max=15; $, = " "; print map { int(rand($max-$min))+$min } 1..$n'
#Generates 10 random numbers between 5 and 15.

perl -le 'print map { ("a".."z",”0”..”9”)[rand 36] } 1..8'
#Generates a 8 character password from a to z and number 0 – 9.

perl -le 'print map { ("a",”t”,”g”,”c”)[rand 4] } 1..20'
#Generates a 20 nucleotide long random residue.

perl -le 'print "a"x50'
#generate a string of ‘x’ 50 character long

perl -le 'print join ", ", map { ord } split //, "hello world"'
#Will print the ascii value of the string hello world.

perl -le '@ascii = (99, 111, 100, 105, 110, 103); print pack("C*", @ascii)'
#converts ascii values into character strings.

perl -le '@odd = grep {$_ % 2 == 1} 1..100; print "@odd"'
#Generates an array of odd numbers.

perl -le '@even = grep {$_ % 2 == 0} 1..100; print "@even"'
#Generate an array of even numbers

perl -lpe 'y/A-Za-z/N-ZA-Mn-za-m/' file
#Convert the entire file into 13 characters offset(ROT13)

perl -nle 'print uc'
#Convert all text to uppercase:

perl -nle 'print lc'
#Convert text to lowercase:

perl -nle 'print ucfirst lc'
#Convert only first letter of first word to uppercas

perl -ple 'y/A-Za-z/a-zA-Z/'
#Convert upper case to lower case and vice versa

perl -ple 's/(\w+)/\u$1/g'
#Camel Casing

perl -pe 's|\n|\r\n|'
#Convert unix new lines into DOS new lines:

perl -pe 's|\r\n|\n|'
#Convert DOS newlines into unix new line

perl -pe 's|\n|\r|'
#Convert unix newlines into MAC newlines:

perl -pe '/regexp/ && s/foo/bar/'
#Substitute a foo with a bar in a line with a regexp.

Reference/Sources:

http://genomics-array.blogspot.in/2010/11/some-unixperl-oneliners-for.html

http://genomespot.blogspot.com/2013/08/a-selection-of-useful-bash-one-liners.html

http://biowize.wordpress.com/2012/06/15/command-line-magic-for-your-gene-annotations/

http://genomics-array.blogspot.com/2010/11/some-unixperl-oneliners-for.html

http://bioexpressblog.wordpress.com/2013/04/05/split-multi-fasta-sequence-file/

Common methods to discover tandem repeats

BioStar — Thu, 09 Mar 2023 02:40:52 -0600

Tandem repeats are DNA sequences that are repeated in a contiguous manner in the genome. These sequences are often used as genetic markers and are important in many areas of genetics and genomics research. Here are some methods for discovering tandem repeats in genomes:

Tandem Repeat Finder: Tandem Repeat Finder is a software tool that identifies tandem repeats in DNA sequences. It is available for free download and can be used on both nucleotide and protein sequences. The tool uses a statistical algorithm to identify repeats based on their length, copy number, and overall composition.
RepeatMasker: RepeatMasker is another software tool that can identify tandem repeats in DNA sequences. It works by comparing the input sequence to a database of known repeats and then identifies any tandem repeats that match those in the database.
PCR-based methods: Polymerase chain reaction (PCR) can be used to amplify and detect tandem repeats in genomic DNA. PCR primers are designed to flank the tandem repeat region, and amplification of the target DNA fragment can be visualized on a gel. This method can be useful for detecting novel tandem repeats and for genotyping.
Southern blotting: Southern blotting is a classic method for detecting DNA fragments in a sample. It can be used to detect tandem repeats by digesting genomic DNA with a restriction enzyme, separating the fragments by gel electrophoresis, and then probing the blot with a tandem repeat-specific probe.

Overall, a combination of these methods can be used to comprehensively identify tandem repeats in genomes.

Stephen Friend: The hunt for "unexpected genetic heroes"

Sat, 31 May 2014 14:31:47 -0500

What can we learn from people with the genetics to get sick — who don't? With most inherited diseases, only some family members will develop the disease, while others who carry the same genetic risks dodge it. Stephen Friend suggests we start studying those family members who stay healthy. Hear about the Resilience Project, a massive effort to collect genetic materials that may help decode inherited disorders. TEDTalks is a daily video podcast of the best talks and performances from the TED Conference, where the world's leading thinkers and doers give the talk of their lives in 18 minutes (or less). Look for talks on Technology, Entertainment and Design -- plus science, business, global issues, the arts and much more. Find closed captions and translated subtitles in many languages at http://www.ted.com/translate Follow TED news on Twitter: http://www.twitter.com/tednews Like TED on Facebook: https://www.facebook.com/TED Subscribe to our channel: http://www.youtube.com/user/TEDtalksDirector

List of motif discovery tools !

Neel — Tue, 20 Nov 2018 03:54:26 -0600

In genetics, a sequence motif is a nucleotide or amino-acid sequence pattern that is widespread and has, or is conjectured to have, a biological significance. For proteins, a sequence motif is distinguished from a structural motif, a motif formed by the three-dimensional arrangement of amino acids which may not be adjacent.

Following are the list of tools for motif discovery:

2Dsweep -- protein annotation by secondary structure elements

Perform secondary structure predictions on protein sequences.

3D-footprint -- database of DNA-binding protein structures

Find binding specificity information about DNA-protein complexes.

3D-footprint: DNA-binding protein database

Find information about the binding specificity of DNA-binding proteins.

3D-partner -- a web server to infer interacting partners and binding models

Predict interacting partners and binding models.

3MOTIF -- a protein structure visualization system for conserved sequence motifs

Use this web-based sequence motif visualization system to display sequence motif information in its appropriate three-dimensional (3D) context.

AFAWE -- Automatic functional annotation in a distributed Web Services Environment

Protein function prediction and annotation in an integrated environment powered by web service.

ANCHOR -- Prediction of Protein Binding Regions in Disordered Proteins

Find information about protein binding.

ANNIE -- ANNotation and Interpretation Environment for Protein Sequences

Use to predict function from de novo protein sequences.

Active Sequences Collection (ASC) database -- A new tool to assign functions to protein sequences

Search for short active protein sequences with demonstrated biological activities.

Blocks -- Ungapped segments in conserved protein sequences

Search for ungapped segments corresponding to the most highly conserved regions of proteins.

CASTp -- computed atlas of surface topography of proteins with structural and topographical mapping of functionally annotated residues

Identify and measure surface accessible pockets as well as interior inaccessible cavities, for proteins and other molecules.

CSA -- The Catalytic Site Atlas

To search for catalytic residue annotation for enzymes in the Protein Data Bank.

ConFunc -- Conserved residue Protein Function Prediction Server

Predict protein function using Gene Ontology.

ConSurf-DB -- evolutionary conservation profiles of protein structures database

Automatically calculate evolutionary conservation scores of key amino acid residues and map them on protein structures.

DBAli -- A Database of Structure Alignments

Mine the protein structure space.

DILIMOT -- discovery of linear motifs in proteins

Predict short linear motifs (3-8 residues) in a set of protein sequences.

Dasty2 -- an Ajax protein DAS client

A web client for visualizing protein sequence feature information using DAS.

DomainSweep -- protein annotation by domain analysis

Identify the domain architecture within a protein sequence.

E1DS -- catalytic site prediction based on 1D signatures of concurrent conservation

Predict enzyme catalytic site.

ELM -- Eukarotic Linear Motif Resource

Predict functional sites in eukaryotic proteins.

EXPASY Proteome Tools Collection

Use a collection of tools for protein analyses.

EXPASY-Findmod

Predict potential protein post-translational modifications and find potential single amino acid substitutions in peptides.

EzCatDB -- the Enzyme Catalytic-mechanism Database

Search for information related to the catalytic mechanisms of enzymes.

FFPred -- feature-based function prediction

An integrated feature-based function prediction server for vertebrate proteomes.

FingerPRINT Scan

Identify the closest matching PRINTS sequence motif fingerprints in a protein sequence.

FireDB -- a database of functionally important residues from proteins of known structure

Search for functional annotation of important sites in proteins with known structures.

Frog2 -- a FRee Online druG 3D conformation generator

Produce 3D conformations of small drug compounds.

HGPD -- Human Gene and Protein Database

A database presenting experiment-based results in human proteomics.

HHsenser -- exhaustive transitive profile search using HMMx96HMM comparison

Conduct exhaustive intermediate profile searches of a set of homologous protein sequences.

HotSpot Wizard -- Substrate Specificity Hot Spot Identification web server

Design protein mutations in site-directed mutagenesis.

INTREPID -- INformation-theoretic TREe traversal for Protein functional site IDentification

Use for protein functional site identification.

Integrating protein annotation resources through the Distributed Annotation System

Annotate protein using this integrated annotation resource.

InterProScan -- protein domains identifier

Identify protein family (and DNA) domains, patterns, motifs, protein families, and functional sites.

KFC -- Knowledge-based FADE and Contacts

Interactive forecasting of protein interaction hot spots.

MAGIIC-PRO -- detecting functional signatures by efficient discovery of long patterns in protein sequences

Discover long patterns in protein sequences.

MALISAM -- Manual ALIgnments for Structurally Analogous Motifs

Database containing pairs of structural analogs and their alignments.

MEME -- discovering and analyzing DNA and protein sequence motifs

Find sequence patterns in DNA and protein sequences.

MODPROPEP -- a program for knowledge-based modeling of protein-peptide complexes

A web server for knowledge-based modeling of protein-peptide complexes, specifically peptides in complex with major histocompatibility complex (MHC) proteins and kinases.

MeMo -- a web tool for prediction of protein methylation modifications

Predict protein methylation sites.

MegaMotifBase -- a database of structural motifs in protein families and superfamilies

Find structural segments or motifs for protein structures.

Minimotif Miner -- a tool for investigating protein function

Find motifs in a protein sequence.

Motif3D -- Relating protein sequence motifs to 3D structure

Visualize protein sequence motifs on the 3D protein structures.

MotifScan

Find presence of any known protein motif (Prosite and Pfam) in a protein sequence.

MultiBind -- Multiple Alignment of Protein Binding Sites

Recognize spatial chemical binding patterns common to a set of protein structures.

NMT -- The MYR Predictor

Analyze proteins for the presence of N-terminal N-myristoylation site.

NetNGlyc -- N-Glycosylation sites prediction tool

Find the presence of N-Glycosylation sites in human proteins.

NetOGly 3.1 -- O-glycosylation sites prediction tool

Find the presence of O-GalNAc (mucin type) glycosylation sites in mammalian proteins.

NetPhos 2.0 -- Phosphorylation sites predictions

Analyze eukaryotic proteins for the presence of serine, threonine and tyrosine phosphorylation sites.

NetPhosK 1.0 Server -- kinase specific eukaryotic protein phosphorylation sites prediction tool

Find possible kinase specific phosphorylation sites in eukaryotic proteins.

NetworKIN -- a resource for exploring cellular phosphorylation networks

NeuroPred -- a tool to predict cleavage sites in neuropeptide precursors and provide the masses of the resulting peptides

Predict cleavage sites at basic amino acid locations in neuropeptide precursor sequences.

Non-Redundant Patent Sequences - Patented Sequence Database

Find information about patented nucleotide and protein sequences.

O-GLYCBASE

Search for information about glycoproteins with O-linked and C-linked glycosylation sites.

PANDORA -- Protein ANnotation Diagram ORiented Analysis

Find information about protein sequence annotations.

PAR-3D -- Protein Active site Residue - 3D structural motif

A server to predict protein active site residues.

PDBSite -- a database of the 3D structure of protein functional sites

Search for structural and functional information on the protein functional sites.

PDBSiteScan -- A program for searching for active, binding and posttranslational modification sites in the 3D structures of proteins

Search 3D protein fragments similar in structure to known active, binding and posttranslational modification sites.

PEDANT -- Protein Extraction, Description and ANalysis Tool

Conduct genome wide functional and structural analysis.

PHOSIDA -- Phosphorylation site database

Search for phosphorylation data of any protein of interest.

PHOSPHORYLATION SITE DATABASE

Search for information on prokaryotic proteins that undergo serine, threonine, or tyrosine phosphorylation.

PNU -- Protein Naming Utility

Determine correct names for proteins.

POODLE-S -- Predicition Of Order and Disorder by machine LEarning

Web application for predicting protein disorder by using physicochemical features and reduced amino acid set of a position-specific scoring matrix.

PPISearch -- Protein-Protein Interaction Search

Find homologous protein-protein interactions across multiple species.

PPSearch

Search your query sequence against PROSITE pattern database for protein motifs.

PRIDB -- Protein-RNA Interface DataBase

Find information about protein-RNA complexes from the Protein Data Bank (PDB).

PRINTS and its automatic supplement, prePRINTS -- A compendium of protein fingerprints

Search for protein fingerprints.

PROSITE

Identify protein families and domains for a given protein sequence.

PRRDB -- Pattern Recognition Receptor Database

A comprehensive database of pattern-recognition receptors and their ligands.

PatMatch -- a program for finding patterns in peptide and nucleotide sequences

Search for short nucleotide or peptide sequences such as cis-elements in nucleotide sequences or small domains and motifs in protein sequences.

PepCyber:P~PEP -- a database of human protein protein interactions mediated by phosphoprotein-binding domains

Database specialized in documenting human PPBD-containing proteins and PPBD-mediated interactions.

PeptideCutter -- protein cleavage sites prediction tool

Predicts potential protease cleavage sites and sites cleaved by chemicals in a given protein sequence.

Phobius -- A combined transmembrane topology and signal peptide predictor

Predict combined transmembrane topology and signal peptides.

Phospho.ELM -- a database of phosphorylation sites

Search for eukaryotic phosphorylation sites.

Phospho3D -- a database of three-dimensional structures of protein phosphorylation sites

Search for 3D structure and functional annotation of phosphorylation sites in proteins.

PhosphoSite -- A bioinformatics resource dedicated to physiological protein phosphorylation.

Search the database of in vivo phosphorylation sites of human and mouse proteins

PolyQ -- Polyglutamine Database

Find information about polyglutamine (polyQ) repeats.

Pratt Protein motif and pattern discovery

Find the presence of protein motifs and patterns in an amino acid sequence.

PrediSi -- Prediction of Signal Peptides and their Cleavage Positions

Predict signal peptide sequences and their cleavage positions in bacterial and eukaryotic amino acid sequences.

ProFunc -- a server for predicting protein function from 3D structure

Predict protein functions based on known structures.

ProMateus--an open research approach to protein-binding sites analysis

Predict the location of potential protein-protein binding sites for unbound proteins.

ProTeus -- identifying signatures in protein termini

Identify short linear signatures in protein termini.

ProtSweep -- protein annotation by homology

Analyze and identify newly obtained protein sequences.

Protemot -- prediction of protein binding sites with automatically extracted geometrical templates

Predict protein binding sites in a protein sequence based on geometrical analysis of protein tertiary substructures.

QuasiMotiFinder -- protein annotation by searching for evolutionarily conserved motif-like patterns

Search for evolutionarily conserved motif-like patterns in protein sequences.

RNABindR -- software for prediction of RNA binding residues in proteins

Web-based server for analyzing and predicting RNA binding sites in proteins.

SCANMOT -- searching for similar sequences using a simultaneous scan of multiple sequence motifs

Search for similarities between proteins by simultaneous matching of multiple motifs.

SDPpred -- A Tool for Prediction of Amino Acid Residues that Determine Differences in Functional Specificity of Homologous Proteins

Predict residues in protein sequences that determine the proteins' functional specificity.

SDR -- Specificity Determining Residues Database

Predict specificity-determining residues in protein families.

SLiMDisc -- Short, Linear Motif Discovery

Find shared motifs in proteins with a common attribute.

SUMOsp -- a web server for sumoylation site prediction

Conduct in silico sumoylation sites prediction.

SWAKK -- a web server for detecting positive selection in proteins using a sliding window substitution rate analysis

Detect protein sequence section under positive evolution selection.

ScanProsite

Search for motifs and patterns within protein sequences.

ScanProsite -- detection of PROSITE signature matches and ProRule-associated functional and structural residues in proteins

Detect patterns, profiles and motifs in a protein sequence.

ScanSite 2.0 -- Proteome-wide prediction of cell signaling interactions using short sequence motifs

Search for motifs within proteins that are likely to be phosphorylated by specific protein kinases or bind to domains such as SH2 domains, 14-3-3 domains or PDZ domains.

SePreSA -- SErver for the PREdiction of populations susceptible to Serious Adverse drug reaction

Find information about populations carrying polymorphisms within protein binding pockets that make them susceptible to serious adverse drug reaction (SADR).

Sequence Motif Search

Search the presence of a motif in either amino acid sequence or nucleotide sequence.

Signal-3L -- A 3-layer approach for predicting signal peptides

Predict signal peptides.

SignalP -- Machine learning approaches to the prediction of signal peptides, their cleavage sites, and other protein sorting signals

Predict signal peptides and their cleavage sites.

Sulfinator -- tyrosine sulfation sites prediction tool

Predict the presence of tyrosine sulfation sites in protein sequences

SuperSite -- Ligand Binding Site Database

Look at protein structure from a ligand and binding site perspective.

Swiss EMBnet node web server

Use a collection of bioinformatics tools at this portal site.

T-REKS -- identification of Tandem REpeats in sequences with a K-meanS based algorithm

Find information about tandem repeats in proteins that carry fundamental biological functions and are related to a number of human diseases.

TMFunction -- The Functional Database of Membrane Proteins

Find information about functional residues in alpha-helical and beta-barrel membrane proteins.

TOPDOM -- Conservatively Located Domains and Motifs in Transmembrane Proteins

Database of domains and motifs with conservative location in transmembrane proteins.

The EMOTIF database

Search for highly conserved and specific protein sequence motifs.

TreeDet -- Predicting Functional Residues in Protein Sequence Alignments

Predict functional sites in protein sequence alignments use different methodologies.

W-ChIPMotifs -- ChIP-based protein Motif discovery web server

Find de novo protein motifs from chromatin immunoprecipitation data.

WebFEATURE -- an interactive web tool for identifying and visualizing functional sites on macromolecular structures

Scan query structures for functional sites in both proteins and nucleic acids.

WebProAnalyst -- an interactive tool for analysis of quantitative structurex96activity relationships in protein families

Analyze quantitative structure-activity relationship of related protein families.

eBLOCKs -- enumerating conserved protein blocks to achieve maximal sensitivity and specificity

Search for ungapped alignments of highly conserved regions among a protein family or superfamily.

eF-seek -- prediction of the functional sites of proteins by searching for similar electrostatic potential and molecular surface shape

Predict the functional sites of proteins.

firestar -- prediction of functionally important residues using structural templates and alignment reliability

An expert system for predicting ligand-binding residues in protein structures.

iMOTdb -- a comprehensive collection of spatially interacting motifs in proteins

Automatically identify spatially interacting motifs among distantly related proteins sharing similar folds and possessing common ancestral lineage.

INSPIRE Faculty Scheme: a component of “Assured Opportunity for Research Career (AORC)” under INSPIRE.

Sat, 19 Jul 2014 14:59:30 -0500

Ministry of Science and Technology, Department of Science and Technology

7th ADVERTISEMENT – 2014 (2)

INSPIRE Faculty Scheme: a component of “Assured Opportunity for Research Career (AORC)” under INSPIRE.

The Department of Science and Technology, Government of India, has launched the “Innovation in Science Pursuit for Inspired Research (INSPIRE)” [http://www.inspire-dst.gov.in] program in 2008.

The program aims to attract talent for study of science and careers with research. INSPIRE includes many components. The importance of Assured Career Opportunity in R&D sector has been recognized.

INSPIRE Faculty Scheme opens up an “Assured Opportunity for Research Career (AORC)” for young researchers in the age group of 27-32 years. It offers a contractual research awards to young achievers and opportunity for independent research in the near term and emerge as a future leader in the long term.

Eligibility

Essential Indian citizens and people of Indian origin including NRI/PIO status with PhD (in science, mathematics, engineering, pharmacy, medicine, and agriculture related subjects) from any recognized university in the world,

Those who have submitted their PhD Theses and are awaiting award of the degree are also
eligible. However, the award will be conveyed only after confirmation of the awarding the
PhD degree.

The upper age limit as on 1st July 2014 should be 32 years for considering support for a
period of 5 years. However, for SC and ST candidates, upper age limit will be 35 years.

Publication(s) in highly reputed Journals demonstrating research potential of the candidate.

Desirable

Candidates who are within top 1% at the School Leaving Examination, IIT-JEE rank, 1st Rank Holder either in graduation or post-graduation level university examination (which are used presently for identifying INSPIRE Scholars at under-graduate level and INSPIRE Fellows for doctoral degree)

More at http://www.inspire-dst.gov.in/faculty_scheme.html

Flye: Fast and accurate de novo assembler for single molecule sequencing reads

BioJoker — Tue, 02 Apr 2019 21:54:55 -0500

Flye is a de novo assembler for single molecule sequencing reads, such as those produced by PacBio and Oxford Nanopore Technologies. It is designed for a wide range of datasets, from small bacterial projects to large mammalian-scale assemblies. The package represents a complete pipeline: it takes raw PB / ONT reads as input and outputs polished contigs. Flye also includes a special mode for metagenome assembly.

Address of the bookmark: https://github.com/fenderglass/Flye

Commercial and public next-gen-seq (NGS) software

Surabhi Chaudhary — Tue, 03 Jun 2014 20:45:11 -0500

Integrated solutions
CLCbio Genomics Workbench - de novo and reference assembly of Sanger, Roche FLX, Illumina, Helicos, and SOLiD data. Commercial next-gen-seq software that extends the CLCbio Main Workbench software. Includes SNP detection, CHiP-seq, browser and other features. Commercial. Windows, Mac OS X and Linux.
Galaxy - Galaxy = interactive and reproducible genomics. A job webportal.
Genomatix - Integrated Solutions for Next Generation Sequencing data analysis.
JMP Genomics - Next gen visualization and statistics tool from SAS. They are working with NCGR to refine this tool and produce others.
NextGENe - de novo and reference assembly of Illumina, SOLiD and Roche FLX data. Uses a novel Condensation Assembly Tool approach where reads are joined via "anchors" into mini-contigs before assembly. Includes SNP detection, CHiP-seq, browser and other features. Commercial. Win or MacOS.
Partek - Commercial software for NGS, microarray, and qPCR data analysis. Streamlined analysis workflows for: ChIP-Seq, RNA-Seq, DNA-Seq, DNA Methylation, Gene Expression, Exon, miRNA Expression, Copy Number, Allele-Specific Copy Number, LOH, Association, Trio Analysis, and Tiling. Supports all commercial sequencing and microarray technologies.
SeqMan Genome Analyser - Software for Next Generation sequence assembly of Illumina, Roche FLX and Sanger data integrating with Lasergene Sequence Analysis software for additional analysis and visualization capabilities. Can use a hybrid templated/de novo approach. Commercial. Win or Mac OS X.
SHORE - SHORE, for Short Read, is a mapping and analysis pipeline for short DNA sequences produced on a Illumina Genome Analyzer. A suite created by the 1001 Genomes project. Source for POSIX.
SlimSearch - Fledgling commercial product.
Synamatix has SXOligoSearch (http://synasite.mgrc.com.my:8080/sxo...ligoSearch.php)
The SWIFT suit is a software collection for fast index-based sequence comparison. It contains the following programs: SWIFT — fast local alignment search, guaranteeing to find epsilon-matches between two sequences; SWIFT BALSAM — a very fast program to find semiglobal non-gapped alignments based on k-mer seeds. http://bibiserv.techfak.uni-bielefeld.de/swift/
biolib.is library and a set of script targeted to NGS. There are modules to: clean sequences (sanger, 454, ilumina), parse caf, ace and bowtie map files, clean and filter contigs, look for snps and indels., filter snps, do statistics for: reads, contigs and snps.

Align/Assemble to a reference
BFAST - Blat-like Fast Accurate Search Tool. Written by Nils Homer, Stanley F. Nelson and Barry Merriman at UCLA.
Bowtie - Ultrafast, memory-efficient short read aligner. It aligns short DNA sequences (reads) to the human genome at a rate of 25 million reads per hour on a typical workstation with 2 gigabytes of memory. Uses a Burrows-Wheeler-Transformed (BWT) index. Link to discussion thread here. Written by Ben Langmead and Cole Trapnell. Linux, Windows, and Mac OS X.
BWA - Heng Lee's BWT Alignment program - a progression from Maq. BWA is a fast light-weighted tool that aligns short sequences to a sequence database, such as the human reference genome. By default, BWA finds an alignment within edit distance 2 to the query sequence. C++ source.
ELAND - Efficient Large-Scale Alignment of Nucleotide Databases. Whole genome alignments to a reference genome. Written by Illumina author Anthony J. Cox for the Solexa 1G machine.
Exonerate - Various forms of pairwise alignment (including Smith-Waterman-Gotoh) of DNA/protein against a reference. Authors are Guy St C Slater and Ewan Birney from EMBL. C for POSIX.
GenomeMapper - GenomeMapper is a short read mapping tool designed for accurate read alignments. It quickly aligns millions of reads either with ungapped or gapped alignments. A tool created by the 1001 Genomes project. Source for POSIX.
GMAP - GMAP (Genomic Mapping and Alignment Program) for mRNA and EST Sequences. Developed by Thomas Wu and Colin Watanabe at Genentec. C/Perl for Unix.
gnumap - The Genomic Next-generation Universal MAPper (gnumap) is a program designed to accurately map sequence data obtained from next-generation sequencing machines (specifically that of Solexa/Illumina) back to a genome of any size. It seeks to align reads from nonunique repeats using statistics. From authors at Brigham Young University. C source/Unix.
MAQ - Mapping and Assembly with Qualities (renamed from MAPASS2). Particularly designed for Illumina with preliminary functions to handle ABI SOLiD data. Written by Heng Li from the Sanger Centre. Features extensive supporting tools for DIP/SNP detection, etc. C++ source
MOSAIK - MOSAIK produces gapped alignments using the Smith-Waterman algorithm. Features a number of support tools. Support for Roche FLX, Illumina, SOLiD, and Helicos. Written by Michael Strömberg at Boston College. Win/Linux/MacOSX
MrFAST and MrsFAST - mrFAST & mrsFAST are designed to map short reads generated with the Illumina platform to reference genome assemblies; in a fast and memory-efficient manner. Robust to INDELs and MrsFAST has a bisulphite mode. Authors are from the University of Washington. C as source.
MUMmer - MUMmer is a modular system for the rapid whole genome alignment of finished or draft sequence. Released as a package providing an efficient suffix tree library, seed-and-extend alignment, SNP detection, repeat detection, and visualization tools. Version 3.0 was developed by Stefan Kurtz, Adam Phillippy, Arthur L Delcher, Michael Smoot, Martin Shumway, Corina Antonescu and Steven L Salzberg - most of whom are at The Institute for Genomic Research in Maryland, USA. POSIX OS required.
Novocraft - Tools for reference alignment of paired-end and single-end Illumina reads. Uses a Needleman-Wunsch algorithm. Can support Bis-Seq. Commercial. Available free for evaluation, educational use and for use on open not-for-profit projects. Requires Linux or Mac OS X.
PASS - It supports Illumina, SOLiD and Roche-FLX data formats and allows the user to modulate very finely the sensitivity of the alignments. Spaced seed intial filter, then NW dynamic algorithm to a SW(like) local alignment. Authors are from CRIBI in Italy. Win/Linux.
RMAP - Assembles 20 - 64 bp Illumina reads to a FASTA reference genome. By Andrew D. Smith and Zhenyu Xuan at CSHL. (published in BMC Bioinformatics). POSIX OS required.
SeqMap - Supports up to 5 or more bp mismatches/INDELs. Highly tunable. Written by Hui Jiang from the Wong lab at Stanford. Builds available for most OS's.
SHRiMP - Assembles to a reference sequence. Developed with Applied Biosystem's colourspace genomic representation in mind. Authors are Michael Brudno and Stephen Rumble at the University of Toronto. POSIX.
Slider- An application for the Illumina Sequence Analyzer output that uses the probability files instead of the sequence files as an input for alignment to a reference sequence or a set of reference sequences. Authors are from BCGSC. Paper is here.
SOAP - SOAP (Short Oligonucleotide Alignment Program). A program for efficient gapped and ungapped alignment of short oligonucleotides onto reference sequences. The updated version uses a BWT. Can call SNPs and INDELs. Author is Ruiqiang Li at the Beijing Genomics Institute. C++, POSIX.
SSAHA - SSAHA (Sequence Search and Alignment by Hashing Algorithm) is a tool for rapidly finding near exact matches in DNA or protein databases using a hash table. Developed at the Sanger Centre by Zemin Ning, Anthony Cox and James Mullikin. C++ for Linux/Alpha.
SOCS - Aligns SOLiD data. SOCS is built on an iterative variation of the Rabin-Karp string search algorithm, which uses hashing to reduce the set of possible matches, drastically increasing search speed. Authors are Ondov B, Varadarajan A, Passalacqua KD and Bergman NH.
SWIFT - The SWIFT suit is a software collection for fast index-based sequence comparison. It contains: SWIFT — fast local alignment search, guaranteeing to find epsilon-matches between two sequences. SWIFT BALSAM — a very fast program to find semiglobal non-gapped alignments based on k-mer seeds. Authors are Kim Rasmussen (SWIFT) and Wolfgang Gerlach (SWIFT BALSAM)
SXOligoSearch - SXOligoSearch is a commercial platform offered by the Malaysian based Synamatix. Will align Illumina reads against a range of Refseq RNA or NCBI genome builds for a number of organisms. Web Portal. OS independent.
Vmatch - A versatile software tool for efficiently solving large scale sequence matching tasks. Vmatch subsumes the software tool REPuter, but is much more general, with a very flexible user interface, and improved space and time requirements. Essentially a large string matching toolbox. POSIX.
Zoom - ZOOM (Zillions Of Oligos Mapped) is designed to map millions of short reads, emerged by next-generation sequencing technology, back to the reference genomes, and carry out post-analysis. ZOOM is developed to be highly accurate, flexible, and user-friendly with speed being a critical priority. Commercial. Supports Illumina and SOLiD data.
NCGR uses GMAP (http://www.gene.com/share/gmap/) to alignment Solexa reads. GMAP is free, though.
Exonerate (http://www.ebi.ac.uk/~guy/exonerate/)
MUMmer (http://mummer.sourceforge.net/)
The mapping short reads called gnumap (http://dna.cs.byu.edu/gnumap/) made to increase the accuracy with duplicate matches. Open source, creates viewable output (with Affy's Integrated Genome Browser), and produces results very similar to novocraft's.
SOCS (short oligonucleotides in color space)
BFAST https://secure.genome.ucla.edu/index.php/BFAST

De novo Align/Assemble
ABySS - Assembly By Short Sequences. ABySS is a de novo sequence assembler that is designed for very short reads. The single-processor version is useful for assembling genomes up to 40-50 Mbases in size. The parallel version is implemented using MPI and is capable of assembling larger genomes. By Simpson JT and others at the Canada's Michael Smith Genome Sciences Centre. C++ as source.
ALLPATHS - ALLPATHS: De novo assembly of whole-genome shotgun microreads. ALLPATHS is a whole genome shotgun assembler that can generate high quality assemblies from short reads. Assemblies are presented in a graph form that retains ambiguities, such as those arising from polymorphism, thereby providing information that has been absent from previous genome assemblies. Broad Institute.
Edena - Edena (Exact DE Novo Assembler) is an assembler dedicated to process the millions of very short reads produced by the Illumina Genome Analyzer. Edena is based on the traditional overlap layout paradigm. By D. Hernandez, P. François, L. Farinelli, M. Osteras, and J. Schrenzel. Linux/Win.
EULER-SR - Short read de novo assembly. By Mark J. Chaisson and Pavel A. Pevzner from UCSD (published in Genome Research). Uses a de Bruijn graph approach.
MIRA2 - MIRA (Mimicking Intelligent Read Assembly) is able to perform true hybrid de-novo assemblies using reads gathered through 454 sequencing technology (GS20 or GS FLX). Compatible with 454, Solexa and Sanger data. Linux OS required.
SEQAN - A Consistency-based Consensus Algorithm for De Novo and Reference-guided Sequence Assembly of Short Reads. By Tobias Rausch and others. C++, Linux/Win.
SHARCGS - De novo assembly of short reads. Authors are Dohm JC, Lottaz C, Borodina T and Himmelbauer H. from the Max-Planck-Institute for Molecular Genetics.
SSAKE - The Short Sequence Assembly by K-mer search and 3' read Extension (SSAKE) is a genomics application for aggressively assembling millions of short nucleotide sequences by progressively searching for perfect 3'-most k-mers using a DNA prefix tree. Authors are René Warren, Granger Sutton, Steven Jones and Robert Holt from the Canada's Michael Smith Genome Sciences Centre. Perl/Linux.
SOAPdenovo - Part of the SOAP suite. See above.
VCAKE - De novo assembly of short reads with robust error correction. An improvement on early versions of SSAKE.
Velvet - Velvet is a de novo genomic assembler specially designed for short read sequencing technologies, such as Solexa or 454. Need about 20-25X coverage and paired reads. Developed by Daniel Zerbino and Ewan Birney at the European Bioinformatics Institute (EMBL-EBI).
SOAP (http://soap.genomics.org.cn) by Ruiqiang Li, as has been pointed by ECO.
Euler-SR (Euler-Short Reads Assembly, http://euler-assembler.ucsd.edu/portal/) by Mark J. Chaisson and Pavel A. Pevzner from UCSD. (published in Genome Research)
RMAP (A program for mapping Solexa reads, http://rulai.cshl.edu/rmap/) by Andrew D. Smith and Zhenyu Xuan at CSHL. (published in BMC Bioinformatics)
Short read aligner called Bowtie (http://bowtie-bio.sourceforge.net/) designed for fast mapping of Illumina reads

SNP/Indel Discovery
ssahaSNP - ssahaSNP is a polymorphism detection tool. It detects homozygous SNPs and indels by aligning shotgun reads to the finished genome sequence. Highly repetitive elements are filtered out by ignoring those kmer words with high occurrence numbers. More tuned for ABI Sanger reads. Developers are Adam Spargo and Zemin Ning from the Sanger Centre. Compaq Alpha, Linux-64, Linux-32, Solaris and Mac
PolyBayesShort - A re-incarnation of the PolyBayes SNP discovery tool developed by Gabor Marth at Washington University. This version is specifically optimized for the analysis of large numbers (millions) of high-throughput next-generation sequencer reads, aligned to whole chromosomes of model organism or mammalian genomes. Developers at Boston College. Linux-64 and Linux-32.
PyroBayes - PyroBayes is a novel base caller for pyrosequences from the 454 Life Sciences sequencing machines. It was designed to assign more accurate base quality estimates to the 454 pyrosequences. Developers at Boston College.
Maq is also able to find SNPs with its own alignment. It has a graphical viewer, but again for its own alignment format.
SSAHA has been optimized for short-reads, too. But yes, SSAHASNP appears in your "SNP/INDEL discovery" category.

Genome Annotation/Genome Browser/Alignment Viewer/Assembly Database
EagleView - An information-rich genome assembler viewer. EagleView can display a dozen different types of information including base quality and flowgram signal. Developers at Boston College.
LookSeq - LookSeq is a web-based application for alignment visualization, browsing and analysis of genome sequence data. LookSeq supports multiple sequencing technologies, alignment sources, and viewing modes; low or high-depth read pileups; and easy visualization of putative single nucleotide and structural variation. From the Sanger Centre.
MapView - MapView: visualization of short reads alignment on desktop computer. From the Evolutionary Genomics Lab at Sun-Yat Sen University, China. Linux.
SAM - Sequence Assembly Manager. Whole Genome Assembly (WGA) Management and Visualization Tool. It provides a generic platform for manipulating, analyzing and viewing WGA data, regardless of input type. Developers are Rene Warren, Yaron Butterfield, Asim Siddiqui and Steven Jones at Canada's Michael Smith Genome Sciences Centre. MySQL backend and Perl-CGI web-based frontend/Linux.
STADEN - Includes GAP4. GAP5 once completed will handle next-gen sequencing data. A partially implemented test version is available here
XMatchView - A visual tool for analyzing cross_match alignments. Developed by Rene Warren and Steven Jones at Canada's Michael Smith Genome Sciences Centre. Python/Win or Linux.

Counting e.g. CHiP-Seq, Bis-Seq, CNV-Seq
BS-Seq - The source code and data for the "Shotgun Bisulphite Sequencing of the Arabidopsis Genome Reveals DNA Methylation Patterning" Nature paper by Cokus et al. (Steve Jacobsen's lab at UCLA). POSIX.
CHiPSeq - Program used by Johnson et al. (2007) in their Science publication
CNV-Seq - CNV-seq, a new method to detect copy number variation using high-throughput sequencing. Chao Xie and Martti T Tammi at the National University of Singapore. Perl/R.
FindPeaks - perform analysis of ChIP-Seq experiments. It uses a naive algorithm for identifying regions of high coverage, which represent Chromatin Immunoprecipitation enrichment of sequence fragments, indicating the location of a bound protein of interest. Original algorithm by Matthew Bainbridge, in collaboration with Gordon Robertson. Current code and implementation by Anthony Fejes. Authors are from the Canada's Michael Smith Genome Sciences Centre. JAVA/OS independent. Latest versions available as part of the Vancouver Short Read Analysis Package
MACS - Model-based Analysis for ChIP-Seq. MACS empirically models the length of the sequenced ChIP fragments, which tends to be shorter than sonication or library construction size estimates, and uses it to improve the spatial resolution of predicted binding sites. MACS also uses a dynamic Poisson distribution to effectively capture local biases in the genome sequence, allowing for more sensitive and robust prediction. Written by Yong Zhang and Tao Liu from Xiaole Shirley Liu's Lab.
PeakSeq - PeakSeq: Systematic Scoring of ChIP-Seq Experiments Relative to Controls. a two-pass approach for scoring ChIP-Seq data relative to controls. The first pass identifies putative binding sites and compensates for variation in the mappability of sequences across the genome. The second pass filters out sites that are not significantly enriched compared to the normalized input DNA and computes a precise enrichment and significance. By Rozowsky J et al. C/Perl.
QuEST - Quantitative Enrichment of Sequence Tags. Sidow and Myers Labs at Stanford. From the 2008 publication Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data. (C++)
SISSRs - Site Identification from Short Sequence Reads. BED file input. Raja Jothi @ NIH. Perl.
SeqMap (http://biogibbs.stanford.edu/~jiangh/SeqMap/) - work like ELand, can do 3 or more bp mismatches and also insdel
ChIPSeq analysis is: http://dir.nhlbi.nih.gov/papers/lmi/epigenomes/sissrs/

See also this thread for ChIP-Seq, until I get time to update this list.

Alternate Base Calling
Rolexa - R-based framework for base calling of Solexa data. Project publication
Alta-cyclic - "a novel Illumina Genome-Analyzer (Solexa) base caller"

Transcriptomics
ERANGE - Mapping and Quantifying Mammalian Transcriptomes by RNA-Seq. Supports Bowtie, BLAT and ELAND. From the Wold lab.
G-Mo.R-Se - G-Mo.R-Se is a method aimed at using RNA-Seq short reads to build de novo gene models. First, candidate exons are built directly from the positions of the reads mapped on the genome (without any ab initio assembly of the reads), and all the possible splice junctions between those exons are tested against unmapped reads. From CNS in France.
MapNext - MapNext: A software tool for spliced and unspliced alignments and SNP detection of short sequence reads. From the Evolutionary Genomics Lab at Sun-Yat Sen University, China.
QPalma - Optimal Spliced Alignments of Short Sequence Reads. Authors are Fabio De Bona, Stephan Ossowski, Korbinian Schneeberger, and Gunnar Rätsch. A paper is available.
RSAT - RSAT: RNA-Seq Analysis Tools. RNASAT is developed and maintained by Hui Jiang at Stanford University.
TopHat - TopHat is a fast splice junction mapper for RNA-Seq reads. It aligns RNA-Seq reads to mammalian-sized genomes using the ultra high-throughput short read aligner Bowtie, and then analyzes the mapping results to identify splice junctions between exons. TopHat is a collaborative effort between the University of Maryland and the University of California, Berkeley
NGS-Trex: Next Generation Sequencing Transcriptome profile explorer http://www.biomedcentral.com/1471-2105/14/S7/S10

Reference

Illumina has a software list: http://www.illumina.com/pagesnrn.ilmn?ID=245.

Some softwares in his blog (http://www.fejes.ca/labels/DNA.html)

http://seqanswers.com/wiki/Software

NCBI Webinar

Jit — Sun, 08 Jun 2014 02:47:01 -0500

In less than two weeks, NCBI will offer a webinar entitled "Introducing 3 NCBI Resources to Navigate Testing for Disease Linked Variants: MedGen, GTR and ClinVar". This webinar will delve into the lifecycle of genetic testing and teach attendees how to navigate the NIH Genetic Testing Registry, ClinVar, and MedGen resources. These resources can be used to prepare for clinical cases, access detailed information about orderable genetic tests, interpret test results, and more.

More at https://attendee.gotowebinar.com/register/8452228815737989634