BOL: Related items

A fast package to parse BLAST

Jitendra Narayan — Tue, 10 Sep 2013 16:58:56 -0500

In current era, we are handling huge amount of genomics data, and analysing it to make some biological sense out of it. Large-scale sequence studies requiring BLAST-based analysis produce huge amounts of data to be parsed. There are several BLAST parsers are available, but they are often missing some important features, such as keeping all information from the raw BLAST output, allowing direct access to single results, and performing logical operations over them.

Massimiliano Orsini and Simone Carcangiu develope a new and fast fast package "BlaSTorage" to parse and store BLAST results. BlaSTorage shows comparable speed of more basic parser written in compiled languages as C++ and can be easily integrated into web applications or software pipelines.

Find more @ http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3571973/

http://biowiki.crs4.it/biowiki/MassimilianoOrsini

BLAST Ring Image Generator (BRIG)

Anjana — Fri, 30 Sep 2016 09:18:50 -0500

BRIG is a free cross-platform (Windows/Mac/Unix) application that can display circular comparisons between a large number of genomes, with a focus on handling genome assembly data. The application is available at: http://sourceforge.net/projects/brig

If you have any questions or comments, post them on one of the trackers on BRIG’s SourceForge page: http://sourceforge.net/tracker/?group_id=328245.

Features:

Images show similarity between a central reference sequence and other sequences as concentric rings.
BRIG will perform all BLAST comparisons and file parsing automatically via a simple GUI.
Contig boundaries and read coverage can be displayed for draft genomes; customized graphs and annotations can be displayed.
Using a user-defined set of genes as input, BRIG can display gene presence, absence, truncation or sequence variation in a set of complete genomes, draft genomes or even raw, unassembled sequence data.
BRIG also accepts SAM-formatted read-mapping files enabling genomic regions present in unassembled sequence data from multiple samples to be compared simultaneously

Address of the bookmark: http://brig.sourceforge.net/

Basic command-line to run BLAST

Shruti Paniwala — Wed, 14 Mar 2018 05:10:34 -0500

The goal of this tutorial is to run you through a demonstration of the command line, which you may not have seen or used much before.

All of the commands below can copy/pasted.

Install software

Copy and paste the following commands

sudo apt-get update && sudo apt-get -y install python ncbi-blast+

This updates the software list and installs the Python programming language and NCBI BLAST+.

Get Data

Grab some data to play with. Grab some cow and human RefSeq proteins:

wget ftp://ftp.ncbi.nih.gov/refseq/B_taurus/mRNA_Prot/cow.1.protein.faa.gz
wget ftp://ftp.ncbi.nih.gov/refseq/H_sapiens/mRNA_Prot/human.1.protein.faa.gz

This is only the first part of the human and cow protein files - there are 24 files total for human.

The database files are both gzipped, so lets unzip them

gunzip *gz
ls

Take a look at the head of each file:

head cow.1.protein.faa
head human.1.protein.faa

These are protein sequences in FASTA format. FASTA format is something many of you have probably seen in one form or another – it’s pretty ubiquitous. It’s just a text file, containing records; each record starts with a line beginning with a ‘>’, and then contains one or more lines of sequence text.

Note that the files are in fasta format, even though they end if ”.faa” instead of the usual ”.fasta”. This NCBI’s way of denoting that this is a fasta file with amino acids instead of nucleotides.

How many sequences are in each one?

grep -c '^>' cow.1.protein.faa
grep -c '^>' human.1.protein.faa

This grep command uses the c flag, which reports a count of lines with match to the pattern. In this case, the pattern is a regular expression, meaning match only lines that begin with a >.

This is a bit too big, lets take a smaller set for practice. Lets take the first two sequences of the cow proteins, which we can see are on the first 6 lines

head -6 cow.1.protein.faa > cow.small.faa

BLAST

Now we can blast these two cow sequences against the set of human sequences. First, we need to tell blast about our database. BLAST needs to do some pre-work on the database file prior to searching. This helps to make the software work a lot faster. Because you installed your own version of the sotware, you need to tell the shell where the software is located. Use the full path and the makeblastdb command:

makeblastdb -in human.1.protein.faa -dbtype prot
ls

Note that this makes a lot of extra files, with the same name as the database plus new extensions (.pin, .psq, etc). To make blast work, these files, called index files, must be in the same directory as the fasta file.

blastp [-h] [-help] [-import_search_strategy filename]
[-export_search_strategy filename] [-task task_name] [-db database_name]
[-dbsize num_letters] [-gilist filename] [-seqidlist filename]
[-negative_gilist filename] [-negative_seqidlist filename]
[-entrez_query entrez_query] [-db_soft_mask filtering_algorithm]
[-db_hard_mask filtering_algorithm] [-subject subject_input_file]
[-subject_loc range] [-query input_file] [-out output_file]
[-evalue evalue] [-word_size int_value] [-gapopen open_penalty]
[-gapextend extend_penalty] [-qcov_hsp_perc float_value]
[-max_hsps int_value] [-xdrop_ungap float_value] [-xdrop_gap float_value]
[-xdrop_gap_final float_value] [-searchsp int_value]
[-sum_stats bool_value] [-seg SEG_options] [-soft_masking soft_masking]
[-matrix matrix_name] [-threshold float_value] [-culling_limit int_value]
[-best_hit_overhang float_value] [-best_hit_score_edge float_value]
[-window_size int_value] [-lcase_masking] [-query_loc range]
[-parse_deflines] [-outfmt format] [-show_gis]
[-num_descriptions int_value] [-num_alignments int_value]
[-line_length line_length] [-html] [-max_target_seqs num_sequences]
[-num_threads int_value] [-ungapped] [-remote] [-comp_based_stats compo]
[-use_sw_tback] [-version]

Now we can run the blast job. We will use blastp, which is appropriate for protein to protein comparisons.

blastp -query cow.small.faa -db human.1.protein.faa

This gives us a lot of information on the terminal screen. But this is difficult to save and use later - Blast also gives the option of saving the text to a file.

    blastp -query cow.small.faa -db human.1.protein.faa -out cow_vs_human_blast_results.txt
ls

Take a look at the results using less. Note that there can be more than one match between the query and the same subject. These are referred to as high-scoring segment pairs (HSPs).

less cow_vs_human_blast_results.txt

So how do you know about all the options, such as the flag to create an output file? Lets also take a look at the help pages. Unfortunately there are no man pages (those are usually reserved for shell commands, but some software authors will provide them as well), but there is a text help output

blastp -help

To scroll through slowly

blastp -help | less

To quit the less screen, press the q key.

Parameters of interest include the -evalue (Default is 10?!?) and the -outfmt

Lets filter for more statistically significant matches with a different output format:

blastp \
-query cow.small.faa \
-db human.1.protein.faa \
-out cow_vs_human_blast_results.tab \
-evalue 1e-5 \
-outfmt 7

I broke the long single command into many lines with by “escaping” the newline. That forward slash tells the command line “Wait, I’m not done yet!”. So it waits for the next line of the command before executing.

Check out the results with less.

Lets try a medium sized data set next

head -199 cow.1.protein.faa > cow.medium.faa

What size is this db?

grep -c '^>' cow.medium.faa

Lets run the blast again, but this time lets return only the best hit for each query.

blastp \
-query cow.medium.faa \
-db human.1.protein.faa \
-out cow_vs_human_blast_results.tab \
-evalue 1e-5 \
-outfmt 6 \
-max_target_seqs 1

Summary

Review:

command line programs such as blast use flags to get information about how and what to do
blast options can be found by typing blastp -help
break a command up over many lines by using `` to “escape” the new line

Blastn

blastn [-h] [-help] [-import_search_strategy filename]
[-export_search_strategy filename] [-task task_name] [-db database_name]
[-dbsize num_letters] [-gilist filename] [-seqidlist filename]
[-negative_gilist filename] [-negative_seqidlist filename]
[-entrez_query entrez_query] [-db_soft_mask filtering_algorithm]
[-db_hard_mask filtering_algorithm] [-subject subject_input_file]
[-subject_loc range] [-query input_file] [-out output_file]
[-evalue evalue] [-word_size int_value] [-gapopen open_penalty]
[-gapextend extend_penalty] [-perc_identity float_value]
[-qcov_hsp_perc float_value] [-max_hsps int_value]
[-xdrop_ungap float_value] [-xdrop_gap float_value]
[-xdrop_gap_final float_value] [-searchsp int_value]
[-sum_stats bool_value] [-penalty penalty] [-reward reward] [-no_greedy]
[-min_raw_gapped_score int_value] [-template_type type]
[-template_length int_value] [-dust DUST_options]
[-filtering_db filtering_database]
[-window_masker_taxid window_masker_taxid]
[-window_masker_db window_masker_db] [-soft_masking soft_masking]
[-ungapped] [-culling_limit int_value] [-best_hit_overhang float_value]
[-best_hit_score_edge float_value] [-window_size int_value]
[-off_diagonal_range int_value] [-use_index boolean] [-index_name string]
[-lcase_masking] [-query_loc range] [-strand strand] [-parse_deflines]
[-outfmt format] [-show_gis] [-num_descriptions int_value]
[-num_alignments int_value] [-line_length line_length] [-html]
[-max_target_seqs num_sequences] [-num_threads int_value] [-remote]
[-version]

DESCRIPTION
Nucleotide-Nucleotide BLAST 2.7.0+

New Layout for BLAST ftp Database Site

Jit — Tue, 21 Jan 2020 11:57:11 -0600

As announced previously, the new default database version for BLAST+ is dbV5. To complete this transition, the ftp database site will be updated to support this change. We expect this change to happen around February 4^th, please adjust your scripts or procedures accordingly.

Here is a list of what is changing:

All databases at the root level will be dbV5.
The dbV5 file naming, “_v5” will be removed. Databases with no “_vX” descriptor will be dbV5.
dbV4 tarballs will be renamed with "_v4", files included in tarball will not be renamed.
dbV4 databases will be moved to a v4 subdirectory.
As of 1/13/20 the Cloud directory will be frozen with no more new entries.
The will be no more updates to dbV4 databases.
The FASTA directory will contain nr, nt, swissprot, and pdbaa files.

If you have any questions or concerns, please contact blast-help@ncbi.nlm.nih.gov

New BLAST Core Nucleotide Database (core_nt)

LEGE — Tue, 13 Aug 2024 07:12:53 -0500

The Core Nucleotide Database (core_nt) is now the default nucleotide BLAST database. Core_nt is also available on the command line. You get faster searches & more focused results.

Core_nt contains the same eukaryotic transcript and gene-related sequences as nt. The core_nt database is nt without most eukaryotic chromosome sequences. Most nucleotide BLAST searches with core_nt will be similar to the nt database. However, core_nt is better than nt for accomplishing your most common BLAST search goals, such as identifying gene-related sequences like transcript sequences and complete bacterial chromosomes. This is because, in recent years, nt has acquired more low-relevance, non-annotated, and non-gene content.

Learn more: https://ncbiinsights.ncbi.nlm.nih.gov/2024/07/18/new-blast-core-nucleotide-database/

RepeatMasker compatible blast tool

Neel — Fri, 07 Dec 2018 08:13:03 -0600

RMBlast is a RepeatMasker compatible version of the standard NCBI blastn program. The primary difference between this distribution and the NCBI distribution is the addition of a new program "rmblastn" for use with RepeatMasker and RepeatModeler.

RMBlast supports RepeatMasker searches by adding a few necessary features to the stock NCBI blastn program. These include:

Support for custom matrices ( without KA-Statistics ).
Support for cross_match-like complexity adjusted scoring. Cross_match is Phil Green's seeded smith-waterman search algorithm.
Support for cross_match-like masklevel filtering.

https://anaconda.org/bioconda/rmblast

Address of the bookmark: http://www.repeatmasker.org/RMBlast.html

Visualise blast results !

Abhi — Tue, 11 Oct 2022 03:15:10 -0500

Kablammo helps you create interactive visualizations of BLAST results from your web browser. Find your most interesting alignments, list detailed parameters for each, and export a publication-ready vector image, all without installing any software.

Address of the bookmark: https://kablammo.wasmuthlab.org/

A Step-by-Step Guide to Running BLAST Offline

LEGE — Sat, 07 Dec 2024 22:32:37 -0600

BLAST (Basic Local Alignment Search Tool) is a powerful algorithm used to compare nucleotide or protein sequences to sequence databases, identifying regions of similarity. Running BLAST offline provides more control, ensures data security, and allows customization for specific research needs. Here’s a detailed guide to set up and run BLAST locally on your system.

Step 1: Install BLAST

Download BLAST:
- Visit the NCBI BLAST+ download page to download the appropriate version for your operating system (Windows, macOS, or Linux).
Install BLAST:
- Extract the downloaded archive. For Linux/Mac, use:
```
tar -xvzf ncbi-blast-*.tar.gz
cd ncbi-blast-*
```
- Add the BLAST binary folder to your system PATH for easier access:
```
export PATH=$PATH:/path/to/ncbi-blast-*/bin
```
Verify Installation:
Run the following command to ensure BLAST is installed correctly:
```
blastn -version
```

Step 2: Prepare a Local Database

To run BLAST offline, you’ll need a sequence database.

Download a Pre-Built Database (Optional):
- NCBI provides ready-to-use databases such as nt, nr, and Swiss-Prot. Use the update_blastdb.pl script (bundled with BLAST) to download these:
```
update_blastdb.pl --decompress nt
```
Create a Custom Database:
If you have specific sequences to use as a database:
- Prepare a FASTA file containing the sequences.
- Use makeblastdb to create a database:
```
makeblastdb -in your_sequences.fasta -dbtype [nucl|prot] -out custom_db
```
  Replace [nucl|prot] with nucl for nucleotide sequences or prot for protein sequences.

Step 3: Prepare the Query Sequence

Save your query sequence(s) in FASTA format.
Ensure the file is properly formatted, with a header line starting with > followed by the sequence name, and the sequence on subsequent lines.

Example:

>query_sequence
ATGCGTAGCTAGCGTAGCTAGCTAGCTA

Step 4: Run BLAST

Choose the Appropriate BLAST Tool:
Depending on your data type:
- blastn: For nucleotide-nucleotide searches.
- blastp: For protein-protein searches.
- blastx: Translates nucleotide sequences into proteins and compares them to a protein database.
- tblastn: Compares protein sequences to a nucleotide database.
- tblastx: Translates both nucleotide query and database sequences.
Run the Command:
Example command for blastn:
```
blastn -query query.fasta -db custom_db -out results.txt -outfmt 6 -evalue 1e-5
```
Explanation of Parameters:
- -query: Specifies the query file.
- -db: Points to the local database.
- -out: Output file name.
- -outfmt: Output format (e.g., 6 for tabular format).
- -evalue: E-value cutoff for significance.

Step 5: Interpret Results

Output Formats:
- Default (outfmt 0): Human-readable format.
- Tabular (outfmt 6): Includes fields like query ID, subject ID, percent identity, alignment length, etc.
Analyze Results:
Use tools like grep, Python, or R to parse and filter results for downstream analysis.

Step 6: Optimize Performance

For large datasets, BLAST can be resource-intensive. To improve performance:

Multithreading:
Use the -num_threads option to leverage multiple CPU cores:

blastn -query query.fasta -db custom_db -out results.txt -num_threads 4

Database Subsetting:
Split large databases into smaller chunks for faster searches.
Adjust Parameters:
- Lower the -evalue threshold for stricter matches.
- Use -max_target_seqs to limit the number of results per query.

Step 7: Update Databases (Optional)

If using NCBI databases, regularly update them to ensure the inclusion of the latest sequences:

update_blastdb.pl --decompress nt

Conclusion

Running BLAST offline is a straightforward process that offers flexibility and security for bioinformaticians working with sensitive data. By following this guide, you can harness the power of BLAST to analyze sequences efficiently and gain valuable biological insights.

For advanced use cases, explore BLAST’s customization options, such as custom scoring matrices, filtering, and iterative searches with tools like PSI-BLAST. Happy BLASTing!

mScaffolder: A comparative genome scaffolding tool

Jit — Fri, 15 Jun 2018 04:48:01 -0500

A comparative genome scaffolding tool based on MUMmer

mScaffolder scaffolds a genome using an existing high quality genome as the reference. It aligns the two genomes using nucmer utility from MUMmer and then orders and orients the contigs of the candidate genome guided by their alignments to the reference genome. Please send your questions and comments to mchakrab@uci.edu.

Citation https://www.nature.com/articles/s41588-017-0010-y

Address of the bookmark: https://github.com/mahulchak/mscaffolder

Gap filling or Contigs extensions tools !

Rahul Nayak — Fri, 01 Jun 2018 08:07:32 -0500

There are many tools to perform gap filling using Illumina short reads, for example "GapFiller: a de novo assembly approach to fill the gap within paired reads" or "Toward almost closed genomes with GapFiller". There are also some tools like GAPresolution that can help to perform local re-assemblies using 454 reads. We used GAPresolution but it is not a very good software, it is useful only in some specific situations.

Take a look at the PRICE software from the DeRisi lab. Its meant to do something very similar. http://derisilab.ucsf.edu/index.php?page=software

You could also look at SSPACE (http://www.baseclear.com/landingpages/basetools-a-wide-range-of-bioinformatics-solutions/sspacev12/), ATLAS tools (http://www.hgsc.bcm.tmc.edu/content/bcm-hgsc-software), and SCARPA (http://compbio.cs.toronto.edu/hapsembler/scarpa.html).

See the PAGIT protocol: http://www.sanger.ac.uk/resources/software/pagit/

In particular, take a look at the IMAGE tool: http://genomebiology.com/2010/11/4/R41

Also SOAPdenovo has ha function for scaffolding. Not sure about ABYSS

Here there is a useful explanation of several tools.

https://bioinformaticsonline.com/search?q=scaffolding&entity_type=object&entity_subtype=bookmarks&offset=0&search_type=entities

I could be wrong, but the above answers to your hypothetical scenario appear to miss the point that you aren't interested in assembling the full genome, just the 100 kb part you're interested in. I suggest the following algorithm:

1. Start with the initial assembly C0 of the contigs you have identified as overlapping your region of interest, and the set S of reads those contigs contain. Let C = C0.

2. Repeat:
a. Identify paired-end reads (not in C) for which one or both ends align within, or extending, contigs in C.
b. Identify unpaired reads that align extending these new paired-end reads.
c. Construct a new assembly C' from C and the new reads identified in (a) and (b).
d. Trim C' so it does not extend more than 100 kb to either end of C0. Set C = C'.
e. Let S' denote the reads that contribute to C'. If S' does not contain any reads not present in S, stop. Otherwise, Set S = S'.

3. If you don't have a complete assembly of the region of interest, generate an STS for each end of each contig, probe a library for clones including these STSes, subclone these clones into a paired-end sequencing vector, and generate paired-end reads for this library; then try steps (1) and (2) again, adding these new sequencing reads to what you had before.

4. If your average sequencing depth for the region of interest exceeds 25 or so without filling all gaps, it is likely that the remaining gaps represent sequences that are not getting cloned in your sequencing vectors. Try different sequencing vectors.