<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/17176?offset=1290</link>
	<atom:link href="https://bioinformaticsonline.com/related/17176?offset=1290" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	
<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/4728/3-days-intensive-course-on-understanding-omics-data-in-basel-switzerland-19-21st-november</guid>
  <pubDate>Mon, 23 Sep 2013 10:46:57 -0500</pubDate>
  <link></link>
  <title><![CDATA[3 days intensive course on Understanding 'omics data in Basel, Switzerland, 19-21st November]]></title>
  <description><![CDATA[
<p>Benefits for the participants</p>

<p>- Plan more efficient experiments<br />- Correctly interpret results<br />- Communicate results in publications more effectively</p>

<p>The course focus is on methodologies, not on particular software tools. After the course participants should be able to apply the methods in their respective environment. However, during the course, hands-on sessions will be performed using the Genedata Expressionist® software, which enables participants to quickly apply the discussed methods and visualize results. No previous knowledge on Expressionist® is required; access to the software is free of charge during the course.</p>

<p>More @ http://www.dixa-fp7.eu/dixa-training/dixa-training-agenda/genedata-academy#!</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/44516/16srna-database-download</guid>
	<pubDate>Wed, 24 Apr 2024 04:33:15 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/44516/16srna-database-download</link>
	<title><![CDATA[16sRNA Database Download]]></title>
	<description><![CDATA[<p>Downloading 16S rRNA databases can be crucial for various bioinformatics analyses, especially in microbiome research. However, it's important to note that databases can vary based on your specific needs, such as the taxonomic coverage you require or the type of analysis you're performing. Here's a general guideline on how you can obtain 16S rRNA databases:</p><ol>
<li>
<p><span>NCBI (National Center for Biotechnology Information)</span>:</p>
<ul>
<li>NCBI provides various databases related to genetic information, including 16S rRNA sequences.</li>
<li>You can access the 16S ribosomal RNA sequences from NCBI's Nucleotide database (<a href="https://www.ncbi.nlm.nih.gov/nucleotide/" target="_new">https://www.ncbi.nlm.nih.gov/nucleotide/</a>).</li>
<li>Perform a search using keywords like "16S rRNA" or specific bacterial names to find relevant sequences.</li>
<li>You can download sequences individually or in batches using the provided tools.</li>
</ul>
</li>
<li>
<p><span>GreenGenes</span>:</p>
<ul>
<li>GreenGenes is a widely used 16S rRNA gene sequence database.</li>
<li>You can access it at <a target="_new">http://greengenes.secondgenome.com/</a>.</li>
<li>GreenGenes provides precompiled databases for various purposes, including classification, alignment, and phylogenetic analysis.</li>
</ul>
</li>
<li>
<p><span>SILVA</span>:</p>
<ul>
<li>SILVA (<a href="https://www.arb-silva.de/" target="_new">https://www.arb-silva.de/</a>) is another comprehensive database for ribosomal RNA (rRNA) sequences.</li>
<li>It covers not only 16S rRNA but also other ribosomal RNA sequences.</li>
<li>SILVA provides precompiled databases for various purposes, including taxonomic classification and alignment.</li>
</ul>
</li>
<li>
<p><span>Ribosomal Database Project (RDP)</span>:</p>
<ul>
<li>RDP (<a target="_new">http://rdp.cme.msu.edu/</a>) is a curated database that offers 16S rRNA sequences.</li>
<li>It provides tools for sequence analysis and classification.</li>
<li>You can download sequences and taxonomy information from their website.</li>
</ul>
</li>
<li>
<p><span>QIIME (Quantitative Insights Into Microbial Ecology)</span>:</p>
<ul>
<li>QIIME (<a href="https://qiime2.org/" target="_new">https://qiime2.org/</a>) is a widely used bioinformatics platform for microbiome analysis.</li>
<li>It provides tools for analyzing microbial communities, including processing 16S rRNA sequences.</li>
<li>QIIME often includes its own preprocessed 16S rRNA databases that can be used for analysis within the platform.</li>
</ul>
</li>
</ol><p>Before downloading any database, make sure to read the terms of use and citation requirements, as some databases may have specific usage policies. Additionally, consider the compatibility of the database with your analysis pipeline and software tools.</p><p>&nbsp;</p><p>NCBI 16s RNA database location&nbsp;ftp://ftp.ncbi.nih.gov/blast/db/16SMicrobial.tar.gz</p>]]></description>
	<dc:creator>LEGE</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/2422/bioinformatics-codes-search</guid>
	<pubDate>Thu, 15 Aug 2013 11:08:52 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/2422/bioinformatics-codes-search</link>
	<title><![CDATA[Bioinformatics Codes Search]]></title>
	<description><![CDATA[<p>I bet, this website will be your best friend in near future. This helps us to explore the existing open source codes and learn from it.</p>
<p>You can find some useful open source bioinformatics codes for your analysis work. You can use the left bar options to filtere out or narrow down your search result. This webpage can be an useful resource for a beginners bioinformatician as it contain several bioinformatics basics script that are commonly used by biological programmers and biologist.</p>
<p>Stand on the slumped, dandruff-covered shoulders of millions of computer nerds. _/\_</p>
<p>Enjoy the code and research work.</p>
<p>http://code.ohloh.net/search?s=bioinformatics</p><p>Address of the bookmark: <a href="http://code.ohloh.net/search?s=bioinformatics" rel="nofollow">http://code.ohloh.net/search?s=bioinformatics</a></p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/4946/crcri-bioinfomatics-walk-in-on-08102013</guid>
  <pubDate>Fri, 27 Sep 2013 10:59:53 -0500</pubDate>
  <link></link>
  <title><![CDATA[CRCRI Bioinfomatics Walk In on 08.10.2013]]></title>
  <description><![CDATA[
<p>Walk-in-Interview for recruitment of one Project Fellow for a period of 10 months purely on temporary basis is proposed to be held at Central Tuber Crops Research Institute, Sreekariyam, Thiruvananthapuram for a KSCSTE funded project entitled “PARTICIPATORY DEVELOPMENT OF A WEB BASED USER FRIENDLY CASSAVA EXPERT SYSTEM”</p>

<p>Salary: Rs. 10,000/- per month.</p>

<p>Age limit: 35 for men and 40 for women &amp; SC/ST.</p>

<p>Qualification: First class in M. Sc (Agriculture)/MCA/M.Sc (IT)/ M. Sc (Computer Application)/M.Sc (Bioinformatics)/M.Sc (Geoinformatics).</p>

<p>Desirable: Two years experience in web design and web programming.</p>

<p>Date &amp; time of interview: 08.10.2013, 10 am</p>

<p>Interested candidates may appear for an interview at this institute along with their application in plain paper containing the following particulars viz. (1) Name (2) Father/Husband/Guardian’s Name (3) date of birth &amp; age as on 01.10.2013 (4) Permanent address (5) Address for communication (6) Email address and Telephone No. with code (7) Qualification (8) National fellowship like ICAR/CSIR/UGC etc. if any (9) Whether SC/ST/OBC (10) Details of experience (Attested copies of degree certificate, proof of age, mark sheets). Original certificates should be produced for verification.</p>

<p>No TA/DA will be admissible to the candidates attending the test. The selected candidate will have to join immediately.</p>

<p>Advertisement: http://www.ctcri.org/careers/mithra_SRF.doc</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/5191/programming-language-to-build-synthetic-dna</guid>
	<pubDate>Mon, 30 Sep 2013 16:37:24 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/5191/programming-language-to-build-synthetic-dna</link>
	<title><![CDATA[Programming language to build synthetic DNA]]></title>
	<description><![CDATA[<p style="color: #333333; font-size: 13px; font-style: normal; font-weight: normal; text-align: start;">A team led by <a href="http://homes.cs.washington.edu/~seelig/index.html">Georg Seelig</a>&nbsp;(<a href="http://homes.cs.washington.edu/~seelig/index.html">http://homes.cs.washington.edu/~seelig/index.html</a>) at&nbsp;University of Washington has developed a programming language for chemistry that it hopes will streamline efforts to design a network that can guide the behavior of chemical-reaction mixtures in the same way that embedded electronic controllers guide cars, robots and other devices. In medicine, such networks could serve as &ldquo;smart&rdquo; drug deliverers or disease detectors at the cellular level.</p><p style="color: #333333; font-size: 13px; font-style: normal; font-weight: normal; text-align: start;">Reference &amp; More @</p><p style="color: #333333; font-size: 13px; font-style: normal; font-weight: normal; text-align: start;"><a href="http://www.nature.com/nnano/journal/vaop/ncurrent/full/nnano.2013.189.html">http://www.nature.com/nnano/journal/vaop/ncurrent/full/nnano.2013.189.html</a></p><p style="color: #333333; font-size: 13px; font-style: normal; font-weight: normal; text-align: start;"><a href="http://www.washington.edu/news/2013/09/30/uw-engineers-invent-programming-language-to-build-synthetic-dna/">http://www.washington.edu/news/2013/09/30/uw-engineers-invent-programming-language-to-build-synthetic-dna/</a></p><p style="color: #333333; font-size: 13px; font-style: normal; font-weight: normal; text-align: start;">Image source:&nbsp;washington.edu</p><p style="color: #333333; font-size: 13px; font-style: normal; font-weight: normal; text-align: start;"><img src="http://www.washington.edu/news/files/2013/09/Programmable-chemistry-2.jpg" alt="image" style="border: 0px; border: 0px;"></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/22938/research-assistant-in-computational-biology</guid>
  <pubDate>Wed, 24 Jun 2015 07:55:16 -0500</pubDate>
  <link></link>
  <title><![CDATA[Research assistant in computational biology]]></title>
  <description><![CDATA[
<p>http://www.au.dk/en/about/vacant-positions/scientific-positions/stillinger/Vacancy/show/743161/5283/</p>

<p>Qualifications:<br />MSc degree in computer science, engineering, genetics or similar field with a strong emphasis on computational methods.</p>

<p>Deadline<br />01.08.2015</p>
]]></description>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/5255/walk-in-interview-indian-agricultural-statistics-research-institute</guid>
  <pubDate>Wed, 02 Oct 2013 15:40:17 -0500</pubDate>
  <link></link>
  <title><![CDATA[Walk-in-Interview @ Indian Agricultural Statistics Research Institute]]></title>
  <description><![CDATA[
<p>Indian Agricultural Statistics Research Institute<br />Library Avenue, Pusa, New Delhi – 110012</p>

<p>Walk-in-Interview</p>

<p>Walk-in-interview will be held on October 5, 2013 at 10:00 A.M. at IASRI, New Delhi for a project “A New Distributed Computing Framework for Data Mining” funded by Department of Electronics and Information Technology, Government of India for the following posts. The appointment will be on contractual basis upto 14th October, 2015 or till the termination of the project whichever is earlier and the incumbent shall not have any claim for regular appointment under ICAR.</p>

<p>Research Associate</p>

<p>    Ph.D. in Bioinformatics/ Agricultural Statistics/ Statistics/ Computer Science/ Computer Application or equivalent or</p>

<p>    Post-Graduation in Bioinformatics/ Agricultural Statistics/ Statistics/ Computer Science/ Computer Application or equivalent with 1st Division and at least two years of research experience</p>

<p>     Knowledge of Statistical Analysis /Bioinformatics tools for computational genomics.</p>

<p>     Knowledge of R/Perl programming language</p>

<p>Research Associate</p>

<p>    Ph.D. in Computer Science/ Computer Application / Bioinformatics/ Agricultural<br />    Statistics/ Statistics or equivalent or</p>

<p>    Post-Graduation in Computer Science/ Computer Application /Bioinformatics/ Agricultural Statistics/ Statistics or equivalent with 1st Division and at least two years of research experience</p>

<p>     Expertise in Java programming.<br />     Knowledge of system administration and networking under Linux environment.<br />     Knowledge of parallel programming and cluster computing.</p>

<p>Emoluments for Research Associate: Consolidated Rs:24000/- per month + HRA (for Ph.D. Degree holders) and Rs:23000/- per month + HRA (for Master’s Degree holders)</p>

<p>Age Limit: Age should be not more than 40 years (5 years relaxation for  SC/ST/women candidates and 3 years for OBC candidates as on date of interview).</p>

<p>Interested candidates are requested to appear for Walk-in-Interview on the date and time as specified above in Room No. 106, Training Cum Administrative Block of the Institute along with their application giving bio-data with attested copies of certificates, degrees, testimonials, etc. and one passport size photograph.</p>

<p>Original certificates/ Degrees are needed to be produced at the time of interview.</p>

<p>No T.A. /D.A. will be paid for appearing in the interview.</p>

<p>Advertisement: http://www.iasri.res.in/employment/employment.htm</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/44601/free-resources-to-learn-statistics</guid>
	<pubDate>Sat, 06 Jul 2024 10:30:50 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/44601/free-resources-to-learn-statistics</link>
	<title><![CDATA[Free resources to learn statistics]]></title>
	<description><![CDATA[<p><span>Welcome to the course notes for&nbsp;</span><span>STAT 414: Introduction to Probability Theory</span><span>. These notes are designed and developed by Penn State's&nbsp;</span><a href="https://science.psu.edu/stat">Department of Statistics</a><span>&nbsp;and offered as open educational resources. These notes are free to use under Creative Commons license&nbsp;</span><a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC 4.0</a><span>.</span></p>
<p>&nbsp;</p>
<p>A free online version of the second edition of the book based on Stat 110,&nbsp;<em>Introduction to Probability</em>&nbsp;by Joe Blitzstein and Jessica Hwang,&nbsp;is now available at&nbsp;<a href="http://probabilitybook.net/" title="http://probabilitybook.net">http://probabilitybook.net</a></p>
<p>Print copies are available via&nbsp;<a href="https://www.crcpress.com/Introduction-to-Probability-Second-Edition/Blitzstein-Hwang/p/book/9781138369917" title="">CRC Press</a>,&nbsp;<a href="https://amzn.to/2Ubh7D8" title="">Amazon</a>, and elsewhere.&nbsp;</p>
<p>Stat110x is also available as an&nbsp;edX course.&nbsp;Free signup at&nbsp;<a href="https://www.edx.org/course/introduction-to-probability-0" title="https://www.edx.org/course/introduction-to-probability-0">https://www.edx.org/course/introduction-to-probability-0</a></p>
<p>The edX course focuses on animations, interactive features, readings, and problem-solving, and&nbsp;is&nbsp;<strong>complementary</strong>&nbsp;to the Stat 110 lecture videos on YouTube, which are available at&nbsp;<a href="https://goo.gl/i7njSb" title="https://goo.gl/i7njSb">https://goo.gl/i7njSb</a></p>
<p>The Stat110x animations are available within the course and at&nbsp;<a href="https://goo.gl/g7pqTo" title="">https://goo.gl/g7pqTo</a></p>
<p><a href="https://projects.iq.harvard.edu/stat110/home">https://projects.iq.harvard.edu/stat110/home</a>&nbsp;</p><p>Address of the bookmark: <a href="https://online.stat.psu.edu/stat414/" rel="nofollow">https://online.stat.psu.edu/stat414/</a></p>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/34702/run-miniasm-assembler-on-nanopore-reads</guid>
	<pubDate>Mon, 18 Dec 2017 04:07:50 -0600</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/34702/run-miniasm-assembler-on-nanopore-reads</link>
	<title><![CDATA[Run miniasm assembler on nanopore reads !]]></title>
	<description><![CDATA[<p>Miniasm is a very fast OLC-based&nbsp;<em>de novo</em>&nbsp;assembler for noisy long reads. It takes all-vs-all read self-mappings (typically by&nbsp;<a href="https://github.com/lh3/minimap">minimap</a>) as input and outputs an assembly graph in the&nbsp;<a href="https://github.com/pmelsted/GFA-spec/blob/master/GFA-spec.md">GFA</a>&nbsp;format. Different from mainstream assemblers, miniasm does not have a consensus step. It simply concatenates pieces of read sequences to generate the final&nbsp;<a href="http://wgs-assembler.sourceforge.net/wiki/index.php/Celera_Assembler_Terminology">unitig</a>&nbsp;sequences. Thus the per-base error rate is similar to the raw input reads.</p><p>Find the detail of the reads repeats:</p><blockquote><p>fq2fa ONT_A.fastq ONT_A.fasta&nbsp;<br /><br />minimap2 -xava-ont ONT_A.fasta ONT_A.fasta -t10 -X &gt; AONT.paf&nbsp;<br /><br />awk '{if($1==$6){print}}' AONT.paf &gt; AONTself.paf&nbsp;<br /><br />awk '$5=="-"' AONTself.paf | awk '{print $1}'| sort|uniq &gt; invertedrepeat.list</p></blockquote><p>Generated a few palindrome and repeats plots (highlighting only repeats largest than 10, 20 and 30 kb)</p><blockquote><p>minidot -f 5 -m 30000 AONTself.paf &gt; AONTself30000.eps&nbsp;<br />sed 's/_template_pass_FAH31515//' AONTself30000.eps &gt; AONTself30000final.eps&nbsp;<br /><br />minidot -f 5 -m 20000 AONTself.paf &gt; AONTself20000.eps&nbsp;<br />sed 's/_template_pass_FAH31515//' AONTself20000.eps &gt; AONTself20000final.eps&nbsp;<br /><br />minidot -f 5 -m 10000 AONTself.paf &gt; AONTself10000.eps&nbsp;<br />sed 's/_template_pass_FAH31515//' AONTself10000.eps &gt; AONTself10000final.eps&nbsp;</p></blockquote><p>Assemble with miniasm:</p><blockquote><p>miniasm -f ONT_A.fasta AONT.paf &gt; AONT.gfa&nbsp;</p><p>grep '^S' AONT.gfa |awk '{print "&gt;"$2"\n"$3}' &gt; AONT_miniasm.fasta&nbsp;<br /><br />minimap2 -xasm10 AONT_miniasm.fasta AONT_miniasm.fasta -t1 -X &gt; AONT_miniasm.paf&nbsp;<br /><br />awk '{if($1==$6){print}}' AONT_miniasm.paf &gt; AONT_miniasm_self.paf&nbsp;<br /><br />minidot -f 5 -m 10000 AONT_miniasm_self.paf &gt; AONT_miniasm_self10000.eps&nbsp;</p></blockquote><p>Njoy the assembly !</p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/35923/basic-command-line-to-run-blast</guid>
	<pubDate>Wed, 14 Mar 2018 05:10:34 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/35923/basic-command-line-to-run-blast</link>
	<title><![CDATA[Basic command-line to run BLAST]]></title>
	<description><![CDATA[<p>&nbsp;</p><p>The goal of this tutorial is to run you through a demonstration of the command line, which you may not have seen or used much before.</p><p>All of the commands below can copy/pasted.</p><div id="install-software"><h2>Install software<a href="http://angus.readthedocs.io/en/2016/running-command-line-blast.html#install-software" title="Permalink to this headline"></a></h2><p>Copy and paste the following commands</p><div><div><pre>sudo apt-get update &amp;&amp; sudo apt-get -y install python ncbi-blast+
</pre></div></div><p>This updates the software list and installs the Python programming language and NCBI BLAST+.</p></div><div id="get-data"><h2>Get Data<a href="http://angus.readthedocs.io/en/2016/running-command-line-blast.html#get-data" title="Permalink to this headline"></a></h2><p>Grab some data to play with. Grab some cow and human RefSeq proteins:</p><div><div><pre>wget ftp://ftp.ncbi.nih.gov/refseq/B_taurus/mRNA_Prot/cow.1.protein.faa.gz
wget ftp://ftp.ncbi.nih.gov/refseq/H_sapiens/mRNA_Prot/human.1.protein.faa.gz
</pre></div></div><p>This is only the first part of the human and cow protein files - there are 24 files total for human.</p><p>The database files are both gzipped, so lets unzip them</p><div><div><pre>gunzip *gz
ls
</pre></div></div><p>Take a look at the head of each file:</p><div><div><pre>head cow.1.protein.faa
head human.1.protein.faa
</pre></div></div><p>These are protein sequences in FASTA format. FASTA format is something many of you have probably seen in one form or another &ndash; it&rsquo;s pretty ubiquitous. It&rsquo;s just a text file, containing records; each record starts with a line beginning with a &lsquo;&gt;&rsquo;, and then contains one or more lines of sequence text.</p><p>Note that the files are in fasta format, even though they end if &rdquo;.faa&rdquo; instead of the usual &rdquo;.fasta&rdquo;. This NCBI&rsquo;s way of denoting that this is a fasta file with amino acids instead of nucleotides.</p><p>How many sequences are in each one?</p><div><div><pre>grep -c '^&gt;' cow.1.protein.faa
grep -c '^&gt;' human.1.protein.faa
</pre></div></div><p>This grep command uses the c flag, which reports a count of lines with match to the pattern. In this case, the pattern is a regular expression, meaning match only lines that begin with a &gt;.</p><p>This is a bit too big, lets take a smaller set for practice. Lets take the first two sequences of the cow proteins, which we can see are on the first 6 lines</p><div><div><pre>head -6 cow.1.protein.faa &gt; cow.small.faa
</pre></div></div></div><div id="blast"><h2>BLAST<a href="http://angus.readthedocs.io/en/2016/running-command-line-blast.html#blast" title="Permalink to this headline"></a></h2><p>Now we can blast these two cow sequences against the set of human sequences. First, we need to tell blast about our database. BLAST needs to do some pre-work on the database file prior to searching. This helps to make the software work a lot faster. Because you installed your own version of the sotware, you need to tell the shell where the software is located. Use the full path and the makeblastdb command:</p><div><div><pre>makeblastdb -in human.1.protein.faa -dbtype prot
ls
</pre></div></div><p>Note that this makes a lot of extra files, with the same name as the database plus new extensions (.pin, .psq, etc). To make blast work, these files, called index files, must be in the same directory as the fasta file.</p><p><br /> blastp [-h] [-help] [-import_search_strategy filename]<br /> [-export_search_strategy filename] [-task task_name] [-db database_name]<br /> [-dbsize num_letters] [-gilist filename] [-seqidlist filename]<br /> [-negative_gilist filename] [-negative_seqidlist filename]<br /> [-entrez_query entrez_query] [-db_soft_mask filtering_algorithm]<br /> [-db_hard_mask filtering_algorithm] [-subject subject_input_file]<br /> [-subject_loc range] [-query input_file] [-out output_file]<br /> [-evalue evalue] [-word_size int_value] [-gapopen open_penalty]<br /> [-gapextend extend_penalty] [-qcov_hsp_perc float_value]<br /> [-max_hsps int_value] [-xdrop_ungap float_value] [-xdrop_gap float_value]<br /> [-xdrop_gap_final float_value] [-searchsp int_value]<br /> [-sum_stats bool_value] [-seg SEG_options] [-soft_masking soft_masking]<br /> [-matrix matrix_name] [-threshold float_value] [-culling_limit int_value]<br /> [-best_hit_overhang float_value] [-best_hit_score_edge float_value]<br /> [-window_size int_value] [-lcase_masking] [-query_loc range]<br /> [-parse_deflines] [-outfmt format] [-show_gis]<br /> [-num_descriptions int_value] [-num_alignments int_value]<br /> [-line_length line_length] [-html] [-max_target_seqs num_sequences]<br /> [-num_threads int_value] [-ungapped] [-remote] [-comp_based_stats compo]<br /> [-use_sw_tback] [-version]</p><p>Now we can run the blast job. We will use blastp, which is appropriate for protein to protein comparisons.</p><div><div><pre>blastp -query cow.small.faa -db human.1.protein.faa
</pre></div></div><p>This gives us a lot of information on the terminal screen. But this is difficult to save and use later - Blast also gives the option of saving the text to a file.</p><div><div><pre>    blastp -query cow.small.faa -db human.1.protein.faa -out cow_vs_human_blast_results.txt
ls
</pre></div></div><p>Take a look at the results using less. Note that there can be more than one match between the query and the same subject. These are referred to as high-scoring segment pairs (HSPs).</p><div><div><pre>less cow_vs_human_blast_results.txt
</pre></div></div><p>So how do you know about all the options, such as the flag to create an output file? Lets also take a look at the help pages. Unfortunately there are no man pages (those are usually reserved for shell commands, but some software authors will provide them as well), but there is a text help output</p><div><div><pre>blastp -help
</pre></div></div><p>To scroll through slowly</p><div><div><pre>blastp -help | less
</pre></div></div><p>To quit the less screen, press the q key.</p><p>Parameters of interest include the -evalue (Default is 10?!?) and the -outfmt</p><p>Lets filter for more statistically significant matches with a different output format:</p><div><div><pre>blastp \
-query cow.small.faa \
-db human.1.protein.faa \
-out cow_vs_human_blast_results.tab \
-evalue 1e-5 \
-outfmt 7
</pre></div></div><p>I broke the long single command into many lines with by &ldquo;escaping&rdquo; the newline. That forward slash tells the command line &ldquo;Wait, I&rsquo;m not done yet!&rdquo;. So it waits for the next line of the command before executing.</p><p>Check out the results with less.</p><p>Lets try a medium sized data set next</p><div><div><pre>head -199 cow.1.protein.faa &gt; cow.medium.faa
</pre></div></div><p>What size is this db?</p><div><div><pre>grep -c '^&gt;' cow.medium.faa
</pre></div></div><p>Lets run the blast again, but this time lets return only the best hit for each query.</p><div><div><pre>blastp \
-query cow.medium.faa \
-db human.1.protein.faa \
-out cow_vs_human_blast_results.tab \
-evalue 1e-5 \
-outfmt 6 \
-max_target_seqs 1
</pre></div></div></div><div id="summary"><h2>Summary<a href="http://angus.readthedocs.io/en/2016/running-command-line-blast.html#summary" title="Permalink to this headline"></a></h2><p>Review:</p><ul>
<li>command line programs such as blast use flags to get information about how and what to do</li>
<li>blast options can be found by typing&nbsp;<cite>blastp -help</cite></li>
<li>break a command up over many lines by using&nbsp;<a href="http://angus.readthedocs.io/en/2016/running-command-line-blast.html#id1">`</a>` to &ldquo;escape&rdquo; the new line</li>
</ul><p>&nbsp;</p><p>Blastn</p><p>blastn [-h] [-help] [-import_search_strategy filename]<br /> [-export_search_strategy filename] [-task task_name] [-db database_name]<br /> [-dbsize num_letters] [-gilist filename] [-seqidlist filename]<br /> [-negative_gilist filename] [-negative_seqidlist filename]<br /> [-entrez_query entrez_query] [-db_soft_mask filtering_algorithm]<br /> [-db_hard_mask filtering_algorithm] [-subject subject_input_file]<br /> [-subject_loc range] [-query input_file] [-out output_file]<br /> [-evalue evalue] [-word_size int_value] [-gapopen open_penalty]<br /> [-gapextend extend_penalty] [-perc_identity float_value]<br /> [-qcov_hsp_perc float_value] [-max_hsps int_value]<br /> [-xdrop_ungap float_value] [-xdrop_gap float_value]<br /> [-xdrop_gap_final float_value] [-searchsp int_value]<br /> [-sum_stats bool_value] [-penalty penalty] [-reward reward] [-no_greedy]<br /> [-min_raw_gapped_score int_value] [-template_type type]<br /> [-template_length int_value] [-dust DUST_options]<br /> [-filtering_db filtering_database]<br /> [-window_masker_taxid window_masker_taxid]<br /> [-window_masker_db window_masker_db] [-soft_masking soft_masking]<br /> [-ungapped] [-culling_limit int_value] [-best_hit_overhang float_value]<br /> [-best_hit_score_edge float_value] [-window_size int_value]<br /> [-off_diagonal_range int_value] [-use_index boolean] [-index_name string]<br /> [-lcase_masking] [-query_loc range] [-strand strand] [-parse_deflines]<br /> [-outfmt format] [-show_gis] [-num_descriptions int_value]<br /> [-num_alignments int_value] [-line_length line_length] [-html]<br /> [-max_target_seqs num_sequences] [-num_threads int_value] [-remote]<br /> [-version]</p><p>DESCRIPTION<br /> Nucleotide-Nucleotide BLAST 2.7.0+</p></div>]]></description>
	<dc:creator>Shruti Paniwala</dc:creator>
</item>

</channel>
</rss>