<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/42143?offset=90</link>
	<atom:link href="https://bioinformaticsonline.com/related/42143?offset=90" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36267/pspairwise-sequentially-markovian-coalescent-psmc-model</guid>
	<pubDate>Thu, 19 Apr 2018 05:29:23 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36267/pspairwise-sequentially-markovian-coalescent-psmc-model</link>
	<title><![CDATA[PSPairwise Sequentially Markovian Coalescent (PSMC) model]]></title>
	<description><![CDATA[<p><span>Implementation of the Pairwise Sequentially Markovian Coalescent (PSMC) model</span></p><p>Address of the bookmark: <a href="https://github.com/lh3/psmc" rel="nofollow">https://github.com/lh3/psmc</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/38649/ngs-platforms-launched-by-bgi%E2%80%99s-mgi-tech</guid>
	<pubDate>Thu, 10 Jan 2019 04:42:06 -0600</pubDate>
	<link>https://bioinformaticsonline.com/news/view/38649/ngs-platforms-launched-by-bgi%E2%80%99s-mgi-tech</link>
	<title><![CDATA[NGS Platforms launched by BGI’s MGI Tech]]></title>
	<description><![CDATA[<p>MGI Tech Co., Ltd. (MGI), a subsidiary of BGI Group, is committed to enabling effective and affordable healthcare solutions for all. Based on its proprietary technology, MGI produces sequencing devices, equipment, consumables and reagents to support life science research, medicine and healthcare. MGI's multi-omics platforms include genetic sequencing, mass spectrometry and medical imaging. Providing real-time, comprehensive, life-long solutions, its mission&nbsp;is to&nbsp;develop and promote advanced life science tools for future healthcare.</p><p>MGI, a subsidiary of global genomics leader BGI Group, announced pricing and its first early access customer for the new ultra high-throughput sequencer, MGISEQ-T7, saying it has driven down sequencing cost to&nbsp;$5&nbsp;per gigabyte, with exceptionally high accuracy. Such innovations are helping more people to realize the benefits of genomic information.</p><p>In October, MGI launched the MGISEQ-T7, a highly flexible production-scale platform that is the most powerful sequencer to date. It can produce as many as 60 whole human genomes in one day. The instrument sells for&nbsp;$1 million.</p><p>The T7 enables simultaneous but independent operation of up to four flow cells, which means different applications such as single-cell RNA sequencing, whole exome sequencing and whole genome sequencing can be run in different flow cells at the same time. This helps to reduce costs, allowing MGI to offer the most competitive sequencing price in the market.</p><p><span>Powered by DNBseq&trade;, MGISEQ delivers quality data with accuracy for SNP and Indel calling rate of 99.9% and 99%, respectively, along with decreased duplication rate down to less than 2 percent, and almost zero Index mis-assignment rate.</span></p><p><span><span>SOURCE MGI</span></span></p><p>https://www.bgi.com/global/company/news/bgis-mgi-tech-launches-two-new-ngs-platforms/</p><p>http://en.mgitech.cn/</p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/41394/ngsymposium-in-computational-biology</guid>
  <pubDate>Mon, 09 Mar 2020 06:00:30 -0500</pubDate>
  <link></link>
  <title><![CDATA[NGSymposium in Computational Biology]]></title>
  <description><![CDATA[
<p>We have a great pleasure to invite you to the NGSymposium in Computational Biology to celebrate the 5th anniversary of the NGSchool Summer Schools. This international conference will make way for exchanging knowledge and experiences between experienced and early-stage researchers as well as bioinformaticians. The meeting will be held on 31.07 - 1.08.2020 in Warsaw. It will be a satellite event to the #NGSchool2020: Statistical Learning in Genomics. It will cover a wide range of topics from basic and applied biomedical sciences: bioinformatics, genomics, transcriptomics, computational biology, Machine Learning.</p>

<p>Registration of active participants will be open from February, 27 12 PM CET to April 17, 23:59 CET. In registration forms you will be asked for providing us with some basic information about yourself. You will also be able to submit your abstract. You can save your registration form after filling it partially and come back later to supply more data e.g. upload an abstract. Your registration will be completed only with the payment of the registration fee reaching our accounts - please make sure to transfer the money in advance!</p>

<p>Registration of passive participants will be open after closing of registration of active participants.</p>

<p>Details an registration: https://ngschool.eu/conference/</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/4574/tools-to-detect-synteny-blocks-regions-among-multiple-genomes</guid>
	<pubDate>Mon, 16 Sep 2013 17:12:02 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/4574/tools-to-detect-synteny-blocks-regions-among-multiple-genomes</link>
	<title><![CDATA[Tools to detect synteny blocks regions among multiple genomes]]></title>
	<description><![CDATA[<p>The synteny block (which etymologically means &ldquo;on the same ribbon&rdquo;) is a collection of contiguous genes located on the same chromosome. These block regions have mostly been preserved by genome rearrangements, and so synteny blocks from two related species (e.g., humans and mice) will be roughly similar but flipped around on the respective genomes. Ovcharenko et. al. define it as &lsquo;any conserved sequence blocks, regardless of whether it encompasses multiple genes, an area containing single genes, or areas devoid of known genes to be considers as synteny block as long as there is conservation at the sequence level. Today, however, biologists usually refer to synteny as the conservation of blocks of order within two sets of chromosomes that are being compared with each other. This concept can also be referred to as shared synteny. The NHBLI/NCBI Glossary define synteny as &ldquo;Two genes which occur on the same chromosome are syntenic; however, syntenic genes may or may not be "linked."</p><p>Now a day, geneticists have developed a language of their own. They are pouring lots of money and energy to read the entire genomic text and understand the gods own code ATGC. It is somewhat fascinating, not only for geneticist but also for non-biologist to know that there are several conserved blocks in genome which remain conserved over hundreds of millions of years. There have been several researches on conserved blocks and non-conserved regions to understand the mechanism and importance of all these regions (http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2675965/). The finding indicates conservation and rearrangements of certain evolutionary important genes play an important role in evolution/adaptive changes (http://www.nature.com/nature/journal/v491/n7424/abs/nature11622.html https://academic.oup.com/gbe/article/8/8/2442/2198198/Novel-Insights-into-Chromosome-Evolution-in-Birds , http://science.sciencemag.org/content/346/6215/1311).</p><p>But the puzzle remains open, how to correctly define the synteny (presence of two or more genes on the same chromosome) and conserved synteny (presence of two or more genes on chromosome of each of the two species) on several genomes.</p><p><img src="http://bioinformaticsonline.com/mod/photo/syntenyImg.jpg" alt="image" width="720" height="179" style="border: 0px; border: 0px;"></p><p>Figure: Image generated with Evolution Highway (EH) tool http://eh-demo.ncsa.illinois.edu/&nbsp;</p><p>Keeping the new approach to define conserved synteny in mind there have been various algorithms developed to identify the conserved homologous synteny blocks (HSB) amongst species. Some of them which were commonly used for synteny detections are:</p><p>SyntenyTracker ( http://www-app.igb.uiuc.edu/labs/lewin/donthu/Synteny_assign/html/),</p><p>SyntenyTracker was shown to be an efficient and accurate automated tool for defining HSBs using datasets that may contain minor errors resulting from limitations in map construction methodologies.</p><p>CoGe (http://genomevolution.org/CoGe/SynFind.pl )</p><p>Satsuma (http://evomics.org/learning/genomics/satsuma/)</p><p>Cinteny (http://cinteny.cchmc.org/) ,</p><p>Cinteny server can be used for finding regions syntenic across multiple genomes and measuring the extent of genome rearrangement using reversal distance as a measure.</p><p>OrthoCluster (http://krono.act.uji.es/noticias/orthocluster-a-new-tool-for-mining-syntenic-blocks)</p><p>A new tool for mining syntenic blocks in comparative genomics</p><p>SynMap (http://genomevolution.org/wiki/index.php/SynMap),</p><p>SyMAP (http://www.symapdb.org/)</p><p>SyMAP (Synteny Mapping and Analysis Program) v4.0 is an automated system for identifying and displaying genome synteny alignments. The genomes may be represented by sequenced chromosomes (pseudomolecules), by draft sequence contigs, or by FPC physical maps (with BAC-end or marker sequence).</p><p>http://genomevolution.org/CoGe/SynMap.pl</p><p>RegionMiner (http://www.genomatix.de/online_help/help_regionminer/orthologous.html)</p><p>SyntenyMiner is being developed as an application to visualize and interrogate comparisons among multiple complete genome sequences. http://syntenyminer.sourceforge.net/</p><p>AutoGRAPH ( http://autograph.genouest.org/),</p><p>AutoGRAPH is an integrated web server for multi-species comparative genomic analysis. It is designed for constructing and visualizing synteny maps between two or three species, determination and display of macrosynteny and microsynteny relationships among species, and for highlighting evolutionary breakpoints.</p><p>SynChro(http://www.lgm.upmc.fr/CHROnicle/SynChro.html)</p><p>SynChro is a tool designed to define conserved synteny blocks. It reconstructs synteny blocks between pairwise comparison of multiple genomes. The reconstructed synteny blocks may overlap each other, be included in one another or duplicated due to micro-rearrangements.</p><p>SyntenyView ( http://www.cbs.dtu.dk/dtucourse/cookbooks/nikob/exercises/gf1_output_5.html),</p><p>Ensembl 'SyntenyView' shows conservation of large-scale gene order between species pairs. A brief summary of the calculation method appears at the bottom of this help page.&nbsp; The left of a 'SyntenyView' page displays a diagram of chromosomes with blocks of conserved synteny. The right of a page shows homology matches between individual genes within syntenic blocks.</p><p>SynBrowse ( http://www.synbrowse.org/),</p><p>SynBrowse (Synteny Browser) is a generic sequence comparison tool for visualizing genome alignments both within and between species. It is intended to help scientists study and analyze synteny, homologous genes and other conserved elements between sequences. This software is useful in studying genome duplication and evolution. It can also aid in identifying uncharacterized genes, putative regulatory elements and novel structural features of study species by comparing to a well annotated reference sequence, thus enabling genome curators to refine and edit annotations of species that have incomplete genome annotations.</p><p>Sibelia (http://arxiv.org/abs/1307.7941).</p><p>A comparative genomic tool: It assists biologists in analysing the genomic variations that correlate with pathogens, or the genomic changes that help microorganisms adapt in different environments. Sibelia will also be helpful for the evolutionary and genome rearrangement studies for multiple strains of microorganisms.</p><p>GSV (http://cas-bioinfo.cas.unt.edu/gsv/homepage.php)</p><p>Genome Synteny Viewer allows users to upload files which contain synteny regions between two or more genomes and interactively visualize the synteny between them. GSV also allows users to upload annotation files to visualize annotated regions in addition to synteny regions.</p><p>MicroSyn (http://www.lgm.upmc.fr/CHROnicle/SynChro.html)</p><p>MicroSyn software as a means of detecting microsynteny in adjacent genomic regions surrounding genes in gene families. MicroSyn searches for conserved, flanking colinear homologous gene pairs between two genomic fragments to determine the relationship between two members in a gene family.</p><p>SynOrth (http://synorth.genereg.net/)</p><p>Synorth [s n &ocirc;rth], named in combination of "synteny" and "ortholog", is designed for the study of evolutionary changes of genomic regulatory blocks (GRBs) in vertebrate genomes, and especially the changes following the whole-genome duplication in teleost fish, by tracing the ortholog genes gain and loss in ancient synteny blocks.</p><p>SyDiG (http://www.ncbi.nlm.nih.gov/pubmed/21441096)</p><p>Uncovering Synteny in Distant Genomes.</p><p>MapSynteny&nbsp; (http://www.automatizacionysistemas.com/download.html)</p><p>MapSynteny is a macro in MS Excel&reg; able to create images to show the relationship between genetic maps and large sequences (scaffolds, chromosomes, BACs, etc.). Based on tab &ndash; delimited BLAST results and some formulas, a suitable image of syntenic relationships or physical mapping can be obtained. http://www.automatizacionysistemas.com/Poster_MapSynteny.pdf</p><p>One of the best synteny tutorial for beginer @&nbsp;http://www.nature.com/scitable/topicpage/synteny-inferring-ancestral-genomes-44022</p><p>Reference:</p><p><a href="http://www.nature.com/scitable/topicpage/synteny-inferring-ancestral-genomes-44022">http://www.nature.com/scitable/topicpage/synteny-inferring-ancestral-genomes-44022</a></p><p><a href="http://www.nature.com/nature/journal/v491/n7424/full/nature11622.html">http://www.nature.com/nature/journal/v491/n7424/full/nature11622.html</a></p><p><a href="http://en.wikipedia.org/wiki/Synteny">http://en.wikipedia.org/wiki/Synteny</a></p><p><a href="http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2675965/">http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2675965/</a></p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/29992/spines</guid>
	<pubDate>Mon, 28 Nov 2016 05:33:26 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/29992/spines</link>
	<title><![CDATA[Spines]]></title>
	<description><![CDATA[<p><a href="https://www.broadinstitute.org/ftp/distribution/software/spines/"><em>Spines</em></a>&nbsp;is a collection of software tools, developed and used by the Vertebrate Genome Biology Group at the Broad Institute. It provides basic data structures for efficient data manipulation (mostly genomic sequences, alignments, variation etc.), as well as specialized tool sets for various analyses. It also features three sequence alignment packages:&nbsp;<em>Satsuma,</em>&nbsp;a highly parallelized program for high-sensitivity, genome-wide synteny;&nbsp;<em>Papaya,</em>&nbsp;an all-purpose alignment tool for less diverged sequences; and&nbsp;<em>SLAP,</em>&nbsp;a context-sensitive local aligner for diverged sequences with large gaps.</p>
<p>Access&nbsp;<em>Spines</em>&nbsp;<a href="https://www.broadinstitute.org/ftp/distribution/software/spines/">here</a>.</p><p>Address of the bookmark: <a href="https://www.broadinstitute.org/genome-sequencing-and-analysis/spines" rel="nofollow">https://www.broadinstitute.org/genome-sequencing-and-analysis/spines</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/43062/jcvi-utility-libraries</guid>
	<pubDate>Sat, 08 May 2021 22:04:02 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/43062/jcvi-utility-libraries</link>
	<title><![CDATA[JCVI utility libraries]]></title>
	<description><![CDATA[<p><span>Collection of Python libraries to parse bioinformatics files, or perform computation related to assembly, annotation, and comparative genomics.</span></p><p>Address of the bookmark: <a href="https://github.com/tanghaibao/jcvi" rel="nofollow">https://github.com/tanghaibao/jcvi</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/14191/scalpel</guid>
	<pubDate>Wed, 20 Aug 2014 02:07:58 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/14191/scalpel</link>
	<title><![CDATA[Scalpel]]></title>
	<description><![CDATA[<p>A team from Cold Spring Harbor Laboratory has released an algorithm, called Scalpel, for finding insertions and deletions in next generation sequencing data sets. Scalpel, which is open source and <a href="http://scalpel.sourceforge.net/" title="available for download">available for download</a> on SourceForge,&nbsp;<span>outperformed the popular tools GATK HaplotypeCaller and SOAPindel in test runs on both simulated and real whole human exomes.</span></p><p>Like other indel callers, Scalpel works by performing <em>de novo</em>&nbsp;assembly of regions of interest, so that misalignment to the reference genome cannot obscure the presence of an insertion or deletion. Scalpel's innovation is to repeatedly check its assembly before comparing to the reference genome, to account for simple sequence repeats that are a regular source of error in indel calling. When Scalpel assembles an exon, it collects reads that map to that exon (including partial matches), splits them into k-mers, and creates a de Bruijn graph to span the exon; however, if it detects repeats in the map, it iteratively increases the size of the k-mers by one base until the repeats are eliminated. This ensures that the final assembly of the exon is highly accurate while minimizing compute time.</p><p>The Cold Spring Harbor team's validation of Scalpel, <a href="http://www.nature.com/nmeth/journal/vaop/ncurrent/full/nmeth.3069.html" title="published over the weekend in Nature Methods">published over the weekend in <em>Nature Methods</em></a>, compares Scalpel's performance on a live whole exome against HaplotypeCaller and SOAPindel. The donor is an individual with serious neurological disorders, which may be linked to a high incidence of indels. One thousand indels from this individual's exome, called by one or more of the informatics pipelines, were selected for focused resequencing. This resequencing revealed a 77% true positive rate for Scalpel calls, dramatically better than the rates for either of the competing tools; Scalpel performed especially well with indels longer than five base pairs, a traditional weak point for indel callers.</p><p>Finally, the authors demonstrate Scalpel's use on a large set of genetic data from nearly 600 families who donated samples to the Simons Simplex Collection, a project of the Simons Foundation Autism Research Initiative. Scalpel found a very high enrichment for indels in children affected by autism, compared with their unaffected siblings, a pattern that persisted even after excluding common variants.</p>]]></description>
	<dc:creator>Shruti Paniwala</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/27113/picard</guid>
	<pubDate>Fri, 29 Apr 2016 08:21:54 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/27113/picard</link>
	<title><![CDATA[Picard]]></title>
	<description><![CDATA[<p>Picard is a set of command line tools for manipulating high-throughput sequencing (HTS) data and formats such as SAM/BAM/CRAM and VCF. These file formats are defined in the <a href="http://samtools.github.io/hts-specs/">Hts-specs</a> repository. See especially the <a href="http://samtools.github.io/hts-specs/SAMv1.pdf">SAM specification</a> and the <a href="http://samtools.github.io/hts-specs/VCFv4.3.pdf">VCF specification</a>.</p>
<p>Note that the information on this page is targeted at end-users. For developers, the source code, building instructions and implementation/development resources are available on <a href="https://github.com/broadinstitute/picard">GitHub</a>.</p>
<p>The Picard toolkit is open-source under the <a href="https://tldrlegal.com/license/mit-license">MIT license</a> and free for all uses.</p>
<p>Enjoy!</p><p>Address of the bookmark: <a href="http://broadinstitute.github.io/picard/" rel="nofollow">http://broadinstitute.github.io/picard/</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/27331/andi</guid>
	<pubDate>Fri, 13 May 2016 05:16:35 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/27331/andi</link>
	<title><![CDATA[Andi]]></title>
	<description><![CDATA[<p>This is the <code>andi</code> program for estimating the evolutionary distance between closely related genomes. These distances can be used to rapidly infer phylogenies for big sets of genomes. Because <code>andi</code> does not compute full alignments, it is so efficient that it scales even up to thousands of bacterial genomes.</p>
<p>This readme covers all necessary instructions for the impatient to get <code>andi</code> up and running. For extensive instructions please consult the <a href="https://github.com/EvolBioInf/andi/blob/master/andi-manual.pdf">manual</a>.</p>
<p>More at https://github.com/evolbioinf/andi/</p><p>Address of the bookmark: <a href="http://bioinformatics.oxfordjournals.org/content/early/2015/01/13/bioinformatics.btu815.full" rel="nofollow">http://bioinformatics.oxfordjournals.org/content/early/2015/01/13/bioinformatics.btu815.full</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/27799/bbmapbbtools-package-multipurpose-tool-designed-for-converting-reads-or-other-nucleotide-data-between-different-formats</guid>
	<pubDate>Mon, 13 Jun 2016 05:47:21 -0500</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/27799/bbmapbbtools-package-multipurpose-tool-designed-for-converting-reads-or-other-nucleotide-data-between-different-formats</link>
	<title><![CDATA[BBMap/BBTools package: Multipurpose tool designed for converting reads or other nucleotide data between different formats.]]></title>
	<description><![CDATA[<div id="post_message_148585"><a href="https://sourceforge.net/projects/bbmap/" target="_blank">Reformat</a>is a member of the <a href="https://sourceforge.net/projects/bbmap/" target="_blank">BBMap/BBTools package</a>. It is a multipurpose tool designed for converting reads or other nucleotide data between different formats. It supports, and can inter-convert:<br /> <br /> fastq<br /> fasta<br /> fasta+qual<br /> sam<br /> scarf (an old Illumina format)<br /> bam (if samtools is installed)<br /> gzip<br /> zip<br /> ascii-33 (sanger)<br /> ascii-64 (old Illumina)<br /> paired files<br /> interleaved files<br /> <br /> It is multithreaded and can process data at over 500 megabytes per second, and can accept streams from standard in and write to standard out, allowing it to be easily dropped into the middle of a pipeline for format conversion. Reformat autodetects formats based on file extensions and content, making it very easy to use; and the autodetection can be overridden, allowing flexibility for people who don't like to follow naming conventions, or out-of-spec fastq files with qualities values like -17 or 120.<br /> <br /> The program has been gradually expanded, and can now perform various other functions. None of these will break pairing, if the input is paired.<br /> <br /> Quality trimming (either or both ends)<br /> Quality filtering<br /> Fixed-length trimming<br /> Generation of histograms (base composition, quality, etc)<br /> Subsampling (to a fraction of input reads, or an exact number of reads or bases)<br /> Changing fasta line-wrapping length<br /> Reverse-complementing (all reads or only read 2)<br /> Adding /1 and /2 suffix to read names<br /> GC-content filtering<br /> Length-filtering<br /> Testing for corrupted interleaved files<br /> <br /> Reformat is compatible with any platform that supports Java 1.7 or higher. It also has a bash shellscript for simpler invocation. Typical usage examples:<br /> <br /> Reformat fastq into fasta:<br /> <strong>reformat.sh in=x.fq out=y.fa</strong><br /> <br /> Interleave paired reads:<br /> <strong>reformat.sh in1=x1.fq in2=x2.fq out=y.fq</strong><br /> <br /> Note - you can actually use a shortcut if paired read files have the same name with a 1 and a 2. This is equivalent to the above command:<br /> <strong>reformat.sh in=x#.fq out=y.fq</strong><br /> <br /> De-interleave reads:<br /> <strong>reformat.sh in=x.fq out1=y1.fq out2=y2.fq</strong><br /> <br /> Verify that interleaving appears correct, assuming Illumina namimg conventions:<br /> <strong>reformat.sh in=x.fq vint</strong><br /> <br /> Convert ASCII-33 to ASCII-64:<br /> <strong>reformat.sh in=x.fq out=y.fq qin=33 qout=64</strong><br /> <br /> Quality-trim paired reads to Q10 on the left and right ends and discard reads shorter than 50bp after trimming:<br /> <strong>reformat.sh in1=x1.fq in2=x2.fq out1=y1.fq out2=y2.fq outsingle=singletons.fq qtrim=rl trimq=10 minlength=50</strong><br /> <br /> Subsample 10% of the first 20000 pairs in an interleaved file:<br /> <strong>reformat.sh in=x.fq out=y.fq reads=20000 samplerate=0.1 int=t</strong><br /> (in this case "int=t" overrides interleaving autodetection, to ensure reads are treated as pairs)<br /> <br /> Pipe in a gzipped sam file and pipe out fasta:<br /> <strong>reformat.sh in=stdin.sam.gz out=stdout.fa</strong><br /> <br /> Reverse-complement reads:<br /> <strong>reformat.sh in=x.fq out=y.fq rcomp</strong><br /> <br /> For reformatting a file with very long sequences, Reformat will need more memory; just add the additional flag "-Xmx2g". For example, to change the line-wrapping length on the human genome (which has individual sequences over 200Mbp long) to 70 characters:<br /> <strong>reformat.sh -Xmx2g in=HG19.fa.gz out=HG19_wrapped.fa.gz fastawrap=70</strong><br /> <br /> For additional functions, please run the shellscript with no arguments, or just read it with a text editor. If you have any questions, please post them in this thread.<br /> <br /> For people using a non-bash terminal, you may need to type "bash reformat.sh" instead of just "reformat.sh".<br /> For users of Windows or other platforms that do not support bash shellscripts, replace "reformat.sh" with "java -ea -Xmx200m /path/to/bbmap/current/ jgi.ReformatReads"<br /> for example,<br /> <strong>java -ea -Xmx200m C:\bbmap\current\ jgi.ReformatReads in=x.fq out=y.fa</strong><br /> <br /> Reformat can be downloaded with BBTools here:<br /> <a href="https://sourceforge.net/projects/bbmap/" target="_blank">https://sourceforge.net/projects/bbmap/</a></div>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

</channel>
</rss>