<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/31012?offset=300</link>
	<atom:link href="https://bioinformaticsonline.com/related/31012?offset=300" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/40208/ragoo-fast-reference-guided-scaffolding-of-genome-assembly-contigs</guid>
	<pubDate>Sun, 27 Oct 2019 00:57:23 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/40208/ragoo-fast-reference-guided-scaffolding-of-genome-assembly-contigs</link>
	<title><![CDATA[RaGOO: Fast Reference-Guided Scaffolding of Genome Assembly Contigs]]></title>
	<description><![CDATA[<p>Alonge M, Soyk S, Ramakrishnan S, Wang X, Goodwin S, Sedlazeck FJ, Lippman ZB, Schatz MC:&nbsp;<a href="https://www.biorxiv.org/content/early/2019/01/13/519637">Fast and accurate reference-guided scaffolding of draft genomes</a>.&nbsp;<em>bioRxiv</em>&nbsp;2019.</p>
<p>RaGOO is a tool for coalescing genome assembly contigs into pseudochromosomes via minimap2 alignments to a closely related reference genome. The focus of this tool is on practicality and therefore has the following features:</p>
<ol>
<li>Good performance. On a MacBook Pro using Arabidopsis data, pseudochromosome construction takes less than a minute and the whole pipeline with SV calling takes ~2 minutes.</li>
<li>Intact ordering and orienting of contigs.</li>
<li><a href="https://github.com/malonge/RaGOO/wiki/Misassembly-Correction">Misassembly correction</a></li>
<li><a href="https://github.com/malonge/RaGOO/wiki/GFF-File-Lift-Over">GFF lift-over</a></li>
<li><a href="https://github.com/malonge/RaGOO/wiki/Calling-Structural-Variants">Structural variant calling with and integrated version of Assemblytics</a></li>
<li>Confidence scores associated with the grouping, localization, and orientation for each contig.</li>
</ol><p>Address of the bookmark: <a href="https://github.com/malonge/RaGOO" rel="nofollow">https://github.com/malonge/RaGOO</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45358/the-variant-everyone-ignored</guid>
	<pubDate>Mon, 05 Oct 2026 12:14:20 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45358/the-variant-everyone-ignored</link>
	<title><![CDATA[The Variant Everyone Ignored]]></title>
	<description><![CDATA[<p>Consider a scenario in which a patient's genome has been sequenced. Among billions of DNA bases, a structural alteration may explain the patient's disease. Multiple advanced algorithms analyze the data, yet only one detects the variant, while the others do not. In standard bioinformatics workflows, such a solitary result is often regarded as unreliable and subsequently discarded. Although the solution exists within the data, prevailing computational protocols may overlook it.</p><p>A recent study published in Genome Biology (https://link.springer.com/article/10.1186/s13059-026-04280-y) addressed this challenge by introducing dicast (https://github.com/burgshrimps/dicast), a machine-learning approach for detecting structural variants in short-read sequencing data. Structural variants, such as large deletions, insertions, duplications, and inversions, can have significant biological and clinical implications, yet they are challenging to identify with short-read technologies. Because different detection methods frequently yield divergent results, researchers commonly employ consensus calling, considering a variant valid only if multiple tools detect it. While this approach reduces false positives, it relies on the potentially flawed assumption that the majority is always correct.</p><p>The researchers explored the impact of evaluating the supporting evidence for each variant, rather than simply tallying the number of algorithms that identified it. To establish a ground truth, they analyzed nine genomes using multiple sequencing technologies and 15 detection methods, initially identifying approximately 35 million potential variants. Through extensive filtering, evidence integration, and manual review of over 11,500 variants, they developed a robust benchmark comprising more than 236,000 structural variants. The findings underscored the complexity of the problem: short-read methods detected fewer than half of deletions and less than 10 percent of insertions, with performance declining markedly in repetitive genomic regions. In contrast, long-read technologies demonstrated superior detection capabilities. However, replacing the substantial volume of existing short-read data in clinical and research settings is not immediately feasible. Consequently, the researchers questioned whether short-read data might harbor more information than conventional analytical pipelines currently extract.</p><p>This line of inquiry led to the development of dicast. Rather than merely confirming agreement among multiple tools, dicast identifies patterns in sequencing data, including split and clipped reads, discordant read pairs, alignment characteristics, and the surrounding genomic context. An XGBoost machine-learning model evaluates which combinations of these signals are indicative of genuine structural variants. Thus, the approach shifts from tallying algorithmic consensus to interpreting the underlying evidence.</p><p>The researchers subsequently conducted a targeted evaluation by examining structural variants detected by only a single short-read tool, which are typically missed by consensus-based approaches. dicast successfully recovered approximately 81% of these single-caller deletions, insertions, and duplications. The signals for these variants were present in the data, but conventional filtering methods failed to integrate them effectively.</p><p>The utility of dicast was further demonstrated in cohorts with rare diseases, including congenital limb malformations, atrial fibrillation, and neuromuscular disorders. In one instance, dicast achieved a deletion recall rate of approximately 0.96, compared to 0.74 using consensus calling. The median number of variants requiring manual review was 29 per sample. Among 31 experimentally validated variants that standard filters would have missed, dicast identified 12, whereas consensus calling detected only one.</p><p>Overall, dicast identified approximately 20 percent more potential disease-causing deletions than consensus-based methods. While a 20 percent increase may appear modest, in clinical genomics such improvements can have significant practical implications. Missing a deletion may leave a case unresolved, whereas detecting a structural variant can provide critical diagnostic insights.</p><p>The study does not claim that machine learning has rendered short-read sequencing superior to long-read approaches. Instead, the results underscore the effectiveness of long-read sequencing for structural variant detection. However, dicast highlights a more nuanced perspective: substantial biological information may still be recoverable from the extensive short-read datasets already available.</p><p>The principal lesson extends beyond the detection of structural variants. For many years, bioinformatics pipelines have relied on threshold-based criteria, such as minimum coverage, quality scores, or support from multiple tools. While these rules are useful, biological phenomena do not always conform to rigid checklists; multiple weak signals, when considered collectively, can provide compelling evidence.</p><p>This perspective prompts consideration of the solitary variant: one algorithm identifies it, while several others do not. Traditional consensus techniques might have dismissed it, yet machine learning approaches evaluate the available evidence to determine whether the variant is plausible.</p><p>Occasionally, the most significant variant within a genome is the one that is almost universally overlooked.</p><p>Read more at&nbsp;https://link.springer.com/article/10.1186/s13059-026-04280-y</p>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/24297/bioinformatics-walkin-at-nii</guid>
  <pubDate>Fri, 04 Sep 2015 21:48:15 -0500</pubDate>
  <link></link>
  <title><![CDATA[Bioinformatics WalkIn at NII]]></title>
  <description><![CDATA[
<p>ADVERTISEMENT OF WALK-IN-INTERVIEW</p>

<p>NAME OF THE POST : Bioinformatician (Part time 3 days in a week) (One Position only)</p>

<p>DURATION : One Year</p>

<p>NAME OF THE PROJECT : Next generation sequencing facility</p>

<p>EDUCATIONAL QUALIFICATIONS : At least a Masters degree in Bioinformatics and Bachelors degree in any stream of life sciences</p>

<p>REQUIREMENTS :</p>

<p>Around 5 years of experience and proven track record in next generation sequence data analysis (supported by publications in peer-reviewed journals), ability to analyze transcriptomics, Chip-seq, and small RNA –seq data.</p>

<p>: Should have the ability to analyze raw primary data generated by Illumina next generation sequencing platforms and create / troubleshoot custom analysis Pipelines.</p>

<p>Should have ability to handle all downstream secondary and tertiary data analysis using commercially available as well as open source softwares (transcriptomics, ChIP-seq, small RNA-seq)</p>

<p>Apart from these, the applicant should have knowledge of the following: Programming: Perl and Python. Operating system:</p>

<p>Linux and Windows. NGS Analysis tools: Maq, BWA, Bowtie, SAM tools, BEDTools, MACS, Galaxy, FastQC, Bismark, MEDIPS, Tophat, Cufflinks, AvadisNGS, CLC Genomics Workbench, Galaxy, BaseSpace, Trinity Statistics: Microsoft Excel and R. Database: MySQL Genome Browser: UCSC, Ensemble, IGV, IGB Motif Analysis Tools: MEME Suite, Transfac and RSAT Functional Annotation Tools: DAVID, GeneCodis, Gene Cards Networking Tools: Cytoscape</p>

<p>EMOLUMENTS : The incumbent will be paid a fee of Rs. 2000/- per sitting/ per day.</p>

<p>SCIENTIST NAME : Dr. Arnab Mukhopadhyay,</p>

<p>Staff Scientific V Next generation sequencing facility</p>

<p>SCIENTIST’S E-MAIL ID : arnab@nii.ac.in</p>

<p>WALK IN INTERVIEW ON : 18th September, 2015</p>

<p>REGISTRATION OF CANDIDATES: 10.30 AM to 11.00 AM</p>

<p>PLEASE NOTE- 1. CANDIDATE MAY FILL UP APPLICATION IN THE PRECRIBED FORMAT ALONG WITH NECESSARY DOCUMENTS FOR VERIFICATION. 2. APPLICATIONS CONTAINING INCOMPLETE INFORMATION SHALL NOT BE ENTERTAINED. 3. DATE OF PASSING THE EXAMINATIONS MUST BE INDICATED CLEARLY. 4. ONLY REGISTERED CANDIDATES WILL BE INTERVIEWED. 5. NO TA/DA WILL BE PAID FOR ATTENDING THE INTERVIEW PRESCRIBED FORM 1. NAME 2. FATHER’S NAME 3. MOTHER’S NAME 4. DATE OF BIRTH 5. SEX (MALE/FEMALE) 6. CATEGORY (SC/ ST/ OBC/ PH) 7. ADDRESS a. (CORRSPONDENCE) b. (PERMANENT) 8. E MAIL, TELEPHONE NO. &amp; MOBILE No (if any) 9. ACADEMIC &amp; PROFESSIONAL QUALIFICATIONS NAME OF EXAMINATION PASSED WITH SUBJECTS YEAR OF PASSING BOARD/ UNIVERSITY PERCENTAGE/ DIVISION REMARKS 10. PAST EXPERIENCE &amp; PRESENT EMPLOYMENT, IF ANY 11. CANDIDATES SHOULD STATE CLEARLY WHETHER THEY HAVE BEEN AWARDED PH.D DEGREE OR THESIS HAS BEEN SUBMITTED. 12. HAVE YOU APPLIED FOR A POSITION EARLIER IN THE INSTITUTE? IF SO:- (1) THE DETAILS OF THE PROJECT AND PROJECT INVESTIGATOR (2) IF CALLED FOR INVERVIEW, RESULTS THEREOF</p>

<p>More at http://www1.nii.res.in/sites/default/files/walkininterview-18sept2015.pdf</p>
]]></description>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/24762/postdoctoral-fellowship-in-bioinformatics-at-pesolelab</guid>
  <pubDate>Thu, 01 Oct 2015 07:20:48 -0500</pubDate>
  <link></link>
  <title><![CDATA[Postdoctoral Fellowship in Bioinformatics at pesolelab]]></title>
  <description><![CDATA[
<p>Job Description: Bioinformatics postdoc positions are available in the area of genomics with main focus on exome and RNAseq technologies by ultra high-throughput sequencing platforms. Successful applicants should have the following qualities:</p>

<p>1) demonstrated experience in Bioinformatics research,<br />2) programing experience (python and/or R, C and C++ are very welcome),<br />3) knowledge of Linux/Unix environment,<br />4) experience in handling deep-seq data,<br />5) highly motivated and hard working, and<br />6) interested to work with a multi-disciplinary team combining bioinformatics, genomics, computational biology approaches with experimental biology.</p>

<p>Our research interest covers different areas of bioinformatics and genomics in order to achieve a deeper understanding of gene and genome structure and function (please look at our PubMed publications for more details about our research http://www.ncbi.nlm.nih.gov/pubmed/?term=pesole+g).</p>

<p>Interested applicants should email the curriculum vitae to Prof. Graziano Pesole at graziano.pesole@uniba.it or Dr. Ernesto Picardi at Ernesto.picardi@uniba.it.</p>

<p>Start date: immediate</p>

<p>Duration: up to 24 months<br />Contact Person (Referent): Ernesto Picardi<br />Ref. E-Mail: ernesto.picardi@uniba.it<br />Tel: +390805443308<br />Fax: +390805443317</p>

<p>Group Web Page: http://www.pesolelab.it/</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/26325/crossmap</guid>
	<pubDate>Mon, 08 Feb 2016 15:47:00 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/26325/crossmap</link>
	<title><![CDATA[CrossMap]]></title>
	<description><![CDATA[<p>CrossMap is a program for convenient conversion of genome coordinates (or annotation files) between <em>different assemblies</em> (such as Human <a href="http://www.ncbi.nlm.nih.gov/assembly/2928/">hg18 (NCBI36)</a> &lt;&gt; <a href="http://www.ncbi.nlm.nih.gov/assembly/2758/">hg19 (GRCh37)</a>, Mouse <a href="http://www.ncbi.nlm.nih.gov/assembly/165668/">mm9 (MGSCv37)</a> &lt;&gt; <a href="http://www.ncbi.nlm.nih.gov/assembly/327618/">mm10 (GRCm38)</a>).</p>
<p>It supports most commonly used file formats including SAM/BAM, Wiggle/BigWig, BED, GFF/GTF, VCF.</p>
<p>CrossMap is designed to liftover genome coordinates between assemblies. It&rsquo;s <em>not</em> a program for aligning sequences to reference genome.</p>
<p>We <em>do not</em> recommend using CrossMap to convert genome coordinates between species.</p>
<p>More at http://crossmap.sourceforge.net/</p><p>Address of the bookmark: <a href="http://crossmap.sourceforge.net/" rel="nofollow">http://crossmap.sourceforge.net/</a></p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/26409/ucsc-genome-browser-and-blat-software</guid>
	<pubDate>Thu, 18 Feb 2016 03:18:57 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/26409/ucsc-genome-browser-and-blat-software</link>
	<title><![CDATA[UCSC Genome Browser and Blat software !]]></title>
	<description><![CDATA[<p>This directory contains Genome Browser and Blat application binaries built for standalone <br>command-line use on various supported Linux and UNIX platforms. To determine which set of binaries <br>to download, type "uname -a" on the command line to display your machine type. In most cases the <br>usage statement for the application can be viewed by running the binary with no arguments. <br><br>The UCSC Genome Browser and Blat software are free for academic, nonprofit, and personal use. A <br>license is required for commercial download and installation of these binaries, with the exception <br>of items built from the following source code directories, which are freely available for all uses:<br><br>&nbsp;- kent/src/utils (includes big* tools)<br>&nbsp;- kent/src/lib<br>&nbsp;- kent/src/hg/autoSql<br>&nbsp;- kent/src/hg/autoXml<br><br>For information about commercial licensing of the Genome Browser software, see <br>http://genome.ucsc.edu/license/. The Blat and In-Silico PCR software may be commercially<br>licensed through Kent Informatics (http://www.kentinformatics.com).</p>
<p>More at http://hgdownload.cse.ucsc.edu/admin/exe/</p><p>Address of the bookmark: <a href="http://hgdownload.cse.ucsc.edu/admin/exe/" rel="nofollow">http://hgdownload.cse.ucsc.edu/admin/exe/</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/researchlabs/view/26456/the-mills-lab</guid>
  <pubDate>Wed, 24 Feb 2016 16:18:38 -0600</pubDate>
  <link></link>
  <title><![CDATA[The Mills lab]]></title>
  <description><![CDATA[
<p>The laboratory is focused on the discovery and analysis of structural variation (SVs) from genomic sequence data. As part of the 1000 Genomes Project and other endeavors, we have helped produce initial fine-scale maps using a variety of SV discovery approaches including: (i) paired-end mapping (or read pair analysis) based on abnormally mapped pairs of clone ends; (ii) read-depth analysis, which detects deletions and duplications through analysis of the read depth-of-coverage; (iii) split read analysis, which detects SVs by evaluating gapped sequence alignments; and (iv) sequence assembly, which enables the discovery of novel (non-reference) sequence insertions.</p>

<p>http://millslab.org/research.html</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/26909/sequence-assembly-with-mira-4</guid>
	<pubDate>Wed, 06 Apr 2016 08:21:22 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/26909/sequence-assembly-with-mira-4</link>
	<title><![CDATA[Sequence assembly with MIRA 4]]></title>
	<description><![CDATA[<p>MIRA is a multi-pass DNA sequence data assembler/mapper for whole genome and EST/RNASeq projects. MIRA assembles/maps reads gained by</p>
<div>
<ul>
<li>
<p>electrophoresis sequencing (aka Sanger sequencing)</p>
</li>
<li>
<p>454 pyro-sequencing (GS20, FLX or Titanium)</p>
</li>
<li>
<p>Ion Torrent</p>
</li>
<li>
<p>Solexa (Illumina) sequencing</p>
</li>
<li>
<p>(in development) Pacific Biosciences sequencing</p>
</li>
</ul>
</div>
<p>into contiguous sequences (called <span><em>contigs</em></span>). One can use the sequences of different sequencing technologies either in a single assembly run (a <span><em>true hybrid assembly</em></span>) or by mapping one type of data to an assembly of other sequencing type (a <span><em>semi-hybrid assembly (or mapping)</em></span>) or by mapping a data against consensus sequences of other assemblies (a <span><em>simple mapping</em></span>).</p>
<p>The MIRA acronym stands for <span><strong>M</strong></span>imicking <span><strong>I</strong></span>ntelligent <span><strong>R</strong></span>ead <span><strong>A</strong></span>ssembly and the program pretty well does what its acronym says (well, most of the time anyway). It is the Swiss army knife of sequence assembly that I've used and developed during the past 14 years to get assembly jobs I work on done efficiently - and especially accurately. That is, without me actually putting too much manual work into it.</p>
<p>More at http://mira-assembler.sourceforge.net/docs/DefinitiveGuideToMIRA.html</p><p>Address of the bookmark: <a href="http://mira-assembler.sourceforge.net/docs/DefinitiveGuideToMIRA.html" rel="nofollow">http://mira-assembler.sourceforge.net/docs/DefinitiveGuideToMIRA.html</a></p>]]></description>
	<dc:creator>Priya Singh</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/26972/understanding-fastqc-output</guid>
	<pubDate>Fri, 15 Apr 2016 05:47:40 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/26972/understanding-fastqc-output</link>
	<title><![CDATA[Understanding Fastqc Output]]></title>
	<description><![CDATA[<p>Understanding Following table and graphs</p>
<ol>
<li>Duplication level</li>
<li>kmer profile</li>
<li>per base GC content</li>
<li>per base N content</li>
<li>per base quality</li>
<li>per base sequence content</li>
<li>per sequence GC content</li>
<li>per sequence quality</li>
<li>sequence length distribution</li>
</ol>
<p>More at http://www.bioinformatics.babraham.ac.uk/projects/fastqc/Help/3%20Analysis%20Modules/</p><p>Address of the bookmark: <a href="http://www.bioinformatics.babraham.ac.uk/projects/fastqc/Help/3%20Analysis%20Modules/" rel="nofollow">http://www.bioinformatics.babraham.ac.uk/projects/fastqc/Help/3%20Analysis%20Modules/</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/27104/gatb-genome-analysis-toolbox-with-de-bruijn-graph</guid>
	<pubDate>Thu, 28 Apr 2016 11:16:51 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/27104/gatb-genome-analysis-toolbox-with-de-bruijn-graph</link>
	<title><![CDATA[GATB : Genome Analysis Toolbox with de-Bruijn graph]]></title>
	<description><![CDATA[<p>The&nbsp;<strong><strong>Genome Analysis Toolbox with de-Bruijn graph</strong> (GATB)</strong> provides a set of <a href="https://gatb.inria.fr/gatb-global-architecture/">highly efficient algorithms to analyse NGS data sets</a>. These methods enable the analysis of data sets of any size on multi-core desktop computers, including very huge amount of reads data coming from any kind of organisms such as bacteria, plants, animals and even complex samples (<em>e.g.</em> metagenomes).</p>
<p>More at https://gatb.inria.fr/</p><p>Address of the bookmark: <a href="https://gatb.inria.fr/" rel="nofollow">https://gatb.inria.fr/</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

</channel>
</rss>