<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/44387?offset=50</link>
	<atom:link href="https://bioinformaticsonline.com/related/44387?offset=50" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/44852/what-is-data-science-%E2%80%94-a-bioinformatics-perspective</guid>
	<pubDate>Mon, 16 Jun 2025 01:44:34 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/44852/what-is-data-science-%E2%80%94-a-bioinformatics-perspective</link>
	<title><![CDATA[What is Data Science? — A Bioinformatics Perspective]]></title>
	<description><![CDATA[<p>In today&rsquo;s era of big biology, we&rsquo;re generating more data than ever before&mdash;genomes, transcriptomes, proteomes, metabolomes, microbiomes&hellip; you name it. But raw biological data doesn&rsquo;t speak for itself. Making sense of it requires more than traditional biology. This is where data science steps in.</p><p><strong>So, What Is Data Science?</strong><br />At its core, data science is the interdisciplinary field that extracts knowledge and insights from data using programming, statistics, and domain expertise. In bioinformatics, data science enables us to turn gigabytes of sequence data into biological meaning.</p><p>Imagine trying to understand gene regulation in cancer by analyzing thousands of RNA-seq samples, or predicting antibiotic resistance from bacterial genomes&mdash;these challenges are not solvable through wet lab experiments alone. They require data-driven thinking.</p><p><strong>Data Science Meets Bioinformatics</strong><br />Bioinformatics is inherently a data science domain. From genomics to systems biology, every field in modern biology relies on data science techniques to:</p><p>Clean and process massive datasets</p><p>Discover patterns in high-dimensional data</p><p>Build predictive models (e.g., for disease classification)</p><p>Visualize complex biological networks and trends</p><p>Integrate diverse data types (e.g., transcriptomic + epigenomic data)</p><p><strong>The Bioinformatics Toolkit</strong><br />Here&rsquo;s what data science typically looks like in bioinformatics:</p><p>Task Data Science Role<br />Sequence alignment Efficient algorithms, indexing, parallel processing<br />Gene expression analysis Statistical modeling (e.g., DESeq2, limma)<br />Variant calling Data filtering, probabilistic models<br />Clustering of cells in single-cell data Unsupervised learning<br />Protein structure prediction Deep learning models (e.g., AlphaFold)<br />Metagenomics Data integration, classification, dimensionality reduction</p><p>Common tools include Python, R, Bioconductor, scikit-learn, Pandas, Seurat, and TensorFlow&mdash;often working together in reproducible workflows.</p><p><strong>It's Not Just About Coding</strong><br />A common misconception is that bioinformatics is just programming or scripting. But being a data scientist in bioinformatics also means:</p><p>Understanding experimental design</p><p>Asking biologically meaningful questions</p><p>Choosing the right statistical or machine learning models</p><p>Communicating findings effectively (e.g., plots, dashboards, papers)</p><p>In other words, data science in bioinformatics is where biology, statistics, and computer science converge.</p><p><strong>Why It Matters</strong><br />The real power of data science in bioinformatics is its ability to scale discovery.</p><p>Instead of studying one gene, we can study thousands.</p><p>Instead of analyzing one species, we can explore entire ecosystems.</p><p>Instead of waiting months for lab results, we can generate hypotheses in days.</p><p>From personalized medicine and cancer diagnostics to agricultural genomics and pandemic surveillance, data science is at the heart of the bioinformatics revolution.</p><p><strong>Final Thoughts</strong><br />If you&rsquo;re a biologist who&rsquo;s curious about code, or a data enthusiast fascinated by life sciences, bioinformatics is your playground&mdash;and data science is your toolkit.</p><p>In bioinformatics, data science isn&rsquo;t just useful. It&rsquo;s essential.</p><p>&nbsp;</p>]]></description>
	<dc:creator>Abhi</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/27348/ngago-challenge-crispr</guid>
	<pubDate>Tue, 17 May 2016 03:31:32 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/27348/ngago-challenge-crispr</link>
	<title><![CDATA[NgAgo challenge CRISPR !!]]></title>
	<description><![CDATA[<p><a href="http://www.nature.com/nbt/journal/vaop/ncurrent/full/nbt.3547.html" target="_blank" title="A recent Nature Biotechnology paper"><strong>A recent Nature Biotechnology paper</strong></a>&nbsp;from Chunyu Han&rsquo;s lab,&nbsp;DNA-guided genome editing using the&nbsp;<em>Natronobacterium gregoryi&nbsp;</em>Argonaute,&nbsp;is a must-read for genome editing folks who want to learn about NgAgo. Their team sums up NgAgo&rsquo;s potential pluses this way (<strong>emphasis</strong>&nbsp;mine):</p><blockquote><p>&ldquo;The useful features of NgAgo for genome editing include the following.<strong>First, it has a low tolerance to guide&ndash;target mismatch</strong>. A single nucleotide mismatch at each position of the gDNA impaired the cleavage efficiency of NgAgo, and mismatches at three positions completely blocked cleavage in our experiments.&nbsp;<strong>Second, 5&prime; phosphorylated short ssDNAs are rare in mammalian cells, which minimizes the possibility of cellular oligonucleotides misguiding NgAgo</strong>.<strong>Third, NgAgo follows a &lsquo;one-guide-faithful&rsquo; rule,</strong>&nbsp;that is, a guide can only be loaded when NgAgo protein is in the process of expression, and, once loaded, NgAgo cannot swap its gDNA with other free ssDNA at 37 &deg;C. All of these features could minimize off-target effects.&nbsp;<strong>Finally, it is easy to design and synthesize ssDNAs and to adjust their concentration</strong>, which is difficult with the Cas9-sgRNA system, if the sgRNA is expressed from a plasmid and the normal dosage of an ssDNA guide is only ~1/10 of that of a sgRNA expression plasmid.</p></blockquote><p>NgAgo might be a more orderly way and perhaps even simpler way to go about genome editing than CRISPR, but the jury is still out on that until there are more papers and data. The NgAgo edit efficiency at this preliminary stage of technology development seems very strong. See the pics below</p><p><img src="http://i1.wp.com/www.ipscell.com/wp-content/uploads/2016/05/NgAgo1.jpg" alt="image" width="1311" height="559" style="border: 0px; border: 0px;"></p><p>&nbsp;</p><p>Reference:&nbsp;http://www.nature.com/nbt/journal/vaop/ncurrent/full/nbt.3547.html</p>]]></description>
	<dc:creator>Abhimanyu Singh</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/39380/mgert-mobile-genetic-elements-retrieving-tool</guid>
	<pubDate>Sat, 18 May 2019 08:58:01 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/39380/mgert-mobile-genetic-elements-retrieving-tool</link>
	<title><![CDATA[MGERT: Mobile Genetic Elements Retrieving Tool]]></title>
	<description><![CDATA[<p><em>MGERT</em><span>&nbsp;is a computational pipeline for easy retrieving of MGE's coding sequences of a particular family from genome assemblies.&nbsp;</span><em>MGERT</em><span>&nbsp;utilizes several established bioinformatic tools combined into single pipeline which hides different technical quirks from an inexperienced user.</span></p><p>Address of the bookmark: <a href="https://github.com/andrewgull/MGERT" rel="nofollow">https://github.com/andrewgull/MGERT</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/42923/flanker</guid>
	<pubDate>Sat, 27 Feb 2021 22:04:53 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/42923/flanker</link>
	<title><![CDATA[Flanker]]></title>
	<description><![CDATA[<p><span>Flanker, a Python package which performs alignment-free clustering of gene flanking sequences in a consistent format, allowing investigation of&nbsp;<span>mobile genetic elements (</span>MGEs) without prior knowledge of their structure.&nbsp;<span>Flanker can be flexibly parameterised to finetune outputs by characterising upstream and downstream regions separately and investigating variable lengths of flanking sequence.</span></span></p>
<p><span><img src="https://github.com/wtmatlock/flanker/raw/main/docs/frontpage.png" alt="image" style="border: 0px;"></span></p><p>Address of the bookmark: <a href="https://github.com/wtmatlock/flanker" rel="nofollow">https://github.com/wtmatlock/flanker</a></p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/38443/genoplotr-plot-gene-and-genome-maps-project</guid>
	<pubDate>Wed, 12 Dec 2018 08:33:41 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/38443/genoplotr-plot-gene-and-genome-maps-project</link>
	<title><![CDATA[genoPlotR - plot gene and genome maps project!]]></title>
	<description><![CDATA[<p>genoPlotR is a R package to produce reproducible, publication-grade graphics of gene and genome maps. It allows the user to read from usual format such as protein table files and blast results, as well as home-made tabular files.</p>
<h3>Features</h3>
<ul>
<li>Linear representation of several segments of DNA</li>
<li>Comparisons represented by areas between the segments (like Artemis, for example)</li>
<li>Reads from common formats: Genbank, EMBL, blast, Mauve, and from user-generated tab files</li>
<li>Plot several subsegments of the same segment on the same line, separated by a //</li>
<li>Automatic or manual placement of the segments on the plot</li>
<li>Add annotations to all the lines</li>
<li>Create smart, automatic annotations for genomes, based on gene names</li>
<li>Add a user-generated tree</li>
<li>Add a global scale or a scale to each line</li>
<li>Use user-defined graphical functions to represent genes</li>
<li></li>
</ul><p>Address of the bookmark: <a href="http://genoplotr.r-forge.r-project.org/" rel="nofollow">http://genoplotr.r-forge.r-project.org/</a></p>]]></description>
	<dc:creator>Abhimanyu Singh</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/41920/liftoff-an-accurate-tool-that-maps-annotations-in-gff-or-gtf-between-assemblies</guid>
	<pubDate>Tue, 30 Jun 2020 21:40:52 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/41920/liftoff-an-accurate-tool-that-maps-annotations-in-gff-or-gtf-between-assemblies</link>
	<title><![CDATA[Liftoff: an accurate tool that maps annotations in GFF or GTF between assemblies]]></title>
	<description><![CDATA[<p><span>&nbsp;Liftoff, an accurate tool that maps annotations in GFF or GTF between assemblies of the same, or closely-related species. Unlike current coordinate lift-over tools which require a pre-generated &ldquo;chain&rdquo; file as input, Liftoff is a standalone tool that takes two genome assemblies and a reference annotation as input and outputs an annotation of the target genome.&nbsp;</span></p><p>Address of the bookmark: <a href="https://github.com/agshumate/Liftoff" rel="nofollow">https://github.com/agshumate/Liftoff</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/43427/ogdraw-draw-organelle-genome-maps</guid>
	<pubDate>Tue, 05 Oct 2021 03:34:35 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/43427/ogdraw-draw-organelle-genome-maps</link>
	<title><![CDATA[OGDRAW - Draw Organelle Genome Maps]]></title>
	<description><![CDATA[<p>OrganellarGenomeDRAW converts annotations in the&nbsp;<a href="https://www.ncbi.nlm.nih.gov/genbank/">GenBank</a>&nbsp;or&nbsp;<a href="https://www.ebi.ac.uk/ena">EMBL/ENA</a>&nbsp;format into graphical maps. The input has to be a&nbsp;<a href="https://www.ncbi.nlm.nih.gov/Sitemap/samplerecord.html">GenBank&nbsp;</a>or&nbsp;<a href="https://www.ebi.ac.uk/ena/submit/flat-file">EMBL/ENA flat file</a>&nbsp;wherase the output can vary among several types of files. The application is optimized to create detailed high-quality maps of organellar genomes (plastid and mitochondria). Nevertheless, you can upload most<span style="color: #0b0118;">&nbsp;database</span>&nbsp;entries.</p>
<p>&nbsp;</p>
<p>Please take a look at our&nbsp;<a href="https://chlorobox.mpimp-golm.mpg.de/OGDraw-FAQ.html">FAQ section</a>&nbsp;and do not hesitate to report bugs or suggestions for improvements by&nbsp;<a href="mailto:chlorobox@mpimp-golm.mpg.de?subject=OGDRAW">email</a>.</p><p>Address of the bookmark: <a href="https://chlorobox.mpimp-golm.mpg.de/OGDraw.html" rel="nofollow">https://chlorobox.mpimp-golm.mpg.de/OGDraw.html</a></p>]]></description>
	<dc:creator>Abhi</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/34396/pore-an-r-package-for-the-visualization-and-analysis-of-nanopore-sequencing-data</guid>
	<pubDate>Thu, 23 Nov 2017 09:55:57 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/34396/pore-an-r-package-for-the-visualization-and-analysis-of-nanopore-sequencing-data</link>
	<title><![CDATA[poRe: an R package for the visualization and analysis of nanopore sequencing data]]></title>
	<description><![CDATA[<p><strong>Motivation:</strong>&nbsp;The Oxford Nanopore MinION device represents a unique sequencing technology. As a mobile sequencing device powered by the USB port of a laptop, the MinION has huge potential applications. To enable these applications, the bioinformatics community will need to design and build a suite of tools specifically for MinION data.</p>
<p><strong>Results:</strong>&nbsp;Here we present poRe, a package for R that enables users to manipulate, organize, summarize and visualize MinION nanopore sequencing data. As a package for R, poRe has been tested on Windows, Linux and MacOSX. Crucially, the Windows version allows users to analyse MinION data on the Windows laptop attached to the device.</p>
<p><strong>Availability and implementation:</strong>&nbsp;poRe is released as a package for R at&nbsp;<a href="http://sourceforge.net/projects/rpore/" target="">http://sourceforge.net/projects/rpore/</a>&nbsp;. A tutorial and further information are available at&nbsp;<a href="https://sourceforge.net/p/rpore/wiki/Home/" target="">https://sourceforge.net/p/rpore/wiki/Home/</a></p>
<p><strong>Contact:</strong><a href="mailto:mick.watson@roslin.ed.ac.uk" target="">mick.watson@roslin.ed.ac.uk</a></p><p>Address of the bookmark: <a href="https://academic.oup.com/bioinformatics/article/31/1/114/2365693" rel="nofollow">https://academic.oup.com/bioinformatics/article/31/1/114/2365693</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36833/bfc-a-standalone-high-performance-tool-for-correcting-sequencing-errors-from-illumina-sequencing-data</guid>
	<pubDate>Thu, 31 May 2018 09:35:23 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36833/bfc-a-standalone-high-performance-tool-for-correcting-sequencing-errors-from-illumina-sequencing-data</link>
	<title><![CDATA[BFC: a standalone high-performance tool for correcting sequencing errors from Illumina sequencing data]]></title>
	<description><![CDATA[BFC is a standalone high-performance tool for correcting sequencing errors from Illumina sequencing data. It is specifically designed for high-coverage whole-genome human data, though also performs well for small genomes.

The BFC algorithm is a variant of the classical spectrum alignment algorithm introduced by Pevzner et al (2001). It uses an exhaustive search to find a k-mer path through a read that minimizes a heuristic objective function jointly considering penalties on correction, quality and k-mer support. This algorithm was first implemented in my fermi assembler and then refined a few times in fermi, fermi2 and now in BFC. In the k-mer counting phase, BFC uses a blocked bloom filter to filter out most singleton k-mers and keeps the rest in a hash table (Melsted and Pritchard, 2011). The use of bloom filter is how BFC is named, though other correctors such as Lighter and Bless actually rely more on bloom filter than BFC.

https://github.com/lh3/bfc<p>Address of the bookmark: <a href="https://github.com/lh3/bfc" rel="nofollow">https://github.com/lh3/bfc</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37527/nanopack-visualizing-and-processing-long-read-sequencing-data</guid>
	<pubDate>Fri, 10 Aug 2018 18:41:34 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37527/nanopack-visualizing-and-processing-long-read-sequencing-data</link>
	<title><![CDATA[NanoPack: visualizing and processing long-read sequencing data]]></title>
	<description><![CDATA[<p>The NanoPack tools are written in Python3 and released under the GNU GPL3.0 License. The source code can be found at&nbsp;<a href="https://github.com/wdecoster/nanopack" target="">https://github.com/wdecoster/nanopack</a>, together with links to separate scripts and their documentation. The scripts are compatible with Linux, Mac OS and the MS Windows 10 subsystem for Linux and are available as a graphical user interface, a web service at&nbsp;<a href="http://nanoplot.bioinf.be/" target="">http://nanoplot.bioinf.be</a>&nbsp;and command line tools.</p>
<p>&nbsp;https://academic.oup.com/bioinformatics/article/34/15/2666/4934939</p><p>Address of the bookmark: <a href="https://github.com/wdecoster/nanoQC" rel="nofollow">https://github.com/wdecoster/nanoQC</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

</channel>
</rss>