BOL: Related items

LoVis4u: Locus Visualisation tool for comparative genomics

LEGE — Tue, 17 Sep 2024 02:30:57 -0500

Description

LoVis4u is a bioinformatics tool for Loci Visualisation.

LoVis4u, a command-line tool and Python API designed for highly customizable and fast visualisation of multiple genomic loci. LoVis4u generates vector images in PDF format based on annotation data from GenBank or GFF files. It is capable of visualising entire genomes of bacteriophages as well as plasmids and user-defined regions of longer prokaryotic genomes. Additionally, LoVis4u offers optional data processing steps to identify and highlight accessory and core genes in input sequences.

https://art-egorov.github.io/lovis4u/

Address of the bookmark: https://github.com/art-egorov/lovis4u

dbCAN: a web server and DataBase for automated Carbohydrate-active enzyme ANnotation

Jit — Mon, 29 May 2017 05:39:29 -0500

dbCAN is a web server and DataBase for automated Carbohydrate-active enzyme ANnotation, funded by the BioEnergy Science Center of the DOE. Similar resources on the web include CAZy database and CAT. All data in dbCAN are generated based on the family classification from CAZy database while it has the following unique features compared with CAZy database and CAT:

dbCAN provides the capability of automated and comprehensive CAZyme annotation of a given genome submitted by the user;
dbCAN provides an explicitly defined signature domain for each and every CAZyme family along with its location in all the relevant full-length CAZyme proteins in all sequenced genomes;
dbCAN provides the most complete set of metagenomic CAZyme genes published so far and represents the first step towards discovering novel CAZyme catalysts in metagenomes;
dbCAN provides a subfamily classification of the existing CAZyme families based on sequence similarities;
dbCAN make all pre-computed data freely available to the public, including sequence alignments, hidden markov models (HMMs) and phylogenies of the signature domain regions in each and every CAZyme family and subfamily.

dbCAN is updated regularly when CAZy database created new families based on latest literature.

Address of the bookmark: http://csbl.bmb.uga.edu/dbCAN/index.php

Bokeh: An interactive visualization library that targets modern web browsers for presentation

Jit — Fri, 10 Aug 2018 18:43:08 -0500

Bokeh is an interactive visualization library that targets modern web browsers for presentation. Its goal is to provide elegant, concise construction of versatile graphics, and to extend this capability with high-performance interactivity over very large or streaming datasets. Bokeh can help anyone who would like to quickly and easily create interactive plots, dashboards, and data applications.

To get started using Bokeh to make your visualizations, see the User Guide.

To see examples of how you might use Bokeh with your own data, check out the Gallery.

A complete API reference of Bokeh is at Reference Guide.

If you are interested in contributing to Bokeh, or extending the library, see the Developer Guide.

Address of the bookmark: https://bokeh.pydata.org/en/latest/

plumber:An R package that converts your existing R code to a web API

BioJoker — Wed, 13 Mar 2019 19:20:10 -0500

plumber allows you to create a REST API by merely decorating your existing R source code with special comments. Take a look at an example.

# plumber.R

#* Echo back the input
#* @param msg The message to echo
#* @get /echo
function(msg=""){
  list(msg = paste0("The message is: '", msg, "'"))
}

#* Plot a histogram
#* @png
#* @get /plot
function(){
  rand <- rnorm(100)
  hist(rand)
}

#* Return the sum of two numbers
#* @param a The first number to add
#* @param b The second number to add
#* @post /sum
function(a, b){
  as.numeric(a) + as.numeric(b)
}

Address of the bookmark: https://www.rplumber.io/

CSAR-web: a web server of contig scaffolding using algebraic rearrangements

BioStar — Fri, 10 Apr 2020 04:39:36 -0500

CSAR-web is a web-based tool that allows the users to efficiently and accurately scaffold (i.e. order and orient) the contigs of a target draft genome based on a complete or incomplete reference genome from a related organism.

CSAR-web can serve as a convenient and useful scaffolding tool allowing the users to efficiently and accurately scaffold their draft genomes according to a complete or incomplete reference genome.

Address of the bookmark: http://genome.cs.nthu.edu.tw/CSAR-web

Enrichr: a comprehensive gene set enrichment analysis

Jit — Thu, 27 Apr 2017 05:42:09 -0500

Enrichment analysis is a popular method for analyzing gene sets generated by genome-wide experiments. Here we present a significant update to one of the tools in this domain called Enrichr. Enrichr currently contains a large collection of diverse gene set libraries available for analysis and download. In total, Enrichr currently contains 180 184 annotated gene sets from 102 gene set libraries. New features have been added to Enrichr including the ability to submit fuzzy sets, upload BED files, improved application programming interface and visualization of the results as clustergrams. Overall, Enrichr is a comprehensive resource for curated gene sets and a search engine that accumulates biological knowledge for further biological discoveries. Enrichr is freely available at: http://amp.pharm.mssm.edu/Enrichr.

https://academic.oup.com/nar/article-lookup/doi/10.1093/nar/gkw377

Address of the bookmark: http://amp.pharm.mssm.edu/Enrichr/

List of cancer genomics research web resources !

biogeek — Wed, 27 Dec 2017 20:33:09 -0600

Major web resources for cancer genomics research

CGHub
https://cghub.ucsc.edu/
Comprehensive data repository; huge data size

EGA
https://www.ebi.ac.uk/ega/
Comprehensive data repository; huge data size

COSMIC
http://cancer.sanger.ac.uk
Largest somatic mutation database; genome sequencing paper curation

CPRG
http://www.broadinstitute.org/software/cprg
Interface for cancer program resources

GDAC
http://gdac.broadinstitute.org/
Data analysis; automatic pipelines; user-friendly reports

SNP500Cancer
http://snp500cancer.nci.nih.gov
Sequence and genotype verification of SNPs

canEvolve
www.canevolve.org/
Comprehensive analysis of tumor profile; Data from 90 studies involving more than 10,000 patients

MethyCancer
http://methycancer.psych.ac.cn
Relationship among DNA methylation, gene expression and cancer

SomamiR
http://compbio.uthsc.edu/SomamiR/
Correlation between somatic mutation and microRNA; genome-wide displaying

cBioPortal
http://www.cbioportal.org/public-portal/
Graphical summaries; gene alteration; processed data; visualization

UCSC Cancer Genomics Browser
https://genome-cancer.soe.ucsc.edu/
Clinical information; gene expression; copy number variation; visualization

CGWB
https://cgwb.nci.nih.gov/
Visualization; gene mutation and variation; automated analysis pipeline

GDSC
http://www.cancerrxgene.org
Drug sensitivity information; drug response information

canSAR
https://cansar.icr.ac.uk/
Multidisciplinary information; drug discovery

NONCODE
http://www.noncode.org/ ncRNAs;
lncRNAs; up-to-date and comprehensive resource

Omega2: metagenome assembly pipeline

Jit — Mon, 10 Jul 2017 05:56:07 -0500

Omega found overlaps between reads using a prefix/suffix hash table. The overlap graph of reads was simplified by removing transitive edges and trimming short branches. Unitigs were generated based on minimum cost flow analysis of the overlap graph and then merged to contigs and scaffolds using mate-pair information. In comparison with three de Bruijn graph assemblers (SOAPdenovo, IDBA-UD and MetaVelvet), Omega provided comparable overall performance on a HiSeq 100-bp dataset and superior performance on a MiSeq 300-bp dataset. In comparison with Celera on the MiSeq dataset, Omega provided more continuous assemblies overall using a fraction of the computing time of existing overlap-layout-consensus assemblers. This indicates Omega can more efficiently assemble longer Illumina reads, and at deeper coverage, for metagenomic datasets.

Address of the bookmark: http://omega.omicsbio.org/

miniasm: very fast OLC-based de novo assembler for noisy long reads

Jit — Mon, 27 Nov 2017 07:58:49 -0600

Miniasm is a very fast OLC-based de novo assembler for noisy long reads. It takes all-vs-all read self-mappings (typically by minimap) as input and outputs an assembly graph in the GFA format. Different from mainstream assemblers, miniasm does not have a consensus step. It simply concatenates pieces of read sequences to generate the final unitig sequences. Thus the per-base error rate is similar to the raw input reads.

So far miniasm is in early development stage. It has only been tested on a dozen of PacBio and Oxford Nanopore (ONT) bacterial data sets. Including the mapping step, it takes about 3 minutes to assemble a bacterial genome. Under the default setting, miniasm assembles 9 out of 12 PacBio datasets and 3 out of 4 ONT datasets into a single contig. The 12 PacBio data sets are PacBio E. coli sample, ERS473430, ERS544009, ERS554120, ERS605484, ERS617393, ERS646601, ERS659581, ERS670327, ERS685285, ERS743109 and a deprecated PacBio E. coli data set. ONT data are acquired from the Loman Lab.

For a C. elegans PacBio data set (only 40X are used, not the whole dataset), miniasm finishes the assembly, including reads overlapping, in ~10 minutes with 16 CPUs. The total assembly size is 105Mb; the N50 is 1.94Mb. In comparison, the HGAP3produces a 104Mb assembly with N50 1.61Mb. This dotter plot gives a global view of the miniasm assembly (on the X axis) and the HGAP3 assembly (on Y). They are broadly comparable. Of course, the HGAP3 consensus sequences are much more accurate. In addition, on the whole data set (assembled in ~30 min), the miniasm N50 is reduced to 1.79Mb. Miniasm still needs improvements.

Miniasm confirms that at least for high-coverage bacterial genomes, it is possible to generate long contigs from raw PacBio or ONT reads without error correction. It also shows that minimap can be used as a read overlapper, even though it is probably not as sensitive as the more sophisticated overlapers such as MHAP and DALIGNER. Coupled with long-read error correctors and consensus tools, miniasm may also be useful to produce high-quality assemblies.

Minimap and miniasm are ultrafast tools for (i) mapping and (ii) assembly. Designed for long, noisy reads, they do not have a correction or consensus step, and therefore the resulting assemblies are contiguous (i.e. long) but very noisy (i.e. full of errors)

We start with an all against all comparison:

minimap -Sw5 -L100 -m0 -t8 reads.fq reads.fq | gzip -1 > reads.paf.gz

Then we can assemble

miniasm -f reads.fq reads.paf.gz > reads.gfa

Convert GFA to FASTA:

awk '/^S/{print ">"$2"\n"$3}' reads.gfa | fold > reads.fa

And then count how many contigs:

grep ">" reads.fa | wc -l

# Download sample PacBio from the PBcR website
wget -O- http://www.cbcb.umd.edu/software/PBcR/data/selfSampleData.tar.gz | tar zxf -
ln -s selfSampleData/pacbio_filtered.fastq reads.fq
# Install minimap and miniasm (requiring gcc and zlib)
git clone https://github.com/lh3/minimap && (cd minimap && make)
git clone https://github.com/lh3/miniasm && (cd miniasm && make)
# Overlap
minimap/minimap -Sw5 -L100 -m0 -t8 reads.fq reads.fq | gzip -1 > reads.paf.gz
# Layout
miniasm/miniasm -f reads.fq reads.paf.gz > reads.gfa

Address of the bookmark: https://github.com/lh3/miniasm

MashMap: a fast and approximate software for mapping long reads (PacBio/ONT) or assembly to reference genome(s)

Jit — Tue, 12 Dec 2017 17:23:31 -0600

MashMap is a fast and approximate software for mapping long reads (PacBio/ONT) or assembly to reference genome(s). It maps a query sequence against a reference region if and only if its estimated alignment identity is above a specified threshold. It does not compute the alignments explicitly, but rather estimates a k-mer based Jaccard similarity using a combination of Winnowing and MinHash. This is then converted to an estimate of sequence identity using the Mash distance. An appropriate k-mer sampling rate is automatically determined given minimum local alignment length and identity thresholds. The efficiency of the algorithm improves as both of these thresholds are increased.

Address of the bookmark: https://github.com/marbl/MashMap