BOL: Related items

Purge Haplotigs: Pipeline to help with curating heterozygous diploid genome assemblies

Rahul Nayak — Mon, 17 Dec 2018 03:17:20 -0600

Some parts of a genome may have a very high degree of heterozygosity. This causes contigs for both haplotypes of that part of the genome to be assembled as separate primary contigs, rather than as a contig and an associated haplotig. This can be an issue for downstream analysis whether you're working on the haploid or phased-diploid assembly.

Identify pairs of contigs that are syntenic and move one of them to the haplotig 'pool'. The pipeline uses mapped read coverage and Minimap2 alignments to determine which contigs to keep for the haploid assembly. Dotplots are optionally produced for all flagged contig matches, juxtaposed with read-coverage, to help the user determine the proper assignment of any remaining ambiguous contigs. The pipeline will run on either a haploid assembly (i.e. Canu, FALCON or FALCON-Unzip primary contigs) or on a phased-diploid assembly (i.e. FALCON-Unzip primary contigs + haplotigs). Here are two examples of how Purge Haplotigs can improve a haploid and diploid assembly.

Address of the bookmark: https://bitbucket.org/mroachawri/purge_haplotigs

LTR_Finder: an efficient program for finding full-length LTR retrotranspsons in genome sequences.

Neel — Sun, 13 Jan 2019 07:05:53 -0600

LTR_Finder is an efficient program for finding full-length LTR retrotranspsons in genome sequences.

The Program first constructs all exact match pairs by a suffix-array based algorithm and extends them to long highly similar pairs. Then Smith-Waterman algorithm is used to adjust the ends of LTR pair candidates to get alignment boundaries. These boundaries are subject to re-adjustment using supporting information of TG..CA box and TSRs and reliable LTRs are selected. Next, LTR_FINDER tries to identify PBS, PPT and RT inside LTR pairs by build-in aligning and counting modules. RT identification includes a dynamic programming to process frame shift. For other protein domains, LTR_FINDER calls ps_scan (from PROSITE, http://www.expasy.org/prosite/) to locate cores of important enzymes if they occur.

Address of the bookmark: https://github.com/xzhub/LTR_Finder

Apollo: First instantaneous, collaborative genomic annotation editor available on the Web

Jit — Fri, 31 May 2019 19:55:39 -0500

Apollo is a plug-in for the JBrowse Genome Viewer.
In addition to genes and pseudogenes, users can annotate ncRNAs (snRNA, snoRNA, tRNA, rRNA), miRNAs, repeat regions, and transposable elements; each annotation type has its own configuration of the ‘Information Editor’.
History tracking with undo/redo functions is available.
Users are able to directly set an annotation to a specific state, choosing from the ‘History’ display.
Adding and updating PubMed IDs will prompt users with a publication title to confirm their submission.
Gene Ontology (GO) terms are supported and GO ID auto-completion has been incorporated.
Users may access a ‘Recent Changes’ page.
Help page with Apollo specific content is available.

Address of the bookmark: http://genomearchitect.github.io/

simuG: a general-purpose genome simulator

BioStar — Thu, 28 Nov 2019 04:33:18 -0600

Simulated genomes with pre-defined and random genomic variants can be very useful for benchmarking genomic and bioinformatics analyses. Here we introduce simuG, a lightweight tool for simulating the full-spectrum of genomic variants (single nucleotide polymorphisms, Insertions/Deletions, copy number variants, inversions and translocations) for any organisms (including human). The simplicity and versatility of simuG make it a unique general-purpose genome simulator for a wide-range of simulation-based applications.

Address of the bookmark: https://github.com/yjx1217/simuG

mutatrix: a population genome simulator which generates simulated genomes.

Jit — Tue, 28 Jan 2020 04:06:58 -0600

genome simulation across a population with zeta-distributed allele frequency, snps, insertions, deletions, and multi-nucleotide polymorphisms

More at https://github.com/ekg/mutatrix

./mutatrix -S sample -P test/ -p 2 -n 10 reference.fasta

Address of the bookmark: https://github.com/ekg/mutatrix

shinyChromosome:a GUI for the interactive creation of non-circular whole genome diagrams

Jit — Sat, 29 Feb 2020 00:39:50 -0600

shinyChromosome is a graphical user interface for interactive creation of non-circular whole genome diagrams developed using the R Shiny package.

To create single-genome plot by aligning genome data along all chromosomes of a single genome, go to the Single-genome plot menu.

To cretae two-genome plot for comparison of data across two genomes, go to the Two-genome plot menu.

For the detail format of input data, check the Input data format submenu of the Help menu.

shinyChromosome is deployed at http://150.109.59.144:3838/shinyChromosome/, http://shinyChromosome.ncpgr.cn, and https://yimingyu.shinyapps.io/shinyChromosome for online use. The source code and manual of shinyChromosome are freely available at https://github.com/venyao/shinyChromosome.

https://yimingyu.shinyapps.io/shinychromosome/

https://www.sciencedirect.com/science/article/pii/S1672022919301883

Address of the bookmark: https://yimingyu.shinyapps.io/shinychromosome/

merqury: Evaluate genome assemblies with k-mers

Jit — Fri, 03 Jul 2020 19:29:34 -0500

Often, genome assembly projects have illumina whole genome sequencing reads available for the assembled individual. The k-mer spectrum of this read set can be used for independently evaluating assembly quality without the need of a high quality reference. Merqury provides a set of tools for this purpose.

More at https://www.biorxiv.org/content/10.1101/2020.03.15.992941v1.full

Address of the bookmark: https://github.com/marbl/merqury

Flanker

Rahul Nayak — Sat, 27 Feb 2021 22:04:53 -0600

Flanker, a Python package which performs alignment-free clustering of gene flanking sequences in a consistent format, allowing investigation of mobile genetic elements (MGEs) without prior knowledge of their structure. Flanker can be flexibly parameterised to finetune outputs by characterising upstream and downstream regions separately and investigating variable lengths of flanking sequence.

Address of the bookmark: https://github.com/wtmatlock/flanker

ncbi-datasets-cli -- Quickstart: command line tools !

Jit — Tue, 07 Dec 2021 02:51:26 -0600

Install and use the NCBI Datasets command line tools

The NCBI Datasets datasets command line tools are datasets and dataformat .

Use datasets to download biological sequence data across all domains of life from NCBI.

Use dataformat to convert metadata from JSON Lines format to other formats.

Conda download:

https://anaconda.org/conda-forge/ncbi-datasets-cli

Buld Download

https://www.ncbi.nlm.nih.gov/datasets/builder/?tax_id=29979

Address of the bookmark: https://www.ncbi.nlm.nih.gov/datasets/docs/v1/quickstarts/command-line-tools/

Smudgeplot: Inference of ploidy and heterozygosity structure using whole genome sequencing data

Neel — Fri, 25 Feb 2022 04:42:09 -0600

This tool extracts heterozygous kmer pairs from kmer count databases and performs gymnastics with them. We are able to disentangle genome structure by comparing the sum of kmer pair coverages (CovA + CovB) to their relative coverage (CovB / (CovA + CovB)). Such an approach also allows us to analyze obscure genomes with duplications, various ploidy levels, etc.

Smudgeplots are computed from raw or even better from trimmed reads and show the haplotype structure using heterozygous kmer pairs. For example:

Address of the bookmark: https://github.com/KamilSJaron/smudgeplot