BOL: Related items

GenomeMapper: Simultaneous alignment of short reads against multiple genomes

Jit — Fri, 25 May 2018 09:29:44 -0500

GenomeMapper is a short read mapping tool designed for accurate read alignments. It quickly aligns millions of reads either with ungapped or gapped alignments. It can be used to align against multiple genomes simulanteously or against a single reference. If you are unsure which one is the appropriate GenomeMapper, you might want to use the latter https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2768987/

Address of the bookmark: http://1001genomes.org/software/genomemapper.html

MSAProbs - Parallel and accurate multiple sequence alignment

Neel — Tue, 09 Jul 2019 23:58:44 -0500

MSAProbs is a well-established state-of-the-art multiple sequence alignment algorithm for protein sequences. The design of MSAProbs is based on a combination of pair hidden Markov models and partition functions to calculate posterior probabilities. Assessed using the popular benchmarks: BAliBASE, PREFAB, SABmark and OXBENCH, MSAProbs achieves statistically significant accuracy improvements over the existing top performing aligners, including ClustalW, MAFFT, MUSCLE, ProbCons and Probalign. In addition, MSAProbs is optimized for shared-memory CPUs by employing a multi-threaded design, and further parallelized for distributed-memory systems using MPI to overcome high memory overhead barrier and achieve good parallel and data-size scalability.

Address of the bookmark: http://msaprobs.sourceforge.net/homepage.htm#latest

VG: variation graph data structures, interchange formats, alignment, genotyping, and variant calling methods

Jit — Tue, 28 Jan 2020 03:53:24 -0600

Variation graphs provide a succinct encoding of the sequences of many genomes. A variation graph (in particular as implemented in vg) is composed of:

nodes, which are labeled by sequences and ids
edges, which connect two nodes via either of their respective ends
paths, describe genomes, sequence alignments, and annotations (such as gene models and transcripts) as walks through nodes connected by edges

Address of the bookmark: https://github.com/vgteam/vg

UniAligner: a parameter-free framework for fast sequence alignment

Abhi — Fri, 08 Mar 2024 23:36:12 -0600

UniAligner (formerly, TandemAligner) is the first parameter-free algorithm for sequence alignment that introduces a sequence-dependent alignment scoring that automatically changes for any pair of compared sequences. Classical alignment approaches, such as the Smith-Waterman algorithm, that work well for most sequences, fail to construct biologically adequate alignments of extra-long tandem repeats (ETRs), such as human centromeres and immunoglobulin loci. This limitation was overlooked in the previous studies since the sequences of the centromeres and other ETRs across multiple genomes only became available recently.

More at https://www.nature.com/articles/s41592-023-01970-4

Address of the bookmark: https://github.com/seryrzu/unialigner

Seal: SEquence ALignment evaluation suite

Jit — Wed, 03 Jan 2018 05:05:46 -0600

Seal is a comprehensive sequencing simulation and alignment tool evaluation suite. This software (implemented in Java) provides several utilities that can be used to evaluate alignment algorithms, including:

Reading a pre-existing reference genome from one or more FASTA files.
Alternatively, generating an artificial reference genome based on input parameters (length, repeat count, repeat length, repeat variability rate).
Simulating reads from random locations in the genome based on input parameters of read length, coverage, sequencing error rate, and indel rate.
Applying alignment tools to the genome and the reads through a standardized interface.
Parsing the output of the alignment tool and calculating the number of reads that were correctly or incorrectly mapped.
Computing run times and measures of accuracy.

Seal has interfaces to evaluate the following software packages:

Bowtie
BWA
MAQ
mrFAST
mrsFAST
Novoalign
SHRiMP
SOAPv2

Address of the bookmark: http://compbio.case.edu/seal/

BamView: a free interactive display of read alignments in BAM data files

Neel — Fri, 09 Nov 2018 13:43:22 -0600

To run the application on UNIX from the downloaded jar file run the UNIX:

java -mx512m -jar BamView.jar

and extra command line options are given when '-h' is used:

java -jar BamView.jar -h

BAM files can be specified on the command line with the '-a' option:

java -mx512m -jar BamView.jar -a pathToFile/sorted.bam

If a BAM filename is not given on the command line BamView will prompt for a file to be entered. The BAM index file should have the same name as the BAM file but with a '.bai' suffix. Multiple BAM files can be loaded and overlaid in the viewer. To make this easier BamView will read in files that contain a list of filenames.

Address of the bookmark: http://bamview.sourceforge.net/

Quip: Aggressive compression of FASTQ, SAM and BAM files.

Neel — Tue, 24 May 2022 06:31:48 -0500

This will help us to reduce the amount of drive space we take up and decrease data transfer times

Quip compresses next-generation sequencing data with extreme prejudice. It supports input and output in the FASTQ and SAM/BAM formats, compressing large datasets to as little as 15% of their original size.

Address of the bookmark: https://github.com/dcjones/quip

SeQuiLa-cov: A fast and scalable library for depth of coverage calculations

Jit — Sun, 15 Dec 2019 10:19:35 -0600

The Docker image is available at https://hub.docker.com/r/biodatageeks/. Supplementary information on benchmarking procedure as well as test data are publicly accessible at the project documentation site http://biodatageeks.org/sequila/benchmarking/benchmarking.html#depth-of-coverage. An archival copy of the code and supporting data is also available via the GigaScience database GigaDB

• Project name: SeQuiLa-cov

• Project home page: http://biodatageeks.org/sequila/

• Source code repository: https://github.com/ZSI-Bio/bdg-sequila

• Operating system: Platform independent

• Programming language: Scala

• Other requirements: Docker

• License: Apache License 2.0

Address of the bookmark: https://academic.oup.com/gigascience/article/8/8/giz094/5543653

The Human Genome Project Video 3D Animation Introduction Low)

Sat, 24 Aug 2013 19:01:19 -0500

The genome factory !!!

Madhvan Reddy — Thu, 16 Jan 2014 02:09:31 -0600

Illumina, Inc. announced Tuesday that its new HiSeq X Ten Sequencing System has broken the “sound barrier” of human genomics by enabling the $1,000 genome. “This platform includes dramatic technology breakthroughs that enable researchers to undertake studies of unprecedented scale by providing the throughput to sequence tens of thousands of human whole genomes in a single year in a single lab,” Illumina stated.

Initial customers for the HiSeq X Ten System, which will ship in Q1 2014, include Macrogen, based in Seoul, South Korea and its CLIA laboratory in Rockville, Maryland, the Broad Institute in Cambridge, Massachusetts, and the Garvan Institute of Medical Research in Sydney, Australia.

“For the first time, it looks like it will be possible to deliver the $1,000 genome, which is tremendously exciting,” said Eric Lander, founding director of the Broad Institute and a professor of biology at MIT. “The HiSeq X Ten should give us the ability to analyze complete genomic information from huge sample populations. Over the next few years, we have an opportunity to learn as much about the genetics of human disease as we have learned in the history of medicine.”

“The HiSeq X Ten is an ideal platform for scientists and institutions focused on the discovery of genotypic variation to enable a deeper understanding of human biology and genetic disease,” Illumina stated. “It can sequence tens of thousands of samples annually with high-quality, high-coverage sequencing, delivering a comprehensive catalog of human variation within and outside coding regions.”

HiSeq X Ten utilizes a number of advanced design features to generate massive throughput. Patterned flow cells, which contain billions of nanowells at fixed locations, combined with a new clustering chemistry deliver a significant increase in data density (6 billion clusters per run). Using state-of-the art optics and faster chemistry, HiSeq X Ten can process sequencing flow cells more quickly than ever before — generating a 10x increase in daily throughput when compared to current HiSeq 2500 performance.

The HiSeq X Ten is sold as a set of 10 or more ultra-high throughput sequencing systems, each generating up to 1.8 terabases (Tb) of sequencing data in less than three days or up to 600 gigabases (Gb) per day, per system, providing the throughput to sequence tens of thousands of high-quality, high-coverage genomes per year. Illumina says the $1,000 includes typical instrument depreciation, DNA extraction, library preparation, and estimated labor.