BOL: Related items

LAMSA: fast split read alignment with long approximate matches

Jit — Tue, 15 May 2018 04:44:42 -0500

LAMSA (Long Approximate Matches-based Split Aligner) is a novel split alignment approach with faster speed and good ability of handling SV events. It is well-suited to align long reads (over thousands of base-pairs). LAMSA takes takes the advantage of the rareness of SVs to implement a specifically designed two-step strategy. That is, LAMSA initially splits the read into relatively long fragments and co-linearly align them to solve the small variations or sequencing errors, and mitigate the effect of repeats. The alignments of the fragments are then used for implementing a sparse dynamic programming (SDP)-based split alignment approach to handle the large or non-co-linear variants. We benchmarked LAMSA with simulated and real datasets having various read lengths and sequencing error rates, the results demonstrate that it is substantially faster than the state-of-the-art long read aligners; mean-while, it also has good ability to handle various categories of SVs. LAMSA is open source and free for non-commercial use. LAMSA is mainly designed by Bo Liu & Yan Gao and developed by Yan Gao in Center for Bioinformatics, Harbin Institute of Technology, China.

Address of the bookmark: https://github.com/hitbc/LAMSA

nanofilt: Filtering and trimming of long read sequencing data

Jit — Mon, 30 Jul 2018 12:01:52 -0500

Filtering on quality and/or read length, and optional trimming after passing filters.
Reads from stdin, writes to stdout.

Intended to be used:

directly after fastq extraction
prior to mapping
in a stream between extraction and mapping

https://github.com/wdecoster/nanofilt

Address of the bookmark: https://github.com/wdecoster/nanofilt

rHAT: a seed-and-extension-based noisy long read alignment tool

Abhimanyu Singh — Sun, 23 Sep 2018 05:12:22 -0500

rHAT is a seed-and-extension-based noisy long read alignment tool. It is suitable for aligning 3rd generation sequencing reads which are in large read length with relatively high error rate, especially Pacbio's Single Molecule Read-time (SMRT) sequencing reads.

Address of the bookmark: https://github.com/dfguan/rHAT

URMAP, an ultra-fast read mapper

Jit — Thu, 29 Oct 2020 23:03:54 -0500

URMAP, a new read mapping algorithm. URMAP is an order of magnitude faster than BWA with comparable accuracy on several validation tests. On a Genome in a Bottle (GIAB) variant calling test with 30× coverage 2×150 reads, URMAP achieves high accuracy (precision 0.998, sensitivity 0.982 and F-measure 0.990) with the strelka2 caller. However, GIAB reference variants are shown to be biased against repetitive regions which are difficult to map and may therefore pose an unrealistically easy challenge to read mappers and variant callers.

More at https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7320720/

Address of the bookmark: https://github.com/rcedgar/urmap

Statistics Using R with Biological Examples

Neel — Thu, 03 Nov 2016 04:55:41 -0500

This book is a manifestation of my desire to teach researchers in biology a bit more about statistics than an ordinary introductory course covers and to introduce the utilization of R as a tool for analyzing their data. My goal is to reach those with little or no training in higher level statistics so that they can do more of their own data analysis, communicate more with statisticians, and appreciate the great potential statistics has to offer as a tool to answer biological questions.

This is necessary in light of the increasing use of higher level statistics in biomedical research. I hope it accomplishes this mission and encourage its free distribution and use as a course text or supplement.

K Seefeld, May 2007

Modern Statistics with R

LEGE — Thu, 22 Aug 2024 04:44:06 -0500

This is the online version of the second edition of Modern Statistics with R. It is free to use, and always will be. Printed copies are available from CRC Press.

Live online courses on statistics with R based on this book, led by the author, are offered regularly; see this page for more information and dates.

The past decades have transformed the world of statistical data analysis, with new methods, new types of data, and new computational tools. The aim of Modern Statistics with R is to introduce you to key parts of the modern statistical toolkit. It teaches you:

Data wrangling - importing, formatting, reshaping, merging, and filtering data in R.
Exploratory data analysis - using visualisations and multivariate techniques to explore datasets.
Statistical inference - modern methods for testing hypotheses and computing confidence intervals.
Predictive modelling - regression models and machine learning methods for prediction, classification, and forecasting.
Simulation - using simulation techniques for sample size computations and evaluations of statistical methods.
Ethics in statistics - ethical issues and good statistical practice.
R programming - writing code that is fast, readable, and (hopefully!) free from bugs.

The book includes plenty of examples and more than 200 exercises with worked solutions. The datasets used for the examples and the exercises can be downloaded here.

Address of the bookmark: https://www.modernstatisticswithr.com/

biostarhandbook

Neel — Fri, 27 Aug 2021 01:31:01 -0500

Nice book collection for bioinformatician ... highly recommended.

Address of the bookmark: https://www.biostarhandbook.com/

Biotechnology Eligibility Test (BET) Answer set !

Abhimanyu Singh — Wed, 17 Apr 2019 04:57:58 -0500

DBT-BET JRF 2019 Exam was held successfully on 14th April 2019. Official Question Paper & Answer key of DBT-BET 2019 Exam has been released.

The best part about the DBT-BET Exam is both the question paper & answer key is released immediately after the exam so you get the upper hand here to check how well you did in the DBT-BET exam, how much you are going to score and if you can pass the Exam with good score. This helps you in planning your career ahead.

DBT Exam is held once in a year by Department of Biotechnology – Govt of India. Candidates willing to get DBT-Junior Research Fellowship – DBT JRF fellowship should take up the DBT-BET Exam, DBT – Biotechnology Eligibility Test (BET) 2020 to be held on somewhere in the month of April. Official Notification for DBT-BET 2020 will be released in Feb 2020.

Question at https://bioinformaticsonline.com/file/view/39260/biotechnology-eligibility-test-bet-question-paper

Grinder / Biogrinder - A versatile omics shotgun and amplicon sequencing read simulator

Jit — Wed, 24 May 2017 08:41:41 -0500

Grinder is a versatile program to create random shotgun and amplicon sequence libraries based on DNA, RNA or proteic reference sequences provided in a FASTA file.

Grinder can produce genomic, metagenomic, transcriptomic, metatranscriptomic, proteomic, metaproteomic shotgun and amplicon datasets from current sequencing technologies such as Sanger, 454, Illumina. These simulated datasets can be used to test the accuracy of bioinformatic tools under specific hypothesis, e.g. with or without sequencing errors, or with low or high community diversity. Grinder may also be used to help decide between alternative sequencing methods for a sequence-based project, e.g. should the library be paired-end or not, how many reads should be sequenced.

Address of the bookmark: https://sourceforge.net/projects/biogrinder/files/biogrinder/

LRCstats: Long Read Correction Statistics

Jit — Fri, 05 Jan 2018 04:04:20 -0600

LRCstats is an open-source pipeline for benchmarking DNA long read correction algorithms for long reads outputted by third generation sequencing technology such as machines produced by Pacific Biosciences. The reads produced by third generation sequencing technology, as the name suggests, are longer in length than reads produced by next generation sequencing technologies, such as those produced by Illumina. However, long reads are plagued by high error rates, which can cause issues in downstream analysis. Long read correction algorithms reduce the error rate of long reads either through self-correcting methods or using accurate, short reads outputted by next generation sequencing technologies to correct long reads.

Of course, some long read correction algorithms are better than others, and developers of long read correction algorithms will wish to compare their algorithm with others currently available. LRCstats benchmarks long read correction algorithms using long reads produced by simulators (such as SimLoRD or PBSim) where the two-way alignments between the uncorrected long reads (uLR) and the corresponding sequences in the reference genome (Ref) are given in some sort of alignment file and then aligning the corrected long reads (cLR) to the Ref-uLR two-way alignments to create three-way alignments using a dynamic programming algorithm. Statistics on these three-way alignments are then collected, such as the overall error rates of the corrected long reads.

https://www.healthcare.uiowa.edu/labs/au/LSC/

Address of the bookmark: https://github.com/cchauve/lrcstats