BOL: Related items

SNP Analysis: Unlocking the Secrets in Our DNA

Abhi — Wed, 16 Jul 2025 01:31:45 -0500

Single Nucleotide Polymorphisms (SNPs) are the most common type of genetic variation in humans—and many other organisms. A single base change in the DNA sequence (for example, an A instead of a G) can influence everything from our eye color to our risk of developing diseases. Analyzing these tiny changes has become central to modern genetics, medicine, agriculture, and evolutionary biology.

What are SNPs?
SNPs (pronounced "snips") are positions in the genome where individuals differ by a single nucleotide. For example:

Reference: ...A T G C A T G A...
Variant: ...A T G T A T G A...

Here, the C in the reference genome has been replaced by a T in the variant.

SNPs occur roughly every 300–1,000 bases in the human genome, meaning there are millions of them scattered throughout our DNA. Most SNPs have no effect on health, but some are linked to disease susceptibility, drug response, and other traits.

Why Do We Analyze SNPs?
1. Medical Genetics

Identify disease-associated variants (e.g., BRCA1/2 in breast cancer).

Predict drug response (pharmacogenomics).

Enable precision medicine by tailoring treatments.

2. Population Genetics & Ancestry

Trace human migration and ancestry.

Study genetic diversity within and between populations.

3. Agriculture & Animal Breeding

Select for desirable traits (drought resistance, yield, disease resistance).

Improve breeding efficiency in livestock.

4. Evolutionary Biology

Track natural selection.

Study adaptation in wild populations.

How is SNP Analysis Performed?
SNP analysis can be broadly divided into three steps:

SNP Detection
Genotyping arrays: Chips that test hundreds of thousands of known SNP positions simultaneously. Fast and affordable, widely used in consumer ancestry testing.

Whole-genome or whole-exome sequencing: Can detect known and novel SNPs across the genome.

Targeted sequencing or PCR: For focused analysis of specific regions.

Variant Calling
Sequencing data is aligned to a reference genome. Bioinformatics tools (e.g., GATK, bcftools) identify positions where the sequenced sample differs from the reference.

Annotation and Interpretation
Tools (e.g., SnpEff, VEP) predict the functional impact of SNPs.

Are the SNPs in coding regions? Do they cause amino acid changes? Are they known to be pathogenic?

Databases like dbSNP, ClinVar, and GWAS Catalog provide information on known associations.

Common Tools for SNP Analysis
Alignment: BWA, Bowtie2

Variant Calling: GATK, FreeBayes

Visualization: IGV, UCSC Genome Browser

Annotation: SnpEff, VEP

Statistical Analysis: PLINK, SNPTEST

Challenges in SNP Analysis
False positives/negatives: Sequencing errors, alignment issues.

Population stratification: Confounding in association studies.

Interpretation: Many SNPs have unknown or complex effects.

Researchers address these with rigorous quality control, large datasets, and increasingly sophisticated statistical models.

The Future of SNP Analysis
With advances in sequencing technology and AI-driven analysis, SNP studies are expanding:

Polygenic risk scores predict disease risk based on thousands of SNPs.

Large-scale biobanks (e.g., UK Biobank, All of Us) enable powerful genome-wide association studies (GWAS).

CRISPR and functional assays help validate SNP effects in the lab.

SNP analysis is at the heart of the genomic revolution, promising insights into biology, health, and evolution at unprecedented scale.

Conclusion
From diagnosing rare diseases to designing better crops, SNP analysis is a foundational tool in modern science. As our ability to sequence and interpret genomes improves, so will our understanding of these tiny—but mighty—variations in DNA.

Biological file format tutorial

Jit — Sun, 17 Dec 2017 18:13:03 -0600

This section explains some of the commonly used file formats in bioinformatics. The information provided here is basic and designed to help users to distinguish the difference between different formats. Please refer user manual or other information resources on web for more details.

Address of the bookmark: https://bioinformatics.uconn.edu/resources-and-events/tutorials/file-formats-tutorial/

Pangolin tutorial !

Abhi — Fri, 10 Dec 2021 05:58:59 -0600

This is a tutorial for using the Pangolin Web Application. For information on using the command line tool, please visit the command line tool usage page.

https://cov-lineages.org/resources/pangolin/tutorial.html

Address of the bookmark: https://cov-lineages.org/resources/pangolin/tutorial.html

Comparison of Short Read De Novo Alignment Algorithms

Rahul Agarwal — Wed, 21 Aug 2013 07:56:01 -0500

Excellent article to introduce different sequencing methods along with tools for de novo assembly of sequencing reads and their relevant references.

Title: Comparison of Short Read De Novo Alignment Algorithms

Author: Nikhil Gopal

Address of the bookmark: http://biochem218.stanford.edu/Projects%202011/Gopal%202011.pdf

3d-dna: 3D de novo assembly (3D DNA) pipeline

Jit — Thu, 28 Dec 2017 10:09:37 -0600

This code is designed to enable anyone to reproduce the Hs2-HiC and the AaegL4 genomes reported in: Dudchenko et al., De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science, 2017.

Unless otherwise noted, all terminology below is consistent with this paper, and all references to figures and tables in this readme refer to this paper. Specifically, some of the terminology used below is outlined in Figure S2. The assembly procedure is described in detail in the Supporting Online Materials, specifically in the section labelled “Pipeline description”.

In addition, the pipeline uses tools and methods from Juicer (Durand & Shamim et al., Cell Systems, 2016) and Juicebox (Durand & Robinson et al., Cell Systems, 2016), as well as additional dependencies noted below.

Feel free to post your questions and comments at: http://www.aidenlab.org/forum.html

http://aidenlab.org/documentation.html

Address of the bookmark: https://github.com/theaidenlab/3d-dna

PERGA: A Paired-End Read Guided De Novo Assembler for Extending Contigs Using SVM and Look Ahead Approach

Rahul Nayak — Tue, 05 Jun 2018 09:57:11 -0500

PERGA - Paired End Reads Guided Assembler PERGA is a novel sequence reads guided de novo assembly approach which adopts greedy-like prediction strategy for assembling reads to contigs and scaffolds. Instead of using single-end reads to construct contig, PERGA uses paired-end reads and different read overlap sizes from O ≥ Omax to Omin to resolve the gaps and branches. Moreover, by constructing a decision model using machine learning approach based on branch features, PERGA can determine the correct extension in 99.7% of cases. PERGA will try to extend the contigs by all feasible nucleotides and determine if these multiple extensions due to sequencing errors or repeats by using looking ahead technology, and it also try to separate the different repeats of nearby genomic regions to make the assembly result more longer and accurate. The simulated E.coli paired-end reads data are generated using GemSim (KE McElroy, F Luciani, T Thomas. Gemsim: General, Error-Model Based Simulator of Next-Generation Sequencing Data. BMC Genomics 2012, 13:74), with coverage 50x, 60x, 100x, read lengths 100-bp, and can be downloaded from https://github.com/zhuxiao/data_PERGA.

Address of the bookmark: https://github.com/hitbio/PERGA

ASplice: a scalable and memory-efficient algorithm for de novo transcriptome assembly

Rahul Nayak — Tue, 03 Jul 2018 04:09:46 -0500

With increased availability of de novo assembly algorithms, it is feasible to study entire transcriptomes of non-model organisms. While algorithms are available that are specifically designed for performing transcriptome assembly from high-throughput sequencing data, they are very memory-intensive, limiting their applications to small data sets with few libraries. Texas A&M University researchers develop a transcriptome assembly algorithm that recovers alternatively spliced isoforms and expression levels while utilizing as many RNA-Seq libraries as possible that contain hundreds of gigabases of data. New techniques are developed so that computations can be performed on a computing cluster with moderate amount of physical memory. Availability – A software program that implements the algorithm is available at: http://faculty.cse.tamu.edu/shsze/asplice. Sze SH, Pimsler ML, Tomberlin JK, Jones CD, Tarone AM. (2017) A scalable and memory-efficient algorithm for de novo transcriptome assembly of non-model organisms. BMC Genomics 18(Suppl 4):387.

Address of the bookmark: http://faculty.cse.tamu.edu/shsze/asplice/

MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph

Jit — Wed, 14 Nov 2018 04:50:27 -0600

MEGAHIT is a single node assembler for large and complex metagenomics NGS reads, such as soil. It makes use of succinct de Bruijn graph (SdBG) to achieve low memory assembly. MEGAHIT can optionally utilize a CUDA-enabled GPU to accelerate its SdBG contstruction. The GPU-accelerated version of MEGAHIT has been tested on NVIDIA GTX680 (4G memory) and Tesla K40c (12G memory) with CUDA 5.5, 6.0 and 6.5. MEGAHIT v1.0 or greater also supports IBM Power PC and has been tested on IBM POWER8.

https://academic.oup.com/bioinformatics/article/31/10/1674/177884

Address of the bookmark: https://github.com/voutcn/megahit

StringTie Transcript assembly and quantification for RNA-Seq

Jit — Tue, 09 Jun 2020 05:21:11 -0500

StringTie is a fast and highly efficient assembler of RNA-Seq alignments into potential transcripts. It uses a novel network flow algorithm as well as an optional de novo assembly step to assemble and quantitate full-length transcripts representing multiple splice variants for each gene locus. Its input can include not only alignments of short reads that can also be used by other transcript assemblers, but also alignments of longer sequences that have been assembled from those reads. In order to identify differentially expressed genes between experiments, StringTie's output can be processed by specialized software like Ballgown, Cuffdiff or other programs (DESeq2, edgeR, etc.).

Address of the bookmark: https://ccb.jhu.edu/software/stringtie/

3D de novo assembly (3D DNA) pipeline

Jit — Sun, 02 Feb 2020 13:41:55 -0600

For a detailed description of the pipeline and how it integrates with other tools designed by the Aiden Lab see Genome Assembly Cookbook on http://aidenlab.org/assembly.

For the original version of the pipeline and to reproduce the Hs2-HiC and the AaegL4 genomes reported in (Dudchenko et al., Science, 2017) see the original commit.

For the detailed description of the merge section see https://github.com/theaidenlab/AGWG-merge.

Address of the bookmark: https://github.com/theaidenlab/3d-dna