BOL: Related items

SNP Analysis: Unlocking the Secrets in Our DNA

Abhi — Wed, 16 Jul 2025 01:31:45 -0500

Single Nucleotide Polymorphisms (SNPs) are the most common type of genetic variation in humans—and many other organisms. A single base change in the DNA sequence (for example, an A instead of a G) can influence everything from our eye color to our risk of developing diseases. Analyzing these tiny changes has become central to modern genetics, medicine, agriculture, and evolutionary biology.

What are SNPs?
SNPs (pronounced "snips") are positions in the genome where individuals differ by a single nucleotide. For example:

Reference: ...A T G C A T G A...
Variant: ...A T G T A T G A...

Here, the C in the reference genome has been replaced by a T in the variant.

SNPs occur roughly every 300–1,000 bases in the human genome, meaning there are millions of them scattered throughout our DNA. Most SNPs have no effect on health, but some are linked to disease susceptibility, drug response, and other traits.

Why Do We Analyze SNPs?
1. Medical Genetics

Identify disease-associated variants (e.g., BRCA1/2 in breast cancer).

Predict drug response (pharmacogenomics).

Enable precision medicine by tailoring treatments.

2. Population Genetics & Ancestry

Trace human migration and ancestry.

Study genetic diversity within and between populations.

3. Agriculture & Animal Breeding

Select for desirable traits (drought resistance, yield, disease resistance).

Improve breeding efficiency in livestock.

4. Evolutionary Biology

Track natural selection.

Study adaptation in wild populations.

How is SNP Analysis Performed?
SNP analysis can be broadly divided into three steps:

SNP Detection
Genotyping arrays: Chips that test hundreds of thousands of known SNP positions simultaneously. Fast and affordable, widely used in consumer ancestry testing.

Whole-genome or whole-exome sequencing: Can detect known and novel SNPs across the genome.

Targeted sequencing or PCR: For focused analysis of specific regions.

Variant Calling
Sequencing data is aligned to a reference genome. Bioinformatics tools (e.g., GATK, bcftools) identify positions where the sequenced sample differs from the reference.

Annotation and Interpretation
Tools (e.g., SnpEff, VEP) predict the functional impact of SNPs.

Are the SNPs in coding regions? Do they cause amino acid changes? Are they known to be pathogenic?

Databases like dbSNP, ClinVar, and GWAS Catalog provide information on known associations.

Common Tools for SNP Analysis
Alignment: BWA, Bowtie2

Variant Calling: GATK, FreeBayes

Visualization: IGV, UCSC Genome Browser

Annotation: SnpEff, VEP

Statistical Analysis: PLINK, SNPTEST

Challenges in SNP Analysis
False positives/negatives: Sequencing errors, alignment issues.

Population stratification: Confounding in association studies.

Interpretation: Many SNPs have unknown or complex effects.

Researchers address these with rigorous quality control, large datasets, and increasingly sophisticated statistical models.

The Future of SNP Analysis
With advances in sequencing technology and AI-driven analysis, SNP studies are expanding:

Polygenic risk scores predict disease risk based on thousands of SNPs.

Large-scale biobanks (e.g., UK Biobank, All of Us) enable powerful genome-wide association studies (GWAS).

CRISPR and functional assays help validate SNP effects in the lab.

SNP analysis is at the heart of the genomic revolution, promising insights into biology, health, and evolution at unprecedented scale.

Conclusion
From diagnosing rare diseases to designing better crops, SNP analysis is a foundational tool in modern science. As our ability to sequence and interpret genomes improves, so will our understanding of these tiny—but mighty—variations in DNA.

Biological file format tutorial

Jit — Sun, 17 Dec 2017 18:13:03 -0600

This section explains some of the commonly used file formats in bioinformatics. The information provided here is basic and designed to help users to distinguish the difference between different formats. Please refer user manual or other information resources on web for more details.

Address of the bookmark: https://bioinformatics.uconn.edu/resources-and-events/tutorials/file-formats-tutorial/

Machine learning training and courses in bioinformatics !

Rahul Nayak — Tue, 31 Dec 2019 19:33:07 -0600

Machine learning techniques have been successful in analyzing biological data because of their capabilities in handling randomness and uncertainty of data noise and in generalization. In this class, we will learn basics about probabilistic models and machine learning techniques. We will focus on probabilistic models (Markov models, Hidden Markov models, and Bayesian networks) for biological sequence analysis and systems biology. Other machine learning techniques, such as Naive bayes, neural networks and SVMs will only be covered briefly.

More at http://homes.sice.indiana.edu/yye/lab/teaching/spring2017-I529/

Pangolin tutorial !

Abhi — Fri, 10 Dec 2021 05:58:59 -0600

This is a tutorial for using the Pangolin Web Application. For information on using the command line tool, please visit the command line tool usage page.

https://cov-lineages.org/resources/pangolin/tutorial.html

Address of the bookmark: https://cov-lineages.org/resources/pangolin/tutorial.html

HiCdat

Jit — Fri, 12 Feb 2016 05:23:44 -0600

HiCdat: a fast and easy-to-use Hi-C data analysis tool

HiCdat is easy-to-use and provides solutions starting from aligned reads up to in-depth analyses. Importantly, HiCdat is focussed on the analysis of larger structural features of chromosomes, their correlation to genomic and epigenomic features, and on comparative studies. It uses simple input and output formats and can therefore easily be integrated into existing workflows or combined with alternative tools.

More at http://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-015-0678-x

Address of the bookmark: https://github.com/MWSchmid/HiCdat

KisSplice

Jit — Tue, 16 Aug 2016 08:34:19 -0500

KisSplice is a software that enables to analyse RNA-seq data with or without a reference genome. It is an exact local transcriptome assembler that allows to identify SNPs, indels and alternative splicing events. It can deal with an arbitrary number of biological conditions, and will quantify each variant in each condition. It has been tested on Illumina datasets of up to 1G reads. Its memory consumption is around 5Gb for 100M reads.

KisSplice is not a full-length transcriptome assembler. This means that it will output the variable regions of the transcripts, not reconstruct them entirely.

KisSplice comes as a workflow, with several possible post-treatments meant to facilitate the analysis of the results. The choice of the post-treatment depends on the availability of a reference genome/transcriptome and on the need to perform a differential analysis, as summarised in the following table.

Address of the bookmark: http://kissplice.prabi.fr/

Machine Learning !!!

Gudiya Pal — Fri, 01 Jul 2016 12:57:12 -0500

In machine learning, computers apply statistical learning techniques to automatically identify patterns in data. These techniques can be used to make highly accurate predictions.

Keep scrolling. Using a data set about homes, we will create a machine learning model to distinguish homes in New York from homes in San Francisco.

Address of the bookmark: http://www.r2d3.us/visual-intro-to-machine-learning-part-1/

WiseScaffolder

Poonam Mahapatra — Wed, 13 Jul 2016 08:08:57 -0500

Function

WiseScaffolder is a stand-alone semi-automatic application for genome scaffolding of pre-assembled contigs using mate-pair data. It also produces editable scaffold maps, allowing either to build gapped scaffolds or usable as a common thread for the manual improvement of scaffolds.

Description

WiseScaffolder includes 4 subcommands: dumpconfig generates a configuration file that notably specifies the average insert size of the mate-pair library preprocess allows the detection and correction of chimerae, the estimation of contigs copy number and produces valuable outputs for the manual improvement of scaffolds scaffold constitutes the central scaffold-builder and comprises two modules:

i) the interative_scaffold_extender, which works with big, unambiguous contigs, or when they run out, single copy contigs, and

ii) the small_contig_inserter, which inserts the small contigs within scaffolds buildfasta converts the scaffold(s) map(s) into Fasta sequences.

Address of the bookmark: http://abims.sb-roscoff.fr/wisescaffolder

CrossMap

Abhimanyu Singh — Mon, 05 Sep 2016 04:07:38 -0500

CrossMap is a program for convenient conversion of genome coordinates (or annotation files) between different assemblies (such as Human hg18 (NCBI36) <> hg19 (GRCh37), Mouse mm9 (MGSCv37) <> mm10 (GRCm38)).
It supports most commonly used file formats including SAM/BAM, Wiggle/BigWig, BED, GFF/GTF, VCF.
CrossMap is designed to liftover genome coordinates between assemblies. It’s not a program for aligning sequences to reference genome.
We do not recommend using CrossMap to convert genome coordinates between species.

Address of the bookmark: http://crossmap.sourceforge.net/

TEannot

Jit — Thu, 18 Aug 2016 10:02:03 -0500

We advise to run first the TEdenovo pipeline but it is not compulsory. We suppose you begin by running the TEannot pipeline on the example provided in the directory "db/" rather than directly on your own genomic sequences. Thus, from now on, the project name is "DmelChr4".

Address of the bookmark: https://urgi.versailles.inra.fr/Tools/REPET/TEannot-tuto