BOL: Related items

The Genome 10K Project

Tue, 29 Jul 2014 09:11:04 -0500

https://genome10k.soe.ucsc.edu The Genome 10K project aims to assemble a genomic zoo—a collection of DNA sequences representing the genomes of 10,000 vertebrate species, approximately one for every vertebrate genus. The trajectory of cost reduction in DNA sequencing suggests that this project will be feasible within a few years. Capturing the genetic diversity of vertebrate species would create an unprecedented resource for the life sciences and for worldwide conservation efforts. The growing Genome 10K Community of Scientists (G10KCOS), made up of leading scientists representing major zoos, museums, research centers, and universities around the world, is dedicated to coordinating efforts in tissue specimen collection that will lay the groundwork for a large-scale sequencing and analysis project.

SNP Analysis: Unlocking the Secrets in Our DNA

Abhi — Wed, 16 Jul 2025 01:31:45 -0500

Single Nucleotide Polymorphisms (SNPs) are the most common type of genetic variation in humans—and many other organisms. A single base change in the DNA sequence (for example, an A instead of a G) can influence everything from our eye color to our risk of developing diseases. Analyzing these tiny changes has become central to modern genetics, medicine, agriculture, and evolutionary biology.

What are SNPs?
SNPs (pronounced "snips") are positions in the genome where individuals differ by a single nucleotide. For example:

Reference: ...A T G C A T G A...
Variant: ...A T G T A T G A...

Here, the C in the reference genome has been replaced by a T in the variant.

SNPs occur roughly every 300–1,000 bases in the human genome, meaning there are millions of them scattered throughout our DNA. Most SNPs have no effect on health, but some are linked to disease susceptibility, drug response, and other traits.

Why Do We Analyze SNPs?
1. Medical Genetics

Identify disease-associated variants (e.g., BRCA1/2 in breast cancer).

Predict drug response (pharmacogenomics).

Enable precision medicine by tailoring treatments.

2. Population Genetics & Ancestry

Trace human migration and ancestry.

Study genetic diversity within and between populations.

3. Agriculture & Animal Breeding

Select for desirable traits (drought resistance, yield, disease resistance).

Improve breeding efficiency in livestock.

4. Evolutionary Biology

Track natural selection.

Study adaptation in wild populations.

How is SNP Analysis Performed?
SNP analysis can be broadly divided into three steps:

SNP Detection
Genotyping arrays: Chips that test hundreds of thousands of known SNP positions simultaneously. Fast and affordable, widely used in consumer ancestry testing.

Whole-genome or whole-exome sequencing: Can detect known and novel SNPs across the genome.

Targeted sequencing or PCR: For focused analysis of specific regions.

Variant Calling
Sequencing data is aligned to a reference genome. Bioinformatics tools (e.g., GATK, bcftools) identify positions where the sequenced sample differs from the reference.

Annotation and Interpretation
Tools (e.g., SnpEff, VEP) predict the functional impact of SNPs.

Are the SNPs in coding regions? Do they cause amino acid changes? Are they known to be pathogenic?

Databases like dbSNP, ClinVar, and GWAS Catalog provide information on known associations.

Common Tools for SNP Analysis
Alignment: BWA, Bowtie2

Variant Calling: GATK, FreeBayes

Visualization: IGV, UCSC Genome Browser

Annotation: SnpEff, VEP

Statistical Analysis: PLINK, SNPTEST

Challenges in SNP Analysis
False positives/negatives: Sequencing errors, alignment issues.

Population stratification: Confounding in association studies.

Interpretation: Many SNPs have unknown or complex effects.

Researchers address these with rigorous quality control, large datasets, and increasingly sophisticated statistical models.

The Future of SNP Analysis
With advances in sequencing technology and AI-driven analysis, SNP studies are expanding:

Polygenic risk scores predict disease risk based on thousands of SNPs.

Large-scale biobanks (e.g., UK Biobank, All of Us) enable powerful genome-wide association studies (GWAS).

CRISPR and functional assays help validate SNP effects in the lab.

SNP analysis is at the heart of the genomic revolution, promising insights into biology, health, and evolution at unprecedented scale.

Conclusion
From diagnosing rare diseases to designing better crops, SNP analysis is a foundational tool in modern science. As our ability to sequence and interpret genomes improves, so will our understanding of these tiny—but mighty—variations in DNA.

Research Associate at Indian Institute of Chemical Technology (IICT), Hyderabad

Thu, 07 Aug 2014 01:55:21 -0500

Indian Institute of Chemical Technology (IICT), Hyderabad, a constituent of CSIR is a leading research Institute in the area of chemical sciences. The core strength of IICT lies in Organic Chemistry, and it continues to excel in this field for over six decades. The research efforts during these years have resulted in the development of several innovative processes for a variety of products necessary for human welfare such as drugs, agrochemicals, food, organic intermediates, adhesives etc. More than 150 technologies developed by IICT are now in commercial production.

CSIR-IICT is conducting Walk in Interview for the following position on a purely temporary basis for the sponsored project "GENESIS (BSC-0121) at 10.00 AM on 19th August 2014 at IICT, Hyderabad

Position : Research Associate
No of Post : One
Desired Profile : PhD in computation biology or M.Tech in Computational Biology with three years experience in relevant subject and atleast one research paper in SCI journal

Experience : Knowledge in vector and vector borne disease, disease modeling, GIS mapping and modeling.
Age : 35 years
Stipend : Rs 22000/- + HRA

Eligible candidate may download the application form from our website http://www.iictindia.org and appear for interview along with the duly filled in application form supported by bio-data and one set of attested photo copies of certificates of educational qualification, age, experience, caste, latest photograph and the cadndidates are required to bring all the original certificates for verification

Walk in Interview : 19.08.14

Single Cell RNAseq data analysis tutorial !!

Robert M Willioms — Mon, 27 Nov 2017 16:24:29 -0600

A major breakthrough (replaced microarrays) in the late 00’s and has been widely used since
Measures the average expression level for each gene across a large population of input cells
Useful for comparative transcriptomics, e.g. samples of the same tissue from different species
Useful for quantifying expression signatures from ensembles, e.g. in disease studies
Insufficient for studying heterogeneous systems, e.g. early development studies, complex tissues (brain)
Does not provide insights into the stochastic nature of gene expression

Following are the useful links:

Single Cell RNAseq data analysis Tutorial

A step-by-step workflow for low-level analysis of single-cell RNA-seq data

A step-by-step workflow for low-level analysis of single-cell RNA-seq data with Bioconductor

SCell: single-cell RNA-seq analysis software

https://github.com/diazlab/SCell

Beta-Poisson model for single-cell RNA-seq data analyses

https://github.com/nghiavtr/BPSC

Sincera: A Computational Pipeline for Single Cell RNA-Seq Profiling Analysis

https://research.cchmc.org/pbge/sincera.html

SC3 – consensus clustering of single-cell RNA-Seq data

http://biorxiv.org/content/early/2016/09/02/036558

Citrus: A toolkit for single cell sequencing analysis

http://biorxiv.org/content/early/2016/09/14/045070

Single-Cell Resolution of Temporal Gene Expression during Heart Development

http://www.cell.com/developmental-cell/fulltext/S1534-5807(16)30682-7

Scalable latent-factor models applied to single-cell RNA-seq data separate biological drivers from confounding effects

http://biorxiv.org/content/early/2016/11/15/087775

Single cell transcriptomes identify human islet cell signatures and reveal cell-type-specific expression changes in type 2 diabetes

http://genome.cshlp.org/content/early/2016/11/18/gr.212720.116.abstract

SCODE: An efficient regulatory network inference algorithm from single-cell RNA-Seq during differentiation

http://biorxiv.org/content/early/2016/11/21/088856

SCOUP is a probabilistic model to analyze single-cell expression data during differentiation

https://github.com/hmatsu1226/SCOUP

scLVM is a modelling framework for single-cell RNA-seq data

https://github.com/PMBio/scLVM

Selective Locally linear Inference of Cellular Expression Relationships (SLICER) algorithm for inferring cell trajectories

https://github.com/jw156605/SLICER

SinQC: A Method and Tool to Control Single-cell RNA-seq Data Quality

http://www.morgridge.net/SinQC.html

TSCAN: Pseudo-time reconstruction and evaluation in single-cell RNA-seq analysis

https://github.com/zji90/TSCAN

Visualization and cellular hierarchy inference of single-cell data using SPADE

http://www.nature.com/nprot/journal/v11/n7/full/nprot.2016.066.html

OEFinder: Identify ordering effect genes in single cell RNA-seq data

https://github.com/lengning/OEFinder

JRF position in the Faculty of Life Sciences & Biotechnology at Sauth Asian University

Wed, 13 Aug 2014 07:16:30 -0500

Opening for a Project-JRF position in the Faculty of Life Sciences & Biotechnology

Applications are invited for the post of Junior Research Fellow (JRF) in a DBT funded IYBA project entitled “Generatingaprotein-ncRNA interactome for Dorsal mediated gene regulation and dorso-ventral patterning genes in Drosophila” in the Lab. Of Molecular Biology at the Faculty of Life Sciences and Biotechnology, South Asian University, New Delhi. The project requires extensive use of molecular, genetic and genomic approaches.

POST: Junior Research Fellow (JRF)

NO. OF VACANCIE(S) - (01)

FELLOWSHIP: Rs. 16,000/- plus HRA

PROJECT DURATION: 2014-2016 (Two years)

LAST DATE FOR APPLICATION: Aug 18, 2014.

Eligibility criteria:

M.Sc./M.Tech./ in Biological Sciences/Biotechnology/Bio-Informatics. Candidates with research experience in the field of Drosophila/Yeast genetics will be preferred.

Application Procedure:

A covering letter along with your CV, copy of prior publications (if any) and proof of experience should be e-mailed to lmb_sau@aol.com. Hardcopy of the application should be brought on the day of interview along with other testimonials and marks statements for verification purpose.

IMPORTANT NOTE:

-No TA/DA will be paid for attending the interview.

-SAU may select candidates against the post depending upon qualification and experience of candidates and reserves the right to relax any of the qualifications in case the candidate is found otherwise well qualified by the Selection Committee

-The abovementioned post is temporary and will be initially offered for a period of one year and can be extended, on satisfactory performance.

More at http://www.sau.ac.in/recruitment/vacancy.html

AWK for beginners !

BioJoker — Fri, 26 Apr 2019 16:19:41 -0500

AWK is a standard tool on every POSIX-compliant UNIX system. It’s like flex/lex, from the command-line, perfect for text-processing tasks and other scripting needs. It has a C-like syntax, but without mandatory semicolons (although, you should use them anyway, because they are required when you’re writing one-liners, something AWK excels at), manual memory management, or static typing. It excels at text processing. You can call to it from a shell script, or you can use it as a stand-alone scripting language.

Why use AWK instead of Perl? Readability. AWK is easier to read than Perl. For simple text-processing scripts, particularly ones that read files line by line and split on delimiters, AWK is probably the right tool for the job.

#!/usr/bin/awk -f

# Comments are like this


# AWK programs consist of a collection of patterns and actions.
pattern1 { action; } # just like lex
pattern2 { action; }

# There is an implied loop and AWK automatically reads and parses each
# record of each file supplied. Each record is split by the FS delimiter,
# which defaults to white-space (multiple spaces,tabs count as one)
# You can assign FS either on the command line (-F C) or in your BEGIN
# pattern

# One of the special patterns is BEGIN. The BEGIN pattern is true
# BEFORE any of the files are read. The END pattern is true after
# an End-of-file from the last file (or standard-in if no files specified)
# There is also an output field separator (OFS) that you can assign, which
# defaults to a single space

BEGIN {

    # BEGIN will run at the beginning of the program. It's where you put all
    # the preliminary set-up code, before you process any text files. If you
    # have no text files, then think of BEGIN as the main entry point.

    # Variables are global. Just set them or use them, no need to declare..
    count = 0;

    # Operators just like in C and friends
    a = count + 1;
    b = count - 1;
    c = count * 1;
    d = count / 1; # integer division
    e = count % 1; # modulus
    f = count ^ 1; # exponentiation

    a += 1;
    b -= 1;
    c *= 1;
    d /= 1;
    e %= 1;
    f ^= 1;

    # Incrementing and decrementing by one
    a++;
    b--;

    # As a prefix operator, it returns the incremented value
    ++a;
    --b;

    # Notice, also, no punctuation such as semicolons to terminate statements

    # Control statements
    if (count == 0)
        print "Starting with count of 0";
    else
        print "Huh?";

    # Or you could use the ternary operator
    print (count == 0) ? "Starting with count of 0" : "Huh?";

    # Blocks consisting of multiple lines use braces
    while (a < 10) {
        print "String concatenation is done" " with a series" " of"
            " space-separated strings";
        print a;

        a++;
    }

    for (i = 0; i < 10; i++)
        print "Good ol' for loop";

    # As for comparisons, they're the standards:
    # a < b   # Less than
    # a <= b  # Less than or equal
    # a != b  # Not equal
    # a == b  # Equal
    # a > b   # Greater than
    # a >= b  # Greater than or equal

    # Logical operators as well
    # a && b  # AND
    # a || b  # OR

    # In addition, there's the super useful regular expression match
    if ("foo" ~ "^fo+$")
        print "Fooey!";
    if ("boo" !~ "^fo+$")
        print "Boo!";

    # Arrays
    arr[0] = "foo";
    arr[1] = "bar";

    # You can also initialize an array with the built-in function split()

    n = split("foo:bar:baz", arr, ":");

    # You also have associative arrays (actually, they're all associative arrays)
    assoc["foo"] = "bar";
    assoc["bar"] = "baz";

    # And multi-dimensional arrays, with some limitations I won't mention here
    multidim[0,0] = "foo";
    multidim[0,1] = "bar";
    multidim[1,0] = "baz";
    multidim[1,1] = "boo";

    # You can test for array membership
    if ("foo" in assoc)
        print "Fooey!";

    # You can also use the 'in' operator to traverse the keys of an array
    for (key in assoc)
        print assoc[key];

    # The command line is in a special array called ARGV
    for (argnum in ARGV)
        print ARGV[argnum];

    # You can remove elements of an array
    # This is particularly useful to prevent AWK from assuming the arguments
    # are files for it to process
    delete ARGV[1];

    # The number of command line arguments is in a variable called ARGC
    print ARGC;

    # AWK has several built-in functions. They fall into three categories. I'll
    # demonstrate each of them in their own functions, defined later.

    return_value = arithmetic_functions(a, b, c);
    string_functions();
    io_functions();
}

# Here's how you define a function
function arithmetic_functions(a, b, c,     d) {

    # Probably the most annoying part of AWK is that there are no local
    # variables. Everything is global. For short scripts, this is fine, even
    # useful, but for longer scripts, this can be a problem.

    # There is a work-around (ahem, hack). Function arguments are local to the
    # function, and AWK allows you to define more function arguments than it
    # needs. So just stick local variable in the function declaration, like I
    # did above. As a convention, stick in some extra whitespace to distinguish
    # between actual function parameters and local variables. In this example,
    # a, b, and c are actual parameters, while d is merely a local variable.

    # Now, to demonstrate the arithmetic functions

    # Most AWK implementations have some standard trig functions
    localvar = sin(a);
    localvar = cos(a);
    localvar = atan2(b, a); # arc tangent of b / a

    # And logarithmic stuff
    localvar = exp(a);
    localvar = log(a);

    # Square root
    localvar = sqrt(a);

    # Truncate floating point to integer
    localvar = int(5.34); # localvar => 5

    # Random numbers
    srand(); # Supply a seed as an argument. By default, it uses the time of day
    localvar = rand(); # Random number between 0 and 1.

    # Here's how to return a value
    return localvar;
}

function string_functions(    localvar, arr) {

    # AWK, being a string-processing language, has several string-related
    # functions, many of which rely heavily on regular expressions.

    # Search and replace, first instance (sub) or all instances (gsub)
    # Both return number of matches replaced
    localvar = "fooooobar";
    sub("fo+", "Meet me at the ", localvar); # localvar => "Meet me at the bar"
    gsub("e+", ".", localvar); # localvar => "m..t m. at th. bar"

    # Search for a string that matches a regular expression
    # index() does the same thing, but doesn't allow a regular expression
    match(localvar, "t"); # => 4, since the 't' is the fourth character

    # Split on a delimiter
    n = split("foo-bar-baz", arr, "-"); # a[1] = "foo"; a[2] = "bar"; a[3] = "baz"; n = 3

    # Other useful stuff
    sprintf("%s %d %d %d", "Testing", 1, 2, 3); # => "Testing 1 2 3"
    substr("foobar", 2, 3); # => "oob"
    substr("foobar", 4); # => "bar"
    length("foo"); # => 3
    tolower("FOO"); # => "foo"
    toupper("foo"); # => "FOO"
}

function io_functions(    localvar) {

    # You've already seen print
    print "Hello world";

    # There's also printf
    printf("%s %d %d %d\n", "Testing", 1, 2, 3);

    # AWK doesn't have file handles, per se. It will automatically open a file
    # handle for you when you use something that needs one. The string you used
    # for this can be treated as a file handle, for purposes of I/O. This makes
    # it feel sort of like shell scripting, but to get the same output, the string
    # must match exactly, so use a variable:

    outfile = "/tmp/foobar.txt";

    print "foobar" > outfile;

    # Now the string outfile is a file handle. You can close it:
    close(outfile);

    # Here's how you run something in the shell
    system("echo foobar"); # => prints foobar

    # Reads a line from standard input and stores in localvar
    getline localvar;

    # Reads a line from a pipe (again, use a string so you close it properly)
    cmd = "echo foobar";
    cmd | getline localvar; # localvar => "foobar"
    close(cmd);

    # Reads a line from a file and stores in localvar
    infile = "/tmp/foobar.txt";
    getline localvar < infile; 
    close(infile);
}

# As I said at the beginning, AWK programs consist of a collection of patterns
# and actions. You've already seen the BEGIN pattern. Other
# patterns are used only if you're processing lines from files or standard
# input.
#
# When you pass arguments to AWK, they are treated as file names to process.
# It will process them all, in order. Think of it like an implicit for loop,
# iterating over the lines in these files. these patterns and actions are like
# switch statements inside the loop. 

/^fo+bar$/ {

    # This action will execute for every line that matches the regular
    # expression, /^fo+bar$/, and will be skipped for any line that fails to
    # match it. Let's just print the line:

    print;

    # Whoa, no argument! That's because print has a default argument: $0.
    # $0 is the name of the current line being processed. It is created
    # automatically for you.

    # You can probably guess there are other $ variables. Every line is
    # implicitly split before every action is called, much like the shell
    # does. And, like the shell, each field can be access with a dollar sign

    # This will print the second and fourth fields in the line
    print $2, $4;

    # AWK automatically defines many other variables to help you inspect and
    # process each line. The most important one is NF

    # Prints the number of fields on this line
    print NF;

    # Print the last field on this line
    print $NF;
}

# Every pattern is actually a true/false test. The regular expression in the
# last pattern is also a true/false test, but part of it was hidden. If you
# don't give it a string to test, it will assume $0, the line that it's
# currently processing. Thus, the complete version of it is this:

$0 ~ /^fo+bar$/ {
    print "Equivalent to the last pattern";
}

a > 0 {
    # This will execute once for each line, as long as a is positive
}

# You get the idea. Processing text files, reading in a line at a time, and
# doing something with it, particularly splitting on a delimiter, is so common
# in UNIX that AWK is a scripting language that does all of it for you, without
# you needing to ask. All you have to do is write the patterns and actions
# based on what you expect of the input, and what you want to do with it.

# Here's a quick example of a simple script, the sort of thing AWK is perfect
# for. It will read a name from standard input and then will print the average
# age of everyone with that first name. Let's say you supply as an argument the
# name of a this data file:
#
# Bob Jones 32
# Jane Doe 22
# Steve Stevens 83
# Bob Smith 29
# Bob Barker 72
#
# Here's the script:

BEGIN {

    # First, ask the user for the name
    print "What name would you like the average age for?";

    # Get a line from standard input, not from files on the command line
    getline name < "/dev/stdin";
}

# Now, match every line whose first field is the given name
$1 == name {

    # Inside here, we have access to a number of useful variables, already
    # pre-loaded for us:
    # $0 is the entire line
    # $3 is the third field, the age, which is what we're interested in here
    # NF is the number of fields, which should be 3
    # NR is the number of records (lines) seen so far
    # FILENAME is the name of the file being processed
    # FS is the field separator being used, which is " " here
    # ...etc. There are plenty more, documented in the man page.

    # Keep track of a running total and how many lines matched
    sum += $3;
    nlines++;
}

# Another special pattern is called END. It will run after processing all the
# text files. Unlike BEGIN, it will only run if you've given it input to
# process. It will run after all the files have been read and processed
# according to the rules and actions you've provided. The purpose of it is
# usually to output some kind of final report, or do something with the
# aggregate of the data you've accumulated over the course of the script.

END {
    if (nlines)
        print "The average age for " name " is " sum / nlines;
}

Project Fellow at Institute of Himalayan Bioresource Technology

Fri, 15 Aug 2014 06:50:08 -0500

Research Associate/ Project FellowDate of posting:14 Aug

Eligibility : MSc, M Phil / Phd, BE/B.Tech
Location : Himachal Pradesh-other
Job Category : Govt Jobs, Research, Walkin
Last Date : 20 Aug 2014

Advertisement No.6/2014

Post : Project Fellow
Research Associate/ Project Fellow Jobs opportunity in CSIR-Institute of Himalayan Bioresource Technology
M.Sc. in Bioinformatics/Computer Science with 55% marks and (ii) M.Sc. Bioinformatics/ Computational biology/ P.G. Diploma in Bioinformatics/B.Tech. or higher Degree in Bioinformatics with 55% marks

Date of Interview: 29.08.2014.

More at http://www.ihbt.res.in/recruit/AdvtNo6_2014.pdf

Snakemake Tutorials !

Rahul Nayak — Mon, 09 May 2022 05:20:41 -0500

A lesson introducing the Snakemake workflow system for bioinformatics analysis.

Prerequisites

This is an intermediate lesson and assumes learners have already done some bioinformatics:

Familiarity with the BASH command shell, including concepts like pipes, variables and loops.

Knowledge of bioinformatics fundamentals like the FASTQ file format and transcriptome sequencing, in order to understand the example workflow.

No previous knowledge of Snakemake or workflow systems is required.

https://carpentries-incubator.github.io/snakemake-novice-bioinformatics/index.html

Address of the bookmark: https://carpentries-incubator.github.io/snakemake-novice-bioinformatics/aio/index.html

Biology, Computers Collide in High-Demand Field of Bioinformatics

Mon, 25 Aug 2014 00:56:10 -0500

Dr. Shivas Amin calls bioinformatics a "collision of biology and computers." Students learn how to use computers and skills in math and biology to analyze genome and proteome projects to prepare for high-demand jobs in the life sciences. Learn more about Amin and hear from student Medina Baitemirova and alumnus Lukas Simon about the fast-growing field of bioinformatics.

A comprehensive atlas of human gene activity released !!!

Abhimanyu Singh — Tue, 02 Sep 2014 14:20:24 -0500

A large international consortium of researchers has produced the first comprehensive, detailed map of the way genes work across the major cells and tissues of the human body. The findings describe the complex networks that govern gene activity, and the new information could play a crucial role in identifying the genes involved with disease.

We are able to pinpoint the regions of the genome that can be active in a disease and in normal activity, whether it’s in a brain cell, the skin, in blood stem cells or in hair follicles. This is a major advance that will greatly increase our ability to understand the causes of disease across the body.

The research is outlined in a series of papers published March 27, 2014, two in the journal Nature and 16 in other scholarly journals. The work is the result of years of concerted effort among 250 experts from more than 20 countries as part of FANTOM 5 (Functional Annotation of the Mammalian Genome). The FANTOM project, led by the Japanese institution RIKEN, is aimed at building a complete library of human genes.

Researchers studied human and mouse cells using a new technology called Cap Analysis of Gene Expression (CAGE), developed at RIKEN, to discover how 95% of all human genes are switched on and off. These “switches” — called “promoters” and “enhancers” — are the regions of DNA that manage gene activity. The researchers mapped the activity of 180,000 promoters and 44,000 enhancers across a wide range of human cell types and tissues and, in most cases, found they were linked with specific cell types.

Referene : www.kurzweilai.net/first-comprehensive-atlas-of-human-gene-activity-released