BOL: Related items

SNP Analysis: Unlocking the Secrets in Our DNA

Abhi — Wed, 16 Jul 2025 01:31:45 -0500

Single Nucleotide Polymorphisms (SNPs) are the most common type of genetic variation in humans—and many other organisms. A single base change in the DNA sequence (for example, an A instead of a G) can influence everything from our eye color to our risk of developing diseases. Analyzing these tiny changes has become central to modern genetics, medicine, agriculture, and evolutionary biology.

What are SNPs?
SNPs (pronounced "snips") are positions in the genome where individuals differ by a single nucleotide. For example:

Reference: ...A T G C A T G A...
Variant: ...A T G T A T G A...

Here, the C in the reference genome has been replaced by a T in the variant.

SNPs occur roughly every 300–1,000 bases in the human genome, meaning there are millions of them scattered throughout our DNA. Most SNPs have no effect on health, but some are linked to disease susceptibility, drug response, and other traits.

Why Do We Analyze SNPs?
1. Medical Genetics

Identify disease-associated variants (e.g., BRCA1/2 in breast cancer).

Predict drug response (pharmacogenomics).

Enable precision medicine by tailoring treatments.

2. Population Genetics & Ancestry

Trace human migration and ancestry.

Study genetic diversity within and between populations.

3. Agriculture & Animal Breeding

Select for desirable traits (drought resistance, yield, disease resistance).

Improve breeding efficiency in livestock.

4. Evolutionary Biology

Track natural selection.

Study adaptation in wild populations.

How is SNP Analysis Performed?
SNP analysis can be broadly divided into three steps:

SNP Detection
Genotyping arrays: Chips that test hundreds of thousands of known SNP positions simultaneously. Fast and affordable, widely used in consumer ancestry testing.

Whole-genome or whole-exome sequencing: Can detect known and novel SNPs across the genome.

Targeted sequencing or PCR: For focused analysis of specific regions.

Variant Calling
Sequencing data is aligned to a reference genome. Bioinformatics tools (e.g., GATK, bcftools) identify positions where the sequenced sample differs from the reference.

Annotation and Interpretation
Tools (e.g., SnpEff, VEP) predict the functional impact of SNPs.

Are the SNPs in coding regions? Do they cause amino acid changes? Are they known to be pathogenic?

Databases like dbSNP, ClinVar, and GWAS Catalog provide information on known associations.

Common Tools for SNP Analysis
Alignment: BWA, Bowtie2

Variant Calling: GATK, FreeBayes

Visualization: IGV, UCSC Genome Browser

Annotation: SnpEff, VEP

Statistical Analysis: PLINK, SNPTEST

Challenges in SNP Analysis
False positives/negatives: Sequencing errors, alignment issues.

Population stratification: Confounding in association studies.

Interpretation: Many SNPs have unknown or complex effects.

Researchers address these with rigorous quality control, large datasets, and increasingly sophisticated statistical models.

The Future of SNP Analysis
With advances in sequencing technology and AI-driven analysis, SNP studies are expanding:

Polygenic risk scores predict disease risk based on thousands of SNPs.

Large-scale biobanks (e.g., UK Biobank, All of Us) enable powerful genome-wide association studies (GWAS).

CRISPR and functional assays help validate SNP effects in the lab.

SNP analysis is at the heart of the genomic revolution, promising insights into biology, health, and evolution at unprecedented scale.

Conclusion
From diagnosing rare diseases to designing better crops, SNP analysis is a foundational tool in modern science. As our ability to sequence and interpret genomes improves, so will our understanding of these tiny—but mighty—variations in DNA.

Single Cell RNAseq data analysis tutorial !!

Robert M Willioms — Mon, 27 Nov 2017 16:24:29 -0600

A major breakthrough (replaced microarrays) in the late 00’s and has been widely used since
Measures the average expression level for each gene across a large population of input cells
Useful for comparative transcriptomics, e.g. samples of the same tissue from different species
Useful for quantifying expression signatures from ensembles, e.g. in disease studies
Insufficient for studying heterogeneous systems, e.g. early development studies, complex tissues (brain)
Does not provide insights into the stochastic nature of gene expression

Following are the useful links:

Single Cell RNAseq data analysis Tutorial

A step-by-step workflow for low-level analysis of single-cell RNA-seq data

A step-by-step workflow for low-level analysis of single-cell RNA-seq data with Bioconductor

SCell: single-cell RNA-seq analysis software

https://github.com/diazlab/SCell

Beta-Poisson model for single-cell RNA-seq data analyses

https://github.com/nghiavtr/BPSC

Sincera: A Computational Pipeline for Single Cell RNA-Seq Profiling Analysis

https://research.cchmc.org/pbge/sincera.html

SC3 – consensus clustering of single-cell RNA-Seq data

http://biorxiv.org/content/early/2016/09/02/036558

Citrus: A toolkit for single cell sequencing analysis

http://biorxiv.org/content/early/2016/09/14/045070

Single-Cell Resolution of Temporal Gene Expression during Heart Development

http://www.cell.com/developmental-cell/fulltext/S1534-5807(16)30682-7

Scalable latent-factor models applied to single-cell RNA-seq data separate biological drivers from confounding effects

http://biorxiv.org/content/early/2016/11/15/087775

Single cell transcriptomes identify human islet cell signatures and reveal cell-type-specific expression changes in type 2 diabetes

http://genome.cshlp.org/content/early/2016/11/18/gr.212720.116.abstract

SCODE: An efficient regulatory network inference algorithm from single-cell RNA-Seq during differentiation

http://biorxiv.org/content/early/2016/11/21/088856

SCOUP is a probabilistic model to analyze single-cell expression data during differentiation

https://github.com/hmatsu1226/SCOUP

scLVM is a modelling framework for single-cell RNA-seq data

https://github.com/PMBio/scLVM

Selective Locally linear Inference of Cellular Expression Relationships (SLICER) algorithm for inferring cell trajectories

https://github.com/jw156605/SLICER

SinQC: A Method and Tool to Control Single-cell RNA-seq Data Quality

http://www.morgridge.net/SinQC.html

TSCAN: Pseudo-time reconstruction and evaluation in single-cell RNA-seq analysis

https://github.com/zji90/TSCAN

Visualization and cellular hierarchy inference of single-cell data using SPADE

http://www.nature.com/nprot/journal/v11/n7/full/nprot.2016.066.html

OEFinder: Identify ordering effect genes in single cell RNA-seq data

https://github.com/lengning/OEFinder

Machine learning training and courses in bioinformatics !

Rahul Nayak — Tue, 31 Dec 2019 19:33:07 -0600

Machine learning techniques have been successful in analyzing biological data because of their capabilities in handling randomness and uncertainty of data noise and in generalization. In this class, we will learn basics about probabilistic models and machine learning techniques. We will focus on probabilistic models (Markov models, Hidden Markov models, and Bayesian networks) for biological sequence analysis and systems biology. Other machine learning techniques, such as Naive bayes, neural networks and SVMs will only be covered briefly.

More at http://homes.sice.indiana.edu/yye/lab/teaching/spring2017-I529/

An Introduction to Applied Bioinformatics

Jit — Fri, 02 Mar 2018 04:26:38 -0600

IAB is primarily being developed by Greg Caporaso(GitHub/Twitter: @gregcaporaso) in the Caporaso Lab at Northern Arizona University. You can find information on the courses I teach on my teaching website and information on my research and lab on my lab website.

Address of the bookmark: http://readiab.org/

Bioinformatics in Africa: Part3 - Mali

BioStar — Sat, 06 Feb 2021 13:28:44 -0600

International Center for Excellence in Research (ICER):

The ICER is a research center composed of the following three programs: 1. The Malaria Research and Training Center Parasitology Group, 2. The Malaria Research and Training Center Entomology Group 3. The SEREFO Group.

The first two programs develop biomedical researches in malaria, Filariasis and Leishmaniasis. The third program develops biomedical researches in tuberculosis and HIV.

Bioinformatics was introduced recently to the ICER and is constantly growing. The ICER has one team, headed by Sidy SOUMARE, which supports the three programmes in all their needs in informatics and bioinformatics. This team can beneficiate from some computational facilities (4 blast servers, 15 other servers and around 200 terminals), but the ICER staff needs some training in order to be able to administrate these facilities.

Research Interest and Activities: The following are the present areas of research interest: 1. Functional genomics 2. Analysis of microarray data 3. Interaction between the vector and the parasite.

Bioinformatics in Africa: Part7 - Tunisia

BioStar — Sat, 06 Feb 2021 21:25:09 -0600

Institut Pasteur de Tunis (IPT):
The IPT is a research institution founded in 1883. IPT is under the supervision of the Ministry of Health and is part of the Université El Manar of Tunis (Ministry of high Education). The missions of the institute are: Public Health Laboratory activities (PHL), Research on infectious diseases, and R/D on vaccines. Research programs are mainly oriented towards local health problems such as leishmaniais, viral hepatitis, and scorpion venoms. The group of Bioinformatics and Modelling of the IPT is hosted by the Laboratoire d’Immunopathologie Vaccinologie et Génétique Moléculaire (LIVGM), and exists since the beginning of 2005. Its present research activities include: genome annotation, EST clustering and modelling of the host/parasite response to Leishmania infection. It consists of two senior scientists, two PhD students and one MSc student

Centre de Biotechnology de Sfax (CBS):
Bioinformatics activity started at CBS in 2001 with the settingup of a research and service unit of bioinformatics. This unit currently includes one senior researcher, one engineer and four Phd students. Activities include sequence annotation (service) and three research programs: ab initio prediction of short eukaryote genes, statistical modelling by Bayesian networks approach of signal transduction pathways and statistical analysis of human sequence variation data (haplotype reconstruction and linkage disequilibrium). Activities of the Bioinformatics unit could be found at the website: http://www.cbs.rnrt.tn/ and the research activity report is available under request to Bioinformatics@cbs.rnrt.tn. Although the computing facilities are good, there is still a need for trained human resources to strengthen bioinformatics capacities at CBS, particularly in structural bioinformatics.

Web site and links: http://www.cbs.rnrt.tn

Basic Structure of Snakemake Pipeline Run !

Abhi — Thu, 14 Oct 2021 07:01:38 -0500

/user/snakemake-demo$ ls

config.json data envs scripts slurm-240702.out Snakefile

data = mock data for the snakefile to use
Snakefile = name of the snakemake “formula” file
- Note: The default file that snakemake looks for in the current working directory is the Snakefile. If you would like to override that you can specify it following the -s
  - snakemake -s snakefile.py
envs = directory for storing the conda environments that the workflow will use.
scripts = directory for storing python scripts called by the snakemake formula.
config.json = json format file with extra parameters for our snakemake file to use.
cluster.json = json format file with specification for running on the HPC
samples.txt = file we will use later relating to the config.json file.

Run the snakemake file as a dry run (the example workflow shown above).

This will build a DAG of the jobs to be run without actually executing them.
snakemake --dry-run

User can execute rules of interest.

snakemake --dry-run all VS. snakemake --dry-run call VS. snakemake --dry-run bwa

Run the snakemake file in order to produce an image of the DAG of jobs to be run.

snakemake --dag | dot -Tsvg > dag.svg OR snakemake --dag | dot -Tsvg > dag.svg

Run the snakemake (this time not as a dry run)

snakemake --use-conda

Prepare for Coding Interview !

Abhi — Tue, 11 Jan 2022 06:14:41 -0600

This is a comprehensive guide to prepare for your next coding interview. It's great for recent graduates and has questions and practice materials structured from traditional big tech interview formats.

While it does not include the latest developments in programming since 2019, it nails the core fundamentals in a very comprehensive and accessible way!

Credits to Kaiyu Zhang, with additional material in the appendix sourced from Reddit.

People say that interviews at Google will cover as much ground as possible. As a new college graduate, the ground that I must capture are the following. Part of the list is borrowed from a reddit post: https://www. reddit.com/r/cscareerquestions/comments/206ajq/my_onsite_interview_experience_at_google/ #bottom-comments.

1. Data structures

2. Trees and Graph algorithms

3. Dynamic Programming

4. Recursive algorithms

5. Scheduling algorithms (Greedy)

6. Caching 1

7. Sorting

8. Files

9. Computability

10. Bitwise operators

11. System design

Online resources on must-read papers in evolutionary biology

BioStar — Fri, 26 Jul 2024 01:39:14 -0500

Online resources on must-read papers in evolutionary biology, for a literature club.

Below is a summary of all answers that we received.

All the best,

Jana and Xiaoyan

1.       *Nick Barton:*

- The textbook "Evolution" by Nick Barton, with resources for
  exploring the literature: Barton, N. H., Briggs, D. E. G., Eisen, J.
  A., Goldstein, D. B., & Patel, N. H. (2007). Evolution. Cold Spring
  Harbor Laboratory Press.

- Papers from a course named "Classics in Evolutionary Biology":

Evolutionary Synthesis
1. Haldane, J. B. S. 1932. The causes of evolution. Longmans. New York.
   (esp. Ch. IV).
2. Fisher, R. A. 1930. The genetical theory of natural selection. Oxford
   University Press, Oxford. Selected Sections - Fundamental Theorem.

Genetic Variation
1a. Lewontin, R. C., and J. L. Hubby. 1966. A molecular approach to
the study of genic heterozygosity in natural populations. II. Amount
of variation and degree of heterozygosity in natural populations of
Drosophila pseudoobscura. Genetics. 54:595-609.

1b. Sachidandam et al. 2001. A map of human genome sequence variation
containing 1.42 million single nucleotide polymorphisms. 409: 928-33.

2. Wright S., Dobzhansky T., Hovanitz W. 1942 Genetics of natural
populations VII The allelism of lethals in the third chromosome of
Drosophila pseudoobscura. Genetics 27: 363-394.

Recombination and evolution
1. Hill, W. G., and A. Robertson. 1966. The effect of linkage on limits
to artificial selection. Genet. Res. 8:269-294.

2. Maynard Smith and Haigh. 1974. The hitch-hiking effect of a favourable
gene. Genet. Res. 23: 23-35.

Understanding sequence variation
1. Begun D. J., Aquadro C. F., 1992 Levels of naturally occurring DNA
polymorphism correlate with recombination rate in Drosophila melanogaster.
Nature 356: 519-520.

2. Green R. E., Reich D., Pääbo S., 2010 A draft sequence of the
Neandertal genome. Science 328: 710-722.

Quantitative Genetics:  variation in complex traits
1. Galton F., 1877 Typical laws of heredity. Nature 15: 492-495-
512-514- 532-533.

2. Turelli M., 1984 Heritable genetic variation via
mutation-selection balance: Lerch's Zeta meets the abdominal
bristle. Theor. Popul. Biol. 25: 138-193.

Quantitative Genetics:  finding the genes
1. Shrimpton A. E., Robertson A., 1988 The Isolation of polygenic factors
controlling bristle score in Drosophila melanogaster II Distribution of
third chromosome bristle effects within chromosome sections. Genetics
118: 445-459.

2. Boyle E. A., Li Y. I., Pritchard J. K., 2017 An expanded view of
complex traits: from polygenic to omnigenic. Cell 169: 1177-1186.

Neutral Evolution
1. Kimura, M. 1968. Evolutionary rate at the molecular level. Science.
217:624-626.

2a. Kern A. D., Hahn M. W., 2018 The Neutral Theory in Light of Natural
Selection. Molecular Biology and Evolution 110: 21077-6.

2b. Jensen J. D., Payseur B. A., Stephan W., Aquadro C. F., Lynch M.,
Charlesworth D., Charlesworth B., 2018 The importance of the Neutral Theory
in 1968 and 50 years on: a response to Kern and Hahn 2018. Evolution 112:
2109-4.

2c. Ellegren & Galtier. 2016. Determinants of genetic diversity. Nature
Reviews Genetics.

Mutation and Genetic Variability
1. Luria, S. E., and M. Delbrück. 1943. Mutations of Bacteria from Virus
Sensitivity to Virus Resistance. Genetics. 28(6):491-511.

2. Hill, W G. 1982. "Rates of Change in Quantitative Traits From Fixation
of New Mutations." Proceedings of the National Academy of Sciences (U.S.A.)
79: 142-45.

Testing for selection
1. McDonald & Kreitman. 1991. Adaptive protein evolution at the Adh locus
in Drosophila. Nature.

2. Begun, et al. Mol. Biol. Evol. 16, 1816-1819 (1999).

3. Siddiq et al. 2016. Experimental test and refutation of a classic case
of molecular adaptation in Drosophila melanogaster.  Nature Ecology &
Evolution.

The shifting balance
1. Wright, S. 1932. The roles of mutation, inbreeding, crossbreeding and
selection in evolution. Proceedings of the VI International Congress of
Genetics: 1. pp 356-366.

2. Coyne, J.A., N.H. Barton, and M. Turelli. 1997. A critique of Wright's
shifting balance theory of evolution.  Evolution 51: 643-671.

3. Barton. 2016. Sewall Wright on Evolution in Mendelian Populations and
the "Shifting Balance". Genetics.

Evolution of Sex
1.  Muller, H.J. 1964. The relation of recombination to mutational advance.
Mutation Res. 1(1):2-9

2. McDonald et al. 2016. Sex speeds adaptation by altering the dynamics of
molecular evolution. Nature.

Kin Selection, Cooperation, and Conflict
1. Hamilton, W. D. 1964. The genetical evolution of social behaviour I.
Journal of Theoretical Biology. 7:1-52.

2. Trivers, R. L. 1974 Parent-offspring conflict. American Zoologist.
14(1):249-264.

Sexual Selection
1. Zahavi, A. 1975. Mate selection - a selection of a handicap. J. Theor.
Biol. 53:205-214.

2. Kirkpatrick, M., and Ryan, M.J. 1991. The evolution of mating
preferences and the paradox of the lek. Nature. 350:33-38.

Fitness Landscapes
1. Dean, A. 1995. A Molecular Investigation of Genotype by Environment
Interactions. Genetics. 139:19-33.

2. Costanzo et al. 2010. The Genetic Landscape of a Cell. Science.

Speciation
1. Coyne, J. A., and H. A. Orr. 1989. Patterns of speciation in Drosophila.
Evolution. 43:362-381.

2. Corbett-Detig et al. 2013. Genetic incompatibilities are widespread
within species. Nature.

2.       *Marcos Antezana:*

Valen, L. v. 1975. Energy and Evolution. University of Chicago, Department
of Biology.

3.       *Remco Folkertsma:*

1. The work by Hopi Hoekstra on local adaptation and oldfield mice

2. Poelstra, J. W., Vijay, N., Bossu, C. M., Lantz, H., Ryll, B., Müller,
I., ... & Wolf, J. B. (2014). The genomic landscape underlying phenotypic
integrity in the face of gene flow in crows. Science, 344(6190), 1410-1414.

4.       *Joshka Kaufmann and Leslie Turner*

They offer us a link to 'papers every evolutionary biologist should read',
the papers are collected by Leslie Turner.
https://static1.squarespace.com/static/53e8cb7ce4b02c4bc3aeeee4/t/5ab8fcb670a6ad55c67fcdf4/1522072758665/EvoBioClassicsRefList.pdf

5.       *Sarah Stockwell*

Matt Ridley collected classic papers in evolutionary biology and printed
part of these papers in his book Evolution (see Matt Ridley. Evolution
(Univ. of Oxford Press, 2nd edition, 2004))

Awk for Bioinformatician and computational biologist

Poonam Mahapatra — Tue, 06 Feb 2018 14:54:35 -0600

Awk is a programming language which allows easy manipulation of structured data and is mostly used for pattern scanning and processing. It searches one or more files to see if they contain lines that match with the specified patterns and then perform associated actions. The basic syntax is:

awk '/pattern1/ {Actions}
/pattern2/ {Actions}' file

The working of Awk is as follows
Awk reads the input files one line at a time.
For each line, it matches with given pattern in the given order, if matches performs the corresponding action.
If no pattern matches, no action will be performed.
In the above syntax, either search pattern or action are optional, But not both.
If the search pattern is not given, then Awk performs the given actions for each line of the input.
If the action is not given, print all that lines that matches with the given patterns which is the default action.
Empty braces with out any action does nothing. It wont perform default printing operation.
Each statement in Actions should be delimited by semicolon.
Say you have data.tsv with the following contents:

$ cat data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3 ACTTATATATATATA
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
By default Awk prints every line from the file.

$ awk '{print;}' data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3 ACTTATATATATATA
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
We print the line which matches the pattern contig3

$ awk '/contig3/' data/test.tsv
contig3 ACTTATATATATATA
Awk has number of builtin variables. For each record i.e line, it splits the record delimited by whitespace character by default and stores it in the $n variables. If the line has 5 words, it will be stored in $1, $2, $3, $4 and $5. $0 represents the whole line. NF is a builtin variable which represents the total number of fields in a record.

$ awk '{print $1","$2;}' data/test.tsv
contig1,ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2,ACTTTATATATT
contig3,ACTTATATATATATA
contig4,ACTTATATATATATA
contig5,ACTTTATATATT

$ awk '{print $1","$NF;}' data/test.tsv
contig1,ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2,ACTTTATATATT
contig3,ACTTATATATATATA
contig4,ACTTATATATATATA
contig5,ACTTTATATATT

Awk has two important patterns which are specified by the keyword called BEGIN and END. The syntax is as follows:

BEGIN { Actions before reading the file}
{Actions for everyline in the file}
END { Actions after reading the file }

For example,
$ awk 'BEGIN{print "Header,Sequence"}{print $1","$2;}END{print "-------"}' data/test.tsv
Header,Sequence
contig1,ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2,ACTTTATATATT
contig3,ACTTATATATATATA
contig4,ACTTATATATATATA
contig5,ACTTTATATATT
-------
We can also use the concept of a conditional operator in print statement of the form print CONDITION ? PRINT_IF_TRUE_TEXT : PRINT_IF_FALSE_TEXT. For example, in the code below, we identify sequences with lengths > 14:

$ awk '{print (length($2)>14) ? $0">14" : $0"<=14";}' data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG>14
contig2 ACTTTATATATT<=14
contig3 ACTTATATATATATA>14
contig4 ACTTATATATATATA>14
contig5 ACTTTATATATT<=14
We can also use 1 after the last block {} to print everything (1 is a shorthand notation for {print $0} which becomes {print} as without any argument print will print $0 by default), and within this block, we can change $0, for example to assign the first field to $0 for third line (NR==3), we can use:

$ awk 'NR==3{$0=$1}1' data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
You can have as many blocks as you want and they will be executed on each line in the order they appear, for example, if we want to print $1 three times (here we are using printf instead of print as the former doesn't put end-of-line character),

$ awk '{printf $1"\t"}{printf $1"\t"}{print $1}' data/test.tsv
contig1 contig1 contig1
contig2 contig2 contig2
contig3 contig3 contig3
contig4 contig4 contig4
contig5 contig5 contig5
Although, we can also skip executing later blocks for a given line by using next keyword:

$ awk '{printf $1"\t"}NR==3{print "";next}{print $1}' data/test.tsv
contig1 contig1
contig2 contig2
contig3
contig4 contig4
contig5 contig5

$ awk 'NR==3{print "";next}{printf $1"\t"}{print $1}' data/test.tsv
contig1 contig1
contig2 contig2

contig4 contig4
contig5 contig5
You can also use getline to load the contents of another file in addition to the one you are reading, for example, in the statement given below, the while loop will load each line from test.tsv into k until no more lines are to be read:

$ awk 'BEGIN{while((getline k <"data/test.tsv")>0) print "BEGIN:"k}{print}' data/test.tsv
BEGIN:contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
BEGIN:contig2 ACTTTATATATT
BEGIN:contig3 ACTTATATATATATA
BEGIN:contig4 ACTTATATATATATA
BEGIN:contig5 ACTTTATATATT
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3 ACTTATATATATATA
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
You can also store data in the memory with the syntax VARIABLE_NAME[KEY]=VALUE which you can later use through for (INDEX in VARIABLE_NAME) command:

$ awk '{i[$1]=1}END{for (j in i) print j"<="i[j]}' data/test.tsv
contig1<=1
contig2<=1
contig3<=1
contig4<=1
contig5<=1