BOL: Related items

Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data

Rahul Nayak — Mon, 13 Nov 2017 05:10:23 -0600

Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data
AfterQC can simply go through all fastq files in a folder and then output three folders: good, bad and QC folders, which contains good reads, bad reads and the QC results of each fastq file/pair.
Currently it supports processing data from HiSeq 2000/2500/3000/4000, Nextseq 500/550, MiniSeq...and other Illumina 1.8 or newer formats

Address of the bookmark: https://github.com/OpenGene/AfterQC

Most Commonly used Awk by Bioinformatician

Neel — Mon, 19 Aug 2013 01:12:38 -0500

Awk is a programming language that is specifically designed for quickly manipulating space delimited data. Although you can achieve all its functionality with Perl, awk is simpler in many practical cases.

Why awk? You can replace a pipeline of 'stuff | grep | sed | cut...' with a single call to awk. For a simple script, most of the timelag is in loading these apps into memory, and it's much faster to do it all with one. This is ideal for something like an openbox pipe menu where you want to generate something on the fly. You can use awk to make a neat one-liner for some quick job in the terminal, or build an awk section into a shell script. You can find a lot of online tutorials, but here I will only show a few examples which cover most of bioinformatician daily uses of awk.

choose rows where column 3 is larger than column 5:

awk '$3>$5' input.txt > output.txt

extract column 2,4,5:

awk '{print $2,$4,$5}' input.txt > output.txt

awk 'BEGIN{OFS="\t"}{print $2,$4,$5}' input.txt

show rows between 20th and 80th:

awk 'NR>=20&&NR<=80' input.txt > output.txt

calculate the average of column 2:

awk '{x+=$2}END{print x/NR}' input.txt

regex (egrep):

awk '/^test[0-9]+/' input.txt

calculate the sum of column 2 and 3 and put it at the end of a row or replace the first column:

awk '{print $0,$2+$3}' input.txt

awk '{$1=$2+$3;print}' input.txt

join two files on column 1:

awk 'BEGIN{while((getline<"file1.txt")>0)l[$1]=$0}$1 in l{print $0"\t"l[$1]}' file2.txt > output.txt

count number of occurrence of column 2 (uniq -c):

awk '{l[$2]++}END{for (x in l) print x,l[x]}' input.txt

apply "uniq" on column 2, only printing the first occurrence (uniq):

awk '!($2 in l){print;l[$2]=1}' input.txt

count different words (wc):

awk '{for(i=1;i!=NF;++i)c[$i]++}END{for (x in c) print x,c[x]}' input.txt

deal with simple CSV:

awk -F, '{print $1,$2}'

substitution (sed is simpler in this case):

awk '{sub(/test/, "no", $0);print}' input.txt

OK now here's where to read this stuff properly explained. roll

Two thorough tutorials:

http://www.gnu.org/software/gawk/manual/gawk.html

http://www.grymoire.com/Unix/Awk.html

A famous list of useful one-liners - though they're short, many are quite tricky:

http://www.pement.org/awk/awk1line.txt

And some nice explanations of those one-liners. After reading this you'll have a pretty good grasp!

http://www.catonmat.net/blog/awk-one-li … -part-one/

http://www.catonmat.net/blog/ten-awk-ti … -pitfalls/

Convert VCF to tab-deilimited table

Seema Singh — Tue, 15 May 2018 07:39:08 -0500

Performed with GATK :

java -Xmx8g -jar GenomeAnalysisTK.jar \
-T VariantsToTable \
-R reference.fa \
-V reference_genomes_GT.vcf \
-F CHROM -F POS -F REF -F ALT -GF GT \
-o reference_genomes_GT.table
multiple_sample.vcf should also be converted to multiple_sample_GT.table using this approach.

Tools for id conversion

LEGE — Sat, 20 May 2023 21:53:32 -0500

g:Convert enables to convert between various gene, protein, microarray probe and numerous other types of namespaces. We provide at least 40 types of IDs for more than 60 species. The 98 different namespaces supported for human include Ensembl, Refseq, Illumina, Entrezgene and Uniprot identifiers. All namespaces are obtained through matching them via Ensembl gene identifiers as a reference.

Address of the bookmark: https://biit.cs.ut.ee/gprofiler/convert

MCAT: Motif Combining and Association Tool

Neel — Sun, 13 Jan 2019 06:27:28 -0600

This is a pipeline for finding motifs in fasta files.
It can be run from the command line as follows:

usage: orange_pipeline_refine.py [-h] [-w W] [--nmotifs NMOTIFS] [--iter ITER] [-c C]
[-s S] [-d] [-ff] [-v V]
positive_seq negative_seq

positional arguments:
positive_seq the fasta file for the positive sequences
negative_seq the fasta file for the negative sequences

Address of the bookmark: https://github.com/yanshen43/MCAT

ALE: a Generic Assembly Likelihood Evaluation Framework for Assessing the Accuracy of Genome and Metagenome Assemblies

Neel — Tue, 26 Apr 2016 03:38:43 -0500

Assembly Likelihood Evaluation (ALE) framework that overcomes these limitations, systematically evaluating the accuracy of an assembly in a reference-independent manner using rigorous statistical methods. This framework is comprehensive, and integrates read quality, mate pair orientation and insert length (for paired-end reads), sequencing coverage, read alignment and k-mer frequency. ALE pinpoints synthetic errors in both single and metagenomic assemblies, including single-base errors, insertions/deletions, genome rearrangements and chimeric assemblies presented in metagenomes. At the genome level with real-world data, ALE identifies three large misassemblies from the Spirochaeta smaragdinae finished genome, which were all independently validated by Pacific Biosciences sequencing. At the single-base level with Illumina data, ALE recovers 215 of 222 (97%) single nucleotide variants in a training set from a GC-rich Rhodobacter sphaeroides genome. Using real Pacific Biosciences data, ALE identifies 12 of 12 synthetic errors in a Lambda Phage genome, surpassing even Pacific Biosciences' own variant caller, EviCons. In summary, the ALE framework provides a comprehensive, reference-independent and statistically rigorous measure of single genome and metagenome assembly accuracy, which can be used to identify misassemblies or to optimize the assembly process.

More at http://www.ncbi.nlm.nih.gov/pubmed/23303509

Address of the bookmark: http://sc932.github.io/ALE/about.html

FQC Dashboard: Integrates FastQC results into a web-based, interactive, and extensible FASTQ quality control tool

Shruti Paniwala — Tue, 10 Nov 2020 01:30:22 -0600

FQC is software that facilitates quality control of FASTQ files by carrying out a QC protocol using FastQC, parsing results, and aggregating quality metrics into an interactive dashboard designed to richly summarize individual sequencing runs. The dashboard groups samples in dropdowns for navigation among the data sets, utilizes human-readable configuration files to manipulate the pages and tabs, and is extensible with CSV data.

Address of the bookmark: https://github.com/pnnl/fqc

WGS Celera Assembler version 8.3rc2

Jit — Mon, 10 Apr 2017 04:45:40 -0500

These are release notes for Celera Assembler version 8.3rc2, which was released on May 24, 2015.

This distribution package provides a stable, tested, documented version of the software. The distribution is usable on most Unix-like platforms, and some platforms have pre-compiled binary distributions ready for installation.

The source code package includes full source code (revision 4627), Makefiles, and scripts. A subset of the kmer package (http://kmer.sourceforge.net/, version r1994), used by some modules of Celera Assembler, is included. This distribution includes [http://samtools.sourceforge.net/ SAMtools], [http://www.cbcb.umd.edu/software/jellyfish/ Jellyfish 2.0], [https://github.com/pbjd/pbutgcns PBUTGCNS], [https://github.com/PacificBiosciences/pbdagcon PBDAGCON], [https://github.com/PacificBiosciences/BLASR BLASR], and parts of the [https://github.com/PacificBiosciences/FALCON/tree/v0.1.3 Falcon assembler].

Full documentation can be found online at http://wgs-assembler.sourceforge.net/.

Interesting scripts within it

urbe@urbo214b[bin] ls []
-rwxrwxr-x 1 urbe urbe 11K Apr 10 11:41 addCNSToStore
-rwxrwxr-x 1 urbe urbe 575K Apr 10 11:41 addReadsToUnitigs
-rwxrwxr-x 1 urbe urbe 128K Apr 10 11:41 analyzeBest
-rwxrwxr-x 1 urbe urbe 257K Apr 10 11:41 analyzePosMap
-rwxrwxr-x 1 urbe urbe 1,5M Apr 10 11:41 analyzeScaffolds
-rwxrwxr-x 1 urbe urbe 224K Apr 10 11:41 asmOutputFasta
-rwxrwxr-x 1 urbe urbe 448K Apr 10 11:41 asmOutputStatistics
-rwxrwxr-x 1 urbe urbe 2,4K Apr 10 11:41 asmToAGP.pl
-rwxrwxr-x 1 urbe urbe 7,6M Apr 10 11:41 blasr
-rwxrwxr-x 1 urbe urbe 1,6M Apr 10 11:41 bogart
-rwxrwxr-x 1 urbe urbe 183K Apr 10 11:41 bogus
-rwxrwxr-x 1 urbe urbe 272K Apr 10 11:41 bogusness
-rwxrwxr-x 1 urbe urbe 247K Apr 10 11:41 buildPosMap
-rwxrwxr-x 1 urbe urbe 213K Apr 10 11:41 buildRefContigs
-rwxrwxr-x 1 urbe urbe 990K Apr 10 11:41 buildUnitigs
-rwxrwxr-x 1 urbe urbe 18K Apr 10 11:41 ca2ace.pl
-rwxrwxr-x 1 urbe urbe 12K Apr 10 11:41 caqc_help.ini
-rwxrwxr-x 1 urbe urbe 61K Apr 10 11:41 caqc.pl
-rwxrwxr-x 1 urbe urbe 23K Apr 10 11:41 cat-corrects
-rwxrwxr-x 1 urbe urbe 24K Apr 10 11:41 cat-erates
-rwxrwxr-x 1 urbe urbe 1,9M Apr 10 11:41 cgw
-rwxrwxr-x 1 urbe urbe 1,4M Apr 10 11:41 cgwDump
-rwxrwxr-x 1 urbe urbe 204K Apr 10 11:41 chimChe
-rwxrwxr-x 1 urbe urbe 201K Apr 10 11:40 chimera
-rwxrwxr-x 1 urbe urbe 220K Apr 10 11:41 classifyMates
-rwxrwxr-x 1 urbe urbe 201K Apr 10 11:41 classifyMatesApply
-rwxrwxr-x 1 urbe urbe 215K Apr 10 11:41 classifyMatesPairwise
-rwxrwxr-x 1 urbe urbe 366K Apr 10 11:41 computeCoverageStat
-rwxrwxr-x 1 urbe urbe 9,8K Apr 10 11:41 convert-fasta-to-v2.pl
-rwxrwxr-x 1 urbe urbe 48K Apr 10 11:41 convertOverlap
-rwxrwxr-x 1 urbe urbe 119K Apr 10 11:41 convertSamToCA
-rwxrwxr-x 1 urbe urbe 20K Apr 10 11:41 convertToPBCNS
-rwxrwxr-x 1 urbe urbe 197K Apr 10 11:41 correct-frags
-rwxrwxr-x 1 urbe urbe 259K Apr 10 11:41 correct-olaps
-rwxrwxr-x 1 urbe urbe 520K Apr 10 11:41 correctPacBio
-rwxrwxr-x 1 urbe urbe 540K Apr 10 11:41 ctgcns
-rwxrwxr-x 1 urbe urbe 162K Apr 10 11:40 deduplicate
-rwxrwxr-x 1 urbe urbe 37K Apr 10 11:41 demotePosMap
-rwxrwxr-x 1 urbe urbe 1,5M Apr 10 11:41 dumpCloneMiddles
-rwxrwxr-x 1 urbe urbe 124K Apr 10 11:41 dumpPBRLayoutStore
-rwxrwxr-x 1 urbe urbe 1,3M Apr 10 11:41 dumpSingletons
-rwxrwxr-x 1 urbe urbe 171K Apr 10 11:41 erate-estimate
-rwxrwxr-x 1 urbe urbe 221K Apr 10 11:40 estimate-mer-threshold
-rwxrwxr-x 1 urbe urbe 1,5M Apr 10 11:41 extendClearRanges
-rwxrwxr-x 1 urbe urbe 1,3M Apr 10 11:41 extendClearRangesPartition
-rwxrwxr-x 1 urbe urbe 205K Apr 10 11:40 extractmessages
-rwxrwxr-x 1 urbe urbe 7,2M Apr 10 11:41 falcon_sense
-rwxrwxr-x 1 urbe urbe 9,8K Apr 10 11:41 fastaToCA
-rwxrwxr-x 1 urbe urbe 124K Apr 10 11:40 fastqAnalyze
-rwxrwxr-x 1 urbe urbe 137K Apr 10 11:40 fastqSample
-rwxrwxr-x 1 urbe urbe 62K Apr 10 11:40 fastqSimulate
-rwxrwxr-x 1 urbe urbe 121K Apr 10 11:40 fastqSimulate-sort
-rwxrwxr-x 1 urbe urbe 246K Apr 10 11:40 fastqToCA
-rwxrwxr-x 1 urbe urbe 140K Apr 10 11:41 filterOverlap
-rwxrwxr-x 1 urbe urbe 341K Apr 10 11:40 finalTrim
-rwxrwxr-x 1 urbe urbe 228K Apr 10 11:41 fixUnitigs
-rwxrwxr-x 1 urbe urbe 147K Apr 10 11:40 fragmentDepth
-rwxrwxr-x 1 urbe urbe 29K Apr 10 11:41 fragsInVars
-rwxrwxr-x 1 urbe urbe 545K Apr 10 11:41 frgs2clones
-rwxrwxr-x 1 urbe urbe 398K Apr 10 11:40 gatekeeper
-rwxrwxr-x 1 urbe urbe 139K Apr 10 11:40 gatekeeperbench
-rwxrwxr-x 1 urbe urbe 167K Apr 10 11:40 gkpStoreCreate
-rwxrwxr-x 1 urbe urbe 147K Apr 10 11:40 gkpStoreDumpFASTQ
-rwxrwxr-x 1 urbe urbe 184K Apr 10 11:41 greedyFragmentTiling
-rwxrwxr-x 1 urbe urbe 1,6K Apr 10 11:41 greedy_layout_to_IUM
-rwxrwxr-x 1 urbe urbe 142K Apr 10 11:40 initialTrim
-rwxrwxr-x 1 urbe urbe 967K Apr 10 11:41 jellyfish
-rwxrwxr-x 1 urbe urbe 219K Apr 10 11:41 markRepeatUnique
-rwxrwxr-x 1 urbe urbe 273K Apr 10 11:40 markUniqueUnique
-rwxrwxr-x 1 urbe urbe 114K Apr 10 11:40 mercy
-rwxrwxr-x 1 urbe urbe 3,8K Apr 10 11:41 mergeqc.pl
-rwxrwxr-x 1 urbe urbe 422K Apr 10 11:40 merTrim
-rwxrwxr-x 1 urbe urbe 125K Apr 10 11:40 merTrimApply
-rwxrwxr-x 1 urbe urbe 376K Apr 10 11:40 meryl
-rwxrwxr-x 1 urbe urbe 176K Apr 10 11:41 metagenomics_ovl_analyses
-rwxrwxr-x 1 urbe urbe 297K Apr 10 11:41 olap-from-seeds
-rwxrwxr-x 1 urbe urbe 275K Apr 10 11:41 outputLayout
-rwxrwxr-x 1 urbe urbe 229K Apr 10 11:41 overlapInCore
-rwxrwxr-x 1 urbe urbe 144K Apr 10 11:40 overlap_partition
-rwxrwxr-x 1 urbe urbe 179K Apr 10 11:41 overlapStats
-rwxrwxr-x 1 urbe urbe 179K Apr 10 11:41 overlapStore
-rwxrwxr-x 1 urbe urbe 153K Apr 10 11:41 overlapStoreBucketizer
-rwxrwxr-x 1 urbe urbe 175K Apr 10 11:41 overlapStoreBuild
-rwxrwxr-x 1 urbe urbe 33K Apr 10 11:41 overlapStoreIndexer
-rwxrwxr-x 1 urbe urbe 48K Apr 10 11:41 overlapStoreSorter
-rwxrwxr-x 1 urbe urbe 604K Apr 10 11:40 overmerry
lrwxrwxrwx 1 urbe urbe 4 Apr 10 11:41 pacBioToCA -> PBcR
-rwxrwxr-x 1 urbe urbe 131K Apr 10 11:41 PBcR
-rwxrwxr-x 1 urbe urbe 2,9M Apr 10 11:41 pbdagcon
-rwxrwxr-x 1 urbe urbe 1,9M Apr 10 11:41 pbutgcns
-rwxrwxr-x 1 urbe urbe 201K Apr 10 11:40 remove_fragment
-rwxrwxr-x 1 urbe urbe 153K Apr 10 11:40 removeMateOverlap
-rwxrwxr-x 1 urbe urbe 2,5K Apr 10 11:41 replaceUIDwithName-fastq
-rwxrwxr-x 1 urbe urbe 1,2K Apr 10 11:41 replaceUIDwithName-posmap
-rwxrwxr-x 1 urbe urbe 1,3M Apr 10 11:41 resolveSurrogates
-rwxrwxr-x 1 urbe urbe 139K Apr 10 11:41 rewriteCache
-rwxrwxr-x 1 urbe urbe 232K Apr 10 11:41 runCA
-rwxrwxr-x 1 urbe urbe 88K Apr 10 11:41 runCA-dedupe
-rwxrwxr-x 1 urbe urbe 14K Apr 10 11:41 runCA-overlapStoreBuild
-rwxrwxr-x 1 urbe urbe 3,6K Apr 10 11:41 run_greedy.csh
-rwxrwxr-x 1 urbe urbe 297K Apr 10 11:40 sffToCA
-rwxrwxr-x 1 urbe urbe 13K Apr 10 11:40 show-corrects
-rwxrwxr-x 1 urbe urbe 557K Apr 10 11:41 splitUnitigs
-rwxrwxr-x 1 urbe urbe 1,4M Apr 10 11:41 terminator
drwxrwxr-x 2 urbe urbe 4,0K Apr 10 11:41 TIGR
-rwxrwxr-x 1 urbe urbe 526K Apr 10 11:41 tigStore
-rwxrwxr-x 1 urbe urbe 35K Apr 10 11:41 tracearchiveToCA
-rwxrwxr-x 1 urbe urbe 35K Apr 10 11:41 tracedb-to-frg.pl
-rwxrwxr-x 1 urbe urbe 44K Apr 10 11:41 trimFastqByQVWindow
-rwxrwxr-x 1 urbe urbe 18K Apr 10 11:40 uidclient
-rwxrwxr-x 1 urbe urbe 589K Apr 10 11:41 unitigger
-rwxrwxr-x 1 urbe urbe 42K Apr 10 11:40 upgrade-v8-to-v9
-rwxrwxr-x 1 urbe urbe 42K Apr 10 11:40 upgrade-v9-to-v10
-rwxrwxr-x 1 urbe urbe 854 Apr 10 11:41 utg2fasta
-rwxrwxr-x 1 urbe urbe 731K Apr 10 11:41 utgcns
-rwxrwxr-x 1 urbe urbe 561K Apr 10 11:41 utgcnsfix

Address of the bookmark: http://wgs-assembler.sourceforge.net/wiki/index.php/Main_Page

AfterQC: Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data

Jit — Fri, 29 Jun 2018 03:26:03 -0500

Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data AfterQC can simply go through all fastq files in a folder and then output three folders: good, bad and QC folders, which contains good reads, bad reads and the QC results of each fastq file/pair. Currently it supports processing data from HiSeq 2000/2500/3000/4000, Nextseq 500/550, MiniSeq...and other Illumina 1.8 or newer formats The author has reimplemented this tool in C++ with multithreading support to make it much faster. The new tool is called fastp and can be found at: https://github.com/OpenGene/fastp . If you prefer a C++ based tool, please use fastp instead. https://github.com/OpenGene/AfterQC

Address of the bookmark: https://github.com/OpenGene/AfterQC

Awk for Bioinformatician and computational biologist

Poonam Mahapatra — Tue, 06 Feb 2018 14:54:35 -0600

Awk is a programming language which allows easy manipulation of structured data and is mostly used for pattern scanning and processing. It searches one or more files to see if they contain lines that match with the specified patterns and then perform associated actions. The basic syntax is:

awk '/pattern1/ {Actions}
/pattern2/ {Actions}' file

The working of Awk is as follows
Awk reads the input files one line at a time.
For each line, it matches with given pattern in the given order, if matches performs the corresponding action.
If no pattern matches, no action will be performed.
In the above syntax, either search pattern or action are optional, But not both.
If the search pattern is not given, then Awk performs the given actions for each line of the input.
If the action is not given, print all that lines that matches with the given patterns which is the default action.
Empty braces with out any action does nothing. It wont perform default printing operation.
Each statement in Actions should be delimited by semicolon.
Say you have data.tsv with the following contents:

$ cat data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3 ACTTATATATATATA
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
By default Awk prints every line from the file.

$ awk '{print;}' data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3 ACTTATATATATATA
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
We print the line which matches the pattern contig3

$ awk '/contig3/' data/test.tsv
contig3 ACTTATATATATATA
Awk has number of builtin variables. For each record i.e line, it splits the record delimited by whitespace character by default and stores it in the $n variables. If the line has 5 words, it will be stored in $1, $2, $3, $4 and $5. $0 represents the whole line. NF is a builtin variable which represents the total number of fields in a record.

$ awk '{print $1","$2;}' data/test.tsv
contig1,ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2,ACTTTATATATT
contig3,ACTTATATATATATA
contig4,ACTTATATATATATA
contig5,ACTTTATATATT

$ awk '{print $1","$NF;}' data/test.tsv
contig1,ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2,ACTTTATATATT
contig3,ACTTATATATATATA
contig4,ACTTATATATATATA
contig5,ACTTTATATATT

Awk has two important patterns which are specified by the keyword called BEGIN and END. The syntax is as follows:

BEGIN { Actions before reading the file}
{Actions for everyline in the file}
END { Actions after reading the file }

For example,
$ awk 'BEGIN{print "Header,Sequence"}{print $1","$2;}END{print "-------"}' data/test.tsv
Header,Sequence
contig1,ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2,ACTTTATATATT
contig3,ACTTATATATATATA
contig4,ACTTATATATATATA
contig5,ACTTTATATATT
-------
We can also use the concept of a conditional operator in print statement of the form print CONDITION ? PRINT_IF_TRUE_TEXT : PRINT_IF_FALSE_TEXT. For example, in the code below, we identify sequences with lengths > 14:

$ awk '{print (length($2)>14) ? $0">14" : $0"<=14";}' data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG>14
contig2 ACTTTATATATT<=14
contig3 ACTTATATATATATA>14
contig4 ACTTATATATATATA>14
contig5 ACTTTATATATT<=14
We can also use 1 after the last block {} to print everything (1 is a shorthand notation for {print $0} which becomes {print} as without any argument print will print $0 by default), and within this block, we can change $0, for example to assign the first field to $0 for third line (NR==3), we can use:

$ awk 'NR==3{$0=$1}1' data/test.tsv
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
You can have as many blocks as you want and they will be executed on each line in the order they appear, for example, if we want to print $1 three times (here we are using printf instead of print as the former doesn't put end-of-line character),

$ awk '{printf $1"\t"}{printf $1"\t"}{print $1}' data/test.tsv
contig1 contig1 contig1
contig2 contig2 contig2
contig3 contig3 contig3
contig4 contig4 contig4
contig5 contig5 contig5
Although, we can also skip executing later blocks for a given line by using next keyword:

$ awk '{printf $1"\t"}NR==3{print "";next}{print $1}' data/test.tsv
contig1 contig1
contig2 contig2
contig3
contig4 contig4
contig5 contig5

$ awk 'NR==3{print "";next}{printf $1"\t"}{print $1}' data/test.tsv
contig1 contig1
contig2 contig2

contig4 contig4
contig5 contig5
You can also use getline to load the contents of another file in addition to the one you are reading, for example, in the statement given below, the while loop will load each line from test.tsv into k until no more lines are to be read:

$ awk 'BEGIN{while((getline k <"data/test.tsv")>0) print "BEGIN:"k}{print}' data/test.tsv
BEGIN:contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
BEGIN:contig2 ACTTTATATATT
BEGIN:contig3 ACTTATATATATATA
BEGIN:contig4 ACTTATATATATATA
BEGIN:contig5 ACTTTATATATT
contig1 ACTGTCTGTCACTGTGTTGTGATGTTGTGTGTG
contig2 ACTTTATATATT
contig3 ACTTATATATATATA
contig4 ACTTATATATATATA
contig5 ACTTTATATATT
You can also store data in the memory with the syntax VARIABLE_NAME[KEY]=VALUE which you can later use through for (INDEX in VARIABLE_NAME) command:

$ awk '{i[$1]=1}END{for (j in i) print j"<="i[j]}' data/test.tsv
contig1<=1
contig2<=1
contig3<=1
contig4<=1
contig5<=1