BOL: Related items

A guide for complete R beginners :- Getting data into R

Archana Malhotra — Tue, 24 Feb 2015 20:15:08 -0600

For a beginner this can be is the hardest part, it is also the most important to get right.

It is possible to create a vector by typing data directly into R using the combine function ‘c’

x

same as

x

creates the vector x with the numbers between 1 and 5.

You can see what is in an object at any time by typing its name;

x

will produce the output ‘[1] 1 2 3 4 5′

Note that names need to be quoted

daysofweek ← c(‘Monday’, ‘Tuesday’, ‘Wednesday’, ‘Thursday’, ‘Friday’);

Usually however you want to input from a file. We have touched on the ‘read.table’ function already.

mydata

Now mydata is a data frame with multiple vectors

each vector can be identified by the default syntax

#if any of these are typed it will print to screen

mydata$V1 mydata$V2 mydata$V3

By default the function assumes certain things from the file

The file is a plain text file (there are function to read excel files: not covered here)
columns are separated by any number of tabs or spaces
there is the same number of data points in each column
there is no header row (labels for the columns)
there is no column with names for the rows** [I’ll explain].

If any of these are false, we need to tell that to the function

If it has a header column

mydata header=T also works

Note that there is a comma between different parts of the functions arguments

If there is one less column in the header row, then R assumes that the 1^st column of data after the header are the row names

Now the vectors (columns) are identified by their name

#if any of these are typed it will print to screen

mydata$A mydata$B mydata$C

# Summary about the whole data frame

summary(mydata)

# Summary information of column A

summary(mydata$A)

We can shortcut having to type the data frame each time by attaching it

attach(mydata)

# summary of column B as ‘mydata’ is attached

summary(B)

Two other important options for read.table

If is is separated only by tabs and has a header

mydata

Really useful if you have spaces in the contents of some columns, so R does not mess up reading the columns . However if the columns or of an uneven length it will tell you.

If you know that the file has uneven columns

mydata

This causes R to fill empty spaces in a columns with ‘NA’ .

The last two examples will still work with our file and give the same result as with only headers=T

Graphs

to get an idea of what R is capable of type

demo(graphics)

steps through the examples, and the code is printed to the screen

We will work with simpler examples that have immediate use to biologists.

Remember to get more information about the options to a function type ‘?function’

Histogram of A

hist(mydata$A)

If there was more data we could increase the number of vertical columns with the option, breaks=50 (or another relevant number).

boxplot(mydata)

We can get rid of the need to type the data frame each time by using the attach function

# if not already done so

attach(mydata)
boxplot(mydata$A, mydata$B, name=c(“Value A”, “Value B”) , ylab=“Count of Something”)

same as

boxplot(A, B, name=c(“Value A”, “Value B”) , ylab=“Count of Something”)

Scatter plot

# if not already done so

attach(mydata)
plot(A,B) # or plot(mydata$A, mydata$B)

SAVING an image

Windows users (Rgui) RIGHT click on image and select which you want.

These instructions work for everyone.

You need to create a new device of the type of file you need, then send the data to that device

to save as a png file (easy to load into the likes of powerpoint, also great for web applications.

png(‘filename’)
boxplot(A, B, name=c(“Value A”, “Value B”) , ylab=“Count of Something”)

or to save as a pdf

pdf(‘filename’)
boxplot(A, B, name=c(“Value A”, “Value B”) , ylab=“Count of Something”)

Note

Nothing will appear on screen, the output is going to the file
Also it may not be saved immediately but will once the device (or R) is turned quit.

To quit R type

q() # If you save your session, next time you start R, you will have your data preloaded.

Or if you want to remain in R

dev.off() #turns of the png (or pdf etc) device, thus forces the data to save

ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data

Jit — Mon, 19 Feb 2018 06:46:15 -0600

ETE v3, featuring numerous improvements in the underlying library of methods, and providing a novel set of standalone tools to perform common tasks in comparative genomics and phylogenetics.

The new features include

(i) building gene-based and supermatrix-based phylogenies using a single command,

(ii) testing and visualizing evolutionary models,

(iii) calculating distances between trees of different size or including duplications, and

(iv) providing seamless integration with the NCBI taxonomy database.

ETE is freely available at http://etetoolkit.org

Address of the bookmark: http://etetoolkit.org

lordFAST: sensitive and Fast Alignment Search Tool for LOng noisy Read sequencing Data

BioJoker — Tue, 27 Nov 2018 04:43:57 -0600

lordFAST is a sensitive tool for mapping long reads with high error rates. lordFAST is specially designed for aligning reads from PacBio sequencing technology but provides the user the ability to change alignment parameters depending on the reads and application.

lordFAST, a novel long-read mapper that is specifically designed to align reads generated by PacBio and potentially other SMS technologies to a reference. lordFAST not only has higher sensitivity than the available alternatives, it is also among the fastest and has a very low memory footprint.

Address of the bookmark: https://github.com/vpc-ccg/lordfast

Genome in a Bottle (GIAB) Consortium

Jit — Sat, 25 Jan 2020 13:50:52 -0600

The Genome in a Bottle (GIAB) Consortium is a public-private-academic consortium hosted by NIST to develop the technical infrastructure (reference standards, reference methods, and reference data) to enable translation of whole human genome sequencing to clinical practice.

https://www.nist.gov/news-events/news/2016/09/nist-releases-new-family-standardized-genomes

Address of the bookmark: https://jimb.stanford.edu/giab/

AutoGluon: AutoML for Text, Image, and Tabular Data

Jit — Thu, 07 Jan 2021 05:33:17 -0600

AutoGluon automates machine learning tasks enabling you to easily achieve strong predictive performance in your applications. With just a few lines of code, you can train and deploy high-accuracy machine learning and deep learning models on text, image, and tabular data.

Address of the bookmark: https://github.com/awslabs/autogluon

NASA Open Science Data Repository

Abhi — Wed, 18 Dec 2024 11:54:47 -0600

The NASA Open Science Data Repository (OSDR) enables access to space-related data from experiments and missions that investigate biological and health responses of terrestrial life to spaceflight. The goal of OSDR is to enable multi-modal and multi-hierarchical fundamental space life science data be reused toward basic science, applied science, and operational outcomes for space exploration and knowledge discovery. These data include ‘omics, phenotypic, physiological, behavioral, hardware, environmental telemetry; raw, processed; tabular, text, code, bioimaging, and video.

https://www.nasa.gov/reference/osdr-data-processing/

Address of the bookmark: https://www.nasa.gov/osdr/

Master Thesis: Trans-membrane topology prediction through Markov based decoders

Rahul Agarwal — Wed, 17 Jul 2013 16:16:17 -0500

Abstract:

Background/Motivation:

The dearth of structural information on alpha helical membrane protein (MPs) has hindered thus far the development of reliable knowledge –based potentials that can be used for automatic prediction of trans-membrane (TM) protein structure. While algorithm for identification of TM segments is available, modelling of the domains of alpha helical MPs involves assembling the segments into a bundle. This requires the correct assignment of the buried and lipid-exposed faces of the TM domains.

Results: In a cross validated test on single sequences, our trans-membrane MM, correctly predicts the entire topology for 77% of the sequences in a standard dataset of 86 proteins with supervised topology. These results compare favorably with existing methods.

Source Code: Matlab

Conclusion/Implementation: Here discriminant data mining approach was used to predict the location and orientation of alpha helices in membrane-spanning proteins. It is based on a first order Markov model (MM) with an architecture that corresponds closely to the biological systems. The model is enriched with three types of states for the loop on the cytoplasmic side (outer loop), loop for the non-cytoplasmic side (inner side), and trans-membrane part. The closed association between the biological and Markov states allows us to infer which part of the model architecture are important to capture the information which encodes the membrane topology, and gain a better understanding of the mechanism and constraints involved. Predictor Model was established by various Markov decoder , and assignment of the membrane helix boundaries was apparent.

Scientist Bioinformatics Positions

Thu, 30 Jan 2020 06:53:40 -0600

Bioinformatics-Multi_Omics_Integration

https://www.researchgate.net/job/939073_Senior_Scientist_Bioinformatics-Multi_Omics_Integration

Senior_Scientist_Bioinformatics-Transcriptomics_Analysis

https://www.researchgate.net/job/939075_Senior_Scientist_Bioinformatics-Transcriptomics_Analysis-Belgium_France_Switzerland_The_Netherlands

Senior Scientist Bioinformatics - Network Analytics

https://www.researchgate.net/job/939070_Senior_Scientist_Bioinformatics-Network_Analytics_Belgium_France_Switzerland_the_Netherlands

Team Leader Bioinformatics Data Sciences - Mechelen, Belgium

https://www.researchgate.net/job/938787_Team_Leader_Bioinformatics_Data_Sciences-Mechelen_Belgium

Automatic Predictive Model Constructor - APMC

Jan Bińkowski — Mon, 16 Sep 2019 09:43:21 -0500

I would like to invite everyone interested in the subject of machine learning in life science, to test APMC module,

it`s a fully automatic tool (created by students) to simply create and develop supervised machine learning models

for classification and regression purposes. Links to tool, instruction and documentation bellow:

APMC: https://gene-calc.pl/apmc
How to use: https://gene-calc.pl/apmc/how-to-use
Documentation: https://gene-calc.pl/apmc/documentation

bacLIFE: an automated genome mining tool for identification of lifestyle associated genes

BioStar — Fri, 15 Mar 2024 04:59:14 -0500

bacLIFE is a streamlined computational workflow that annotates bacterial genomes and performs large-scale comparative genomics to predict bacterial lifestyles and to pinpoint candidate genes, denominated lifestyle-associated genes (LAGs), and biosynthetic gene clusters associated with each lifestyle detected. This whole process is divided into different modules:

Clustering module Predicts, clusters and annotates the genes of every input genome
Lifestyle prediction Employs a machine learning model to forecast bacterial lifestyle or other specified metadata
Analitical module (Shiny app) Results from the previous modules are embedded in a user-friendly interface for comprehensive and interactive comparative genomics.

You can find the complete wiki here [https://github.com/Carrion-lab/bacLIFE/wiki/bacLIFE-wiki]

Address of the bookmark: https://github.com/Carrion-lab/bacLIFE