<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/27479?offset=1070</link>
	<atom:link href="https://bioinformaticsonline.com/related/27479?offset=1070" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	
<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/11035/bioinformatics-jrfsrf-position-at-nii</guid>
  <pubDate>Sun, 25 May 2014 16:54:04 -0500</pubDate>
  <link></link>
  <title><![CDATA[Bioinformatics JRF/SRF position at NII]]></title>
  <description><![CDATA[
<p>NATIONAL INSTITUTE OF IMMUNOLOGY, NEW DELHI-110067</p>

<p>Applications are invited for the position of Senior Research Fellow for the following time-bound sponsored project as per the details given below:</p>

<p>1. BTIS project on, “Bioinformatics Center-National Infrastructural Facility in the Area of Immunology” funded by DBT</p>

<p>Senior Research Fellow (P) (One Position only)</p>

<p>Dr. Debasisa Mohanty<br />Staff Scientist-VI<br />deb@nii.res.in</p>

<p>Qualifications: M.Sc in Biological Sciences or Biotechnology with at least 04 years of Research experience in Bioinformatics or computational Biology after the master’s degree is essential.</p>

<p>Emoluments: The selected candidates will draw consolidated emoluments as per Institute Rules, depending upon qualifications &amp; experience</p>

<p>Rs. 18,000/- per month consolidated plus 30% HRA if Leading to Ph.D/NET/GATE Qualified otherwise Rs. 14,000/- per month + 30% HRA.</p>

<p>Job description: The candidate should be well versed in programming in PERL/C++/HTML/CGI, web server and portal development, computational analysis of<br />protein structure &amp; function, molecular dynamics simulations and use of high performance computing systems.</p>

<p>GENERAL TERMS AND CONDITIONS:-</p>

<p>1. The candidates selected for the above posts will be on contract for one year or duration of the project whichever is shorter, at a time.<br />2. No hostel/ housing facility will be provided.<br />3. Number of posts may vary and shall be need based. Advertisement is no commitment.<br />4. Applicants may clearly mention the category they belong to i.e. SC/ST/OBC/PH and attach documentary proof of the same.<br />5. No TA/DA will be paid for attending the interview, if called for.<br />6. Apart from sending application in the prescribed format given below, candidates should send complete Curriculum Vitae along with the names of three referees. Curriculum Vitae should contain details of the experimental expertise.</p>

<p>HOW TO APPLY Interested candidates may apply directly, STRICTLY IN THE PRESCRIBED FORMAT GIVEN BELOW, through e-mail, to the Investigator of the project, clearly indicating the name of the project along with their complete C.V., e-mail id, fax numbers, telephone numbers. Only Short listed candidates will be called for interview and they required to submit attested copies of all their certificates and a Demand Draft of Rs 100/- drawn on Canara Bank or Indian Bank payable at Delhi/New Delhi in favour of the Director, NII (SC / ST and PH candidates are exempted subject to submission of documentary proof), at the time of interview.</p>

<p>LAST DATE OF RECEIPT OF APPLICATIONS: 06th June, 2014</p>

<p>Advertisement</p>

<p>www1.nii.res.in/sites/default/files/projectappointment-Dr.Mohanty-6June2014.pdf</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/44468/orthoflow-workflow-for-phylogenetic-inference-of-genome-scale-datasets-of-protein-coding-genes</guid>
	<pubDate>Wed, 21 Feb 2024 06:13:08 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/44468/orthoflow-workflow-for-phylogenetic-inference-of-genome-scale-datasets-of-protein-coding-genes</link>
	<title><![CDATA[Orthoflow: workflow for phylogenetic inference of genome-scale datasets of protein-coding genes]]></title>
	<description><![CDATA[<p><span>Orthoflow is a workflow for phylogenetic inference of genome-scale datasets of protein-coding genes. Our goal was to make it straightforward to work from a combination of input sources including annotated contigs in Genbank format and FASTA files containing CDSs. It uses several state of the art inference methods for orthology inference, either based on HMM profiles or de novo inference of orthogroups. Through the use of OrthoSNAP, many additional ortholog alignments can be generated from multi-copy gene families. For phylogenetic inference, users can choose a supermatrix approach and/or gene tree inference followed by supertree reconstruction. Users can specify a range of alignment filtering settings to retain high-quality alignments for phylogenetic inference. The workflow produces a detailed report that, in addition to the phylogenetic results, includes a range of diagnostics to verify the quality of the results.</span></p><p>Address of the bookmark: <a href="https://github.com/rbturnbull/orthoflow" rel="nofollow">https://github.com/rbturnbull/orthoflow</a></p>]]></description>
	<dc:creator>LEGE</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/13014/bioinformatics-jrf-vacancy-at-icgeb-new-delhi</guid>
  <pubDate>Wed, 23 Jul 2014 16:07:15 -0500</pubDate>
  <link></link>
  <title><![CDATA[Bioinformatics JRF vacancy at ICGEB, New Delhi]]></title>
  <description><![CDATA[
<p>Junior Research Fellow for a DBT sponsored project entitled "Computational and experimental characterization of stage specific arginine methylation in P. falciparum proteome". </p>

<p>Candidates should have a 1st class MSc/MTech/BTech degree in Bioinformatics. Please send complete CV, quoting Application for RMETH-JRF-2014, by email to Dr. Dinesh Gupta: dinesh@icgeb.res.in</p>

<p>Closing date for applications: 6 August 2014</p>

<p>More at http://www.icgeb.org/tl_files/Vacancies/JRF.pdf</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/news/view/4408/fourth-branch-of-life</guid>
	<pubDate>Mon, 09 Sep 2013 21:48:37 -0500</pubDate>
	<link>https://bioinformaticsonline.com/news/view/4408/fourth-branch-of-life</link>
	<title><![CDATA[Fourth Branch of Life]]></title>
	<description><![CDATA[<p>Scientist have found the biggest viruses known, pandoraviruses which opened up entirely /completely... new questions questions and raise objections to in science. It even suggesting a fourth domain of life.</p><p>The new visrus are about one micron&mdash;a thousandth of a millimeter&mdash;in length, the newfound genus Pandoravirus dwarfs other viruses, which range in size from about 50 nanometers up to 100 nanometers. A genus is a taxonomic ranking between species and family.</p><p>Find&nbsp; more at @ http://www.nature.com/scitable/blog/viruses101/newly_found_pandoraviruses_hint_at</p><p>http://news.nationalgeographic.co.uk/news/2013/07/130718-viruses-pandoraviruses-science-biology-evolution/</p><p>&nbsp;</p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/11313/linux-sort-commands-for-bioinformatics</guid>
	<pubDate>Sat, 31 May 2014 15:41:16 -0500</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/11313/linux-sort-commands-for-bioinformatics</link>
	<title><![CDATA[Linux Sort Commands for Bioinformatics]]></title>
	<description><![CDATA[<p>Almost all the scripting languages such as Perl, Python etc have built-in sort, but unfortunately none of them are as flexible as sort command. But one when it come to space efficiency GNU sort stands at the top. It can sort a 20Gb file with less than 2Gb memory. It is not trivial to implement so powerful a sort by yourself.</p><p>sort a space-delimited file based on its first column, then the second if the first is the same, and so on:<br />sort input.txt</p><p>sort a huge file (GNU sort ONLY):<br />sort -S 1500M -t $HOME/tmp input.txt &gt; sorted.txt</p><p>sort starting from the third column, skipping the first two columns:<br />sort +2 input.txt</p><p>sort the second column as numbers, descending order; if identical, sort the 3rd as strings, ascending order:<br />sort -k2,2nr -k3,3 input.txt</p><p>sort starting from the 4th character at column 2, as numbers:<br />sort -k2.4n input.txt</p><p>More Linxu sort command information<br /><br />If you have any sort commands you'd like to share, please add them to our comments section below. For more help, you can also type:<br /><br />man sort<br /><br />or<br /><br />sort --help<br /><br />on your Unix/Linux system.</p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/8159/list-of-in-silico-binding-site-prediction-tools</guid>
	<pubDate>Mon, 03 Feb 2014 04:35:01 -0600</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/8159/list-of-in-silico-binding-site-prediction-tools</link>
	<title><![CDATA[List of In-silico Binding Site Prediction Tools]]></title>
	<description><![CDATA[<p>Following are the list of In-silico Binding Site Prediction in Proteins tools</p><p><a href="http://cast.engr.uic.edu/">CASTp</a> : <a href="http://sts.bioengr.uic.edu/castp/">http://sts.bioengr.uic.edu/castp/</a> &nbsp;Computed Atlas of Surface Topography of proteins (CASTp) provides an online resource for locating, delineating and measuring concave surface regions on three-dimensional structures of proteins. These include pockets located on protein surfaces and voids buried in the interior of proteins. The measurement includes the area and volume of pocket or void by solvent accessible surface model (Richards' surface) and by molecular surface model (Connolly's surface), all calculated analytically. CASTp can be used to study surface features and functional regions of proteins. CASTp includes a graphical user interface, flexible interactive visualization, as well as on-the-fly calculation for user uploaded structures. CASTp is updated daily and can be accessed at <a href="http://cast.engr.uic.edu/">http://cast.engr.uic.edu</a>.</p><p><a href="http://www.bigre.ulb.ac.be/Users/benoit/LigASite/index.php?home">LigASite</a>: <a href="http://www.bigre.ulb.ac.be/Users/benoit/LigASite/index.php?home">http://www.bigre.ulb.ac.be/Users/benoit/LigASite/index.php?home</a> is a gold-standard dataset of biologically relevant binding sites in protein structures. It consists of proteins with one unbound structure and at least one structure of the protein-ligand complex. Both a redundant and a non-redundant (sequence identity lower than 25%) version is available. Quaternary structures proposed by PISA <a href="http://www.bigre.ulb.ac.be/Users/benoit/LigASite/index.php?references">(3)</a> are used for all structures in the dataset.</p><p><a href="http://www.ebi.ac.uk/pdbe-site/pdbemotif/">PDBeMotif</a>: <a href="http://www.ebi.ac.uk/pdbe-site/pdbemotif/">http://www.ebi.ac.uk/pdbe-site/pdbemotif/</a> is an extremely fast and powerful search tool that facilitates exploration of the Protein Data Bank (PDB) by combining protein sequence, chemical structure and 3D data in a single search. Currently it is the only tool that offers this kind of integration at this speed. PDBeMotif can be used to examine the characteristics of the binding sites of single proteins or classes of proteins such as Kinases and the conserved structural features of their immediate environments either within the same specie or across different species. For example, it can highlight a conserved activation loop common to protein kinases, which is important in regulating activity and is marked by conserved DFG and APE motifs at the start and end of the loop, respectively. The prediction of the effect of modifications to small molecules that bind to the active and/or regulatory sites of proteins on their efficacy can be based on the outcome of analytic work done using PDBeMotif.</p><p><em><a href="http://pocket.uchicago.edu/fpop/">fPOP</a></em>: <a href="http://pocket.uchicago.edu/fpop/">http://pocket.uchicago.edu/fpop/</a> (footprinting Pockets Of Proteins, http://pocket.uchicago.edu/fpop/) is a database of the protein functional surfaces identified by shape analysis. In this relational database, we collected the spatial patterns of protein binding sites including both holo and apo forms from more than 40,000 structures. To identify protein binding sites, we model the shape of a split pocket induced by a binding ligand(s). Essentially, we use a purely geometric method to extract site-specific spatial patterns of split pockets as templates to match those from unbound structures. To perform an effective shape comparison, we utilize the Smith-Waterman algorithm to footprint an unbound pocket fragment with those selected from the canonical functional surfaces of &gt;19,000 structures in the SplitPocket (http://pocket.uchicago.edu/). The pairwise alignment of the unbound and split-pocket fragments is superimposed to evaluate the local structural similarity for detecting the unbound split characteristic through the RMSD measurement. Furthermore, we conduct a large-scale computation to systematically identify binding sites of proteins. In addition to the geometric measurements, we extensively measure the propensity of surface conservation encapsulated in the evolutionary history.(<a href="http://pocket.uchicago.edu/fpop/intro.html" target="_blank">more</a>)</p><p><a href="http://metapocket.eml.org/">metaPocket</a>: <a href="http://metapocket.eml.org/">http://metapocket.eml.org/</a> &nbsp;is a meta server to identify pockets on protein surface to predict ligand-binding sites. The identification of ligand-binding sites is often the starting point for protein function annotation and structure-based drug design. Many computational methods for the prediction of ligand-binding sites have been developed in recent decades. Here we present a consensus method metaPocket, in which the predicted sites from four methods: LIGSITE<em><sup>cs</sup></em>, PASS, Q-SiteFinder, and SURFNET are combined together to improve the prediction success rate. All these methods are evaluated on two datasets of 48 unbound/bound structures and 210 bound structures. The comparison results show that metaPocket improves the success rate from 70 to 75% at the top 1 prediction. MetaPocket is available at <a href="http://metapocket.eml.org/">http://metapocket.eml.org</a>.</p><p><a href="http://pocketquery.csb.pitt.edu/">PocketQuery</a>: <a href="http://pocketquery.csb.pitt.edu/">http://pocketquery.csb.pitt.edu/</a> &nbsp;is a web service for interactively exploring not only hot spot and anchor residues, but hot <em>regions</em>, defined by clusters of residues, at the interface of protein-protein interactions. An assortment of metrics, including changes in solvent accessible surface area, energy-based scores, and sequence conservation, are available to screen and sort clusters of residues. PocketQuery was developed by <a href="http://www.pitt.edu/%7Edkoes/">David Koes</a> from the <a href="http://smoothdock.ccbb.pitt.edu/">Camacho Lab</a> in the <a href="http://www.csb.pitt.edu/">Department of Computational and System Biology</a> at the <a href="http://www.pitt.edu/">University of Pittsburgh</a>.</p><p><a href="http://www.ncbi.nlm.nih.gov/Structure/ibis/ibis.cgi">IBIS</a>: <a href="http://www.ncbi.nlm.nih.gov/Structure/ibis/ibis.cgi">http://www.ncbi.nlm.nih.gov/Structure/ibis/ibis.cgi</a> is the NCBI Inferred Biomolecular Interactions Server. For a given protein sequence or structure query, IBIS reports physical interactions observed in experimentally-determined structures for this protein. IBIS also infers/predicts interacting partners and binding sites by homology, by inspecting the protein complexes formed by close homologs of a given query. To ensure biological relevance of inferred binding sites, the IBIS algorithm clusters binding sites formed by homologs based on binding site sequence and structure conservation.</p><p><a href="http://www.sbg.bio.ic.ac.uk/%7E3dligandsite/">3DLigandStie</a>: <a href="http://www.sbg.bio.ic.ac.uk/%7E3dligandsite/">http://www.sbg.bio.ic.ac.uk/~3dligandsite/</a> is an automated method for the prediction of ligand binding sites. Users can either submit a sequence or a protein structure. If a sequence is submitted then Phyre is run to predict the structure. The structure is then ussed to search a structural library to identify homologous structures with bound ligands. These ligands are superimposed onto the protein structure to predict a ligand binding site.</p><p><a href="http://www.modelling.leeds.ac.uk/sb/">SitesBase</a>: <a href="http://www.modelling.leeds.ac.uk/sb/">http://www.modelling.leeds.ac.uk/sb/</a> is a database of known ligand binding sites within the PDB which is navigable by PDB identifier or ligand 3 letter code e.g. NAD. Each binding site has a frequently updated register of structurally similar binding sites sharing atomic similarity detected by geometric hashing (Brakoulias and Jackson 2004). Multiple alignments, structural superpositions and links to other structural databases are also available enabling further analysis.</p><p><a href="http://163.43.140.95/top">PROSURFER</a>: <a href="http://163.43.140.95/top">http://163.43.140.95/top</a> contains information about structural similarities with respect to the query surfaces. A pocket search algorithm detected 48,347 potential ligand binding sites from the 9,708 non-redundant protein entries in the PDB database. All-against-all structural comparison was performed for the predicted sites, and the similar sites with the Z-score &ge; 2.5 were selected. These results can be accessed by the PDB code or ligand name.</p><p><a href="http://kbdock.loria.fr/index.php">KBDOCK</a>: <a href="http://kbdock.loria.fr/index.php">http://kbdock.loria.fr/index.php</a> is a 3D database system that defines and spatially clusters protein binding sites for knowledge-based protein docking. KBDOCK integrates protein domain-domain interaction information from <a href="http://3did.irbbarcelona.org/" target="_blank" title="Open in a new tab the 3DID home page">3DID</a> and sequence alignments from <a href="http://pfam.sanger.ac.uk/" target="_blank" title="Open in a new tab the Pfam home page">PFAM</a> together with structural information from the <a href="http://www.rcsb.org/" target="_blank" title="Open in a new tab the PDB home page">PDB</a> in order to analyse the spatial arrangements of DDIs by Pfam family, and to propose structural templates for protein docking. [<a href="http://kbdock.loria.fr/about.php" title="Go to the About page">More</a>]</p><p><a href="http://www.pocketome.org/">Pocketome</a>: <a href="http://www.pocketome.org/">http://www.pocketome.org/</a> The Pocketome is an encyclopedia of conformational ensembles of all druggable binding sites that can be identified experimentally from co-crystal structures in the <a href="http://www.pdb.org/" target="_blank">Protein Data Bank</a>.</p><p><a href="http://cheminfo.u-strasbg.fr:8080/scPDB/2011/db_search/about_scpdb.html">sc-PDB</a>: <a href="http://cheminfo.u-strasbg.fr:8080/scPDB/2011/db_search/about_scpdb.html">http://cheminfo.u-strasbg.fr:8080/scPDB/2011/db_search/about_scpdb.html</a>&nbsp; To assist structure-based approaches in drug design, we have processed the PDB to identify binding sites suitable for the docking of a drug-like ligand and we have so created a database called sc-PDB. The sc-PDB database provides separated MOL2 files for the ligand, its binding site and the corresponding protein chain(s). Ions and cofactors at the vicinity of the ligand are included in the protein. More details about the sc-PDB scope, its content and its evolution during the 2004-2009 period are provided in <a href="http://cheminfo.u-strasbg.fr:8080/scPDB/2011/db_search/txt_files/HDR-scPDB.pdf" target="_blank">a pdf document</a>.</p><p><a href="http://www.reading.ac.uk/bioinf/FunFOLD/FunFOLD_form.html">The FunFOLD Binding Site Residue Prediction Server</a>: BACKGROUND: The accurate prediction of ligand binding residues from amino acid sequences is important for the automated functional annotation of novel proteins. In the previous two CASP experiments, the most successful methods in the function prediction category were those which used structural superpositions of 3D models and related templates with bound ligands in order to identify putative contacting residues. However, whilst most of this prediction process can be automated, visual inspection and manual adjustments of parameters, such as the distance thresholds used for each target, have often been required to prevent over prediction. Here we describe a novel method FunFOLD, which uses an automatic approach for cluster identification and residue selection. The software provided can easily be integrated into existing fold recognition servers, requiring only a 3D model and list of templates as inputs. A simple web interface is also provided allowing access to non-expert users. The method has been benchmarked against the top servers and manual prediction groups tested at both CASP8 and CASP9.RESULTS: The FunFOLD method shows a significant improvement over the best available servers and is shown to be competitive with the top manual prediction groups that were tested at CASP8. The FunFOLD method is also competitive with both the top server and manual methods tested at CASP9. When tested using common subsets of targets, the predictions from FunFOLD are shown to achieve a significantly higher mean Matthews Correlation Coefficient (MCC) scores and Binding-site Distance Test (BDT) scores than all server methods that were tested at CASP8. Testing on the CASP9 set showed no statistically significant separation in performance between FunFOLD and the other top server groups tested. CONCLUSIONS: The FunFOLD software is freely available as both a standalone package and a prediction server, providing competitive ligand binding site residue predictions for expert and non-expert users alike. The software provides a new fully automated approach for structure based function prediction using 3D models of proteins.</p><p><a href="http://probis.cmm.ki.si/index.php">ProBiS</a>: <a href="http://probis.cmm.ki.si/index.php">http://probis.cmm.ki.si/index.php</a> &nbsp;algorithm for detection of structurally similar protein binding sites by local structural alignment. Motivation: Exploitation of locally similar 3D patterns of physicochemical properties on the surface of a protein for detection of binding sites that may lack sequence and global structural conservation. Results: An algorithm, ProBiS is described that detects structurally similar sites on protein surfaces by local surface structure alignment. It compares the query protein to members of a database of protein 3D structures and detects with sub-residue precision, structurally similar sites as patterns of physicochemical properties on the protein surface. Using an efficient maximum clique algorithm, the program identifies proteins that share local structural similarities with the query protein and generates structure-based alignments of these proteins with the query. Structural similarity scores are calculated for the query protein's surface residues, and are expressed as different colors on the query protein surface. The algorithm has been used successfully for the detection of protein&ndash;protein, protein&ndash;small ligand and protein&ndash;DNA binding sites. Availability: The software is available, as a web tool, free of charge for academic users at <a href="http://probis.cmm.ki.si/">http://probis.cmm.ki.si</a></p><p><a href="http://www.scfbio-iitd.res.in/dock/ActiveSite_new.jsp">Active Site prediction</a>: <a href="http://www.scfbio-iitd.res.in/dock/ActiveSite_new.jsp">http://www.scfbio-iitd.res.in/dock/ActiveSite_new.jsp</a> Active Site Prediction of Protein server computes the cavities in a given protein.</p><p><a href="http://mspc.bii.a-star.edu.sg/tankp/run_depth.html">DEPTH</a>: <a href="http://mspc.bii.a-star.edu.sg/tankp/run_depth.html">http://mspc.bii.a-star.edu.sg/tankp/run_depth.html</a> Depth measures the closest distance of a residue/atom to bulk solvent. Accessible surface area is a parameter that is widely used in analyses of protein structure and stability. However accessible surface area does not distinguish between atoms just below the protein surface and those in the core of the protein. In order to differentiate between such buried residues, we describe a computational procedure for calculating the depth of a residue from the protein surface. A detailed description of the computation of depth can be found <a href="http://www.ncbi.nlm.nih.gov/pubmed/10425675">here</a>.</p><p><a href="http://cssb.biology.gatech.edu/findsite">FINDSITE</a>: <a href="http://cssb.biology.gatech.edu/findsite">http://cssb.biology.gatech.edu/findsite</a> &nbsp;FINDSITE is a threading-based binding site prediction/protein functional inference/ligand screening algorithm that detects common ligand binding sites in a set of evolutionarily related proteins. Crystal structures as well as protein models can be used as the target structures.</p><p><a href="http://proline.physics.iisc.ernet.in/pocketdepth/">PocketDepth</a>: <a href="http://proline.physics.iisc.ernet.in/pocketdepth/">http://proline.physics.iisc.ernet.in/pocketdepth/</a>&nbsp; A new depth based algortihm for identification of ligand binding sites. Abstract: Computational methods for identifying and predicting functional sites in protein structures are increasingly becoming important in structural biology and bioinformatics not only for understanding the function of the molecule in detail but also for structure-based design of possible ligands and potential drugs as well as modified protein molecules. While there are a few structure based prediction methods already available, given the complexity and diversity of protein structural types, there is still a great need to explore newer methods and concepts to develop accurate, versatile and efficient binding site prediction algorithms. We have developed a new method PocketDepth, for identification of binding sites in proteins. The method is purely geometry-based and proceeds in two stages, labeling of grid cells with depth factors followed by a depth based clustering that uses neighbourhood information. Depth is an important parameter considered during protein structure visualization and analysis but has been used more often intuitively than systematically. Our current implementation of depth reflects how central a given sub-space is to a putative pocket rather than reflecting merely how far away it is situated from the nearest external surface of the protein. We have tested the algorithm against PDBbind, a large curated set of 1091 proteins obtained from PDB. A prediction was considered a true-positive if the predicted pocket had at-least 10% overlap with the actual ligand. The prediction accuracy using this set was about 96%. Moreover, 87% of the true-positives were identified within the first five ranks for each protein, of which 55% are in the first rank itself. 77% of the predictions had at least 50% overlap with the experimentally observed ligand. High prediction rates were again observed, when the method was tested against a data-set of apo-proteins and compared with their respective ligand complexes. A comparison of our method with four other widely used methods for a chosen representative set is also presented.</p><p><a href="http://strcomp.protein.osaka-u.ac.jp/ghecom/">GHECOM 1.0</a> : <a href="http://strcomp.protein.osaka-u.ac.jp/ghecom/">http://strcomp.protein.osaka-u.ac.jp/ghecom/</a>&nbsp; Grid-based HECOMi finder. A program for finding multi-scale pockets on protein surfaces using mathematical morphology</p><p><a href="http://www.modelling.leeds.ac.uk/pocketfinder/">Pocket-Finder</a>: <a href="http://www.modelling.leeds.ac.uk/pocketfinder/">http://www.modelling.leeds.ac.uk/pocketfinder/</a> is based on the Ligsite algorithm written by Hendlich <em>et al.</em> (1997). Pocket-Finder was written to compare pocket detection with our new ligand binding site detction algorithm <a href="http://www.modelling.leeds.ac.uk/qsitefinder">Q-SiteFinder.</a></p><p><a href="http://luna.bioc.columbia.edu/honiglab/screen2/cgi-bin/screen2.cgi">Screen2</a>: <a href="http://luna.bioc.columbia.edu/honiglab/screen2/cgi-bin/screen2.cgi">http://luna.bioc.columbia.edu/honiglab/screen2/cgi-bin/screen2.cgi</a> &nbsp;is a tool for identifying protein cavities and computing cavity attributes that can be applied for classification and analysis. The original Screen, written by Murad Nayal, was dependent on the obsolete Irix platform and is no longer available. Screen2 was reengineered by Brian Y. Chen for efficiency and compatibility, and made accessible as a web service by Raquel Norel.</p><p><a href="http://compbio.cs.princeton.edu/concavity/">ConCavity</a>: <a href="http://compbio.cs.princeton.edu/concavity/">http://compbio.cs.princeton.edu/concavity/</a> Identifying a protein's functional sites is an important step towards characterizing its molecular function. Numerous structure- and sequence-based methods have been developed for this problem. Here we introduce <em>ConCavity</em>, a small molecule binding site prediction algorithm that integrates evolutionary sequence conservation estimates with structure-based methods for identifying protein surface cavities. In large-scale testing on a diverse set of single- and multi-chain protein structures, we show that <em>ConCavity</em> substantially outperforms existing methods for identifying both 3D ligand binding pockets and individual ligand binding residues. As part of our testing, we perform one of the first direct comparisons of conservation-based and structure-based methods. We find that the two approaches provide largely complementary information, which can be combined to improve upon either approach alone. We also demonstrate that <em>ConCavity</em> has state-of-the-art performance in predicting catalytic sites and drug binding pockets. Overall, the algorithms and analysis presented here significantly improve our ability to identify ligand binding sites and further advance our understanding of the relationship between evolutionary sequence conservation and structural and functional attributes of proteins. Data, source code, and prediction visualizations are available on the <em>ConCavity</em> web site (<a href="http://compbio.cs.princeton.edu/concavity/">http://compbio.cs.princeton.edu/concavit​y/</a>).</p><p><a href="http://bioinfo3d.cs.tau.ac.il/MultiBind/index.html">MultiBind and MAPPIS</a>: <a href="http://bioinfo3d.cs.tau.ac.il/MultiBind/index.html">http://bioinfo3d.cs.tau.ac.il/MultiBind/index.html</a> Web servers for multiple alignment of protein 3D binding sites and their interactions. Analysis of protein&ndash;ligand complexes and recognition of spatially conserved physico-chemical properties is important for the prediction of binding and function. Here, we present two webservers for multiple alignment and recognition of binding patterns shared by a set of protein structures. The first webserver, MultiBind (<a href="http://bioinfo3d.cs.tau.ac.il/MultiBind">http://bioinfo3d.cs.tau.ac.il/MultiBind</a>), performs multiple alignment of protein binding sites. It recognizes the common spatial chemical binding patterns even in the absence of similarity of the sequences or the folds of the compared proteins. The input to the MultiBind server is a set of protein-binding sites defined by interactions with small molecules. The output is a detailed list of the shared physico-chemical binding site properties. The second webserver, MAPPIS (<a href="http://bioinfo3d.cs.tau.ac.il/MAPPIS">http://bioinfo3d.cs.tau.ac.il/MAPPIS</a>), aims to analyze protein&ndash;protein interactions. It performs multiple alignment of protein&ndash;protein interfaces (PPIs), which are regions of interaction between two protein molecules. MAPPIS recognizes the spatially conserved physico-chemical interactions, which often involve energetically important hot-spot residues that are crucial for protein&ndash;protein associations. The input to the MAPPIS server is a set of protein-protein complexes. The output is a detailed list of the shared interaction properties of the interfaces.</p><p><a href="http://bioinfo3d.cs.tau.ac.il/MolAxis/">MolAxis</a>: <a href="http://bioinfo3d.cs.tau.ac.il/MolAxis/">http://bioinfo3d.cs.tau.ac.il/MolAxis/</a>&nbsp; is a tool for the identification of high clearance pathways or <em>corridors</em> which represent molecular channels in the complement space of proteins. It is extremely efficient because it samples the medial axis of the complement of the molecule, reducing the problem dimension to two, since the medial axis is composed of surface patches. It is designed to analyze proteins channels, calculate pore dimensions and analyze atom accessibility. MolAxis reads files in the standard Protein Data Bank format (PDB) containing a single frame or multiple frames generated by molecular dynamics (MD) simulations. MolAxis handles two distinct scenarios: It computes channels that connect a single point (like an inner chamber) to the bulk solvent, and it also computes transmembrane (TM) channels. MolAxis has a friendly web interface (see the <a href="http://bioinfo3d.cs.tau.ac.il/MolAxis/server_channel.html" target="body">Web Server</a> tab). It also has a stand-alone version, statically compiled for linux, which can be downloaded from the <a href="http://bioinfo3d.cs.tau.ac.il/cgi-bin/pdownload/progdownload.pl/?pname=MolAxis" target="body">Download</a> tab.</p><p><a href="http://fpocket.sourceforge.net/">fpocket</a>: <a href="http://fpocket.sourceforge.net/">http://fpocket.sourceforge.net/</a> fpocket is a very fast open source protein pocket (cavity) detection algorithm based on Voronoi tessellation. It was developed in the C programming language and is currently available as command line driven program. A GUI is in development and mdpocket (fpocket on md trajectories) is out now. fpocket includes two other programs (dpocket &amp; tpocket) that allow you to extract pocket descriptors and test own scoring functions respectively. Furthermore a nifty druggability prediction score has been added to fpocket recently. As the algorithm is very fast it can be used on a large scale level (PDB size for instance). If you use fpocket for publication, please cite : <em>Vincent Le Guilloux, Peter Schmidtke and Pierre Tuffery</em>, "Fpocket: An open source platform for ligand pocket detection", BMC Bioinformatics, 2009, 10:168</p><p><a href="http://sumo-pbil.ibcp.fr/cgi-bin/sumo-welcome">SuMo</a>: <a href="http://sumo-pbil.ibcp.fr/cgi-bin/sumo-welcome">http://sumo-pbil.ibcp.fr/cgi-bin/sumo-welcome</a> allows you to screen the <a href="http://www.rcsb.org/" target="_blank">Protein Data Bank</a> (PDB) for finding ligand binding sites matching your protein structure or inversely, for finding protein structures matching a given site in your protein. This method is neither based on aminoacid sequence nor on fold comparisons. Priority is given to biological relevance. SuMo uses its own heuristics for defining ligand binding sites. Automatically selected ligand binding sites are extracted from PDB structure files and stored into <a href="http://sumo-pbil.ibcp.fr/cgi-bin/sumo-database">SuMo's own database</a>.</p><p><a href="http://www.caver.cz/">CAVER</a>: <a href="http://www.caver.cz/">http://www.caver.cz/</a> CAVER is a software tool for analysis and visualization of tunnels and channels in protein structures. Tunnels are void pathways leading from a cavity buried in a protein core to the surrounding solvent. Unlike tunnels, channels lead through the protein structure and their both endings are opened to the surrounding solvent. Studying of these pathways is highly important for drug design and molecular enzymology.</p><p><a href="http://scbx.mssm.edu/sitehound/sitehound-download/download.html">SiteHound</a>: <a href="http://scbx.mssm.edu/sitehound/sitehound-download/download.html">http://scbx.mssm.edu/sitehound/sitehound-download/download.html</a> SiteHound identifies protein regions that are likely to interact with ligands.&nbsp;The only input files required by SITEHOUND are the PDB file of the protein and the Molecular Interaction Field (MIFs) or Affinity Map for that protein structure structure. EasyMIFs is provided as a tool to calculate MIFs, alternatively AutoGrid (part of the AutoDock suite developed by Arthur Olson&rsquo;s group at The Scripps Research Insitute) or the SiteHound-web server can be used to produce Affinity maps or MIFs. A python script named 'auto.py' is provided in the package and can be used to perform binding site identification in a fully automated fashion. The script will prepare the protein PDB file, compute a Molecular Interaction Fields map with EasyMIFs and carry out binding site identification using SiteHound.&nbsp;It is also possible to use EasyMIFs and SiteHound separately.</p><p><a href="http://www.biochem.ucl.ac.uk/%7Eroman/surfnet/surfnet.html">SURFNET</a>: <a href="http://www.biochem.ucl.ac.uk/%7Eroman/surfnet/surfnet.html">http://www.biochem.ucl.ac.uk/~roman/surfnet/surfnet.html</a> The SURFNET program generates surfaces and void regions between surfaces from coordinate data supplied in a PDB file.</p><p><a href="http://appserver.biotec.tu-dresden.de/MSPocket/">MSPocket</a>: <a href="http://appserver.biotec.tu-dresden.de/MSPocket/">http://appserver.biotec.tu-dresden.de/MSPocket/</a> is an orientation independent program for the detection and graphical analysis of protein surface pockets [Zhu2011]. The approach is based on the solvent excluded surfaces generated by <a href="http://mgltools.scripps.edu/packages/MSMS">MSMS</a> [Sanner1996].</p><p><a href="http://pdbfun.uniroma2.it/pfinder/index.html">Pfinder</a> : <a href="http://pdbfun.uniroma2.it/pfinder/index.html">http://pdbfun.uniroma2.it/pfinder/index.html</a>&nbsp; Pfinder is a bioinformatic method for the prediction of phosphate-binding sites in protein structures. Given a protein structure, Pfinder compares it with a set of 215 highly conserved structural motifs known to bind the phosphate moiety of phosphorylated ligands.</p><p><a href="http://xray.bmc.uu.se/cgi-bin/gerard/image_page.pl?image=usf/voodoo.gif">VOIDOO</a>: <a href="http://xray.bmc.uu.se/usf/voidoo.html">http://xray.bmc.uu.se/usf/voidoo.html</a> is a program for detection of cavities in macromolecular structures. It uses an algorithm that makes it possible to detect even certain types of cavities that are connected to "the outside world". Three different types of cavity can be handled by VOIDOO: Vanderwaals cavities (the complement of the molecular Vanderwaals surface), probe-accessible cavities (the cavity volume that can be occupied by the centres of probe atoms) and MS-like probe-occupied cavities (the volume that can be occupied by probe atoms, <em>i.e.</em> including their radii).</p><p><a href="http://gecco.org.chemie.uni-frankfurt.de/pocketpicker/index.html">PocketPicker</a>: <a href="http://gecco.org.chemie.uni-frankfurt.de/pocketpicker/index.html">http://gecco.org.chemie.uni-frankfurt.de/pocketpicker/index.html</a> Background: Identification and evaluation of surface binding-pockets and occluded cavities are initial steps in protein structure-based drug design. Characterizing the active site's shape as well as the distribution of surrounding residues plays an important role for a variety of applications such as automated ligand docking or <em>in situ </em>modeling. Comparing the shape similarity of binding site geometries of related proteins provides further insights into the mechanisms of ligand binding. Results: We present PocketPicker, an automated grid-based technique for the prediction of protein binding pockets that specifies the shape of a potential binding-site with regard to its buriedness. The method was applied to a representative set of protein-ligand complexes and their corresponding <em>apo</em>-protein structures to evaluate the quality of binding-site predictions. The performance of the pocket detection routine was compared to results achieved with the existing methods CAST, LIGSITE, LIGSITE<sup>cs</sup>, PASS and SURFNET. Success rates PocketPicker were comparable to those of LIGSITE<sup>cs </sup>and outperformed the other tools. We introduce a descriptor that translates the arrangement of grid points delineating a detected binding-site into a correlation vector. We show that this shape descriptor is suited for comparative analyses of similar binding-site geometry by examining induced-fit phenomena in aldose reductase. This new method uses information derived from calculations of the buriedness of potential binding-sites. Conclusion: The pocket prediction routine of PocketPicker is a useful tool for identification of potential protein binding-pockets. It produces a convenient representation of binding-site shapes including an intuitive description of their accessibility. The shape-descriptor for automated classification of binding-site geometries can be used as an additional tool complementing elaborate manual inspections.</p><p><a href="http://www.bisb.uni-bayreuth.de/index.php?page=data/mcvol/mcvol">McVol</a>: <a href="http://www.bisb.uni-bayreuth.de/index.php?page=data/mcvol/mcvol">http://www.bisb.uni-bayreuth.de/index.php?page=data/mcvol/mcvol</a>&nbsp; This program was developed to integrate the molecular volume, solven accessible volume an Van der Waals volume of proteins using a Monte carlo algorithm. Based on this calculations, McVol is also able to identify internal cavities as well as surface clefts und fill these cavities with water molecules. Additionally, a membrane of dummy atoms can be placed as a disc atound the protein. The program is available under the Gnu Public Licence. A precompiled binary (X86) can be downloaded free of charge from here (when the associated paper is published).</p><p>&nbsp;</p>]]></description>
	<dc:creator>Shikha Logwani</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/pages/view/11399/next-generation-sequencing-in-r-or-bioconductor-environment</guid>
	<pubDate>Mon, 02 Jun 2014 18:03:09 -0500</pubDate>
	<link>https://bioinformaticsonline.com/pages/view/11399/next-generation-sequencing-in-r-or-bioconductor-environment</link>
	<title><![CDATA[Next generation sequencing in R or bioconductor environment]]></title>
	<description><![CDATA[<p>There are many R software and bioconductor packages for NGS data analysis, some of them are as follows</p><h3><a name="TOC-Biostrings" id="TOC-Biostrings"></a>Biostrings</h3><p>The Biostrings package from Bioconductor provides an advanced environment for efficient sequence management and analysis in R. It contains many speed and memory effective string containers, string matching algorithms, and other utilities, for fast manipulation of large sets of biological sequences. The objects and functions provided by Biostrings form the basis for many other sequence analysis packages. <a href="http://bioconductor.org/packages/release/bioc/html/Biostrings.html">Documentation</a></p><div><div style="text-align: left;"><div style="color: #000000;"><h4><a name="TOC-IRanges-Overview" id="TOC-IRanges-Overview"></a>IRanges Overview</h4><p>IRanges provides the low-level infrastructure and containers for handling sets of integer ranges within Bioconductor's BioC-Seq domain. Its classes and methods provide support for many more high-level packages like GenomicRanges, ShortRead, Rsamtools, etc. <a href="http://bioconductor.org/packages/release/bioc/html/IRanges.html">Documentation</a></p><div style="text-align: right;"><div style="text-align: left;"><h4><a name="TOC-GenomicRanges-Overview" id="TOC-GenomicRanges-Overview"></a>GenomicRanges Overview</h4><p>The <em>GenomicRanges</em> package serves as the foundation for representing genomic locations within the Bioconductor project. It is built upon the <em>IRanges</em> infrastructure and defines three major data containers - <em>GRanges, GRangesList</em> and <em>GappedAlignments</em> - which are supporting other important BioC-Seq packages including <em>ShortRead, Rsamtools, rtracklayer, GenomicFeatures</em> and <em>BSgenome</em>.&nbsp; Compared to the IRanges container, the GRanges/<em>GRangesList</em> classes are more flexible and extensible to store additional information about sequence ranges, such as chromosome identifiers (sequence space), strand information and annotation data. <a href="http://bioconductor.org/packages/release/bioc/html/GenomicRanges.html">Documentation</a></p></div></div></div></div><h3><a name="TOC-Motif-Discovery" id="TOC-Motif-Discovery"></a>Motif Discovery</h3><h4><a name="TOC-cosmo" id="TOC-cosmo"></a>cosmo</h4><p>The cosmo package allows to search a set of unaligned DNA sequences for a shared motif that may function as transcription factor binding site. The algorithm extends the popular motif discovery tool MEME (Bailey and Elkan, 1995) in that it allows the search to be supervised by specifying a set of constraints that the motif to be discovered must satisfy. <a href="http://bioconductor.org/packages/release/bioc/html/cosmo.html">Documentation</a></p></div><div>
<p><span></span><span></span></p>
<div style="color: #0000ff;"><h4><a name="TOC-BCRANK" id="TOC-BCRANK"></a>BCRANK</h4><p>BCRANK is a method that takes a ranked list of genomic regions as input and outputs short DNA sequences that are overrepresented in some part of the list. The algorithm was developed for detecting transcription factor (TF) binding sites in a large number of enriched regions from high-throughput ChIP-chip or ChIP-seq experiments, but it can be applied to any ranked list of DNA sequences. Documentation</p>
<p><a href="http://bioconductor.org/packages/release/bioc/html/BCRANK.html"></a></p>
<p>rGADEM: <a href="http://bioconductor.org/packages/devel/bioc/html/rGADEM.html">Documentation</a></p><p>MotIV: <a href="http://bioconductor.org/packages/devel/bioc/html/MotIV.html">Documentation</a></p></div><h3><a name="TOC-ShortRead" id="TOC-ShortRead"></a>ShortRead</h3><p>The ShortRead package provides input, quality control, filtering, parsing, and manipulation functionality for short read sequences produced by high throughput sequencing technologies. While support is provided for many sequencing technologies, this package is primairly focused on Solexa/Illumina reads. <a href="http://bioconductor.org/packages/release/bioc/html/ShortRead.html">Documentation</a></p><h3><a name="TOC-Rsamtools" id="TOC-Rsamtools"></a>Rsamtools</h3><p>Rsamtools provides functions for parsing and inspecting samtools BAM formatted binary alignment data. SAM/BAM is quickly becoming a universal standard alignment format, and is now supported by a wide variety of alignment tools. <a href="http://bioconductor.org/help/bioc-views/2.7/bioc/html/Rsamtools.html">Documentation</a></p>
<p><a href="http://samtools.sourceforge.net/">Samtools Website</a><br /> <a href="http://bio-bwa.sourceforge.net/">BWA (Burrows-Wheeler Alignment) Website</a><br /><span style="color: #0000ff;"></span></p>
<div style="color: #000000;">&nbsp;</div></div><div>
<p><span style="color: #000000;">Additional tools for SNP analysis:&nbsp;</span></p>
<p><a href="http://bioconductor.org/help/bioc-views/release/bioc/html/snpMatrix.html">snpMatrix</a></p><h3><a name="TOC-BSgenome" id="TOC-BSgenome"></a>BSgenome</h3><p>BSgenome provides an object oriented infrastructure for interacting with a Biostring based genome sequence. BSgenome packages exist for many common genomes, and can be created to represent custom genomes. See the "How to forge a BSgenome data package" Vignette for instructions to create a new BSgenome package if a prebuilt package does not exist for your organism. <a href="http://bioconductor.org/packages/release/bioc/html/BSgenome.html">Documentation</a></p><h3><a name="TOC-rtracklayer" id="TOC-rtracklayer"></a>rtracklayer</h3><p>rtracklayer provides an interface for exporting annotation feature data to various genome browsers and file formats (such as GFF). See the Small RNA Profiling exercise for an example of using rtracklayer to visualize alignment coverage. <a href="http://bioconductor.org/packages/release/bioc/html/rtracklayer.html">Documentation</a></p><h3><a name="TOC-biomaRt" id="TOC-biomaRt"></a>biomaRt</h3><p>The biomaRt package, provides an interface to a growing collection of databases implementing the BioMart software suite (http:// www.biomart.org). The package enables online retrieval of large amounts of data in a uniform way without the need to know the underlying database schemas. This data is retrieved automatically via the Internet, so it's recommended that you cache the data locally, or check versions if your code will be adversely affected by updates to these data. <a href="http://bioconductor.org/packages/release/bioc/html/biomaRt.html">Documentation</a></p><h3><a name="TOC-ChIP-Seq-Analysis-Packages" id="TOC-ChIP-Seq-Analysis-Packages"></a>ChIP-Seq Analysis Packages</h3><p>Bioconductor provides various packages for analyzing and visualizing ChIP-Seq data. Only a small selection of these packages is introduced here. Additional useful introductions to this topic are: <a href="http://www.bioconductor.org/workshops/2009/SeattleJan09/ChIP-seq/">BioC ChIP-seq Case Study</a> and BioC <a href="http://www.bioconductor.org/help/course-materials/2009/SeattleNov09/ChIP-seq/">ChIP-Seq</a>.</p><h4><a name="TOC-chipseq" id="TOC-chipseq"></a>chipseq</h4><p>The chipseq package combines a variety of HT-Seq packages to a pipeline for ChIP-Seq data analysis. <a href="http://bioconductor.org/packages/release/bioc/html/chipseq.html">Documentation</a></p><h4><a name="TOC-BayesPeak" id="TOC-BayesPeak"></a>BayesPeak</h4><p>BayesPeak is a peak calling package for identifying DNA binding sites of proteins in ChIP-Seq experiments. Its algorithm uses hidden Markov models (HMM) and Bayesian statistical methods. The following sample code introduces the identification of peaks with the BayesPeak package as well as the incorporation of read coverage information obtained by the chipseq package. <a href="http://bioconductor.org/packages/release/bioc/html/BayesPeak.html">Documentation</a> [ <a href="http://www.biomedcentral.com/1471-2105/10/299">Publication</a> ]</p><h4><a name="TOC-PICS" id="TOC-PICS"></a>PICS</h4><p>The PICS package applies probabilistic inference to aligned-read ChIP-Seq data in order to identify regions bound by transcription factors. PICS identifies enriched regions by modeling local concentrations of directional reads, and uses DNA fragment length prior information to discriminate closely adjacent binding events via a Bayesian hierarchical t-mixture model. The following sample code uses the test data set from the above BayesPeak package in order to compare the results from both methods by identifying their consensus peak set. <a href="http://www.bioconductor.org/packages/release/bioc/html/PICS.html">Documentation</a> [ <a href="http://www.hubmed.org/display.cgi?uids=20528864">Publication</a> ]</p><h4><a name="TOC-ChIPpeakAnno" id="TOC-ChIPpeakAnno"></a>ChIPpeakAnno</h4><p>The ChIPpeakAnno package provides. batch annotation of the peaks identified from either ChIP-seq or ChIP-chip experiments. It includes functions to retrieve the sequences around peaks, obtain enriched Gene Ontology (GO) terms, find the nearest gene, exon, miRNA or custom features such as most conserved elements and other transcription factor binding sites supplied by users. The package leverages the biomaRt, IRanges, Biostrings, BSgenome, GO.db, multtest and stat packages. <a href="http://bioconductor.org/packages/release/bioc/html/ChIPpeakAnno.html">Documentation</a></p><h4><a name="TOC-Additional-ChIP-Seq-Packages" id="TOC-Additional-ChIP-Seq-Packages"></a>Additional ChIP-Seq Packages</h4><p>DiffBind: <a href="http://www.bioconductor.org/packages/release/bioc/html/DiffBind.html">Documentation</a></p><p>MOSAICS: <a href="http://bioconductor.org/packages/devel/bioc/html/mosaics.html">Documentation</a></p><p>iSeq: <a href="http://bioconductor.org/packages/release/bioc/html/iSeq.html">Documentation</a></p><p>ChIPseqR: <a href="http://bioconductor.org/packages/release/bioc/html/ChIPseqR.html">Documentation</a></p><p>ChiPsim: <a href="http://bioconductor.org/packages/release/bioc/html/ChIPsim.html">Documentation</a></p><p>CSAR: <a href="http://www.bioconductor.org/packages/devel/bioc/html/CSAR.html">Documentation</a></p><p>ChIP-Seq Pipeline: <a href="http://www.bioconductor.org/packages/release/bioc/html/PICS.html">PICS</a>, rGADEM and MotIV (<a href="http://www.rglab.org/pics-and-bioconductor/">developer web site</a>)</p><p>SPP: <a href="http://compbio.med.harvard.edu/Supplements/ChIP-seq/">ChIP-seq processing pipeline</a></p><p><a href="http://compbio.med.harvard.edu/Supplements/ChIP-seq/tutorial.html">SPP Tutorial</a></p><p><a href="http://liulab.dfci.harvard.edu/MACS/index.html">MACS</a></p><p><a href="http://gmdd.shgmo.org/Computational-Biology/ChIP-Seq/download/SIPeS">SIPeS</a></p><h3><a name="TOC-RNA-Seq-Analysis" id="TOC-RNA-Seq-Analysis"></a>RNA-Seq Analysis</h3><h4><a name="TOC-Counting-Reads-that-Overlap-with-Annotation-Ranges-" id="TOC-Counting-Reads-that-Overlap-with-Annotation-Ranges-"></a>Counting Reads that Overlap with Annotation Ranges&nbsp;</h4><p>The GenomicRanges package provides support for importing into R short read alignment data in BAM format (via Rsamtools) and associating them with genomic feature ranges, such as exons or genes. This way one can quantify the number of reads aligning to annotated genomic regions. The package defines general purpose containers for storing genomic intervals as well as more specialized containers for storing alignments against a reference genome. The two main functions for read counting provided by this infrastructure are <span>countOverlaps <span style="color: #000000;"><span>and</span></span> summarizeOverlaps</span>. For their proper usage, it is important to read the corresponding <a href="http://www.bioconductor.org/packages/devel/bioc/vignettes/GenomicRanges/inst/doc/summarizeOverlaps.pdf">PDF manual</a>. <a href="http://bioconductor.org/packages/release/bioc/html/GenomicRanges.html">Documentation</a></p><h4><a name="TOC-Differential-Gene-Expression-Analysis-with-DESeq" id="TOC-Differential-Gene-Expression-Analysis-with-DESeq"></a>Differential Gene Expression Analysis with DESeq</h4><p>The DESeq package contains functions to call differentially expressed genes (DEGs) in count tables based on a model using the negative binomial distribution. It expects as input a data frame with the raw read counts per region/gene of interest (rows) for each test sample (columns).&nbsp; Such a count table can be imported into R or generated from BAM alignment files using the <span>countOverlaps</span> function as introduced above. <a href="http://www.bioconductor.org/packages/release/bioc/html/DESeq.html">Documentation</a></p><h4><a name="TOC-Differential-Gene-Expression-Analysis-with-edgeR" id="TOC-Differential-Gene-Expression-Analysis-with-edgeR"></a>Differential Gene Expression Analysis with edgeR</h4><p>The edgeR package uses empirical Bayes estimation and exact tests based on the negative binomial distribution to call differentially expressed genes (DEGs) in count data.&nbsp;</p>
<p><a href="http://www.bioconductor.org/packages/release/bioc/html/edgeR.html">Documentation</a></p>
<p><span style="color: #000000;">A variety of additional R packages are available for normalizing RNA-Seq read count data and identifying differentially expressed genes (DEG): <br /> </span></p><p><a href="http://bioconductor.org/packages/devel/bioc/html/easyRNASeq.html">easyRNASeq</a> (simplifies read counting per genome feature)</p><p><a href="http://www.bioconductor.org/packages/release/bioc/html/DEXSeq.html">DEXSeq</a> (Inference of differential exon usage);&nbsp;<a href="http://www.bioconductor.org/packages/release/data/experiment/html/parathyroidSE.html">parathyroidSE</a> explains how to generate exon read counts in R</p><p><a href="http://bioconductor.org/packages/release/bioc/html/DEGseq.html">DEGseq</a></p><p><a href="http://www.bioconductor.org/packages/release/bioc/html/baySeq.html">baySeq</a> (also see: <a href="http://www.bioconductor.org/packages/release/bioc/html/segmentSeq.html">segmentSeq</a>)</p><p><a href="http://bioconductor.org/packages/release/bioc/html/Genominator.html">Genominator</a> (<a href="http://www.hubmed.org/display.cgi?uids=20167110">Bullard et al. 2010</a>)</p><div style="text-align: right;"><div style="text-align: left;"><h4><a name="TOC-Detection-of-Alternative-Splice-Junctions" id="TOC-Detection-of-Alternative-Splice-Junctions"></a>Detection of Alternative Splice Junctions</h4>
<p><span style="color: #000000;">Another utility of RNA-Seq experiments is the analysis of splice junctions. The following software suggestions provide this utility:</span></p>
<p><a href="http://woldlab.caltech.edu/rnaseq/">ERANGE<br /> </a><a href="http://tophat.cbcb.umd.edu/">TopHat</a></p><p><a href="http://biogibbs.stanford.edu/%7Ekinfai/SpliceMap/">SpliceMap</a></p><p><a href="http://solidsoftwaretools.com/gf/project/splitseek/">SplitSeek</a></p><h3><a name="TOC-DNA-Methylation-Data-Analysis" id="TOC-DNA-Methylation-Data-Analysis"></a>DNA-Methylation Data Analysis</h3><div><ul>
<li><span style="font-size: 10pt;"><a href="http://www.bioconductor.org/help/course-materials/2012/BiocEurope2012/mattia_pelizzola_methylPipe.pdf">methylPipe</a></span></li>
<li><span style="font-size: 10pt;"><a href="http://www.bioconductor.org/packages/devel/bioc/html/bsseq.html">bsseq</a></span></li>
<li><a href="http://www.bioconductor.org/packages/devel/bioc/html/BiSeq.html">BiSeq</a></li>
<li>Much more under <a href="http://www.bioconductor.org/packages/devel/BiocViews.html#___DNAMethylation">BiocViews</a></li>
</ul></div></div></div><h3><a name="TOC-HT-Seq-Data-Visualization" id="TOC-HT-Seq-Data-Visualization"></a>HT-Seq Data Visualization</h3>
<p><a href="http://www.bioconductor.org/packages/release/bioc/html/ggbio.html">ggbio</a>: ggplot2 extension for genomics data (<a href="http://tengfei.github.com/ggbio/">online manual</a>) <a href="http://www.bioconductor.org/packages/devel/bioc/html/Gviz.html">Gviz</a>:&nbsp;Plotting data and annotation information along genomic coordinates <a href="http://bioconductor.org/packages/release/bioc/html/HilbertVis.html">HilbertVis</a>: Hilbert genome plots</p>
<p><a href="http://bioconductor.org/packages/release/bioc/html/GenomeGraphs.html">GenomeGraphs</a>: Plotting genomic information from Ensembl</p><p><a href="http://www.hubmed.org/display.cgi?uids=18507856">TileQC</a>: Flow Cell Quality Visualization</p><p><a href="http://bioconductor.org/packages/release/bioc/html/rtracklayer.html">rtracklayer</a>: R interface to genome browsers</p><p><a href="http://genoplotr.r-forge.r-project.org/">genoPlotR</a>: Plotting maps of genes and genomes</p><p><a href="http://bioconductor.org/packages/release/bioc/html/Genominator.html">Genominator</a>: Tools for storing, accessing, analyzing and visualizing genomic data.</p><p>&nbsp;</p><p>To install all packages</p><blockquote><p>source("http://bioconductor.org/biocLite.R")<br />biocLite()<br />biocLite(c("ShortRead", "Biostrings", "IRanges", "BSgenome", "rtracklayer", "biomaRt", "chipseq", "ChIPpeakAnno", "Rsamtools", "BayesPeak", "PICS", "GenomicRanges", "DESeq", "edgeR", "leeBamViews", "GenomicFeatures", "BSgenome.Celegans.UCSC.ce2"))</p></blockquote></div>]]></description>
	<dc:creator>John Parker</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37827/genomethreader-gene-prediction-software</guid>
	<pubDate>Wed, 03 Oct 2018 15:34:08 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37827/genomethreader-gene-prediction-software</link>
	<title><![CDATA[GenomeThreader: Gene Prediction Software]]></title>
	<description><![CDATA[<p><em>GenomeThreader</em><span>&nbsp;is a software tool to compute gene structure predictions. The gene structure predictions are calculated using a similarity-based approach where additional cDNA/EST and/or protein sequences are used to predict gene structures via spliced alignments.&nbsp;</span><em>GenomeThreader</em><span>&nbsp;was motivated by disabling limitations in&nbsp;</span><a href="http://bioinformatics.iastate.edu/cgi-bin/gs.cgi"><em>GeneSeqer</em></a><span>, a popular gene prediction program which is widely used for plant genome annotation.</span></p><p>Address of the bookmark: <a href="http://genomethreader.org/" rel="nofollow">http://genomethreader.org/</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/opportunity/view/11494/postdoc-position-at-centre-mediterraneen-de-medecine-moleculaire-nice-france</guid>
  <pubDate>Wed, 04 Jun 2014 07:20:57 -0500</pubDate>
  <link></link>
  <title><![CDATA[Postdoc position at Centre Méditerranéen de Médecine Moléculaire - Nice - France]]></title>
  <description><![CDATA[
<p>The research group of Dr. Michele Trabucchi at the Centre Méditerranéen de Médecine Moléculaire (C3M) at INSERM U1065 (University of Nice Sophia-Antipolis, France) is seeking candidates for a Postdoctoral fellow position to start on October 2014 for 3 years funded by FRM (Fondation pour la Recherche Médicale).<br />The broad interest of the lab is in understanding the expression control and function of small RNAs in activated myeloid cells (visit our webpage to check research interests and publications of the group : http://www.unice.fr/c3m/EN/Equipe10.html ). </p>

<p>The work will focus on the functional studies of small RNAs by using next-generation sequencing approaches.<br /> <br />Candidates should hold a Ph.D. degree and have strong background in bioinformatics.<br />The University of Nice Sophia-Antipolis provides a wide range of facilities and training essential for biomedical research.</p>

<p>Interested applicants should send a PDF with a cover letter stating research interests and qualifications, an updated CV, a summary of previous research experience and contact information for two references to Michele Trabucchi ( mtrabucchi@unice.fr )</p>

<p>Homepage: http://www.unice.fr/c3m/EN/Equipe10.html</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/35899/reference-free-prediction-of-rearrangement-breakpoint-reads</guid>
	<pubDate>Thu, 08 Mar 2018 05:05:25 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/35899/reference-free-prediction-of-rearrangement-breakpoint-reads</link>
	<title><![CDATA[Reference-free prediction of rearrangement breakpoint reads]]></title>
	<description><![CDATA[<p><span>lideSort-BPR (&nbsp;</span><span>b</span><span>&nbsp;reak&nbsp;</span><span>p</span><span>&nbsp;oint&nbsp;</span><span>r</span><span>&nbsp;eads) is based on a fast algorithm for all-against-all comparisons of short reads and theoretical analyses of the number of neighboring reads. When applied to a dataset with a sequencing depth of 100&times;, it finds &sim;88% of the breakpoints correctly with no false-positive reads. Moreover, evaluation on a real prostate cancer dataset shows that the proposed method predicts more fusion transcripts correctly than previous approaches, and yet produces fewer false-positive reads. To our knowledge, this is the first method to detect breakpoint reads without using a reference genome.</span></p>
<p><span>https://github.com/ewijaya/slidesort-bpr</span></p><p>Address of the bookmark: <a href="https://code.google.com/archive/p/slidesort-bpr/" rel="nofollow">https://code.google.com/archive/p/slidesort-bpr/</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>

</channel>
</rss>