<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/38886?offset=290</link>
	<atom:link href="https://bioinformaticsonline.com/related/38886?offset=290" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45289/the-atlas-of-nine-billion-possibilities</guid>
	<pubDate>Wed, 09 Sep 2026 02:07:58 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45289/the-atlas-of-nine-billion-possibilities</link>
	<title><![CDATA[The Atlas of Nine Billion Possibilities]]></title>
	<description><![CDATA[<p>Imagine a book that holds all the instructions for building a human, made up of billions of letters. What if you changed just one letter? Maybe nothing would happen. Or that tiny change could affect how a gene works, quietly shaping a cell&rsquo;s biology or even helping cause disease.</p><p>Scientists face a big challenge with the human genome. They can read its letters, but understanding their roles is much harder. Only about 2 percent of the genome codes for proteins. The rest acts like a huge control panel, deciding when and where genes turn on. With about 9 billion possible single-letter changes, testing them all in a lab just isn&rsquo;t possible.</p><p>So, Google DeepMind asked a new question: what if we could predict what those changes might do?</p><p>This question led to the AlphaGenome Atlas (https://deepmind.google.com/science/alphagenome/atlas?), a detailed map of nearly every possible single-letter change in the human genome. Instead of checking each change one by one, researchers can use the Atlas to spot the ones most likely to matter. The AlphaGenome Variant Impact (AVI) score works like a trail marker, pointing scientists toward the changes worth a closer look.</p><p>This is where the Atlas gets especially useful. Much of the genome lies outside the protein-coding regions, where DNA acts as a switch or controller for genes. AlphaGenome lets researchers explore these areas and see how small changes could affect gene activity.</p><p>An atlas isn't the destination; it's a guide.</p><p>The AlphaGenome Atlas doesn&rsquo;t replace experiments or solve every mystery. Instead, it helps scientists decide where to begin. From billions of possibilities, it turns the vast genetic landscape into something researchers can start to explore.</p><p>There are nine billion possibilities, a vast map, and maybe among them therWith nine billion possibilities and a huge map to explore, there may be clues hidden here to some of medicine&rsquo;s toughest mysteries.</p><p>More at&nbsp;https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/alphagenome-atlas.pdf</p>]]></description>
	<dc:creator>Jitendra Narayan</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45349/finding-the-hidden-switches-the-story-of-kinext-and-protein-kinases</guid>
	<pubDate>Thu, 24 Sep 2026 02:51:28 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45349/finding-the-hidden-switches-the-story-of-kinext-and-protein-kinases</link>
	<title><![CDATA[Finding the Hidden Switches: The Story of KiNext and Protein Kinases]]></title>
	<description><![CDATA[<p>Every newly sequenced genome contains thousands of proteins, but identifying what each protein does is a much harder task. Among these proteins are protein kinases, important molecular regulators that control processes such as cell growth, development, metabolism, stress responses, and signaling. Finding these kinases and determining which families they belong to can reveal important clues about how an organism functions and has evolved.</p><p>This is where KiNext comes into the picture. Introduced in a 2024 study published in BMC Bioinformatics, KiNext is a computational workflow designed to identify and classify protein kinases from predicted protein sequences. Instead of relying on a single search method, it brings together several approaches, including Hidden Markov Models, sequence alignment, phylogenetic analysis, and structural comparison.</p><p>The search begins with a simple question: does a protein contain the characteristics of a kinase? Protein kinases can change considerably during evolution, but important regions of their sequences often retain recognizable patterns. KiNext uses Hidden Markov Models, or HMMs, to detect these patterns. An HMM does not require a protein to be an exact match to a known kinase. Instead, it looks for a statistical sequence signature associated with kinase proteins, making it possible to detect more distant candidates.</p><p>Once potential kinases are identified, KiNext takes the analysis further. It distinguishes conventional eukaryotic protein kinases from atypical protein kinases and then attempts to classify them into different kinase groups and families. This distinction is important because simply identifying a protein as a kinase does not tell the complete story. Different kinase families can have very different evolutionary histories and biological functions.</p><p>The next stage brings evolution into the picture. KiNext can align kinase sequences and construct phylogenetic trees, allowing researchers to examine how newly identified proteins are related to previously characterized kinases. When sequence evidence alone is difficult to interpret, structural information can provide another clue. The workflow can incorporate AlphaFold-predicted structures and Foldseek-based structural comparisons to investigate whether an unusual protein resembles known kinase structures.</p><p>The researchers tested KiNext using two very different organisms: the Pacific oyster, Crassostrea gigas, and the green alga Ostreococcus tauri. In C. gigas, KiNext recovered previously reported kinases while identifying additional candidates. Structural analysis provided further evidence for many of the newly detected proteins. In O. tauri, the workflow similarly recovered most previously reported kinases and identified additional candidates while refining some of their classifications.</p><p>What makes KiNext particularly interesting is not just its ability to find kinases, but how the entire analysis is organized. The workflow uses Nextflow, allowing the different computational steps to be connected into a reproducible pipeline. Containers can also help manage software dependencies, making it easier to run the workflow across different computing environments.</p><p>This reproducibility becomes increasingly important as the number of available genomes continues to grow. A researcher studying one organism may be able to perform an analysis manually, but repeating the same process across hundreds or thousands of genomes quickly becomes impractical. A standardized workflow provides a way to perform the analysis consistently while keeping track of how the results were generated.</p><p>At its core, KiNext demonstrates a broader change taking place in modern genomics. Sequencing a genome provides an enormous amount of information, but the real scientific challenge begins afterward: understanding what all those sequences mean. Protein kinases are only one part of this larger puzzle, yet they are particularly important because they act as molecular switches throughout the cell.</p><p>By combining sequence profiles, evolutionary analysis, and structural evidence within a reproducible computational framework, KiNext provides researchers with a systematic way to uncover these molecular switches. Its real value lies not only in finding more kinases, but in making the process scalable, repeatable, and easier to apply to new genomes.</p><p>As genome sequencing continues to expand across the tree of life, tools such as KiNext can help turn enormous collections of protein sequences into meaningful biological stories&mdash;one kinase at a time.</p><p>Read more about it @</p><p>https://link.springer.com/article/10.1186/s12859-024-05953-w</p>]]></description>
	<dc:creator>LEGE</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36884/halc-high-throughput-algorithm-for-long-read-error-correction</guid>
	<pubDate>Fri, 08 Jun 2018 10:47:41 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36884/halc-high-throughput-algorithm-for-long-read-error-correction</link>
	<title><![CDATA[HALC: High throughput algorithm for long read error correction]]></title>
	<description><![CDATA[HALC, a high throughput algorithm for long read error correction. HALC aligns the long reads to short read contigs from the same species with a relatively low identity requirement so that a long read region can be aligned to at least one contig region, including its true genome region’s repeats in the contigs sufficiently similar to it (similar repeat based alignment approach)

HALC was able to obtain 6.7-41.1% higher throughput than the existing algorithms while maintaining comparable accuracy. The HALC corrected long reads can thus result in 11.4-60.7% longer assembled contigs than the existing algorithms.<p>Address of the bookmark: <a href="https://github.com/lanl001/halc" rel="nofollow">https://github.com/lanl001/halc</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/35059/lrcstats-long-read-correction-statistics</guid>
	<pubDate>Fri, 05 Jan 2018 04:04:20 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/35059/lrcstats-long-read-correction-statistics</link>
	<title><![CDATA[LRCstats: Long Read Correction Statistics]]></title>
	<description><![CDATA[<p>LRCstats is an open-source pipeline for benchmarking DNA long read correction algorithms for long reads outputted by third generation sequencing technology such as machines produced by Pacific Biosciences. The reads produced by third generation sequencing technology, as the name suggests, are longer in length than reads produced by next generation sequencing technologies, such as those produced by Illumina. However, long reads are plagued by high error rates, which can cause issues in downstream analysis. Long read correction algorithms reduce the error rate of long reads either through self-correcting methods or using accurate, short reads outputted by next generation sequencing technologies to correct long reads.</p>
<p>Of course, some long read correction algorithms are better than others, and developers of long read correction algorithms will wish to compare their algorithm with others currently available. LRCstats benchmarks long read correction algorithms using long reads produced by simulators (such as SimLoRD or PBSim) where the two-way alignments between the uncorrected long reads (uLR) and the corresponding sequences in the reference genome (Ref) are given in some sort of alignment file and then aligning the corrected long reads (cLR) to the Ref-uLR two-way alignments to create three-way alignments using a dynamic programming algorithm. Statistics on these three-way alignments are then collected, such as the overall error rates of the corrected long reads.</p>
<p>https://www.healthcare.uiowa.edu/labs/au/LSC/</p><p>Address of the bookmark: <a href="https://github.com/cchauve/lrcstats" rel="nofollow">https://github.com/cchauve/lrcstats</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37645/lsc-improving-pacbio-long-read-accuracy-by-short-read-alignment</guid>
	<pubDate>Thu, 06 Sep 2018 16:27:35 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37645/lsc-improving-pacbio-long-read-accuracy-by-short-read-alignment</link>
	<title><![CDATA[LSC: Improving PacBio Long Read Accuracy by Short Read Alignment]]></title>
	<description><![CDATA[<ul>
<li>Added Command line argument support.</li>
<li>Multi-stage execution modes.</li>
<li>Support for parallelization. Now execution proceeds in batches of long reads the size of which can be set by --long_read_batch_size N.</li>
<li>Better compressed intermediate files.</li>
<li>Added utilities folder.</li>
<li>Added support for multiple short read files.</li>
<li>Removed use of configuration file.</li>
</ul><p>Address of the bookmark: <a href="https://www.healthcare.uiowa.edu/labs/au/LSC/" rel="nofollow">https://www.healthcare.uiowa.edu/labs/au/LSC/</a></p>]]></description>
	<dc:creator>Abhimanyu Singh</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/34445/inc-seq-accurate-single-molecule-reads-using-nanopore-sequencing</guid>
	<pubDate>Mon, 27 Nov 2017 10:38:56 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/34445/inc-seq-accurate-single-molecule-reads-using-nanopore-sequencing</link>
	<title><![CDATA[INC-Seq: accurate single molecule reads using nanopore sequencing]]></title>
	<description><![CDATA[<p><span>INC-Seq reads enabled accurate species-level classification, identification of species at 0.1&nbsp;% abundance and robust quantification of relative abundances, providing a cheap and effective approach for pathogen detection and microbiome profiling on the MinION system.</span></p><p>Address of the bookmark: <a href="https://github.com/CSB5/INC-Seq" rel="nofollow">https://github.com/CSB5/INC-Seq</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36607/tarean-a-computational-tool-for-identification-and-characterization-of-satellite-dna-from-unassembled-short-reads</guid>
	<pubDate>Tue, 15 May 2018 02:53:11 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36607/tarean-a-computational-tool-for-identification-and-characterization-of-satellite-dna-from-unassembled-short-reads</link>
	<title><![CDATA[TAREAN: A computational tool for identification and characterization of satellite DNA from unassembled short reads]]></title>
	<description><![CDATA[<p><strong>TA</strong>ndem&nbsp;<strong>RE</strong>peat&nbsp;<strong>AN</strong>alyzer -TAREAN &ndash; is a computational pipeline for&nbsp;<strong>unsupervised identification of satellite repeats</strong>&nbsp;from unassembled sequence reads. The pipeline uses low-pass whole genome sequence reads and performs their graph-based clustering. Resulting clusters, representing all types of repeats, are then examined for the presence of circular structures and putative satellite repeats are reported.</p>
<p><em><strong>How to use TAREAN</strong></em>:</p>
<ul>
<li>Install a local instance of the pipeline using its source code available from&nbsp;<a href="https://bitbucket.org/petrnovak/repex_tarean" target="_blank" title="TAREAN source code">bitbucket repository</a>.</li>
<li>Use&nbsp; public Galaxy-based server at&nbsp;<a href="https://repeatexplorer-elixir.cerit-sc.cz/" target="_blank">https://repeatexplorer-elixir.cerit-sc.cz/</a>. The server is provided in frame of the&nbsp;<a href="https://www.elixir-czech.cz/" target="_blank">Elixir CZ project</a>&nbsp;and is maintained by&nbsp;<a href="https://www.cesnet.cz/" target="_blank">CESNET</a>&nbsp;and&nbsp;<a href="https://www.cerit-sc.cz/en/index.html" target="_blank">CERIT-SC</a>. Simple registration is required to use this service.</li>
</ul>
<p>Development of TAREAN was supported by&nbsp;<a href="https://www.elixir-czech.cz/" target="_blank" title="ELIXIR-CZ">ELIXIR CZ</a>&nbsp;research infrastructure project (MEYS Grant No: LM2015047).</p>
<p><strong><em>References</em></strong></p>
<p>Novak, P., Avila Robledillo, L., Koblizkova, A., Vrbova, I., Neumann, P., Macas, J. (2017) &ndash;&nbsp;<a href="https://academic.oup.com/nar/article/3574061/" target="_blank">TAREAN: a computational tool for identification and characterization of satellite DNA from unassembled short reads</a>.&nbsp;<em>Nucleic Acids Res.</em>, doi:10.1093/nar/gkx257</p><p>Address of the bookmark: <a href="https://bitbucket.org/petrnovak/repex_tarean" rel="nofollow">https://bitbucket.org/petrnovak/repex_tarean</a></p>]]></description>
	<dc:creator>Surabhi Chaudhary</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36800/genomemapper-simultaneous-alignment-of-short-reads-against-multiple-genomes</guid>
	<pubDate>Fri, 25 May 2018 09:29:44 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36800/genomemapper-simultaneous-alignment-of-short-reads-against-multiple-genomes</link>
	<title><![CDATA[GenomeMapper: Simultaneous alignment of short reads against multiple genomes]]></title>
	<description><![CDATA[GenomeMapper is a short read mapping tool designed for accurate read alignments. It quickly aligns millions of reads either with ungapped or gapped alignments. It can be used to align against multiple genomes simulanteously or against a single reference. If you are unsure which one is the appropriate GenomeMapper, you might want to use the latter

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2768987/<p>Address of the bookmark: <a href="http://1001genomes.org/software/genomemapper.html" rel="nofollow">http://1001genomes.org/software/genomemapper.html</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37574/simlord-a-read-simulator-for-third-generation-sequencing-reads</guid>
	<pubDate>Wed, 22 Aug 2018 10:40:27 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37574/simlord-a-read-simulator-for-third-generation-sequencing-reads</link>
	<title><![CDATA[SimLoRD: A read simulator for third generation sequencing reads]]></title>
	<description><![CDATA[<p>SimLoRD is a read simulator for third generation sequencing reads and is currently focused on the Pacific Biosciences SMRT error model.</p>
<p>Reads are simulated from both strands of a provided or randomly generated reference sequence.</p>
<div id="rst-header-features">
<ul>
<li>The reference can be read from a FASTA file or randomly generated with a given GC content. It can consist of several chromosomes, whose structure is respected when drawing reads. (Simulation of genome rearrangements may be incorporated at a later stage.)</li>
<li>The read lengths can be determined in four ways: drawing from a log-normal distribution (typical for genomic DNA), sampling from an existing FASTQ file (typical for RNA), sampling from a a text file with integers (RNA), or using a fixed length</li>
<li>Quality values and number of passes depend on fragment length.</li>
<li>Provided subread error probabilities are modified according to number of passes</li>
<li>Outputs reads in FASTQ format and alignments in SAM format</li>
</ul>
</div><p>Address of the bookmark: <a href="https://bitbucket.org/genomeinformatics/simlord/" rel="nofollow">https://bitbucket.org/genomeinformatics/simlord/</a></p>]]></description>
	<dc:creator>Aaryan Lokwani</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/39671/flye-fast-and-accurate-de-novo-assembler-for-single-molecule-sequencing-reads</guid>
	<pubDate>Sat, 06 Jul 2019 03:48:22 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/39671/flye-fast-and-accurate-de-novo-assembler-for-single-molecule-sequencing-reads</link>
	<title><![CDATA[Flye: Fast and accurate de novo assembler for single molecule sequencing reads]]></title>
	<description><![CDATA[<p><span>Flye is a de novo assembler for single molecule sequencing reads, such as those produced by PacBio and Oxford Nanopore Technologies. It is designed for a wide range of datasets, from small bacterial projects to large mammalian-scale assemblies. The package represents a complete pipeline: it takes raw PB / ONT reads as input and outputs polished contigs. Flye also includes a special mode for metagenome assembly.</span></p><p>Address of the bookmark: <a href="https://github.com/fenderglass/Flye" rel="nofollow">https://github.com/fenderglass/Flye</a></p>]]></description>
	<dc:creator>Rahul Nayak</dc:creator>
</item>

</channel>
</rss>