<?xml version='1.0'?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:georss="http://www.georss.org/georss" xmlns:atom="http://www.w3.org/2005/Atom" >
<channel>
	<title><![CDATA[BOL: Related items]]></title>
	<link>https://bioinformaticsonline.com/related/34620?offset=370</link>
	<atom:link href="https://bioinformaticsonline.com/related/34620?offset=370" rel="self" type="application/rss+xml" />
	<description><![CDATA[]]></description>
	
	<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45306/the-genes-that-travel-together-a-genomic-detective-story</guid>
	<pubDate>Fri, 11 Sep 2026 13:00:36 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45306/the-genes-that-travel-together-a-genomic-detective-story</link>
	<title><![CDATA[The Genes That Travel Together: A Genomic Detective Story]]></title>
	<description><![CDATA[<p>Imagine looking through thousands of microbial genomes and discovering two genes that repeatedly appear together. The obvious conclusion is that they must somehow be connected&mdash;that perhaps they work together, participate in the same pathway, or depend on each other. But evolution has a way of leaving misleading clues. What if these two genes are found together simply because they were inherited from the same ancient ancestor? In that case, their apparent association may have little to do with their biological function; it may simply be a reflection of shared evolutionary history. This is the intriguing problem addressed by CORGIAS (https://github.com/ynishimuraLv/corgias), a computational framework developed to distinguish genuine gene associations from correlations created by common ancestry.</p><p>Instead of asking only whether two genes occur together across genomes, CORGIAS asks a deeper question: did these genes actually evolve together, repeatedly gaining or losing their presence in concert, or are they merely travelling together because of their inherited history? The framework introduces two complementary approaches, Ancestral State Adjustment (ASA) and the Simultaneous EVolution test (SEV), which incorporate evolutionary information into the search for gene associations. This distinction becomes increasingly important as genome sequencing continues to uncover enormous numbers of microbial genomes, many from organisms that have never been cultured and whose genes remain functionally mysterious. In this growing genomic landscape, simply finding a gene is no longer enough&mdash;we need clues about what that gene might be doing. CORGIAS approaches this problem almost like a genomic detective: rather than treating every association as evidence, it reconstructs the evolutionary story behind the association and asks whether the evidence survives that history. By doing so, it offers a way to separate coincidence from biological connection and potentially uncover functional relationships hidden within the vast microbial genomic landscape. The broader message is beautifully simple: genes do not evolve in isolation, and understanding what they do may require us to look beyond where they are today and reconstruct how they arrived there. Sometimes, the most important clue in a genome is not the gene itself, but the evolutionary journey it has taken.</p><p>More at&nbsp;https://academic.oup.com/nargab/article/7/4/lqaf182/8377910?</p>]]></description>
	<dc:creator>BioStar</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/blog/view/45343/simulating-the-unseen-new-tools-reshaping-metagenomic-research</guid>
	<pubDate>Wed, 23 Sep 2026 09:42:30 -0500</pubDate>
	<link>https://bioinformaticsonline.com/blog/view/45343/simulating-the-unseen-new-tools-reshaping-metagenomic-research</link>
	<title><![CDATA[Simulating the Unseen: New Tools Reshaping Metagenomic Research]]></title>
	<description><![CDATA[<p>When benchmarking a new metagenomic pipeline the first step is generally to face a simple problem. A bioinformatician might have a good deal of real sequencing data, but there's a drawback in that the actual biological composition underlying those reads is almost never known. What organisms really were there? In what quantities? How many sequencing errors occurred? Can the pipeline correctly reconstruct the community? Since the true answers are unknown, it is difficult to deal with these questions.</p><p>Metagenomic simulators are a way of dealing with this issue; they produce artificial sequencing datasets in which both the composition of the microorganisms and the expected results are known. Researchers can then test a pipeline on such a controlled dataset and see precisely how well it functions. However, as metagenomic studies have become larger, more diverse, and more complex, simply generating artificial reads is no longer sufficient. The simulators now have to replicate community variation, sequencing errors, different sequencing platforms, massive data volumes, and even laboratory-based biases.</p><p>MIDASim (2024) was designed with a particular emphasis on microbial community dynamics. Unlike previous tools which regarded each simulated sample as individual, MIDASim is able to efficiently generate realistic multi-sample microbiome datasets, covering the cases where changes over time are being tracked. This feature is particularly useful when investigating how microbial communities vary between people, different experimental setups, or at various times. By overcoming the computational limitations of earlier simulation methods, MIDASim has made it easier to carry out large and complex microbiome simulation studies.</p><p>A major challenge is scale. Since modern metagenomic projects can involve millions or even billions of sequencing reads, generating such large synthetic datasets can become a significant computational problem. In order to tackle this issue, Izzy (2025) was developed. As a high-throughput metagenomic read simulator, Izzy is designed with speed in mind while still reproducing the typical patterns of sequencing errors. The people who developed it stated that it can be as much as 60 times faster than earlier tools, which makes large-scale simulations far more practical. Now, researchers can begin to ask not how long a simulation will take but rather what size dataset they want to test.</p><p>At the same time, the sequencing technology became more diverse. Researchers were not any longer restricted to using Illumina short reads since Oxford Nanopore and PacBio long-read sequencing also became important in the field of metagenomic research. In order to address these new requirements, MeSS (Metagenomic Sequence Simulator, 2025) was developed. The programme is based on a Snakemake framework and is therefore compatible with the Illumina, Oxford Nanopore and PacBio platforms. MeSS has also been designed to operate efficiently and to use less memory than earlier tools; it provides ready-made templates for the microbiomes of different sites in the human body, thus giving researchers a simple means of beginning to generate realistic simulated communities.</p><p>The simulation problem became even more specialized. Suppose that sequencing is not based on uniform sampling of DNA? In the case of targeted sequencing and hybridization-capture experiments, particular sequences are deliberately enriched, and the probability of capturing a target depends on both the probe and the target sequence. A standard model based on uniform sampling might fail to take into account this important feature of real experiments. That is why RAmpSim has been designed specifically to handle simulations for capture-based and targeted sequencing. Instead of depending solely on the assumption of uniform sampling, it includes a thermodynamic nearest-neighbor energy model to represent probe&ndash;target interactions and sequence-dependent enrichment. As a result, it is especially suitable for applications such as hybridization capture in metagenomics, where the efficiency of recovering a sequence can depend strongly on how it relates to the capture probes. RAmpSim thus reflects a wider trend towards simulations that aim not only to reproduce sequencing output but also key aspects of the experimental process itself.</p><p>MIDASim, Izzy, MeSS and RAmpSim all demonstrate how fast metagenomic simulation is evolving. MIDASim is able to deal with realistic variation within a large number of and changing microbial communities. Izzy overcomes the problem of generating very large datasets. MeSS provides support for a number of sequencing technologies in an efficient and repeatable manner. RAmpSim increases the experimental realism for capture-based sequencing. Although each of these tools addresses a separate issue, they all contribute to advancing the field.</p><p>Modern metagenomic simulation therefore aims not only at producing artificial FASTQ files. Researchers nowadays desire simulated datasets that reflect community changes, incorporate the characteristics of the sequencing platform, include real error patterns, take into account computational scale, and account for experimental biases. As metagenomic workflows become more advanced, the simulators have to keep up with this development. The tools are now going beyond that of simple data generators; they produce controlled versions of complex metagenomic experiments and thus provide researchers with something that ordinary sequencing data seldom does: a dataset in which the answer is known before any analysis takes place.</p>]]></description>
	<dc:creator>LEGE</dc:creator>
</item>

<item>
  <guid isPermaLink='true'>https://bioinformaticsonline.com/researchlabs/view/43044/kanthida-lab</guid>
  <pubDate>Wed, 28 Apr 2021 02:27:22 -0500</pubDate>
  <link></link>
  <title><![CDATA[Kanthida Lab !]]></title>
  <description><![CDATA[
<p>Research Interest: </p>

<p>Bioinformatics </p>

<p>High-throughput and high-dimensional data analysis</p>

<p>Microbiome data analysis (Main focus)</p>

<p>Next-generation and third-generation sequencing data analysis for genomics</p>

<p>Gene expression data analysis</p>

<p>Machine learning for biological data</p>

<p>Biomarkers identification </p>

<p>Database and web-application for biological data</p>

<p>More at <br />https://sites.google.com/mail.kmutt.ac.th/kanthida-k/home?authuser=0</p>
]]></description>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/38829/nquire-a-statistical-framework-for-ploidy-estimation-using-ngs-short-read-data</guid>
	<pubDate>Thu, 31 Jan 2019 05:12:19 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/38829/nquire-a-statistical-framework-for-ploidy-estimation-using-ngs-short-read-data</link>
	<title><![CDATA[nQuire: A statistical framework for ploidy estimation using NGS short-read data]]></title>
	<description><![CDATA[<p>nQuire implements a set of commands to estimate ploidy level of individuals from species, where recent polyploidization occurred and intraspecific ploidy variation is observed. Specifically, nQuire uses next-generation sequencing data to distinguish between diploids, triploids and tetraploids, on the basis of frequency distributions at variant sites where only two bases are segregating.</p>
<p>For more background see also the publication at&nbsp;<a href="https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-018-2128-z">BMC Bioinformatics</a>.</p>
<p>https://github.com/clwgg/nQuire</p><p>Address of the bookmark: <a href="https://github.com/clwgg/nQuire" rel="nofollow">https://github.com/clwgg/nQuire</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/30831/fsa-fast-statistical-alignment</guid>
	<pubDate>Mon, 06 Feb 2017 04:26:01 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/30831/fsa-fast-statistical-alignment</link>
	<title><![CDATA[FSA: Fast Statistical Alignment]]></title>
	<description><![CDATA[<p><span>FSA is a probabilistic multiple sequence alignment algorithm which uses a "distance-based" approach to aligning homologous protein, RNA or DNA sequences. Much as distance-based phylogenetic reconstruction methods like Neighbor-Joining build a phylogeny using only pairwise divergence estimates, FSA builds a multiple alignment using only pairwise estimations of homology. This is made possible by the sequence annealing technique for constructing a multiple alignment from pairwise comparisons, developed by Ariel Schwartz in&nbsp;</span><a href="http://www.eecs.berkeley.edu/Pubs/TechRpts/2007/EECS-2007-39.html">"Posterior Decoding Methods for Optimization and Control of Multiple Alignments</a><span>."</span></p>
<p>FSA brings the high accuracies previously available only for small-scale analyses of proteins or RNAs to large-scale problems such as aligning thousands of sequences or megabase-long sequences. FSA introduces several novel methods for constructing better alignments:</p>
<ul>
<li>FSA uses machine-learning techniques to estimate gap and substitution parameters on the fly for each set of input sequences. This "query-specific learning" alignment method makes FSA very robust: it can produce superior alignments of sets of homologous sequences which are subject to very different evolutionary constraints.</li>
<li>FSA is capable of aligning hundreds or even thousands of sequences using a randomized inference algorithm to reduce the computational cost of multiple alignment. This randomized inference can be over ten times faster than a direct approach with little loss of accuracy.</li>
<li>FSA can quickly align very long sequences using the "anchor annealing" technique for resolving anchors and projecting them with transitive anchoring. It then stitches together the alignment between the anchors using the methods described above.</li>
<li>The included GUI, MAD (Multiple Alignment Display), can display the intermediate alignments produced by FSA, where each character is colored according to the probability that it is correctly aligned (see the picture and&nbsp;<a href="http://fsa.sourceforge.net/images/Suchard_SIV.fsa.mov">movie</a>&nbsp;at the top of the page).</li>
</ul>
<p><span>You can see more information on the&nbsp;</span><a href="http://fsa.sourceforge.net/FAQ.html">FAQ</a><span>.&nbsp;</span></p>
<p>&nbsp;</p><p>Address of the bookmark: <a href="http://fsa.sourceforge.net/" rel="nofollow">http://fsa.sourceforge.net/</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/42310/dada2-fast-and-accurate-sample-inference-from-amplicon-data-with-single-nucleotide-resolution</guid>
	<pubDate>Tue, 10 Nov 2020 20:26:00 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/42310/dada2-fast-and-accurate-sample-inference-from-amplicon-data-with-single-nucleotide-resolution</link>
	<title><![CDATA[DADA2: Fast and accurate sample inference from amplicon data with single-nucleotide resolution]]></title>
	<description><![CDATA[<p>The&nbsp;<a href="https://benjjneb.github.io/dada2/tutorial.html">DADA2 tutorial</a>&nbsp;goes through a typical workflow for paired end Illumina Miseq data: raw amplicon sequencing data is processed into the table of exact&nbsp;<strong>amplicon sequence variants (ASVs)</strong>&nbsp;present in each sample.</p>
<p>The&nbsp;<a href="https://benjjneb.github.io/dada2/bigdata.html">DADA2 Workflow on Big Data</a>&nbsp;goes through workflow optimized to run on large datasets (10s of millions to billions of reads).</p>
<p>An&nbsp;<a href="https://benjjneb.github.io/dada2/ITS_workflow.html">ITS-specific version of the DADA2 workflow</a>&nbsp;identifies and verifiably removes primers on both ends of each ITS read, a key step due to the variable length of the ITS region.</p>
<p>Short demonstrations of&nbsp;<a href="https://benjjneb.github.io/dada2/assign.html">assigning taxonomy</a>&nbsp;and&nbsp;<a href="https://benjjneb.github.io/dada2/assign.html">assigning species</a>&nbsp;to sequences.</p><p>Address of the bookmark: <a href="https://benjjneb.github.io/dada2/index.html" rel="nofollow">https://benjjneb.github.io/dada2/index.html</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/36755/minialign-fast-and-accurate-alignment-tool-for-pacbio-and-nanopore-long-reads</guid>
	<pubDate>Thu, 24 May 2018 08:33:26 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/36755/minialign-fast-and-accurate-alignment-tool-for-pacbio-and-nanopore-long-reads</link>
	<title><![CDATA[minialign: fast and accurate alignment tool for PacBio and Nanopore long reads]]></title>
	<description><![CDATA[Minialign is a little bit fast and moderately accurate nucleotide sequence alignment tool designed for PacBio and Nanopore long reads. It is built on three key algorithms, minimizer-based index of the minimap overlapper, array-based seed chaining, and SIMD-parallel Smith-Waterman-Gotoh extension.<p>Address of the bookmark: <a href="https://github.com/ocxtal/minialign" rel="nofollow">https://github.com/ocxtal/minialign</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37496/gsearch-a-fast-and-flexible-general-search-tool-for-whole-genome-sequencing</guid>
	<pubDate>Mon, 06 Aug 2018 17:19:15 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37496/gsearch-a-fast-and-flexible-general-search-tool-for-whole-genome-sequencing</link>
	<title><![CDATA[gSearch: a fast and flexible general search tool for whole-genome sequencing]]></title>
	<description><![CDATA[<p><span>gSearch compares sequence variants in the Genome Variation Format (GVF) or Variant Call Format (VCF) with a pre-compiled annotation or with variants in other genomes. Its search algorithms are subsequently optimized and implemented in a multi-threaded manner.&nbsp;</span></p><p>Address of the bookmark: <a href="http://ml.ssu.ac.kr/gSearch/index.html" rel="nofollow">http://ml.ssu.ac.kr/gSearch/index.html</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/37937/frodock-20-fast-protein%E2%80%93protein-docking-server</guid>
	<pubDate>Wed, 17 Oct 2018 04:31:30 -0500</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/37937/frodock-20-fast-protein%E2%80%93protein-docking-server</link>
	<title><![CDATA[FRODOCK 2.0: fast protein–protein docking server]]></title>
	<description><![CDATA[<p><span>frodock: a&nbsp;user-friendly protein&ndash;protein docking server based on an improved version of FRODOCK that includes a complementary knowledge-based potential. The web interface provides a very effective tool to explore and select protein&ndash;protein models and interactively screen them against experimental distance constraints. The competitive success rates and efficiency achieved allow the retrieval of reliable potential protein&ndash;protein binding conformations that can be further refined with more computationally demanding strategies.</span></p><p>Address of the bookmark: <a href="http://frodock.chaconlab.org/" rel="nofollow">http://frodock.chaconlab.org/</a></p>]]></description>
	<dc:creator>Neel</dc:creator>
</item>
<item>
	<guid isPermaLink="true">https://bioinformaticsonline.com/bookmarks/view/40217/shouji-a-fast-and-efficient-pre-alignment-filter-for-sequence-alignment</guid>
	<pubDate>Mon, 04 Nov 2019 07:09:45 -0600</pubDate>
	<link>https://bioinformaticsonline.com/bookmarks/view/40217/shouji-a-fast-and-efficient-pre-alignment-filter-for-sequence-alignment</link>
	<title><![CDATA[Shouji: a fast and efficient pre-alignment filter for sequence alignment]]></title>
	<description><![CDATA[<p>The ability to generate massive amounts of sequencing data continues to overwhelm the processing capacity of existing algorithms and compute infrastructures. In this work, we explore the use of hardware/software co-design and hardware acceleration to significantly reduce the execution time of short sequence alignment, a crucial step in analyzing sequenced genomes.</p>
<p>&nbsp;<img src="https://github.com/BilkentCompGen/Shoji/raw/master/Figure1-GitHub.png" alt="image" style="border: 0px;"></p>
<p>We introduce Shouji, a highly parallel and accurate pre-alignment filter that remarkably reduces the need for computationally-costly dynamic programming algorithms. The first key idea of our proposed pre-alignment filter is to provide high filtering accuracy by correctly detecting all common subsequences shared between two given sequences. The second key idea is to design a hardware accelerator design that adopts modern FPGA (field-programmable gate array) architectures to further boost the performance of our algorithm.</p>
<p>More at <a href="https://github.com/CMU-SAFARI/Shouji">https://github.com/CMU-SAFARI/Shouji</a></p><p>Address of the bookmark: <a href="https://github.com/CMU-SAFARI/Shouji" rel="nofollow">https://github.com/CMU-SAFARI/Shouji</a></p>]]></description>
	<dc:creator>Jit</dc:creator>
</item>

</channel>
</rss>