Imagine trying to solve a mystery with thousands of clues—but some of the clues are labelled incorrectly.
That is the challenge researchers face when studying how bacteria adapt to different environments. A bacterial gene may appear to be linked to a particular habitat, when the real reason is simply that closely related bacteria happen to live there.
In their 2025 Genome Biology paper, Bujdoš, Walter, and O’Toole introduce aurora, a machine-learning tool designed to tackle this problem.
Aurora identifies potentially mislabeled or unusual bacterial strains before performing genome-wide association studies (GWAS). By cleaning up the dataset and accounting for bacterial evolutionary relationships, it can help researchers find genetic features that are more genuinely connected to habitat adaptation.
The researchers tested aurora using simulated and real bacterial datasets and found that it could recover important genetic associations even when datasets contained misleading labels.
The bigger lesson is simple: better biological discoveries often begin with better data.
Aurora gives researchers a new way to separate real genetic clues from misleading ones—and could help us better understand how bacteria adapt, survive, and evolve in the environments they call home.
Based on Bujdoš et al., “aurora: a machine learning GWAS tool for analyzing microbial habitat adaptation,” Genome Biology (2025).[Read the original paper](https://link.springer.com/article/10.1186/s13059-025-03524-7?)