Story

Attributing the source of Campylobacter infections

How national genomic surveillance and machine learning were used to estimate the origins of human infection in the United States.

Multi-panel source-attribution figure with maps, phylogenies and model results

The problem

Foodborne infection is common; tracing its origin is not.

Campylobacter infections are usually sporadic rather than part of a clearly bounded outbreak. A patient may have encountered the bacterium through food, livestock, water, wildlife or several overlapping routes. Traditional case-control studies remain essential, but they are difficult to repeat continuously at the scale of national surveillance.

Bacterial genomes offer a second line of evidence. Host-associated populations acquire different combinations of genetic variants over time. If those patterns are learned from well-characterised source isolates, they can be used to estimate where human isolates most likely originated.

Building a national genomic comparison

The study assembled 8,856 genomes from human infections and 16,703 genomes from potential sources collected through US surveillance between 2009 and 2019. Poultry, cattle, wild birds and pork were represented. Random-forest and probabilistic approaches were used to learn source-associated genomic patterns, test performance and attribute human cases.

The important shift is from asking whether a source can cause infection to estimating how much it contributes across the population.

What the models suggested

Poultry was the dominant attributed source, accounting for about 68% of infections. Cattle contributed roughly 28%, while wild birds and pork made smaller contributions. These estimates are not permanent biological constants: they reflect the populations, sampling and time period studied. Their value lies in providing a quantitative baseline that can be updated as surveillance changes.

The analysis also showed why attribution and antimicrobial-resistance surveillance belong together. Poultry-associated populations contributed substantially to the burden of infection and included lineages in which resistance had increased over time.

Why this matters for control

Source attribution can help public-health agencies target interventions where the likely return is greatest, then measure whether the population signal changes. It also creates a framework for comparing regions and years, provided that models are retrained with representative data and uncertainty is reported honestly.

The next challenge is contextualisation. Models built from national data may perform differently in local populations or countries with different production systems. That is why projects such as GETcampy and CCC place so much emphasis on collecting local source genomes rather than importing assumptions from elsewhere.