Every year, theoretical physicists dream up hundreds of ideas for how new physics could materialize in the data produced by the Large Hadron Collider. As a physicist on one of the LHC’s biggest experiments, I love the enthusiasm. But testing each one of these ideas takes time—often, the length of an entire PhD thesis. We just don’t have the people power to hack away at all these clever ideas individually. Even more, the data volume generated by the LHC’s collisions is extraordinary; far beyond what we can record and store. We only have microseconds to select the most promising collisions, and if a theorist’s cool new idea doesn’t make the cut, the events that could prove it right are automatically dumped in the trash.

This has always made me and my colleagues on the CMS Experiment at CERN deeply uncomfortable. What if we’re missing out on something extraordinary simply because a new idea didn’t make our “this could be interesting” list?

Illustration by Sandbox Studio, Chicago with Thumy Phan

One idea is that—in addition to our “this could be interesting” list—we also keep a random selection of events that might not seem interesting at first glance, but could later reveal unexpected patterns. But new physics is so rare that catching it in a random sample is like throwing a fishing net into the bay and catching a kraken. Another solution could be to reduce the file sizes of the data we store, like going from high-resolution photographs to pixelated JPEGs. This would let us store a much higher fraction of events. But it would also mean that if we do find evidence of new physics, we would only have a blurry image of it.

So what can we do?

Thanks to a Genesis Mission award from the US Department of Energy, which was granted to my group at the University of Colorado Boulder, and our collaborators at Fermilab, UCSD, and Johns Hopkins, we no longer have to compromise. We can have our selection of high-quality data. We can have our pixelated thumbnails. We can adapt the concept of “random sampling” so that it is no longer random, but optimized to search for the strangest, most amazing events. And we can do all of this while staying within our computing power constraints.

How? By developing ultra-fast artificial intelligence and deploying it at the earliest stages of data collection.

First, we are reimagining something we call an anomaly detection trigger. Triggers are hardware and software tools that automatically sort our data into “this is interesting and should be saved” or “this is boring and can be chucked.”

Traditional trigger systems search for pre-programmed patterns in the data. But for the anomaly detection trigger, we don’t tell it what to look for. We simply ask, “Does this event look different from all the others?” If the answer is yes, we flag and save it. This anomaly trigger will not replace our traditional triggers, but instead add a special “anomalous” data set that preserves the strangest, most amazing events. With this data set, we no longer need to individually test every new physics model; we can simply see if any weird patterns emerge.

Illustration by Sandbox Studio, Chicago

We deployed an initial prototype during LHC Run 3 and proved that the concept works. With the Genesis Mission grant, we can transform this idea from a proof of principle into full production and deploy it during the High Luminosity run of the LHC, which will create some of the most complex data ever seen and at a rate previously unimaginable.

Second, we are pushing a concept called data scouting. In addition to keeping our high-resolution data of the things we want to study (i.e., Higgs bosons), we also want to keep everything else—but in a much more compact, “pixelated” form. The concept is like covering a nature reserve with cheap, low‑resolution wildlife cameras. Sure, we will still have our high-resolution cameras set-up in places where we know we will see something cool, but we will also have the opportunity to catch a blurry tentacle unexpectedly emerging from the bay.

And this brings us to the final part: If we find an anomaly, we don’t just flag it. We use AI to reverse‑engineer what made those events strange and then write new trigger criteria so that we can capture that blurry tentacle in full resolution if it ever appears again. (And thus, maybe finally catch our kraken.)

This three‑part strategy—anomaly detection, data scouting, and AI‑assisted trigger redesign—will give us the flexibility to catch not only the strange things dreamed-up by theorists, but the events nobody saw coming. In this way, these tools are not just new technology: They are a paradigm shift. And for me, that’s the most exciting part of this research. We no longer have to pretend that we know what new physics will look like. We can openly embrace our own ignorance and search for physics signatures that lie beyond the human imagination.