As a physics professor at the University of Alabama, I can tell you something for certain: becoming a physicist is a lot harder today than when I was a student 30 years ago. Our graduate students don’t just learn physics. Before they can get to the good stuff—creative problem-solving and pushing the bounds of human knowledge—they also need to master coding, data science, advanced mathematics, detector engineering, theoretical modeling, computer simulations, machine learning and much more.
And I know what my students do when faced with this massive wall of information: they ask Claude or ChatGPT for help. And who can blame them? I’ll be the first to admit that these tools can provide mentoring that makes the barrier to entry easier. But AI is not all sunshine and rainbows. Not only is there a lot of controversy about using AI in education, but the tools themselves are imperfect and will hallucinate fake information when pushed to the boundary of human knowledge (which is exactly where our research lies).
This puts professors like me in a difficult situation. Should I forbid the use of AI in the classroom and force students to learn all these skills the old-fashioned way? Or should I embrace it and risk students learning the wrong answers? (Or not learning at all?)

I didn’t like either of these options, which is why my colleagues at Fermilab and our team at the University of Alabama decided to pursue a third idea: take commercially available large language models like ChatGPT and Claude and educate them to behave like our teaching assistants.
Why? Because whether we like it or not, our students will be using AI (as well as everybody else, me included). Large language models (LLMs) are good at general reasoning, writing, and coding. And I’d much rather my students spend their time on searches for new physics than searches for bugs in their code. But the big problem is that LLMs lack domain expertise and the ability to say, “I don’t know” when a question is beyond their abilities. (For instance, I recently asked a model if one of my research ideas had already been tested by other scientists. It responded “yes,” and when I asked it for the source, it cited our own conversation back to me.)
The first order of business was to run a test: Could an LLM with physics-specific guardrails outperform an off the shelf frontier model?
First, we ran a commercially available LLM out of the box and asked it all the questions a graduate student might. As expected, we did not get very good results. Then, we provided it with domain expertise: strict definitions, where it is allowed to draw information from, and the circumstances under which it must answer, “I don’t know.” Even with these simple tweaks, we got a much better performance. But it wasn’t good enough to deploy—we still needed something better.

Luckily, a grant from the DOE Genesis Mission is now allowing us to take this project from prototype to production.
Our Phase I Genesis Mission project focuses on the bread-and-butter of particle physics: data simulation and analysis. We are teaching these concepts to our LLM the same way we teach them to our students. (Well, almost the same; you don’t need to tell an LLM “good morning” before diving into a lesson.) Our long-term goal is to gradually expand to more parts of particle physics research and eventually create something that can be used by any experimental collaboration in particle physics and astrophysics.
If we succeed, this project will revolutionize physics education. Instead of spending years lost in code caves and documentation labyrinths, students will be able to focus on the interesting stuff: detector design, data interpretation, and the imaginative act of asking, “What if.” We won’t just need the students who are good at memorization and manual debugging; we can also cultivate those who are better at creative thinking—the quality we actually need if we ever want to break beyond the Standard Model of particle physics. And crucially, we will be able to embrace AI in the classroom rather than go on pretending that our students aren’t already using it.
