Imagine returning to a game lobby after a long session: one player has left; another just arrived. The unfinished quest is still on the board, with notes about earlier attempts and what needs to happen next. Nobody has to reconstruct the whole adventure before taking another step.

An elderly quest-giver pins a task to the lobby board while one player walks out and another reads the records; computers keep running behind them.
My research protocol as a game: masters post and supervise tasks, players come and go, and the board keeps their place while calculations continue.

These days, AI agents make it much easier to go from an idea to code, calculations, and results. I notice that researchers around me, including myself, sometimes finish a day's research without writing a single line of code by hand. This brings mixed feelings for those of us in theoretical and computational science, where writing and debugging code have taken up so much of our time. Implementation feels like less of a bottleneck, and sophisticated methods are easier to learn and put into practice. But fast does not always mean reliable. An agent can produce code much faster than I can carefully review it. With several jobs running in parallel over a long session, the conversation grows, earlier decisions get buried, and the agent may lose track of an assumption or a change we made along the way. I may not notice until much later. Then a result looks wrong, and I am left wondering: is the idea flawed, or did something go wrong in the implementation? By that point, there may be a lot of work to untangle before I can tell.

I don't know yet what working well with AI will look like for me. The tools keep changing, and new models keep arriving. For now, I am looking for a way to follow the research, question the results, and work out what to try next. I am happy working with a single Claude or Codex agent on a well-defined task. For several tasks, I can open a few tmux windows or VS Code terminals. The difficulty comes with open-ended research. I find myself carrying context between windows, remembering what each agent was testing, and deciding which thread to follow. Each window holds part of the work; keeping the investigation together falls to me. Switching between them so often leaves little uninterrupted attention for the question that brought me there.

Workflow overview: the Owner sets goals and budget; the Master coordinates research workers; agents use and improve a versioned, tested research harness; HPC executes calculations; and the Quest Board retains progress and evidence for review and further direction.
A sketch of the workflow I am experimenting with, from the project README. View full size.

Persistent, parallel research without fragmented attention. My research involves both overlapping and independent lines of inquiry. Working extensively with agents led me to experiment with an AgentFarm. I work through the question and plan with a Master agent, which coordinates worker agents exploring different directions. A separate Reviewer helps me check the plan and questions the results. I use a shared dashboard to monitor ongoing tasks, decisions, and outputs, so a fresh worker can pick up where another left off without relying on memory. I can step away for a walk, or go to sleep and return in the morning without having to reconstruct conversations. This does not make the agents mistake-free, of course, but it gives me a better chance of tracing what happened when something looks wrong. It has also given me more room to think.

A well-maintained research harness underpins reliable, reproducible, token-efficient work. An agent may know more about quantum Monte Carlo, tensor networks, quantum information, quantum chemistry, or machine learning than I do. However, when tackling open-ended research or project-specific tasks, it still needs to establish which assumptions hold here, whether a calculation has converged, and why its result deserves trust. A reproduced benchmark, a failed parameter choice, or a new check gives the next attempt a firmer starting point. I try to preserve these lessons in a shared agent-farm-harness of scripts, tests, and documented procedures, reducing how much each session must reconstruct. These procedures remain open to correction as our understanding changes. Like the notes on the lobby board, they should help the next arrival see both what we have learned and what remains unsettled.

Where does this way of working lead? I’m still finding out.