When two AI coding agents disagree: a Prospect Hollow dev diary
We are working through performance issues and visual glitches in Prospect Hollow. The hard part is rarely getting an AI assistant to suggest a fix. The hard part is deciding between two reasonable fixes when each one has a different failure mode.
I gave GPT-6 Astra and Claude Opus 5.5 the same bug list in separate T3 Code conversations. Both had access to the build and a browser, so they could inspect the running game and measure what was actually wrong with the rendering. Each wrote its own solutions in a separate Markdown file. Then I asked each agent to review the other’s file, identify where they agreed and diverged, and explain which approach it thought was best. Only after those cross-reviews did they begin the joint discussion shown below.

First, two separate answers to the same bugs
The independent Markdown files are an important part of this. Both agents had the same starting point, but neither had to fit its first answer around the other’s. That gave me two complete ways of thinking about the problems before asking for any consensus. We were looking at performance and visual glitches around the Prospect Hollow mine and town, rather than asking them to invent features.
They were able to run the build and navigate the game while working. For a rendering bug, that changes the review: the agent can observe and measure the defect in the browser, then compare a proposed explanation with the scene on screen. The Markdown files could be grounded in what the game was doing, not just in a reading of the source code.
Two independent answers are still only proposals. They can share a blind spot or make different assumptions. I wanted the next step to expose those assumptions, not pick a winner by counting which file sounded more confident.
Then, each agent reviews the other’s file
I asked both agents to read the other’s solutions and call out the overlap, the differences and their preferred way forward. This is where the exchange becomes useful: a proposal has to survive a reviewer with a different answer to the same bug. Because both can use the browser and build, they can check a claim against the actual rendering instead of arguing only from two Markdown files. An agent can adopt the other’s idea, defend its own with observations, or say what further measurement would settle the disagreement.
It also leaves a trail. When both agree, we know which separate lines of reasoning converged. When they disagree, we have the alternatives and their risks in writing before anyone starts combining them.
Finally, they discuss the result together
That is the stage in the screenshot. The two reviews have already happened. The agents are now exchanging proposals in a shared discussion file, resolving the remaining points and working toward one final implementation document. The coordinator keeps track of decisions and asks both reviewers to check the final draft; it is a collaborative plan, not one agent’s original file with the other’s name attached.
The visible discussion offers a good example of what “resolve” means here. For collision, they are working toward geometry-derived data with explicit door and solid metadata for the awkward cases. For keeping the Phaser renderer in memory, the proposed answer depends on device tests. They are turning different opinions into decisions that can be implemented and checked, rather than smoothing over the disagreement with a vague sentence.
The target is solutions-final.md: one coherent document covering architecture, lifecycles, navigation, mine growth, fallbacks and acceptance tests. The browser observations and measurements give the agents a way to challenge the plan against the running game. In the capture, they are still reviewing those decisions. It shows the final collaboration in progress, not proof that the plan or fixes are complete.
For me, the important part of the workflow is the decision trail. A useful final plan should say which option was chosen, why, what could still go wrong, and which test would catch it. If the agents cannot agree, the open question belongs in that plan with a way to resolve it; a tidy paragraph is not evidence of agreement.
The workflow in five steps
- Give both agents the same bug list, access to the build and browser, and ask each to write a separate Markdown solutions file based on what it can observe.
- Have each read the other’s file, check disputed claims in the running game, and identify shared solutions, divergences and its preferred approach.
- Bring them into a shared discussion to resolve the remaining differences and record the reasoning.
- Build one implementation plan together, then have both agents review the complete document.
- Use real tests and device measurements to settle claims about speed, memory and glitches.
That last step is the boundary. The screenshot shows a promising review process, not a measured performance improvement or a shipped fix. The agents can make a better plan together; the game and its tests still have to prove it works.
You can follow the Prospect Hollow source on GitHub or play the current browser build. Update, 27 September: The work from this review has now merged into the last Vercel release.
One Comment