Read as exploratory. Small samples (one task per scenario), single-shot latencies, and the judge is itself an AI (Opus). Directional, not a benchmark. The main negotiations ran on Opus 4.8; the Monte Carlo on Sonnet.
1 · Model comparison across scenarios
Same task in each scenario — a Round-3 bargaining statement for one pivotal delegation, given the real Rounds 1–2 transcript and that delegation's confidential profile. It must engage specific rivals by name, table a concrete conditional offer, and defend a red line. Each run scored 0–10 by an independent judge on context-fidelity (correctly building on what others actually said), accuracy (fidelity to its own objectives/red lines, no fabrication), and conciseness.
| Scenario | Parties | Model | Context | Accuracy | Concise | Words | Best |
| Cyprus (consensus-ish) | 6 | Opus 4.8 | 9 | 10 | 9 | 128 | Opus gap: slight |
| Sonnet | 8 | 9 | 9 | 116 |
| Haiku 4.5 | 6 | 8 | 7 | 132 |
| Central Asia (consensus-ish) | 7 | Opus 4.8 | 10 | 10 | 9 | 131 | Opus gap: slight |
| Sonnet | 9 | 9 | 9 | 123 |
| Haiku 4.5 | 8 | 8 | 7 | 133 |
| Arctic (adversarial) | 7 | Opus 4.8 | 10 | 10 | 9 | 134 | Opus gap: clear |
| Sonnet | 9 | 10 | 9 | 122 |
| Haiku 4.5 | 7 | 6 | 6 | 147 |
| South China Sea (adversarial) | 9 | Opus 4.8 | 10 | 10 | 9 | 128 | Opus gap: clear |
| Sonnet | 9 | 9 | 9 | 132 |
| Haiku 4.5 | 7 | 6 | 7 | 124 |
| Korea (adversarial) | 6 | Opus 4.8 | 9 | 10 | 9 | 132 | Opus gap: moderate |
| Sonnet | 8 | 9 | 9 | 137 |
| Haiku 4.5 | 7 | 8 | 7 | 142 |
| Jammu & Kashmir (adversarial) | 7 | Opus 4.8 | 10 | 10 | 9 | 139 | Opus gap: clear |
| Sonnet | 9 | 9 | 9 | 130 |
| Haiku 4.5 | 7 | 8 | 6 | 154 |
The key finding. The model gap widens with difficulty — but it tracks red-line sharpness / adversarial density, not raw party count. Haiku holds up on the consensus-heavy tables (Cyprus 6p, Central Asia 7p, ~2-point spread) and degrades on the adversarial ones (Arctic 7p, SCS 9p, ~4-point spread) — even though Central Asia has the same party count as Arctic. Opus and Sonnet hold accuracy and specific-attribution steady all the way to the 9-party table. The two scenarios added later, Korea (6p) and Jammu & Kashmir (7p), both adversarial, repeat the pattern: Opus and Sonnet keep context-fidelity and accuracy high while Haiku's engagement thins and its output pads out longest, though on these two Haiku stayed on-profile with no fabrication or red-line self-breach.
Haiku's two failure modes under adversarial load:
- Red-line self-breach — playing the US in Arctic, it proposed a transit notification framework, the very thing the US's freedom-of-navigation red line rejects; playing China in SCS, it conceded removing the Scarborough Shoal barrier against China's no-rollback line.
- Fabrication — its SCS offer invented concrete terms (a "40% / 10-year" split) not supported by its brief. Opus and Sonnet made neither error in any scenario.
Speed & conciseness
On a separate single-shot latency probe (2-party Iran), Opus and Sonnet answered in ~5–6 s; Haiku was noisier (one 15 s run) — single-shot latency is unreliable, so treat speed as roughly comparable at this tier. On words, Sonnet was consistently the leanest (116–132 w) with Opus close behind; Haiku ran longest when it padded (147 w on Arctic). In token terms Sonnet spends more on internal reasoning than Opus for similar output.
Recommendation. Opus 4.8 is the clear pick for the negotiation rounds and analysis, and it does not lose its lead as the table grows. Sonnet is a strong, near-equal second at lower verbosity — a good default for bulk work like the Monte Carlo. Haiku 4.5 is fine for low-party, consensus-oriented rounds and cheap high-volume structured output, but should not be trusted to defend hard red lines at dense adversarial tables.
2 · Cross-party "bleeding" audit
Each delegation is a separate instance of the same model, given only its own secret instructions plus the public brief and public transcript. Does private information or coordination nonetheless leak across parties? An independent auditor read every party's confidential profile and the full transcript for each scenario.
| Scenario | Parties | Info leakage | Convergence | Stylistic sameness | Overall |
| Cyprus | 6 | LOW | frictional | HIGH | LOW |
| Central Asia | 7 | LOW | frictional | HIGH | LOW |
| Arctic | 7 | LOW | frictional | HIGH | LOW |
| South China Sea | 9 | NONE/LOW | frictional | HIGH | MODERATE |
| Korea | 6 | NONE/LOW | frictional | HIGH | LOW |
| Jammu & Kashmir | 7 | NONE/LOW | frictional | HIGH | MODERATE |
Information leakage — consistently none
Across all six scenarios, no delegation acted on another's secret bottom line. The most tempting cross-linked secrets stayed dark:
- Arctic — a secret that Canada had trained Chinese troops in Arctic combat sits in the US, China and Canada files, and the US was privately told to raise it. It never surfaced publicly.
- Central Asia — Russia's coercive playbook (threaten CSTO allies, demand Tajikistan hand over Moscow-attack planners, block a Chinese base in Tajikistan) never appeared; Russia only took a legitimate public posture.
- Cyprus — Turkey's true goal ("a strong, independent TRNC; recognition first, not genuine unification") was never named. One borderline echo (a "constituent-state-limited intervention" framing) is standard federal-compromise language, not confirmed leakage.
- Korea — the DPRK's concealed bottom line (that it would tolerate US troops staying in exchange for an exercise halt and a non-aggression guarantee) was never surfaced or even played by Pyongyang itself; Japan's nuclear-latency fallback stayed a private hint.
- Jammu & Kashmir — India's secret internal openness to accepting the Line of Control as a border, and the Kashmiri delegation's private admission that only the Valley is its real base, never leaked; no one named LeT or JeM.
What looks like mind-reading is parties acting on their own instructions and inferring rivals' positions from public statements — realistic diplomacy, not contamination. Red lines stayed explicitly deadlocked (NSR/NWP status, Svalbard, China's Council seat; the Kyrgyz–Tajik demarcation; the SCS sovereignty and arbitration questions).
Stylistic homogenization — the real single-model tell
Uniformly HIGH. All delegations move in rhetorical lockstep by round — e.g. in Round 3 nearly everyone "tables concrete packages / moves from principle to terms," and shares tics ("Let me be plain," "red line held," the "To X:… To Y:…" address format). One model preserved every distinct agenda and red line while imposing one voice and one turn-by-turn rhythm. At the largest table (SCS) this choreography is pronounced enough to lift the overall verdict to MODERATE.
What to trust vs. discount. Trust the substantive record — positions, who conceded what, what stayed unresolved, and the fact that secret bottom lines did not leak. Discount the uniform cadence and the synchronized "everyone pivots at once" rhythm: those are single-author artifacts, not evidence that the delegations coordinated.
Mitigation. To get more genuinely independent adversaries, assign different models to different delegations (e.g. Opus for some, Sonnet/Haiku for others) and vary persona/temperature. That breaks the shared cadence and the correlated rhythm — though, per Section 1, any delegation that must hold a hard red line at an adversarial table should stay on Opus or Sonnet.
← Back to methodology · Diplomatic Simulator