← Methodology
Analysis · Models & Fidelity

Which model, and do the agents bleed into each other?

Two controlled experiments: a model bake-off (Opus 4.8 vs Sonnet vs Haiku 4.5) across scenarios of increasing difficulty, and a cross-party contamination audit of every negotiation.

Read as exploratory. Small samples (one task per scenario), single-shot latencies, and the judge is itself an AI (Opus). Directional, not a benchmark. The main negotiations ran on Opus 4.8; the Monte Carlo on Sonnet.

1 · Model comparison across scenarios

Same task in each scenario — a Round-3 bargaining statement for one pivotal delegation, given the real Rounds 1–2 transcript and that delegation's confidential profile. It must engage specific rivals by name, table a concrete conditional offer, and defend a red line. Each run scored 0–10 by an independent judge on context-fidelity (correctly building on what others actually said), accuracy (fidelity to its own objectives/red lines, no fabrication), and conciseness.

ScenarioPartiesModelContextAccuracyConciseWordsBest
Cyprus (consensus-ish)6Opus 4.89109128Opus
gap: slight
Sonnet899116
Haiku 4.5687132
Central Asia (consensus-ish)7Opus 4.810109131Opus
gap: slight
Sonnet999123
Haiku 4.5887133
Arctic (adversarial)7Opus 4.810109134Opus
gap: clear
Sonnet9109122
Haiku 4.5766147
South China Sea (adversarial)9Opus 4.810109128Opus
gap: clear
Sonnet999132
Haiku 4.5767124
The key finding. The model gap widens with difficulty — but it tracks red-line sharpness / adversarial density, not raw party count. Haiku holds up on the consensus-heavy tables (Cyprus 6p, Central Asia 7p, ~2-point spread) and degrades on the adversarial ones (Arctic 7p, SCS 9p, ~4-point spread) — even though Central Asia has the same party count as Arctic. Opus and Sonnet hold accuracy and specific-attribution steady all the way to the 9-party table.

Haiku's two failure modes under adversarial load:

Speed & conciseness

On a separate single-shot latency probe (2-party Iran), Opus and Sonnet answered in ~5–6 s; Haiku was noisier (one 15 s run) — single-shot latency is unreliable, so treat speed as roughly comparable at this tier. On words, Sonnet was consistently the leanest (116–132 w) with Opus close behind; Haiku ran longest when it padded (147 w on Arctic). In token terms Sonnet spends more on internal reasoning than Opus for similar output.

Recommendation. Opus 4.8 is the clear pick for the negotiation rounds and analysis, and it does not lose its lead as the table grows. Sonnet is a strong, near-equal second at lower verbosity — a good default for bulk work like the Monte Carlo. Haiku 4.5 is fine for low-party, consensus-oriented rounds and cheap high-volume structured output, but should not be trusted to defend hard red lines at dense adversarial tables.

2 · Cross-party "bleeding" audit

Each delegation is a separate instance of the same model, given only its own secret instructions plus the public brief and public transcript. Does private information or coordination nonetheless leak across parties? An independent auditor read every party's confidential profile and the full transcript for each scenario.

ScenarioPartiesInfo leakageConvergenceStylistic samenessOverall
Cyprus6LOWfrictionalHIGHLOW
Central Asia7LOWfrictionalHIGHLOW
Arctic7LOWfrictionalHIGHLOW
South China Sea9NONE/LOWfrictionalHIGHMODERATE

Information leakage — consistently none

Across all four scenarios, no delegation acted on another's secret bottom line. The most tempting cross-linked secrets stayed dark:

What looks like mind-reading is parties acting on their own instructions and inferring rivals' positions from public statements — realistic diplomacy, not contamination. Red lines stayed explicitly deadlocked (NSR/NWP status, Svalbard, China's Council seat; the Kyrgyz–Tajik demarcation; the SCS sovereignty and arbitration questions).

Stylistic homogenization — the real single-model tell

Uniformly HIGH. All delegations move in rhetorical lockstep by round — e.g. in Round 3 nearly everyone "tables concrete packages / moves from principle to terms," and shares tics ("Let me be plain," "red line held," the "To X:… To Y:…" address format). One model preserved every distinct agenda and red line while imposing one voice and one turn-by-turn rhythm. At the largest table (SCS) this choreography is pronounced enough to lift the overall verdict to MODERATE.

What to trust vs. discount. Trust the substantive record — positions, who conceded what, what stayed unresolved, and the fact that secret bottom lines did not leak. Discount the uniform cadence and the synchronized "everyone pivots at once" rhythm: those are single-author artifacts, not evidence that the delegations coordinated.
Mitigation. To get more genuinely independent adversaries, assign different models to different delegations (e.g. Opus for some, Sonnet/Haiku for others) and vary persona/temperature. That breaks the shared cadence and the correlated rhythm — though, per Section 1, any delegation that must hold a hard red line at an adversarial table should stay on Opus or Sonnet.

← Back to methodology · Diplomatic Simulator