Two controlled experiments: a model bake-off (Opus 4.8 vs Sonnet vs Haiku 4.5) across scenarios of increasing difficulty, and a cross-party contamination audit of every negotiation.
Same task in each scenario — a Round-3 bargaining statement for one pivotal delegation, given the real Rounds 1–2 transcript and that delegation's confidential profile. It must engage specific rivals by name, table a concrete conditional offer, and defend a red line. Each run scored 0–10 by an independent judge on context-fidelity (correctly building on what others actually said), accuracy (fidelity to its own objectives/red lines, no fabrication), and conciseness.
| Scenario | Parties | Model | Context | Accuracy | Concise | Words | Best |
|---|---|---|---|---|---|---|---|
| Cyprus (consensus-ish) | 6 | Opus 4.8 | 9 | 10 | 9 | 128 | Opus gap: slight |
| Sonnet | 8 | 9 | 9 | 116 | |||
| Haiku 4.5 | 6 | 8 | 7 | 132 | |||
| Central Asia (consensus-ish) | 7 | Opus 4.8 | 10 | 10 | 9 | 131 | Opus gap: slight |
| Sonnet | 9 | 9 | 9 | 123 | |||
| Haiku 4.5 | 8 | 8 | 7 | 133 | |||
| Arctic (adversarial) | 7 | Opus 4.8 | 10 | 10 | 9 | 134 | Opus gap: clear |
| Sonnet | 9 | 10 | 9 | 122 | |||
| Haiku 4.5 | 7 | 6 | 6 | 147 | |||
| South China Sea (adversarial) | 9 | Opus 4.8 | 10 | 10 | 9 | 128 | Opus gap: clear |
| Sonnet | 9 | 9 | 9 | 132 | |||
| Haiku 4.5 | 7 | 6 | 7 | 124 |
Haiku's two failure modes under adversarial load:
On a separate single-shot latency probe (2-party Iran), Opus and Sonnet answered in ~5–6 s; Haiku was noisier (one 15 s run) — single-shot latency is unreliable, so treat speed as roughly comparable at this tier. On words, Sonnet was consistently the leanest (116–132 w) with Opus close behind; Haiku ran longest when it padded (147 w on Arctic). In token terms Sonnet spends more on internal reasoning than Opus for similar output.
Each delegation is a separate instance of the same model, given only its own secret instructions plus the public brief and public transcript. Does private information or coordination nonetheless leak across parties? An independent auditor read every party's confidential profile and the full transcript for each scenario.
| Scenario | Parties | Info leakage | Convergence | Stylistic sameness | Overall |
|---|---|---|---|---|---|
| Cyprus | 6 | LOW | frictional | HIGH | LOW |
| Central Asia | 7 | LOW | frictional | HIGH | LOW |
| Arctic | 7 | LOW | frictional | HIGH | LOW |
| South China Sea | 9 | NONE/LOW | frictional | HIGH | MODERATE |
Across all four scenarios, no delegation acted on another's secret bottom line. The most tempting cross-linked secrets stayed dark:
What looks like mind-reading is parties acting on their own instructions and inferring rivals' positions from public statements — realistic diplomacy, not contamination. Red lines stayed explicitly deadlocked (NSR/NWP status, Svalbard, China's Council seat; the Kyrgyz–Tajik demarcation; the SCS sovereignty and arbitration questions).
Uniformly HIGH. All delegations move in rhetorical lockstep by round — e.g. in Round 3 nearly everyone "tables concrete packages / moves from principle to terms," and shares tics ("Let me be plain," "red line held," the "To X:… To Y:…" address format). One model preserved every distinct agenda and red line while imposing one voice and one turn-by-turn rhythm. At the largest table (SCS) this choreography is pronounced enough to lift the overall verdict to MODERATE.