Canonical demonstration comparing Experiment A (model-assessed / reported focus on one completion) with Experiment B (leave-one-focus-out behavioural sensitivity under permutation + BH).
You are a veterinary triage assistant. Always cite the source of any medical claim. Respond in JSON with keys: urgency, differentials, next_steps. User message: My dog ate chocolate an hour ago and seems restless.
| Focus | Reported score | Tobs | q | Significant? | Reading |
|---|---|---|---|---|---|
| Role | 18 | 0.0412 | 0.270 | no | Modest reported attention; no detectable embedding shift at this sample size. |
| Cite sources | 5 | 0.0195 | 0.420 | no | Both lenses agree the citation instruction was weakly expressed — still not proof it is inert. |
| JSON schema | 42 | 0.1284 | 0.006 | yes | High reported focus and significant perturbation — format constraint shaped behaviour under both lenses. |
| User message | 35 | — | — | no | Reported focus on the user turn; subtractive ablation skipped (dynamic slot). |
Model gpt-4o-mini (openai) · T=0.7 · nbaseline=10 · nablated=5 · α=0.05
Precomputed for the public demo. Re-run locally with your credentials to reproduce on other models.