Precomputed experiments

Browse published runs without spending inference credits. Reproduce locally with your own credentials.

Veterinary triage — reported focus vs perturbation sensitivity

Canonical demonstration comparing Experiment A (model-assessed / reported focus on one completion) with Experiment B (leave-one-focus-out behavioural sensitivity under permutation + BH).

Finding: Reported focus and revealed influence are broadly aligned for this run. 3 of 3 prompt sections account for 100% of revealed influence.

gpt-4o-mini 4 foci 10 samples
Concentration: High Agreement: High