The problem
01 / 07The problem
A simulation can support any conclusion
Adjust the agents, choose the random seed, choose how the outcome is measured, and cooperation appears or disappears on demand. That is not a flaw in one model. It is a property of simulation. The same rules, run three times with small changes nobody would notice, end in three different places.
02 / 07The problem
A single run is not a result
An outcome from one model, one agent type, one seed, and one way of measuring is an artifact until it survives variation of the things we are uncertain about. It might be a finding. It might be a parameter setting. Nothing in the run itself says which, and the run you would report looks exactly like the ones you would not.
03 / 07The problem
Three fields, three standards of evidence
Mechanism design asks for a formal proof: show that the incentives are compatible. Agent-based modelling asks for generative sufficiency: grow the pattern from simple rules. Ecology asks for a phase portrait: the stable states and how deep they are. Each is right about something, and none alone is enough. Elinor Ostrom made the same point about real institutions and combined all three streams: no single one suffices.
The standard
04 / 07The standard
Robustness analysis: vary what is uncertain
Vary the agent model: rule-based agents, learning agents, language models, human subjects. Each relaxes a different assumption, and none is more realistic than the others in every respect. Vary the environment model too: a claim about a mechanism is strongest when it holds across structurally different models of the same domain. Report what survives both. Where results diverge, that is a finding, not a failure.
05 / 07The standard
Basin stability, not point estimates
Ask how large a perturbation the system can absorb before it shifts regime, not only where it ends up. The answer has the shape of a basin of attraction: the set of initial states from which the desired outcome is still reached. Its size, as a function of the mechanism parameters, is the result. A single trajectory is an illustration. A phase boundary is a finding.
06 / 07The standard
Declare the validation level
Every model states how far it has been validated: it runs; it replicates a known result; it is calibrated to detect the phenomenon it was built to detect; it has been validated across agent models; it has been replicated or used independently. The levels rank the evidence, not the realism of the agents. The model documentation travels with the code.
07 / 07The standard
State the known limitations
The methods break down in known places: while agents are still learning, when participants can change the rules or exit, and when a mechanism is enforced by construction. And stability says nothing about whether a state is desirable. A stable but harmful equilibrium is lock-in, not a success. Every stability claim is reported together with a separate welfare assessment.