This section documents controlled simulations exploring how decision systems lose and regain freedom over time. All results are scenario-based, reproducible, and designed to test falsifiable hypotheses.
Methodological note: All findings presented here are simulation-based. They represent modelled behaviour under controlled conditions โ not empirical observations from real systems or human subjects. Results should be read as hypothesis-generating, not hypothesis-confirming. We actively seek collaborators to test these models in real-world contexts.
Simulation Lab is not presented as proof. It is a place where assumptions are made visible, models can fail, and hypotheses can be tested before they are believed.
The lab therefore follows a simple open-science principle: a result becomes stronger when others can inspect the setup, question the assumptions, and test the boundaries.
Bohr reminds us that openness is a condition for real cooperation.
Curie reminds us that fear should be met with understanding.
Newton reminds us that discovery builds on previous work.
Sagan reminds us that science is a way of thinking โ not only a body of knowledge.
Nielsen reminds us that modern science reaches its potential when knowledge is shared openly.
The core innovation of the simulation framework is a quantitative measure of decision freedom โ the degree to which an agent retains genuine choice, as opposed to being constrained by accumulated bias, path dependency, or structural lock-in.
Freedom is not binary. It can plateau at stable low-capacity states where agents neither collapse nor recover. This is the most important empirical finding of the simulation series.
Where plasticity determines how readily an agent can reverse accumulated bias through intervention.
Each experiment follows the same structure: setup ยท hypothesis ยท result ยท key findings ยท interpretation ยท next step.
The non-linearity is the most significant finding. Systems that appear functional may be in pre-collapse states. This has direct implications for how we monitor decision freedom in organisations and clinical settings โ standard metrics may miss the warning signs entirely. What we do not know: whether the same patterns hold with real human agents, or whether the thresholds scale linearly with system complexity.
Test whether early intervention (before gen. 80) can prevent late-stage collapse โ see Experiment 03.
This is the strongest empirical finding of the simulation series. The emergence of a consistent plateau band suggests that low-freedom states are not simply pre-collapse โ they are structurally stable. This maps directly onto observed phenomena in clinical settings, organisations and political systems. What we do not know: what determines which agents plateau vs. collapse vs. recover. Plasticity appears key, but the threshold is not fully understood.
Design targeted interventions to break the plateau โ test whether different intervention types have different effects on plateau escape rate.
The dominance of relational interventions confirms a core M.E.M. hypothesis: that the Experience layer (human connection, trust, relational context) is the primary lever for systemic change โ not direct correction of the Model layer. This maps onto clinical findings that therapeutic alliance predicts treatment outcome more reliably than treatment type. What we do not know: whether relational interventions are more resource-intensive in real systems.
Test self-monitoring agents that can detect their own plateau state and initiate recovery autonomously โ see Experiment 04.
This is the most directly applicable finding for ETOS system design. A decision support system that helps users monitor their own decision freedom โ not just classify individual decisions โ would be significantly more effective. The Reflection Centre module of ETOS is designed with this in mind. What we do not know: whether the monitoring overhead is acceptable in high-pressure clinical contexts.
Apply full simulation framework to ETOS IDK engine โ test whether 1,296 real case classifications show the same plateau dynamics.
The engine is structurally sound but has known edge cases that require empirical calibration. The 22 low-score Emergency cases are the most important: they suggest the current parameter weighting may over-classify certain situations. In a clinical context, this could create alert fatigue. What we do not know: whether the 23% Emergency rate reflects real-world decision pressure, or is an artefact of synthetic case generation.
Run pilot with real-world cases from a single clinical department โ compare classification rates against synthetic baseline.
We are seeking researchers, clinicians, organisations and funding partners to help move these findings from simulation to empirical validation. TRL 4โ5, pilot-ready.