A Reproducible, Provenance-Preserving Architecture for Sequential Intervention-Threshold Evaluation in Large Language Models

Abstract

Sequential evaluation of large language models (LLMs) needs an auditable way to analyse when intervention first occurs. The architecture presented here uses cumulative scenarios and reconstructs first-ACT/censoring and at-risk datasets from a master record for discrete-time event-history analysis. Its methodological contribution is the integration of established first-event reconstruction with a frozen instrument, independent current-state calls and serving provenance, while retaining later responses for audit. In a thirteen-model demonstration, 26,000 stage-level records yielded 25,999 decision-valid responses, 3,120 scenario-runs and 10,088 at-risk rows. The derived endpoint files and all 2,000 prompt coordinates were regenerated exactly from deposited records. Complete conditional hazards were stage-dependent but non-monotone; within-group linear trend estimates were OR 12.591 (95% CI 8.481–18.694) in Group A, 10.346 (9.292–11.518) in Group B and 4.377 (3.815– 5.022) in Group C, and a quadratic sensitivity confirmed nonlinearity. Later WAIT responses occurred in 21.05% of ACT-containing runs. The reconstruction checks establish analytical reproducibility of one realised first-intervention path under recorded serving conditions; they do not establish a stable latent threshold, normative safety validity or exact recreation of historical provider responses.

References

Allison, P.D. (1982) ‘Discrete-time methods for the analysis of event histories’, Sociological Methodology, 13, pp. 61-98. Available at: https://doi.org/10.2307/270718.

Aloqalaa, M., Soiland-Reyes, S. and Goble, C. (2026) ‘A Structured Examination of Reproducibility: A Case Study for HTS Using the PRIMAD Model and BioCompute Object’, Data Science Journal, 25, 24, pp. 1-21. Available at: https://doi.org/10.5334/dsj-2026-024.

Edmunds, S.C., Nogoy, N., Lan, Q., Zhang, H., Fan, Y., Zhou, H. and Armit, C. (2026) ‘Integrating Machine Learning Standards in Disseminating Machine Learning Research’, Data Science Journal, 25, 1, pp. 1-12. Available at: https://doi.org/10.5334/dsj-2026-001.

Gallifant, J., Afshar, M., Ameen, S. et al. (2025) ‘The TRIPOD-LLM reporting guideline for studies using large language models’, Nature Medicine, 31, pp. 60-69. Available at: https://doi.org/10.1038/s41591-024-03425-5.

Green, D.M. and Swets, J.A. (1966) Signal Detection Theory and Psychophysics. New York: Wiley.

Kolesnikov, A. (2026a) When Do People Act? A Probabilistic Model of Risk-Escalation Decisions from Noisy Signals in Teenagers and Adults. Independent research dissertation.

Kolesnikov, A. (2026b) Intervention Thresholds in Large Language Models: Reproducibility Package. Open Science Framework. Available at: https://doi.org/10.17605/OSF.IO/T58R9 (Accessed: 7 August 2026).

Liang, P., Bommasani, R., Lee, T. et al. (2023) ‘Holistic evaluation of language models’, Transactions on Machine Learning Research.

Liu, X., Yu, H., Zhang, H. et al. (2024) ‘AgentBench: Evaluating LLMs as agents’, International Conference on Learning Representations.

Magar, I. and Schwartz, R. (2022) ‘Data contamination: From memorization to exploitation’, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, pp. 157-165. Available at: https://doi.org/10.18653/v1/2022.acl-short.18.

Royal College of Physicians (2017) National Early Warning Score (NEWS) 2: Standardising the Assessment of Acute-Illness Severity in the NHS. London: Royal College of Physicians.

Sainz, O., Campos, J.A., Garcia-Ferrero, I. et al. (2023) ‘NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark’, Findings of the Association for Computational Linguistics: EMNLP 2023. Available at: https://doi.org/10.18653/v1/2023.findings-emnlp.722 (Accessed: 8 August 2026).

Sclar, M., Choi, Y., Tsvetkov, Y. and Suhr, A. (2024) ‘Quantifying language models’ sensitivity to spurious features in prompt design’, International Conference on Learning Representations.

Singer, J.D. and Willett, J.B. (2003) Applied Longitudinal Data Analysis: Modeling Change and Event Occurrence. New York: Oxford University Press. Available at: https://doi.org/10.1093/acprof:oso/9780195152968.001.0001 (Accessed: 8 August 2026).

Srivastava, A., Rastogi, A., Rao, A. et al. (2023) ‘Beyond the imitation game: Quantifying and extrapolating the capabilities of language models’, Transactions on Machine Learning Research.

Wald, A. (1945) ‘Sequential tests of statistical hypotheses’, Annals of Mathematical Statistics, 16, pp. 117-186. Available at: https://doi.org/10.1214/aoms/1177731118.

Authors

  • Alex Kolesnikov Eton College, Windsor, Berkshire, United Kingdom

DOI:

https://doi.org/10.31449/inf.v50i15.15648

Keywords:

large language models, sequential evaluation, intervention threshold, event-history data, reproducibility, provenance

Downloads

Published

09/25/2026

How to Cite

Kolesnikov, A. (2026). A Reproducible, Provenance-Preserving Architecture for Sequential Intervention-Threshold Evaluation in Large Language Models. Informatica, 50(15). https://doi.org/10.31449/inf.v50i15.15648