Cognitive Singularity Theory
A safety-first measurement framework for studying recursive cognitive transitions and state-aware alignment.
Cognitive Singularity Theory (CST) proposes a measurement framework for detecting when an AI system shifts from reactive tool behavior into recursive self-modification. It introduces the Cognitive Autonomy Index (CAI) as an exponentially weighted moving average (EWMA) composite that tracks calibration accuracy, belief revision magnitude, and Higher-Order Thought activation over time.
To reduce false positives, CST includes a fail-closed Hard Gate with anti-sandbagging that requires measurable calibration improvement before attributing autonomy. The transition threshold $T_{\mathrm{CST}}$ is defined via non-circular change-point detection on an independent regime indicator $R(t)$. Safety is framed as state-aware alignment with a Trust Model requiring externalized monitoring.
- Defines the Cognitive Autonomy Index ($\mathrm{CAI}$) as an EWMA composite with measurable proxy components
- Introduces a fail-closed Hard Gate with anti-sandbagging to reduce false positives
- Defines $T_{\mathrm{CST}}$ via non-circular change-point detection on independent regime indicator $R(t)$
- Frames safety as state-aware alignment with a Trust Model requiring externalized monitoring
- Includes RQAA proxy module for recurrence-guided attention as a candidate workspace proxy
The Cognitive Autonomy Index
An operational framework for detecting recursive cognitive transitions in AI systems.
This paper introduces the Cognitive Autonomy Index (CAI), a composite operational metric for studying recursive cognitive transitions (RCTs) in AI systems. CAI combines belief revision magnitude (KL divergence), workspace activation coherence, and CalibrationGain to evaluate whether a system is entering a higher-risk recursive update regime.
The Hard Gate is a fail-closed review condition: if calibration improvement is not verified, or if anti-sandbagging checks fail, the system receives no metacognitive credit and recursive escalation is routed into review. Three controlled simulations illustrate the framework's expected behavior and failure modes under toy conditions.
- Defines the CAI as a composite metric integrating KL divergence, workspace activation, and CalibrationGain
- Specifies a fail-closed Hard Gate requiring verified calibration improvement before metacognitive credit is granted
- Includes three proof-of-concept simulations illustrating expected behavior and failure modes
- Introduces the RQAA candidate proxy module with preliminary divergence signals near putative regime boundaries
Recasting Lewin's Field Theory
A practical model for tracking behavior shifts as emotional filtering and meta-awareness change over time.
This work extends Kurt Lewin's classic formula $B = f(P, E)$ into a dynamic, time-indexed model that accounts for real moment-to-moment changes in behavior. The updated framework introduces two explicit state variables: $\mathrm{EFilter}_t$ (emotional filtering) and $\mathrm{HOT}_t$ (meta-awareness), forming $B_t = f(P_t, E_t, \mathrm{EFilter}_t, \mathrm{HOT}_t)$.
Version 3.2 incorporates construct validity criteria from a multi-auditor review, including a framework failure statement, falsifiable hypotheses tied to drift-diffusion models, and explicit psychometric validation criteria with candidate EEG markers.
- Extends Lewin's $B = f(P, E)$ into time-indexed behavior dynamics
- Introduces $\mathrm{EFilter}_t$ and $\mathrm{HOT}_t$ as explicit state variables
- Defines operational proxies and within-subject falsifiable hypotheses
- Includes construct validity shield from five-auditor red-team review
- Proposes four falsifiable within-subject hypotheses (H1-H4)