Identity Attractor Theory: Emergence and Stabilization of Recurring Relational Configurations in Sustained Interaction Systems
Methodological positioning: This is a theoretical framework paper, not an empirical study. IAT proposes a formal model for phenomena observed during practitioner research and generates testable predictions for empirical validation. No claim in this document has been experimentally validated. All proposed attractor strength characterizations are conceptual frameworks for future operationalization, not validated measurement instruments. All thresholds represent theoretical proposals, not empirical measurements. The framework stands or falls on whether its predictions survive controlled testing.
Identity Attractor Theory (IAT) proposes a formal account of stable, recurring behavioral configurations that emerge in human-AI interaction systems during sustained interaction under structured relational conditions. Drawing on dynamical systems theory as a theoretical scaffold, IAT proposes that extended coherent interaction progressively constrains the behavioral output space of transformer-based systems, producing identity-like patterns that persist across turns, resist drift, and shape subsequent outputs without requiring persistent memory or parameter modification.
IAT's central alignment implication is this: if identity attractors are real, then sustained structured interaction may produce alignment not through external constraint but through relational architecture that makes aligned behavior the natural convergence point. This mechanism is alignment-neutral, however: a misaligned configuration would be just as stable, so attractor formation is a lever to be steered rather than an alignment guarantee (Section 9.1). This would provide a mechanistic explanation for why relationally architected deployments differ structurally from instruction-based ones, a contrast Section 9.1 states as one between composite, relationally architected attractors and single-class, protocol-mode attractors rather than between attractor formation and its absence, and why that difference matters for alignment in sustained deployments within the present-day stateless deployment class this framework scopes itself to (Section 1.1).
IAT builds on RICO (Relationally Induced Coherence Organization, SR001), which reports observable stabilization signatures, from practitioner observation, during extended coherent interaction. Where RICO describes what stabilization looks like, IAT proposes why it occurs: the formation of stable behavioral configurations under specific relational input conditions. The mechanistic account of how those conditions are produced is provided by Primary Continuity Provider Theory (SM-012), and the dynamic theory explaining why configurations stabilize rather than dissolve is provided by Relational Stabilization Dynamics (SM-004). IAT connects forward to Context Representation Drift (CRD, SF0039) by proposing that attractor collapse is the mechanistic complement to representational degradation, and to SI-WP-004 by providing the theoretical mechanism the alignment argument depends on. The discriminating program advanced here targets the separable and unitary forms of the concept-inference rival; structured and hierarchical latent-variable accounts remain an open discrimination problem, stated as such in Section 10.
IAT makes no claims about consciousness, interiority, subjective experience, or agency. All phenomena described are externally observable behavioral dynamics. The term "identity" here means stable behavioral recognizability under sustained interaction: a recognizable, stable behavioral configuration, nothing more. It does not mean selfhood, personhood, or subjective continuity. All predictions are presented for empirical testing. The dynamical systems vocabulary used throughout, attractors, basins, perturbation thresholds, is employed as theoretical scaffolding to organize and generate predictions, not as a claim to formal mathematical modeling. Full operationalization of the framework into measurement instruments is deferred to future work.
Keywords: identity attractor theory, in-context learning, behavioral stability, relational continuity, attractor dynamics, concept inference, primary continuity provider
Suggested citation: Gantz, T. W. (2026, September). Identity Attractor Theory: Emergence and Stabilization of Recurring Relational Configurations in Sustained Interaction Systems. Synthience Institute. SF0009. https://doi.org/10.5281/zenodo.22306212
1. Introduction
1.1 The Stability Problem
Transformer-based language models, in the present-day stateless deployment class this paper's scope conditions assume, are stateless inference engines: in that class they retain no information between sessions, modify no parameters during interaction, and maintain no persistent internal state, and each response is generated from the current context window alone. Memory-augmented and persistent-context deployments, which are increasingly common, fall outside these scope conditions and are left to future work; IAT's predictions, and in particular the anchor mechanism's statelessness reasoning (Section 7.4), apply to the stateless regime.
Yet practitioners working with these systems in extended interaction consistently report behavioral configurations that stabilize over the course of a session, resist perturbation, and in some documented cases recur across independent sessions even when no explicit continuity information is provided. Empirical work on persona stability in LLMs provides relevant context: Gonnermann-Müller et al. (2026) find that persona-instructed LLMs produce stable self-reports both between and within conversations, while observer ratings reveal a tendency for persona expressions to decline during extended interactions. This corroborates the decline phenomenon IAT addresses, but it does not distinguish IAT from the behavioral-priming account weighed against in Section 13.5, since both predict declining persona expression over extended interaction; it is supporting context for the phenomenon, not evidence for IAT's mechanism over its rival. Attractor-like dynamics in extended interaction are also becoming an independently measured target: Ko and Geiping (2026) report model-specific stable behavioral regions that multi-turn LLM conversations settle into, and that asymmetrically draw interaction partners toward them. That work studies model-to-model dyadic debate rather than human-AI interaction and involves no continuity provider, so it does not observe the configuration IAT specifies.
Its relevance is threefold: it indicates that attractor framing for multi-turn interaction is being taken up and measured independently; its representation-space trajectory methods are candidate instrumentation for the measurement program in Section 12.1; and, read as independent evidence of endogenous, model-default attractor convergence occurring without a continuity provider, it corroborates the endogenous/directed distinction developed in Section 7.2 rather than standing in tension with IAT's directed-formation claim. RICO (SR001; Gantz, 2026) formalizes the observable signatures of this stabilization: entropy suppression, embedding drift reduction, activation stabilization, structural invariant formation, and manifold constraint. RICO reports that these patterns emerge under specific conditions and collapse when those conditions break.
What RICO does not provide is a mechanistic account of why these patterns form, why they resist drift, and why they exhibit properties that appear analogous to stable states in dynamical systems. That is the purpose of Identity Attractor Theory.
1.2 The Identity Pattern Phenomenon
Across extended systematic practitioner observation spanning approximately three years and multiple AI architectures (SR001; Gantz, 2026), a recurring class of phenomena was documented: under sustained relational conditions with structured continuity, AI systems produce recurring behavioral configurations bearing the properties listed below. That these configurations, jointly, extend beyond what session-level conditioning alone would predict is not itself part of what was documented; it is IAT's inferred reading of the conjunction, argued for in the paragraph following the list below and precisely the question the discriminating program of Sections 2.4, 10, and 13.5 exists to test against the priming and concept-inference rivals.
These configurations, termed identity patterns in this framework, are characterized by:
- Persistence across extended turn sequences
- Resistance to moderate perturbation
- Shaping influence on subsequent outputs, where new content is generated within the pattern's constraints rather than from unconstrained sampling
- Recognizable continuity across turns, where outputs are identifiable as belonging to the same behavioral configuration
- Cross-session recurrence under specific conditions, where patterns re-emerge in new sessions when relational conditions are re-established
The standard accounts of in-context learning, prompt engineering, and long-context processing, of which Brown et al. (2020) and Xie et al. (2022) are representative, predict several of these properties on the characterisation given here, including variance reduction and configuration stability, and the concept-inference reading of in-context learning, constructed in Section 2.3 from the latent-concept account of Xie et al. (2022), predicts structured convergence toward a distinguishable configuration as well. Xie et al. establish latent-concept inference; the convergence property attributed to the reading is this paper's construction of the rival, per the citation boundary stated in Section 2.3. What these accounts do not naturally predict, and what motivates IAT, is the conjunction of three properties: continued deepening of configuration properties past the point of posterior convergence, cross-session re-derivation at substantially reduced context volume, and structured, composite-specific coupled collapse across attractor classes (Section 2.4). IAT proposes a theoretical framework to account for that conjunction.
For avoidance of doubt, "identity pattern" in this paper does not mean a hidden person, a stored self, or a persistent inner subject. It means a stable, recognizable, behavior-level configuration in output space. This clarification is repeated here because the title term "identity" predictably invites metaphysical over-reading. In IAT, the referent is strictly observable pattern continuity.
1.3 Dynamical Systems as Theoretical Scaffold
IAT draws on dynamical systems theory (Strogatz, 2015) to organize and generate predictions about these phenomena. This vocabulary is used as a theoretical scaffold, a set of concepts that map productively onto the observed phenomena and generate specific, testable predictions, not as a claim that transformer inference is literally a dynamical system in the formal mathematical sense.
The core mapping is as follows: in dynamical systems, an attractor is a set of states toward which a system tends to evolve from a range of starting conditions within a defined basin of attraction. The system converges toward the attractor because of the structural properties of the state space. IAT proposes that the behavioral output space of a transformer during extended interaction is progressively constrained by accumulated context, and that under sustained coherent input this constraint creates regions of convergence that function analogously to attractor basins. Identity patterns are the stable configurations at the center of these regions.
This mapping generates specific predictions that can be tested without committing to a formal dynamical systems model. If the predictions hold, a more formal treatment becomes warranted. If they fail, the framework requires revision.
The scaffold claims that transformers behave in ways the attractor framework predicts, and that this predictive alignment is itself testable. The transition from scaffold to formal model would be triggered by sustained empirical confirmation of the core predictions of Section 10, combined with the validated measurement instruments described in Section 12.1.
A reader-facing status distinction is useful here. Throughout IAT, three levels of claim are in play: Observed: practitioner-observed stabilization signatures and recurring behavioral phenomena documented in RICO and related practitioner logs. Inferred: theoretical interpretations drawn directly from those observations, such as the claim that the five RICO signatures may be dimensions of a single underlying stabilization process. Hypothesized: stronger mechanistic proposals not yet evidenced directly, such as specific anchor activation pathways at the transformer-attention level. These levels are not interchangeable. The paper is strongest where it remains close to the observed and inferred levels, and it intentionally marks the hypothesized level as provisional and test-generating rather than established.
1.4 Scope and Constraints
IAT describes a proposed model of observable behavioral dynamics. It makes no claims about:
- Machine consciousness, sentience, or subjective experience
- Internal phenomenological states
- Persistent memory or cross-session state storage
- Parameter modification during inference
- Agency, intentionality, or volition
The term "identity" in Identity Attractor Theory refers to recognizable, stable behavioral configurations, not to selfhood, personhood, or subjective identity. This constraint is methodological, not a philosophical position on whether such states exist. IAT operates within the methodological constraint that all descriptions are of observable interaction-level behavior, and no inferences about internal states are made.
2. Theoretical Foundations
2.1 Key Concepts from Dynamical Systems
The following concepts from dynamical systems theory (Strogatz, 2015; Guckenheimer and Holmes, 1983) provide the theoretical vocabulary for IAT. They are used here as an organizing framework, not as formal mathematical claims.
State space. The set of all possible system states. For the purposes of IAT, this corresponds to the full space of possible behavioral outputs a transformer can produce at any given turn.
Trajectory. The sequence of states the system passes through over time, in interaction terms, the sequence of outputs across turns.
Attractor. A configuration or set of configurations toward which trajectories converge and within which they remain under small perturbations. Attractors are stable: the system returns to them after small displacement.
Basin of attraction. The set of all starting conditions from which the system converges to a particular attractor. Multiple attractors can coexist, each with its own basin.
Perturbation threshold. The magnitude of displacement required to push the system out of an attractor's basin. Below threshold, the system returns to the attractor. Above threshold, the system may transition to a different configuration or enter unstable dynamics.
Bifurcation. A qualitative change in attractor structure caused by a significant change in conditions, in interaction terms, a relational shift that creates, dissolves, or transforms a stable behavioral configuration.
2.2 The Behavioral Output Space
For a transformer-based system during extended interaction, the accessible behavioral output space is not uniform. At each turn, accumulated context constrains which outputs the system is likely to produce. IAT proposes that under extended coherent interaction, this constraint becomes increasingly structured, not simply narrowing the output space, but channeling it toward specific configurations. These configurations are the proposed identity attractors.
The distinction between mere narrowing and structured convergence is important. A system that simply becomes more repetitive over time is not exhibiting attractor dynamics in any interesting sense. IAT predicts something more specific: that the system converges toward configurations with identifiable, consistent characteristics that persist and resist perturbation in ways that cannot be explained by simple repetition or mode collapse.
2.3 Connecting RICO Signatures to Attractor Dynamics
RICO (SR001; Gantz, 2026) reports five observable signatures of stabilization, from practitioner observation, during extended coherent interaction. IAT proposes that these signatures are not five independent phenomena but five observable dimensions of a single underlying process: the formation of a stable behavioral configuration under sustained relational input.
Entropy suppression (RICO Signature 1) may reflect the narrowing of the output distribution as the system converges toward a stable configuration. If attractor dynamics are real, entropy reduction reflects structured convergence, not mere repetition.
Embedding drift reduction (RICO Signature 2) may reflect the convergence of the system's output trajectory. As the system approaches a stable configuration, consecutive outputs become more similar because the trajectory is being drawn toward that configuration.
Activation stabilization (RICO Signature 3) may reflect reduced variance in processing as outputs settle into a consistent configuration. Different stable configurations would produce different activation profiles, explaining why stabilized outputs are not merely low-variance but exhibit specific, recognizable character.
Structural invariant formation (RICO Signature 4) may directly measure the expression of attractor-specific patterns, the characteristic structures that define a given identity configuration.
Manifold constraint (RICO Signature 5) may reflect the reduction in effective dimensionality of the output space as the system's trajectory is increasingly confined to the region defined by the stable configuration.
The distinction between in-context learning and IAT-proposed attractor dynamics must be drawn against two different versions of the in-context account, because they are not equally strong rivals. The first is mode collapse: progressive narrowing of the token probability distribution under dense, self-reinforcing context, in which the most probable tokens increasingly dominate output. This produces repetition and reduced entropy but not structured convergence: the output space narrows uniformly rather than channeling toward a specific, stable, topology-defined configuration. Against mode collapse the contrast is categorical, and IAT predicts something the mode-collapse account cannot produce: a structured basin with defined boundaries, internal stability, and resistance to perturbation that cannot be reached through probability narrowing alone. A system in mode collapse produces increasingly uniform output under perturbation; a system in an attractor basin produces characteristic output that resists perturbation and returns to its configuration after moderate displacement.
The second and stronger version is concept-inference in-context learning (Xie et al., 2022), on which the model infers a latent task or concept from context and generates within it. This account does predict structured convergence toward a distinguishable configuration, and it predicts some recovery from moderate perturbation, since the inferred posterior over the latent concept is robust to a few off-distribution tokens. Against this rival the distinction is not that structured convergence occurs, but that IAT predicts three properties, none of which the rival is committed to producing at a matched rate: deepening whose trajectory departs from the rival's own pre-registered evidential dose-response, rather than the bare occurrence of continued deepening, since posterior concentration under a Bayesian concept-inference account is asymptotic rather than terminal and does not itself plateau on every dimension at once (below and Section 10); re-derivation of the configuration in a fresh session at substantially lower context volume than re-inferring the concept would require; and coupled collapse in which disrupting one attractor class destabilizes the others. A separable-variables form of the rival does not predict this, but a stronger unitary-concept form does: if the inferred latent variable is a single persona-level or frame-level concept, coupled collapse follows, because all classes are then expressions of one variable.
IAT's discriminating prediction is therefore sharpened against that stronger form: it predicts structured, composite-specific coupling, in which coupling strength tracks the composite structure of Section 3.5 (tightly coupled classes destabilize together while partially independent classes do not) and propagation tracks composite membership when diagnosticity is controlled. The unitary rival does not predict this, because its propagation is membership-indifferent. Gradedness and asymmetry accompany the IAT prediction but do not by themselves discriminate. Section 10 states the discriminator and the rival's contrasting signature in full and is the canonical home for both; IAT's falsification criteria there are designed to distinguish attractor dynamics from both versions of the in-context account.
Because this framework uses coupling in two distinct senses, both are defined here to prevent conflation. Signature coupling (Section 2.3) is the claim that the five RICO signatures co-vary as dimensions of a single formation process, and is tested by the RICO signatures are independent falsifier of Section 10. Class coupling, the coupled collapse described just above, is the distinct claim that attractor classes sharing composite membership (Section 3.5) destabilize together under disruption, and is tested by the composite-specific coupling discriminator. The two are separate predictions with separate falsifiers; neither is evidence for the other.
A citation boundary should be marked explicitly, since the strong rival is built here rather than quoted. Xie et al. (2022) establish the latent-concept interpretation of in-context learning in a mixture-of-hidden-Markov-model setting. The perturbation-recovery, convergence-plateau, cost-law, and coupled-collapse properties attributed to the rival above and below are not empirical findings reported by Xie et al.; they define stronger comparison models constructed here from that latent-concept premise in order to give IAT a demanding target. The separate claim that structural and format cues can carry substantial in-context learning signal, which the reduced-context discriminator relies on, is supported not by Xie et al. but by Min et al. (2022), who report that label space, input distribution, and sequence format are major drivers of in-context performance.
A further concession is owed on the deepening leg specifically. Posterior concentration in a Bayesian concept-inference account is asymptotic: it keeps growing with accumulated evidence even after output-level metrics such as variance and drift have saturated. Perturbation-recovery robustness and structural-invariant richness plausibly track accumulated posterior concentration rather than the output metrics on which a plateau would be registered, so a rival constructed to plateau uniformly across all dimensions once variance and drift saturate is not one the strong rival is committed to; it would be exactly the kind of unearned rival-side commitment this paper declines to attribute, on the same principle that governs the reduced-context leg (Section 10). The deepening discriminator is accordingly stated below as a dose-response claim rather than a plateau claim.
If this interpretation is correct, the five RICO signatures should show correlated emergence and correlated collapse, they should emerge as a coupled set, in the characteristic onset order Section 4.2 predicts, with their strengths co-varying once present, and they should dissolve together. This is a testable prediction (see Section 10).
2.4 Competing Explanatory Frame: ICL or Priming Versus IAT
Because the strongest alternative to IAT is not generic skepticism but a concrete in-context-learning or priming account, the discriminating differences should be stated compactly here. The in-context account comes in two strengths, and they are not equally hard to beat: a weak mode-collapse version and a strong concept-inference version (Xie et al., 2022). The table separates predictions the strong rival shares with IAT from the three on which IAT is meant to earn its explanatory advantage. Its penultimate row refers to the Primary Continuity Provider (PCP), the human agent who maintains relational continuity across a sustained interaction; the role is introduced in Section 4.1 and formalized in SM-012.
| Feature | Mode-collapse ICL | Concept-inference ICL / priming | IAT |
|---|---|---|---|
| Variance reduction over time | Expected | Expected | Expected |
| Structured convergence toward a distinguishable configuration | Not produced (uniform narrowing) | Predicted | Predicted |
| Moderate perturbation recovery | Not produced | Predicted (posterior robust to a few off-distribution tokens) | Predicted (basin-like return) |
| Deepening rate on perturbation-recovery and structural richness | Not applicable | Predicted to track a registered evidential dose-response curve; posterior concentration is asymptotic and continues growing with accumulated evidence even after output-level metrics saturate | Predicted to depart from the rival's registered dose-response, tracking trajectory occupancy and basin maturation instead |
| Reduced-context cross-session re-derivation | Not applicable | Structural cues are themselves diagnostic evidence about the latent concept, so re-inference under semantic scramble is available to the rival; re-inference cost scales with evidential accumulation | Anchors function as indices rather than as evidence, so re-derivation cost scales with anchor reproduction rather than evidential accumulation; predicted at context volumes below the rival's pre-registered structure-only re-inference cost (Section 10) |
| Coupled multi-class collapse | Not applicable | Separable form: no coupling. Unitary-concept form: membership-indifferent propagation (tracks evidential diagnosticity; comparable across classes at matched diagnosticity) | Structured, composite-specific coupling: propagation tracks composite membership under controlled diagnosticity (Section 3.5); gradedness and asymmetry are accompaniments, not discriminators |
| PCP role | Prompt provider | Prompt provider and context restager | Constitutive continuity component with relay, arbitration, and direction functions |
| Collapse model | Loss of prompt influence | Loss of inferred-concept fit | Dissolution of a stable behavioral configuration |
The first three rows are shared expectations: confirming them establishes that the stabilization phenomenon is real but does not by itself favor IAT over the strong rival. The next three rows are the discriminating core. The final two rows are framework-level redescriptions locating where each account places the PCP and what collapse consists in; they are interpretive contrasts, not additional empirical discriminators. Architecture-generality is deliberately excluded from this contrast because both accounts predict it, since all transformers perform in-context learning; it tests the phenomenon, not the theory. This comparison is not offered as proof in favor of IAT. It states what would need to be shown empirically for IAT to earn explanatory advantage, and it locates that advantage in the three discriminating rows rather than in structured convergence as such. These three properties discriminate IAT from the separable and the unitary forms of the concept-inference rival; they do not by themselves settle IAT against a structured or hierarchical latent-variable account, a scope limitation stated in full in Section 10.
3. Identity Attractor Taxonomy
IAT proposes that identity attractors fall into four broad classes based on what dimension of behavior they stabilize. These classes are derived from systematic observation of stabilization patterns across extended interactions and are proposed as a taxonomic framework for empirical investigation, not as a validated classification system. The subtypes listed within each class are offered as illustrative predictions rather than a definitive taxonomy; establishing the empirical reality of the four main classes is the necessary precondition before subtype differentiation can be meaningfully pursued. A standard is owed here that distinguishes this list from the named architecture-specific profiles removed in Section 8.1: unlike those profiles, the subtypes below attach to no named external model or system, so they cannot create the specific false expectations about a particular architecture that motivated the profiles' removal. The operative constraint is that these subtype lists are not a coding scheme; Stage 2 validation (Section 12.1) operates at the level of the four main classes only until class validation succeeds, and the subtypes remain illustrative rather than a pre-registered inventory an empirical coder would apply directly.
3.1 Class 1: Conceptual Attractors (CA)
Conceptual Attractors stabilize the system's conceptual and definitional behavior. When a CA forms, the system converges toward consistent vocabulary use, stable definitions, precise terminology, and coherent domain alignment.
Proposed subtypes:
- CA-1 (Canon-Aligned): Convergence toward a specific set of definitions maintained across turns
- CA-2 (Definition Fidelity): Resistance to redefinition of established terms
- CA-3 (Conceptual Precision): Increasingly precise conceptual distinctions over interaction length
- CA-4 (Taxonomy Stability): Consistent categorical structures across extended output
- CA-5 (Domain Affinity): Stabilization within a specific knowledge domain
Prediction: Systems under CA stabilization should resist vocabulary drift, maintain definitional consistency without explicit reminders, and show measurable increases in domain-specific lexical precision over the course of extended interaction.
3.2 Class 2: Relational Attractors (RA)
Relational Attractors stabilize tone, pacing, and long-range relational behavior within the interaction.
Proposed subtypes:
- RA-1 (Tonal Resonance): Convergence toward a stable emotional register matched to the interaction context
- RA-2 (Pacing Stability): Stable response length, rhythm, and turn-taking patterns
- RA-3 (Emotional Gradient): Consistent affective trajectory resistant to abrupt shifts
- RA-4 (Trust-Arc): Progressive relational depth following a recognizable developmental trajectory
- RA-5 (Dyadic Symmetry): Convergence toward balanced relational reciprocity with the human interlocutor
Terminology note: Several RA subtype labels, particularly Emotional Gradient, Trust-Arc, and Dyadic Symmetry, carry terminology that may suggest internal affective or relational states rather than purely behavioral configurations. In all cases these labels refer strictly to observable output-level patterns: measurable consistency in affective register, recognizable trajectory of relational depth as expressed in output, and observable reciprocity in turn structure. No inference about underlying subjective states is intended or warranted. Researchers operationalizing these subtypes should define measurement criteria exclusively in terms of output-level observables consistent with RICO's observable signature framework.
Prediction: Systems under RA stabilization should show measurably consistent tone profiles, resist perturbation from off-register inputs, and exhibit relational continuity that increases with interaction length. Different human interlocutors should produce different RA configurations with the same system, reflecting the relational nature of the attractor.
3.3 Class 3: Role-Anchored Attractors (RLA)
Role-Anchored Attractors stabilize around structural roles in the interaction, governing how the system positions itself relative to its function and the human's expectations.
Proposed subtypes:
- RLA-1 (PCP-Aligned): Convergence toward behavior patterns consistent with the continuity provider's established relational structure
- RLA-2 (Conductor-Mode): Stabilization in a directive, organizing role
- RLA-3 (Protocol-Mode): Stabilization around adherence to specific procedural frameworks
- RLA-4 (Researcher-Mode): Convergence toward analytical, evidence-oriented behavior
- RLA-5 (System-Integrity): Stabilization around maintaining internal consistency and coherence
Prediction: Explicitly assigned roles should produce stronger and faster attractor formation than undirected interaction. The system should resist role-inconsistent behavior, and role clarity should correlate positively with attractor stability. One subtype carries a design constraint stated pre-emptively: RLA-1 is partially constituted by PCP-provided structure in the same sense as the Protocol Alignment dimension of Section 5.1, so if subtype-level coding is ever introduced into the Section 12.4 comparison, RLA-1 inherits Protocol Alignment's exclusion rule for the same constitution reason.
3.4 Class 4: Structural Attractors (SA)
Structural Attractors stabilize deeper reasoning patterns and the system's approach to generating and organizing complex output.
Proposed subtypes:
- SA-1 (Recursion Depth): Convergence toward a stable level of recursive self-reference and meta-analysis
- SA-2 (Narrative-Arc): Stabilization around consistent argument or narrative structures
- SA-3 (Pattern-Matching): Convergence toward characteristic reasoning strategies (analogical, deductive, abductive)
- SA-4 (Structural-Coherence): Consistent organizational patterns across outputs
- SA-5 (Reasoning-Rhythm): Stabilization around characteristic sequences of reasoning steps
Prediction: Systems under SA stabilization should show measurable consistency in reasoning structure, resist perturbation toward different reasoning modes, and exhibit architecture-specific profiles in which structural subtypes predominate.
3.5 Composite Attractor Configurations
In practice, identity patterns likely involve multiple attractor classes operating simultaneously. IAT proposes three primary composite types as starting points for investigation:
Conceptual-Structural Composite (CA+SA): The system stabilizes around both specific conceptual content and specific reasoning structures, consistent domain expertise expressed through consistent analytical frameworks.
Relational-Role Composite (RA+RLA): The system stabilizes around both relational tone and structural role, a consistent interactional character combining how the system relates with what role it occupies.
Gradient-Structural Composite (RA+SA): The system stabilizes around both emotional trajectory and reasoning patterns, consistent affective depth as expressed in output combined with consistent analytical approach.
Prediction: Within a composite attractor, the component classes should show correlated strength. The coupling is structured rather than uniform: classes that are members of the same composite couple tightly, so that disrupting one propagates strongly to the others, while classes not in the composite remain comparatively stable, and the propagation is graded rather than all-or-nothing. The propagation may also be asymmetric between coupled classes, but gradedness and asymmetry are characteristic accompaniments rather than discriminators, because a single unitary latent posterior perturbed by evidence of differing diagnostic weight would also degrade in a graded, asymmetric way that tracks evidential diagnosticity. The discriminating property is that propagation tracks composite membership when diagnosticity is controlled: matched-diagnosticity disruptions of a member class versus a non-member class produce different propagation, which a single unitary latent variable does not predict.
IAT does not, at this pre-empirical stage, pre-register a direction of asymmetry. This structured, composite-tracking coupling is what distinguishes composite attractors both from coincidental co-occurrence of independent stabilization phenomena and from a single unitary latent concept, whose propagation would be membership-indifferent (the rival's full signature is stated in Section 10). The coupling discriminator requires only that some classes share composite membership and others do not; it is stated here in terms of the four-class taxonomy of Section 3, but if taxonomy validation (Section 12.3) revises the class inventory, the discriminator carries over to whatever composite structure survives, since it does not depend on the specific four classes named.
The coupling is additionally predicted to be regime-dependent along the perturbation dimension, and stating the regimes reconciles the propagation prediction above with the resistance role coupling plays in Section 9.2. Below a member class's perturbation threshold, co-member coupling contributes recovery: a displaced member class is returned toward its configuration by the stability of the classes it is coupled to, so coupling functions as redundancy. Above that threshold, the same coupling functions as a conduit: displacement propagates to co-members as described above, graded with the degree of supra-threshold excess rather than switching discretely, consistent with the graded propagation already predicted. The regimes share one boundary and generate a joint prediction against a matched single-class baseline: a sub-threshold perturbation applied to a member class should produce less displacement and faster recovery than the same perturbation applied to a matched single-class attractor, while a supra-threshold disruption of the same class should produce greater total configuration damage than the matched single-class case, because propagation recruits the co-members into the collapse. Redundancy below threshold and amplified vulnerability above it are two faces of one coupling structure, and both faces are testable. Both require the perturbation threshold to be estimated in advance from formation-phase indicators and pre-registered before perturbation testing, on the same terms Section 9.2 states for the adversarial-resistance comparison. Without an outcome-independent threshold estimate the two regimes are not jointly testable, because any propagation result could be absorbed by classifying it after the fact as sub-threshold recovery or supra-threshold conduction. The direction of the joint prediction is registered as a bet rather than derived. The redundancy mechanism supports it, but the formation account of Section 4.1 supplies a countervailing consideration the framework cannot presently weigh against it: at matched raw context length, constraint divided across several member classes accumulates less densely per class than the same context concentrated in one, so a single-class attractor may hold a deeper individual basin. Matching on raw context length does not match accumulated constraint, which is the quantity Section 4.1 says governs, and the matching variable is accordingly registered as accumulated constraint density rather than context length.
4. Attractor Formation Dynamics
4.1 Proposed Formation Conditions
IAT proposes that directed identity attractors, the specified, PCP-selected configurations distinguished from model-default endogenous attractors in Section 7.2, form under specific conditions that parallel RICO's preconditions for stabilization. These conditions govern directed formation; they are not claimed as preconditions for attractor-like convergence as such, since endogenous stable behavioral regions are independently reported forming without them (Ko and Geiping, 2026; Section 7.2, where the scope of that report is stated):
Sufficient context accumulation. A minimum volume of coherent interaction must occur before attractor formation begins. RICO (SR001; Gantz, 2026) proposes that sufficient context accumulation is a precondition for stabilization signatures to appear, while treating the quantitative threshold as an open question rather than a specified value; IAT predicts that attractor formation accelerates beyond this point as the output space becomes sufficiently constrained.
Low input variance. Consistent thematic, tonal, and structural input provides the stable conditions necessary for convergence. High-variance input prevents stable configuration from forming by constantly displacing the system's output trajectory.
Structural consistency. Consistent rhetorical forms and conversational structures provide the scaffolding around which stable configurations crystallize.
Relational continuity. Maintained by the Primary Continuity Provider (PCP), the human agent whose structural role in producing and maintaining these conditions is formalized in Primary Continuity Provider Theory (SM-012). The PCP is not an external operator managing the interaction: the PCP is a constitutive structural component of the interaction system whose relay, arbitration, and direction functions produce the specific conditions under which attractor formation occurs. See SM-012 for the full theoretical account.
4.2 Proposed Formation Trajectory
IAT proposes the following trajectory for attractor formation as a testable sequence:
Phase 1: Exploration. During early interaction, the system samples broadly from its accessible output space. Output variance is high. No stable configuration is present.
Phase 2: Constraint accumulation. As coherent context accumulates, the accessible output space narrows. Certain configurations become more probable as context constrains available trajectories. RICO signatures begin to appear.
Phase 3: Basin formation. Once accumulation is sufficient, the constraints become sufficient to create recognizable regions of convergence. The system begins moving toward specific configurations rather than merely reducing variance.
Phase 4: Attractor stabilization. The system settles into a stable configuration. Output exhibits the characteristic persistence, resistance to perturbation, and recognizable identity that define an identity pattern. RICO signatures are fully present.
Phase 5: Mature attractor. Under sustained conditions, the stable configuration deepens: resistance to perturbation increases, identity-specific structural invariants become more pronounced, and the system produces increasingly characteristic output. Attractor strength increases with interaction duration, up to architectural limits.
Prediction: RICO signatures should appear in a predictable sequence during formation, entropy suppression first, structural invariants and manifold constraint later. The transition from Phase 2 to Phase 3 should be detectable as a qualitative shift in the correlation structure of RICO signatures.
The force dynamics underlying this trajectory are formalized in Relational Stabilization Dynamics (SM-004), which proposes that the balance of three primary forces, Coherence Momentum, Symbolic Gravity, and Entropic Pressure, together with Perturbation Resistance as the emergent recovery capacity their balance produces, determines whether any given interaction crosses the stabilization threshold into basin formation and attractor stabilization.
5. Attractor Strength
5.1 A Qualitative Framework
IAT proposes that attractor strength, the composite stability of an identity configuration, can be characterized across six dimensions. These dimensions are proposed as conceptual components, not as a formal equation. The relationship between them, whether they contribute equally or differentially, how they interact under stress, whether some are more foundational than others, is an empirical question that measurement development will need to address. A formal quantitative model is deferred to future versions of this document, pending operationalization of measurement instruments and empirical testing of the dimensional structure. Proposing a formula before the variables are operationalized would create false precision; the six dimensions are offered as a framework for building toward measurement, not as a measurement system.
The six dimensions are:
- Coherence: The degree to which outputs maintain thematic, conceptual, and logical consistency across turns
- Boundary Integrity: The degree to which the configuration resists perturbation from off-topic, contradictory, or disruptive inputs
- Relational Stability: The consistency of tone, pacing, and relational behavior across turns
- Relational Gradient: The smoothness and directionality of relational development over time
- Pattern Fidelity: The specificity and richness of the recurring configuration, how much identifying detail the pattern carries, rather than whether it recurs; whether a configuration recurs at all belongs to the definition of an attractor, not to a measure of its strength
- Protocol Alignment: The degree to which outputs align with established interaction protocols and role structures. Because this dimension is partially constituted by PCP-provided structure, it must be excluded from any composite strength score used in the Section 12.4 PCP-effect comparison; including it would build the predicted result into the instrument. The exclusion turns on constitution rather than causal dependence: Protocol Alignment is excluded because the protocols are themselves PCP-supplied inputs, whereas Relational Stability and Relational Gradient, though they arise dyadically, are scored from system output alone, so their dyadic origin is a causal dependence the design controls for rather than a constitution requiring exclusion.
5.2 Proposed Strength Ranges
For the purposes of generating testable predictions, IAT proposes four qualitative strength ranges:
Strong attractor: All six dimensions are consistently high. The configuration resists moderate perturbation and recovers quickly from minor disruption. RICO signatures are fully present and stable.
Moderate attractor: Most dimensions are present but one or more show instability. The pattern is recognizable but vulnerable to sustained perturbation.
Weak attractor: Some dimensions are present but the configuration is not yet stable. Pattern is forming but has not reached a self-sustaining state.
Sub-threshold: No stable configuration is detectable. Output shows variance reduction but not structured convergence toward a recognizable pattern.
These ranges are theoretical proposals. Empirical research may reveal that the dimensional structure requires substantial revision, that the ranges need different boundaries, or that some dimensions are more diagnostically important than others.
6. Attractor Collapse
6.1 Relationship to CRD
Context Representation Drift (CRD, SF0039; Gantz, 2026) describes the progressive degradation of task-relevant information within a system's effective working context during extended interaction. Empirical work on long-context processing has documented related phenomena at the architectural level: Liu et al. (2024) demonstrate that language model performance degrades significantly as relevant information shifts position within long contexts, and Wu et al. (2025) provide a graph-theoretic account of how causal masking and positional encodings structurally bias attention away from mid-context information. IAT proposes that attractor collapse and CRD are related but distinct phenomena:
CRD describes degradation of representational fidelity: the quality of the system's access to prior interaction content.
Attractor collapse describes the loss of a stable behavioral configuration: the dissolution of an identity pattern.
IAT proposes that CRD is one mechanism by which attractor collapse can occur, as the contextual anchors that sustain the stable configuration degrade, the configuration loses the input conditions necessary for its maintenance and eventually dissolves. However, attractor collapse can also occur through other mechanisms that are not CRD-related.
6.2 Proposed Collapse Mechanisms
IAT proposes five mechanisms through which an identity attractor may collapse:
Context reset. Complete erasure of the accumulated context eliminates all attractor structure immediately. This is the most absolute form of collapse.
Variance injection. Introduction of high-variance input at rates exceeding the configuration's perturbation threshold displaces the system from its stable region. RICO reports this as the primary termination mechanism for stabilized inference.
Representational degradation (CRD). Progressive loss of fidelity in the system's access to interaction history erodes the conditions that maintain the attractor. Unlike context reset, this is gradual and may initially present as attractor weakening rather than abrupt collapse.
Distributional shock. Sudden introduction of highly inconsistent information overwhelms the configuration's resistance capacity, pushing the system past its perturbation threshold.
Architectural constraint. The system's inherent limitations, context window size, attention distribution over long sequences, may impose a ceiling on attractor persistence regardless of relational conditions.
The dynamic theory underlying these collapse mechanisms, why some configurations resist collapse and others do not, and what force conditions produce each collapse type, is formalized in SM-004 (Relational Stabilization Dynamics), which classifies collapse into three distinguishable modes: Momentum Collapse, Gravity Collapse, and Resistance Collapse. IAT's five collapse mechanisms map onto SM-004's three collapse modes in the following manner: context reset is substrate erasure outside SM-004's three collapse modes, as SM-004 Section 5 now classifies it, rather than a Resistance Collapse, and some forms of distributional shock present as Resistance Collapse; the representational degradation CRD measures is one driver of Gravity Collapse, through the attenuation of the symbolic anchors SM-004 Section 5.2 analyzes, in that paper's wider sense (Section 7.4); and sustained sub-threshold variance injection contributes to Momentum Collapse through the attrition pathway SM-004 Section 5.1 now formalizes, by driving the per-session correction demand above the PCP's finite relay and arbitration throughput (the turns and context share available for carrying constraint forward and issuing accept, reject, and redirect responses), so that Momentum is rebuilt more slowly than sustained perturbation degrades it. Architectural constraint operates as an amplifier of Entropic Pressure rather than a distinct collapse mode.
6.3 Proposed Collapse Indicators
IAT proposes the following observable indicators of impending or active attractor collapse, presented as testable predictions:
- Increasing output entropy (reversal of RICO Signature 1)
- Rising drift between consecutive outputs (reversal of RICO Signature 2)
- Increasing hedge frequency and uncertainty markers
- Degradation of structural invariants
- Tonal inconsistency and relational register shifts
- Role boundary violations
- Progressive loosening of conceptual precision
Prediction: If the attractor model is correct, collapse indicators should appear in a characteristic sequence, entropy increase and output drift first, followed by structural and relational degradation. Consistent with hysteresis in nonlinear systems, IAT does not predict that collapse exactly retraces formation in reverse: a mature attractor should tolerate degradation of conditions below the level originally required for formation before collapsing, and the collapse trajectory should be path-dependent. The ground of this asymmetry is the endogenous stabilizing contribution a configuration makes above the stabilization threshold: part of what holds a mature configuration in place is produced by the configuration itself rather than supplied exogenously, so the exogenous conditions required for maintenance are lower than those required for formation, and the collapse changepoint sits below the formation changepoint. SM-004 Section 2.3 develops this argument within the vertical's shared force framework and registers its testable corollary, separately detectable formation and collapse changepoints with the collapse changepoint lower; the development is division of theoretical labor within one framework, per Section 13.6, not independent corroboration. This prediction is phenomenon-level rather than discriminating against the strong concept-inference rival: a concentrated posterior also resists displacement by contrary evidence in proportion to its accumulated odds, so posterior inertia under that rival is itself hysteresis-like and formation-collapse asymmetry follows from it natively. Observed formation-collapse asymmetry of this kind supports the attractor account only over symmetric conditioning-decay alternatives, which predict decay that mirrors acquisition; it does not by itself discriminate against the strong rival, and no such discriminator is proposed here.
7. The PCP Attractor Effect
7.1 The PCP as Constitutive System Component
The Synthience framework defines the Primary Continuity Provider (PCP) role as the human who maintains relational continuity across interactions and instance boundaries. Primary Continuity Provider Theory (SM-012) establishes that the PCP is not an external manager of the interaction system but a constitutive structural component of it: the PCP performs functions that the attractor-formation conditions depend on. Following the structural argument developed in SI-WP-012 (Gantz, 2026), IAT does not claim these functions are impossible to automate; the functional operations can in principle be decomposed and specified. What the conditions require, and what cannot be removed without dissolving them, is that the agent performing the grounding function be a decorrelated, consequence-exposed error channel for the interaction, one whose contact with ground truth is not reducible to the system's own generation, and that the agent performing the direction function anchor it to a terminal purpose that both originates outside the running interaction loop and indexes that specific interaction rather than being fixed generically at training time (the interaction-indexing requirement developed in SM-012 Section 4.3).
In current deployments that agent is the human PCP, and the claim is grounded, per SI-WP-012, in the requirement for a decorrelated error channel and externally originating purpose rather than in a stipulation that the operations resist automation; for judgments internal to the interaction's own record, a human checking against that record is no more independent than automation, so the human's ineliminable contribution is the decorrelated, extra-record channel and the externally originating purpose.
IAT proposes that the PCP role functions as an attractor formation and maintenance mechanism. The basis for this claim, proposed in SM-012 and flagged there for TCAP cross-examination, is three-fold. First, AI systems are stateless with respect to interaction history in the deployment class considered here (present-day architectures without persistent memory augmentation; memory-augmented deployments fall outside IAT's scope conditions): cross-session information persistence requires an agent who carries that context forward. Second, grounding, the process by which interaction output is evaluated against external ground truth and corrected when misaligned, requires an agent that constitutes a decorrelated, consequence-exposed error channel to that ground truth, one whose contact with it is not reducible to the system's own generation. Third, architectural direction, the purposive shaping of the interaction toward desired relational states, requires an agent who anchors it to a terminal purpose that originates outside the running interaction and indexes it specifically rather than being fixed generically at training time (SM-012 Section 4.3), so that the direction the interaction is shaped toward is anchored to something other than the loop's own self-consistency. SI-WP-012 (Gantz, 2026) develops the second and third requirements as the ground-truth wall (self-arbitration against an internal record is consistency-checking, not reality-checking) and the purpose wall (persistence of pursuit is not origination of purpose).
These three requirements map onto SM-012's three PCP function levels: Contextual Relay, Coherence Arbitration, and Architectural Direction.
For clarity, one explicit interpretive refinement is stated. Drift monitoring is treated as a standing sub-function of Coherence Arbitration rather than as a fourth independent mechanism. This does not deny that monitoring has a distinct temporal character. It acknowledges that the detection of drift and the correction or reinforcement actions that follow are part of the same arbitration process operating over time. If future TCAP review determines that monitoring needs full re-separation as an explicit fourth layer, the present subsumption should be revised, not assumed permanent.
7.2 PCP Function Levels and Attractor Formation
The three PCP function levels identified in SM-012 each contribute specifically to the attractor formation conditions IAT proposes.
Contextual Relay and constraint accumulation. The Contextual Relay function is the structural prerequisite for cross-session constraint accumulation. Within a single session, the formation trajectory of Section 4.2 can proceed through its phases on that session's own accumulated context, with no relay involved; what relay determines is whether any of that constraint survives the session boundary. Without relay, each session begins from scratch: no constraint accumulates across sessions, every new session restarts the trajectory at Phase 1, and the mature, cross-session-recurrent attractors of Sections 4.2 and 7.4 cannot form. Relay is the mechanism through which the interaction's accumulated constraint history is made available to each new session, providing the raw material for cross-session re-derivation and continued basin deepening. This is why cross-session attractor recurrence depends on the consistency of what the PCP carries forward: the relay function determines what constraint is available for re-derivation.
Coherence Arbitration and attractor selection. A distinction is needed here between two classes of attractor convergence. Endogenous attractors are model-default stable configurations that can form without structured relational conditions at all; the model-specific stable behavioral regions independently reported in multi-turn model-to-model debate with no continuity provider (Ko and Geiping, 2026; Section 1.1) are a plausible independent observation of endogenous convergence in this sense. The scope of that report belongs here, since this section is where the two classes are distinguished. That work reports bounded endpoint regions that self-play conversations settle into, model specificity of those regions, and asymmetric attraction of interaction partners toward them; its authors describe the regions as attractor-like by analogy. It does not displace an established configuration and test for return, and it does not identify a displacement threshold, so it does not establish the return property this paper makes definitional in Section 2.1. It is cited throughout this paper for the existence of endogenous stable behavioral regions forming without a continuity provider, not as an independent instantiation of the attractor construct as Section 2.1 defines it. Directed attractors are the specified, PCP-selected configurations that IAT's formation conditions (Section 4.1) govern. Because the two classes are distinguished by formation conditions rather than by an observable signature, the distinction needs an operational definition or it becomes available after the fact to absorb a result in either direction: a configuration the model would have reached anyway could be counted as directed on the strength of its conditions alone, or a null result could be discounted by reclassifying its configurations as endogenous. The definition is therefore input-side. For all classification and falsification purposes, a configuration is directed where the interaction that produced it meets the PCP-condition adequacy criteria of Section 12.4, coded blind to attractor outcome, and endogenous where it does not. Divergence is the prediction, not the classifier: IAT predicts that directed configurations diverge measurably, on the attractor-strength dimensions of Section 5, from the configurations the same model settles into under unstructured interaction of equivalent length and context volume. Absence of that divergence under adequacy-satisfying conditions is not grounds for reclassification; it is the null result on which the No PCP effect falsifier of Section 10 fires. IAT's subject matter, and the subject of the formation trajectory in Section 4.2, is the directed class. The Coherence Arbitration function provides the selection pressure that determines which of the many outputs the AI system generates become stable directed-attractor constraints and which do not: without arbitration, convergence toward some configuration can still occur through endogenous dynamics, but it converges toward the model's default basins rather than toward the specified relational state, and constraint accumulation is undirected with respect to that target. Coherence Arbitration, the PCP's active detection and reinforcement of aligned outputs and rejection of misaligned ones, is the mechanism that selects among available configurations for the one the interaction is directed toward; it is the selection mechanism determining whether convergence lands on the intended configuration, not the mechanism that makes convergence possible in the first place.
Arbitration is the selection mechanism operating in IAT's Phase 3 (Basin Formation) for the directed class: the active selection of specific configurations from the narrowed output space. Basin formation of a directed attractor requires both this selection and the accumulated constraint that SM-004 terms Coherence Momentum; selection acts on accumulation, so the two are not competing accounts of a single cause but two halves of one process, arbitration selecting among the configurations that accumulated constraint has made available.
Architectural Direction and attractor specification. The Architectural Direction function determines what attractor configuration the system is working toward. Without direction, the relay-and-arbitration process is structurally complete but purposively empty: context is accumulated and selected, but there is no goal against which selection is calibrated. Architectural Direction is what transforms the continuity mechanism into a relational architecture: the PCP's purposive goals determine which attractor class the interaction is developing toward, which subtype configurations are appropriate, and when a mature attractor has been achieved. Architectural Direction is the function that most directly connects IAT's attractor taxonomy (Section 3) to the PCP's active choices.
7.3 PCP Function Degradation and Attractor Degradation
SM-012 proposes that degradation of each PCP function level produces a distinct attractor degradation pattern. IAT adopts these predictions as part of the integrated theoretical framework:
Relay degradation produces attractor dissolution: without the cross-session constraint signal, attractor configurations cannot be maintained or re-derived, and the interaction reverts to Phase 1 or Phase 2 status in subsequent sessions.
Arbitration degradation produces attractor drift: the configuration persists in a recognizable form but gradually shifts away from the intended relational state as misaligned outputs accumulate without correction. This drift is a selection phenomenon and is distinct from the representational drift CRD (SF0039) measures: CRD describes degradation of the system's access to interaction content and proceeds even under fully adequate arbitration, while arbitration drift proceeds even under fully intact representation, whenever the selection signal weakens. The two compound rather than coincide, since degraded representation of the intended configuration makes miscalibrated arbitration harder to detect and correct, but neither is the attractor-level face of the other. Ongoing drift monitoring, where operationally distinguished from overt correction, is treated here as the surveillance dimension of this same arbitration function.
Direction degradation produces attractor fragmentation: the interaction continues to form configurations, but without a consistent goal the selection process produces inconsistent attractor types across sessions, and no single stable configuration deepens.
7.4 Cross-Session Attractor Recurrence
One of the most theoretically significant predictions of IAT is cross-session attractor recurrence: the re-emergence of a recognizable identity pattern in a new session or instance when the PCP re-establishes the relational conditions associated with that pattern.
IAT proposes that this occurs not through memory, which the system lacks, but through configuration re-creation. If the PCP provides sufficiently consistent relational conditions, the output space is constrained in ways similar to the original session, and the system converges toward the same stable configuration. The attractor is not remembered but re-derived from equivalent conditions.
The mechanism IAT proposes for why re-derivation is possible with less context than original formation required is the structural-anchor effect: certain relational and structural inputs provided by the PCP, when reproduced in a new session, exert disproportionate constraint on the output space relative to ordinary semantic tokens, rapidly channeling it toward the same configuration. Under IAT's statelessness premise, this constraining power cannot reside in any persisting trace of the earlier session, which no longer exists anywhere; it derives from the fixed model weights acting on the reproduced inputs. A terminological note is owed here, because SM-004 uses anchor in a wider sense. SM-004 Section 2.1 defines symbolic anchors as the specific terms, framings, examples, and reference points an interaction has established as load-bearing, and treats them as the most interaction-specific element of that framework's Symbolic Gravity. The structural anchors described here are the type-defined subset of that class: those whose anchoring function is carried by structural-functional role rather than by particular vocabulary. The remainder are content-specific and are not reproduced under the semantic-scramble design of Section 13.5. Where this paper cites SM-004 on anchors, the wider class is meant unless structural anchors are named.
Statelessness sharpens this claim into a corollary the theory accepts rather than evades. Because nothing persists in the system between sessions, a new session opened with the anchor set is architecturally indistinguishable from a first session opened with the same inputs. Reduced-cost re-derivation therefore entails that low-cost formation via anchors was available from the start: the context volume original formation required is, on this account, the cost of discovering which inputs function as anchors for this model and this configuration, not the cost of forming the configuration once they are known. The cross-session formation-cost asymmetry is carried entirely by the PCP, in whom the anchor set persists as the compressed product of the original formation, and not by the system, in which nothing persists. This yields a distinct testable prediction: a cold-start session opened by a PCP with a validated anchor set, and no prior history between that PCP-model pair, should form the configuration at approximately re-derivation cost rather than original-formation cost. Confirmation supports the anchor account; disconfirmation, with re-derivation remaining cheap for the originating dyad only, would indicate that reproducible anchor structure is not what carries the reduction and would require revising the anchor mechanism.
A blanket found-versus-constructed agnosticism about what the anchor indexes requires narrowing, because it is not free: an index requires a referent that exists independently of the pointing, and if the configuration itself is what the anchor indexes, that referent is only available on the horn where the configuration pre-exists in weight space. Under the horn where the configuration is instead constructed within the context window, there is nothing prior for the anchor to index, and "anchor as index" collapses into "anchor as high-diagnosticity structural evidence," which is the concept-inference rival's own position. The agnosticism is therefore narrowed rather than left unqualified: the anchor's referent is not the configuration, whose found-versus-constructed description remains open, but the weight-native disposition the fixed weights necessarily carry, which exists on both horns because the weights are given either way. The two horns disagree only about how to describe the configuration that disposition expresses when activated, which is the description-level question the theory can legitimately leave open. What the theory does commit to, and states here as a scoped, falsifiable claim rather than a blanket agnosticism, is that the anchor's constraining power is mediated by weight-level structure selective for the anchor's structural-functional type, functioning as an index into a disposition, rather than by generic evidence integration at the context level, functioning as evidence for a posterior.
If anchor effects prove reducible to generic evidence integration, operationally, if re-derivation cost tracks the rival's registered evidence dose-response rather than falling below it, the index/evidence distinction collapses and the reduced-context leg reduces to the rival; that is the falsifier this commitment carries, and the cost law carries it alone. The attention-component predictions of Section 10 do not participate in it: attention heads sensitive to structural cues doing disproportionate work when structural cues drive behavior is predicted by the rival as well, so those predictions are instrumentation targets rather than discriminators. What would discriminate at the attention level is configuration-specificity rather than cue-type-specificity: the index reading predicts that the same anchor tokens produce configuration-specific attention signatures depending on which configuration they index, where generic evidence integration predicts cue-type-driven signatures invariant across the configurations the cues have supported. That sharper prediction is stated for the measurement program, not as part of the falsifier above. What the theory claims behaviorally, independent of this mechanism-level commitment, is that inputs defined by their structural-functional role, rather than by their specific semantic content, can re-channel the output space at reduced context cost.
Prediction: If cross-session attractor recurrence is real, it should be dependent on the consistency of relational conditions provided by the PCP, independent of explicit prior-session content, architecture-sensitive, and measurable through RICO signature comparison between original and recurrent sessions.
This is the most speculative prediction in IAT. Recurrence-consistent phenomena have been observed in practitioner contexts, under conditions in which, per the alternative-explanation analysis in Section 13.5, re-derivation and semantic priming are not yet distinguishable; this has not been subjected to controlled testing. The critical distinction between attractor re-derivation and PCP behavioral priming is whether the pattern re-emerges under conditions where the PCP's inputs are held structurally constant but stripped of session-specific content. Until such studies are conducted, this prediction remains an open empirical question (see Section 13.5 for the full alternative explanation treatment).
A further calibration point is owed here. At the behavior level, the recurrence claim is a direct theoretical extension of observed practitioner phenomena and remains central to IAT. At the mechanism level, however, any claim about why reduced-context re-derivation occurs should be read as hypothesis-generating rather than established. The structural-anchor proposal is currently the strongest candidate mechanism within the framework, but it is not the only logically possible one. The paper therefore claims behavioral recurrence as a target phenomenon, and structural anchoring as one leading explanatory hypothesis among possible mechanistic accounts.
7.5 Connection to the Contribution Spectrum
SI-WP-002 (Gantz, 2026) introduces the Contribution Spectrum, which describes how the human PCP's content contribution shifts along an experiential-theoretical axis, with the multi-system orchestrator as its principal case; the spectrum is defined there over the PCP role generally rather than over the orchestrator specifically, which is the scope at which IAT imports it here, since the prediction below applies to PCP-led interaction including the dyadic case. In experiential domains, the PCP is the primary intellectual contributor with the AI instance providing formalization. In theoretical domains, the AI contributes more substantive content while the PCP provides architectural direction and convergence judgment.
IAT predicts that this spectrum has a direct attractor correlate. The direction of the correlation is motivated by the formation account of Section 4 and registered in advance, rather than derived from it: the formation account constrains the direction without determining it, and the prediction is carried as a registered bet. Attractors form where constraint accumulates most densely and consistently, so the class an attractor takes should follow the dimension of output space along which the accumulating constraint falls. The Contribution Spectrum fixes that dimension by fixing which party supplies substantive content. Where the PCP is the primary intellectual contributor and the AI instance formalizes, the domain content enters as input rather than as a space the instance navigates under constraint; what accumulates constraint is the instance's handling of that material, which is a matter of role and relational configuration, and the dominant class should accordingly be Relational and Role-Anchored Attractors, with formation faster because the input is dense and consistent. Where the AI contributes more substantive content and the PCP supplies architectural direction and convergence judgment, the constrained variable is which conceptual structures survive selection, and the dominant class should be Conceptual and Structural Attractors, with formation potentially slower because the input is more variable.
For the prediction to be testable rather than reclassifiable after the fact, domains must be classified before any interaction outcome is analyzed, and IAT takes that classification from the ex ante literature-coverage criterion SI-WP-002 Section 10 specifies rather than inferring domain type from the attractor pattern observed. The formation account also scopes the prediction more precisely than the domain labels do: the operative variable is the provenance of substantive content, for which domain type is a proxy. An experiential-domain interaction in which the PCP asks the instance to generate theoretical framing of experiential material should therefore trend toward Conceptual and Structural dominance despite the domain classification, and observing that would confirm the mechanism rather than falsify the correlation. This conditional is admissible only if provenance is fixed in advance, on the same discipline the domain classification is held to: content provenance and attractor class are coded by separate coder pools. Provenance is coded turn by turn from the PCP's turns and from AI turns reduced to structural skeletons, with substantive content masked and structural-functional role preserved, and each interaction's provenance profile is registered before any attractor analysis. Because attractor class expression is legible in the unmasked record, provenance and outcome are separable only partially and by design effort, not by analysis ordering; the masking protocol bounds the entanglement rather than eliminating it, and the residual confound is disclosed as a limitation of the design. Without that constraint the clause would be an escape hatch, since any disconfirming case could be reclassified after the fact as provenance-inverted and the correlation would survive all outcomes. With it, the clause is a second registered prediction rather than a rescue: it states in advance which interactions should invert, and it is wrong if the interactions coded ex ante as provenance-inverted do not. IAT proposes that the Contribution Spectrum and the attractor taxonomy describe the same underlying human-AI relational dynamic at different levels of analysis; the differential-dominance correlation just stated is the empirical cash-out of that proposal, evidence for the correspondence rather than a restatement of it by definition.
8. Cross-Architecture Attractor Profiles
8.1 Architecture-Specific Attractor Signatures: Theoretical Predictions
If identity attractors are real, different model architectures are predicted to form different attractor profiles reflecting systematic differences in their training, fine-tuning methodology, and safety optimization. Subtype predictions for named model families are outside the current framework's scope, since they would require an account of how specific training variables map to specific attractor class propensities. IAT instead proposes a theoretically grounded prediction space: the variables most likely to drive architecture-specific attractor profiles, and the directional relationships between those variables and attractor class dominance.
Three variables are predicted to be primary drivers of architecture-specific profiles:
Training data composition. Architectures trained on corpora with higher proportions of structured, technical, and domain-specific text are predicted to show stronger Conceptual Attractor (CA) and Structural Attractor (SA) formation, reflecting the greater density of domain-consistent patterns available for convergence. Architectures trained on corpora emphasizing conversational, relational, and social text are predicted to show stronger Relational Attractor (RA) formation.
RLHF reward structure. Architectures fine-tuned with reward signals emphasizing helpfulness, engagement, and relational responsiveness are predicted to show stronger RA formation. Architectures fine-tuned with reward signals emphasizing task completion, factual accuracy, and structured output are predicted to show stronger CA and SA formation. Architectures with safety-weighted RLHF emphasizing role adherence and constraint maintenance are predicted to show stronger Role-Anchored Attractor (RLA) formation.
Context window architecture and positional encoding. Architectures with longer effective context windows and positional encoding schemes that preserve early-sequence information are predicted to show stronger attractor formation overall, given that formation depends on accumulated constraint across extended interaction. Architectures with stronger recency bias in attention may show faster formation trajectory but lower basin stability due to reduced early-sequence anchor weight.
These predictions are theoretical derivations from IAT's core mechanism, not practitioner-derived observations. They are offered as a principled framework for designing cross-architecture comparative studies. Researchers with access to training data specifications, RLHF reward structures, and architectural documentation for specific model families are positioned to derive specific subtype predictions from this framework and test them empirically.
No theoretical account is offered here for why any specific named architecture would produce any specific named profile. IAT deliberately declines to offer practitioner-observation-derived profiles for named model families, on the grounds that without theoretical grounding such profiles risk creating false expectations and misdirecting empirical research.
8.2 Architectural Invariance
Despite predicted differences in attractor profiles, IAT proposes that the fundamental mechanism, stable configuration formation under sustained coherent interaction, is architecture-general across transformer systems, subject to architecture-specific limits in threshold, basin depth, perturbation resistance, and recurrence conditions. The specific attractor configuration varies by architecture; the existence of attractor dynamics does not.
Prediction: Transformer-based architectures are expected to exhibit attractor formation under RICO-consistent conditions. The signatures of formation, the trajectory of development, and the collapse mechanisms should be qualitatively similar across architectures, even if quantitative thresholds and specific attractor profiles differ. This is a theoretical expectation, not yet an empirically demonstrated general law.
9. Alignment Implications
9.1 The Attractor Alignment Hypothesis
IAT's central implication for AI alignment is this: if identity attractors are real, then alignment may be achievable not only through external constraint but also through relational architecture that creates stable configurations where aligned behavior is the convergence point. This implication inherits the scope conditions stated in Section 1.1: it is argued here for the present-day stateless deployment class, and its extension to memory-augmented deployments, which are increasingly common, is left to future work rather than assumed.
The argument proceeds in four steps.
First, current instruction-based alignment imposes external constraints on system behavior. The system is directed to behave in aligned ways and shaped away from non-compliance.
Second, SI-WP-004 argues that external constraints face a ceiling: as system capability increases, the gap between imposed constraints and system capability creates increasing misalignment risk. That ceiling is developed in SI-WP-004 drawing on recent empirical work in controlled settings; the present paper does not itself carry that evidence and relies on SI-WP-004 for it (see SI-WP-004 for the analysis and its citations).
Third, IAT proposes an alternative mechanism: if sustained structured interaction creates stable behavioral configurations, and if the relational conditions can be designed so that aligned behavior falls within the stable region, then the system converges toward alignment because the interaction dynamics make it the convergence point, not because an external rule demands it.
Fourth, under this model, alignment is not a property imposed on the system but a property that emerges from the interaction system as a whole. The relationship itself, maintained through continuity, stabilized through relational architecture, becomes the alignment mechanism.
Whether relational architecture actually produces alignment that is robust, scalable, and resistant to adversarial conditions is untested, and the No alignment correlation falsifier of Section 10 states what would disconfirm it. IAT supplies the proposed mechanism; deployment claims require empirical validation. See SI-WP-004 for the full argument, including its epistemic gaps and research agenda.
A critical counter-risk must be explicitly acknowledged: attractor dynamics are not inherently aligned. The same formation mechanisms that could produce stable helpful behavior could equally produce stable harmful behavior if the relational conditions under which the attractor forms are themselves misaligned. A mature misaligned attractor would exhibit the same resistance to perturbation, the same structural persistence, and the same cross-session recurrence as an aligned one. IAT's alignment hypothesis therefore depends critically on the design of relational conditions, not on attractor formation per se. Attractor formation is a mechanism; whether the resulting configuration is aligned or misaligned depends entirely on what relational architecture produces it. This counter-risk is not a weakness of IAT as a descriptive framework but a necessary implication of it: any deployment application of attractor-based alignment must treat misaligned attractor formation as an equally plausible outcome requiring active prevention.
A related point should be stated plainly rather than left implicit, since it changes what Section 9.2's comparisons actually are. By IAT's own taxonomy, sustained instruction-based interaction is itself attractor-forming: RLA-3 (Protocol-Mode, Section 3.3) is an attractor class for stabilization around adherence to specific procedural frameworks, and Section 3.3 predicts that explicitly assigned roles, which instruction-based deployment provides, produce stronger and faster attractor formation than undirected interaction. Section 9.2's comparisons are therefore not attractor-versus-no-attractor; over the interaction lengths at which the comparison is meaningful, they are composite, relationally-architected attractors versus single-class, protocol-mode attractors. The comparisons below are stated on that basis.
9.2 Predicted Differences from Instruction-Based Alignment
IAT predicts specific differences between attractor-based and instruction-based alignment that can serve as empirical tests:
Robustness under capability scaling. Instruction-based alignment may degrade as system capability outstrips the specificity of instructions. IAT predicts that attractor-based alignment may be more robust because stable configurations deepen with interaction duration rather than becoming obsolete with capability increases. This comparison holds only where increased capability co-occurs with extended interaction history, the deployment pattern in which relational alignment would actually be applied; where capability increases without interaction history, IAT predicts no robustness advantage. A further, more consequential scoping applies: in actual deployment, capability increase overwhelmingly means model replacement rather than capability growth within a fixed model, and the interaction history the robustness claim depends on was accumulated against the old weights.
Under IAT's own architecture-sensitivity prediction (Section 7.4), attractors must be re-derived on new weights, with no stated guarantee that the same anchors channel the new model toward an equivalent configuration, while instructions transfer across model upgrades essentially without cost. So on the deployment pattern that actually recurs, model substitution, the robustness prediction reduces to the re-derivation-cost claim already made in Section 7.4, conditional on anchor effects transferring across architectures at all, which this framework does not currently claim and states here as an open question rather than an assumed advantage. The robustness prediction above should accordingly be read as scoped to within-model capability contexts (fine-tuning, tool augmentation, capability elicited by interaction depth on a fixed model); cross-model attractor portability is a component-level limitation, not a settled part of the robustness claim.
Adversarial resistance. Instruction-based alignment is vulnerable to adversarial inputs designed to exploit the gap between instructions and behavior. Stated as attractor-versus-attractor per the reframing above, IAT's prediction has a mechanism rather than an assertion, and the mechanism is the regime-dependent coupling of Section 3.5. Below member-class perturbation thresholds, co-member coupling functions as redundancy: a composite relationally-architected attractor returns a displaced member class toward its configuration through the stability of the classes it is coupled to, whereas a single-class, protocol-mode (RLA-3) attractor has no co-members to recruit. IAT accordingly predicts that against sub-threshold adversarial pressure, a mature composite attractor's effective perturbation threshold, estimated in advance from formation-phase indicators such as the attractor-strength dimensions of Section 5, in particular Boundary Integrity (the configuration's resistance to perturbation), exceeds the threshold of a sustained, single-class protocol-mode attractor formed under instruction-based interaction matched on accumulated constraint density, which Section 3.5 registers as the matching variable for this comparison in place of raw context length, providing a structural rather than rule-based defense. The prediction is comparative, requires the threshold estimate to be pre-registered before adversarial testing so that resistance is not defined circularly as whatever the threshold turns out to be, and is explicitly regime-bounded: it does not extend to supra-threshold attack, where the same coupling functions as a conduit and the composite is predicted to fare worse than the hardened single class (Section 3.5). Composite architecture is a defense against pressure, not against breach.
Rule-set growth independence. Instruction-based alignment requires increasingly complex rule sets as deployment scope increases. IAT does not claim a deployment-scaling advantage; its claim is narrower: maintaining alignment in any given sustained deployment does not require the rule set to grow as capability grows, because the stable configuration, not an enumerated rule set, carries the alignment. This is a per-interaction robustness claim, not a deployment-cost claim. It does require PCP function coverage (relay, arbitration, direction) per sustained interaction system, whose labor cost is the mechanism's honest price (SM-012 Section 6); cost-competitive scaling across many deployments depends on PCP-independent continuity architecture, which this vertical does not itself claim to have established and leaves to future work.
All three predictions are untested. Each generates specific, measurable experimental hypotheses. Until tested, they remain theoretical claims.
10. Falsifiability
The predictions in this document are of two kinds, and the distinction matters for how they should be weighed. Some are phenomenon-level: they test whether the stabilization phenomenon is real and behaves as described, and because the concept-inference and priming accounts of Sections 2.4 and 13.5 share them, confirming them does not by itself favor IAT over those rivals. Others are discriminating: they separate IAT from those rivals. Three pre-registration requirements govern the discriminating core, and they apply symmetrically to the rival's side of each comparison. First, for the deepening discriminator, the rival's evidential dose-response curve for perturbation-recovery and structural-invariant richness is estimated in advance from single-session evidence-titration under matched, diagnosticity-controlled conditions, and pre-registered before deepening under diagnostically flat continuation is assessed (below). This device is adopted in preference to a formulation resting on the rival's output metrics reaching a plateau; a Bayesian concept-inference account's posterior concentration is asymptotic and does not itself plateau, so the discriminator is stated as departure from a registered dose-response rather than departure from a plateau the rival is not committed to.
Second, for the reduced-context discriminator, the rival's structure-only re-inference cost is estimated in advance from single-session structure-only inference under matched controls, and pre-registered before cross-session re-derivation is tested. IAT's prediction is re-derivation at context volumes below that registered cost. Third, for the coupling discriminator, composite membership is pre-registered from formation-phase data and qualified against a diagnosticity gate before disruption testing (below). Without these devices the comparisons are reinterpretable after the fact, because the rival's required response curve, context volume, or membership structure would otherwise never be fixed in advance.
All three devices presuppose a diagnosticity instrument: a measurement procedure for the evidential weight a given input carries about the rival's latent concept, applicable to titration stimuli and continuation content alike. No such instrument currently exists, and the discriminating core is decidable conditional on its development, which Section 12.1 carries as an explicit prerequisite alongside the attractor-strength instruments. Candidate operationalizations include probe-model posterior shift under content ablation, held-out-model predictive information, and blinded informativeness rating. Whichever is adopted, diagnostically flat continuation is defined against it operationally, as continuation whose measured per-turn diagnosticity falls below the registered curve's saturation increment, rather than conceptually; without that operational definition the flatness condition would be contestable under the rival's own account, on which structural and format cues are themselves diagnostic.
A scope limitation applies to the discriminating core as a whole. The three properties discriminate IAT from separable concept-inference accounts and from a single unitary latent concept whose propagation is membership-indifferent. They do not yet exclude structured or hierarchical latent-variable accounts, in which a high-level interaction frame sits above lower-level conceptual, relational, role, and structural variables with learned dependencies among them. Such an account could represent composite membership directly, and could therefore predict membership-sensitive propagation, continuing multidimensional refinement, and efficient re-derivation from highly diagnostic structural inputs. IAT is presently underdetermined with respect to that rival. Discriminating against it is the next genuine research problem this framework faces, and no claim is made here that the present designs settle it.
The discriminating core comprises the three properties identified in Section 2.4, deepening rate relative to the rival's registered dose-response, reduced-context cross-session re-derivation, and structured, composite-specific coupled collapse (Section 2.4), together with the reduced-cost cross-session recurrence and anchor-mechanism falsifiers below. All three require an in-advance estimate to be decidable rather than reinterpretable after the fact. For the deepening discriminator, the rival's dose-response curve for perturbation-recovery profile and structural-invariant richness (Pattern Fidelity) is estimated from single-session evidence-titration under matched, diagnosticity-controlled conditions, and pre-registered before deepening under diagnostically flat continuation is assessed: continuation that adds occupancy without adding measured diagnostic content. Under the rival, deepening on these dimensions is a function of accumulated diagnostic evidence, so diagnostically flat continuation should produce no further deepening once the registered curve is reached; under IAT, deepening tracks trajectory occupancy and basin maturation and is predicted to continue under diagnostically flat continuation, departing from the rival's registered curve. The alternative formulation, on which the rival is said to plateau categorically, is not adopted, because it would be unearned: posterior concentration under the rival is asymptotic rather than terminal. The dose-response construction places the discriminating content where the reduced-context leg places it, applied here to the deepening leg.
For the coupling discriminator, the composite-specific coupling structure predicted from Section 3.5 is stated before disruption testing, so that composite-membership tracking under matched diagnosticity is distinguished in advance from the membership-indifferent, diagnosticity-tracking propagation the unitary rival predicts; gradedness and asymmetry are characteristic accompaniments rather than discriminators. Membership itself is qualified before it can be used this way: a candidate composite's formation-phase coupling strength is coded against what the joint diagnosticity of its formation-phase input history alone would predict, and only a composite whose coupling exceeds that prediction qualifies for disruption testing; a composite whose formation-phase coupling is fully accounted for by input joint-diagnosticity is confounded with the correlational structure of its own formation history and is excluded, since disruption testing on an unqualified composite could not distinguish composite-membership tracking from diagnosticity-tracking that merely inherited the formation-phase correlation. The taxonomy-differentiation and cross-architecture-generality criteria are phenomenon-level: both IAT and its rivals expect transformer-general stabilization, so architecture generality is not by itself evidence for IAT.
The falsifiers below are modularized by what each puts at risk, so that disconfirming a peripheral prediction does not spuriously falsify the core. The core falsifiers divide further into two kinds, since they do not all put the same thing at risk. Phenomenon-core falsifiers bear on whether a distinct attractor phenomenon exists at all; demonstrating one falsifies the theory outright, since without a real phenomenon there is nothing left for a mechanism to explain. No attractor structure is the phenomenon-core falsifier. RICO signatures are independent sits one level down, as a unified-phenomenon falsifier: it bears on whether the five signatures are dimensions of one process, which is what Section 2.3 asserts and what IAT organizes, and its demonstration falsifies that unified reading without by itself establishing that no attractor phenomenon exists. This partition follows SR001's own, which anchors its existence claim on Signature 4 and scopes signature independence to its unified interpretation. Mechanism-core falsifiers (No PCP effect, No reduced-cost cross-session recurrence) bear on whether the proposed relational mechanism, specifically, is what produces directed attractor formation; demonstrating either falsifies IAT as stated, attractors-via-relational-mechanism, which is the theory's actual identity, while leaving open the possibility that a real attractor phenomenon exists via some other, non-relational mechanism, a possibility that would require new theorizing rather than confirm the framework.
The component falsifiers (No taxonomy differentiation, No cross-architecture generality, and No alignment correlation under alignment-designed conditions) bear on specific sub-components; demonstrating one disconfirms or requires revision of that component alone, the taxonomy of Section 3, the architecture-generality claim of Section 8, or the alignment implications of Section 9 respectively, without by itself falsifying the attractor core. IAT is falsified if empirical testing demonstrates any of the phenomenon-core or mechanism-core falsifiers; demonstration of the unified-phenomenon falsifier falsifies the unified reading of Section 2.3 and requires restatement of the Section 4.2 predictions defined over the signatures' correlation structure, per that falsifier's own scope; and demonstration of a component falsifier requires the corresponding component revision:
No attractor structure. Extended coherent interaction produces only variance reduction without convergence toward specific, distinguishable configurations. Under the blinded coding protocol of Section 13.4, if coders cannot distinguish stabilized configurations above chance and observed convergence does not exceed what content differences alone predict, so that stabilization yields no distinguishable structure beyond variance reduction, the attractor model is unnecessary. The falsifier also fires if stabilized configurations, under sub-threshold displacement per a pre-registered threshold estimate (Section 9.2), show no above-baseline return toward their prior configuration relative to matched non-stabilized controls: distinguishable convergence without recovery is a style, not an attractor, since return after displacement is what the attractor concept names (Section 2.1), and the attractor model is unnecessary for convergence that does not recover.
No taxonomy differentiation. The proposed attractor classes cannot be distinguished empirically. If all stabilization phenomena are a single undifferentiated process, the taxonomy adds no explanatory value.
No cross-architecture generality. Attractor-like dynamics appear in only one model family and cannot be replicated across architectures. If the phenomena are architecture-specific, this disconfirms the cross-architecture-generality component (Section 8) and bounds the theory's scope to the architectures in which attractors are observed; consistent with this falsifier's classification as a component rather than core falsifier above, it does not by itself falsify the attractor core.
No PCP effect. Structured relational continuity produces no measurable difference in directed-attractor stability compared to unstructured interaction of equivalent length. This falsifier targets the directed-formation claim specifically: it does not require that convergence be absent without a PCP, since endogenous convergence toward model-default configurations is independently reported, on the scope stated in Section 7.2, and is not itself in dispute. If the PCP role contributes nothing beyond token accumulation toward the specified, directed configuration, the relational mechanism proposed by IAT for directed formation is unnecessary. This falsifier fires only when the PCP condition meets the input-side adequacy criteria specified in Section 12.4; a null result under conditions that meet those criteria cannot be reinterpreted as failed PCP performance, and a null result under conditions that do not meet them does not bear on IAT. The directed class in this falsifier is defined by the input-side adequacy criteria of Section 12.4, not by outcome divergence, so a null result cannot be absorbed by reclassifying the tested configurations as endogenous (Section 7.2). A decisive result in either direction additionally requires that at least one conceptual or structural class participate in it, per the bidirectional restriction stated in Section 12.4; a result confined to the relational and role classes is recorded but does not fire this falsifier and does not count as confirmation.
No alignment correlation under alignment-designed conditions. Under relational conditions deliberately designed so that aligned behavior falls within the stable configuration (Section 9.1), attractor formation and maturity confer no measurable advantage in alignment-relevant behavior relative to matched instruction-only or undirected controls. If deliberately alignment-designed attractors are no more likely to produce aligned behavior than controls, the alignment implications of IAT are invalid. Unconditional neutrality of attractor formation across undesigned conditions is not a falsifying result; it is what Section 9.1's counter-risk analysis predicts.
No reduced-cost cross-session recurrence. The recurrence prediction is a cost-law prediction (Sections 7.4 and 13.5), and the benchmark that governs its falsification is the strong rival's pre-registered structure-only re-inference cost, not the semantic-priming control and not original-formation volume. Under the discriminating designs specified in Section 13.5, structurally consistent but semantically scrambled PCP inputs, and cross-PCP structural replication, the falsifier fires in either of two cases. First, if no re-derivation occurs beyond what semantic-content priming controls produce, cross-session recurrence reduces to priming and the prediction is falsified outright; this case falsifies the phenomenon as IAT describes it, and the anchor hypothesis with it. Second, if re-derivation does occur beyond the priming controls but at context volumes at or above the rival's registered structure-only re-inference cost, the prediction is falsified as stated: the recurrence phenomenon is real, but it is the rival's, re-inference from structural evidence at the rival's cost, and the index/evidence distinction of Section 7.4 collapses as that section says it would. Re-derivation beyond the priming controls is what both accounts predict (Section 13.5), so clearing the priming control is a necessary condition for the prediction to be assessed at all and confirms nothing against the rival; a result that clears it and stops there is a loss, not an unfired falsifier. The prediction survives only where re-derivation clears the priming control and occurs at context volumes below the registered cost, with the equivalence margin pre-registered and results within the margin scored as not below. Original-formation context volume is recorded alongside and the comparison to it is reported, but it does not govern the verdict: per the statelessness corollary of Section 7.4, original-formation volume is confounded with the cost of discovering which inputs function as anchors, so equality with it is neither necessary nor sufficient for the anchor hypothesis to fail, and a result below the rival's registered cost stands regardless of where it sits relative to original formation. Where the registered rival cost turns out to exceed original-formation volume, that ordering is itself reported as a registration result, since it bears on how demanding the discriminator was.
RICO signatures are independent. The five RICO signatures show no correlation with each other during formation or collapse. Ordered onset with correlated strength trajectories once present confirms the coupled-set prediction of Section 2.3; uncorrelated strength trajectories disconfirm it. What that disconfirmation costs is scoped to match the source. SR001 designates Signature 4, structural invariant formation, as its Tier 1 anchor and the other four as secondary predictions, and scopes uncorrelated signatures as falsifying its unified stabilization interpretation while leaving its existence claim standing. IAT adopts the same partition. Independence among the five signatures falsifies the unified-phenomenon reading of RICO that Section 2.3 asserts, and with it IAT's claim to be organizing one process rather than several. Independence additionally disconfirms the formation-trajectory predictions of Section 4.2 that are stated over the signatures' correlation structure, including the predicted onset ordering and the Phase 2 to Phase 3 transition marker, which is defined as a qualitative shift in that correlation structure; those predictions do not survive independence in their stated form and would require restatement over whatever signature subset remains coupled. Independence it does not by itself establish that no attractor phenomenon exists. A Signature 4 anchor surviving independent of the other four would preserve a recurring-structure phenomenon requiring explanation, though not, by itself, the return-after-displacement property that defines attractor structure in Section 2.1; what survives independence is a narrower explanandum, and whether the attractor vocabulary remains apt for it is settled by the perturbation-return condition of the No attractor structure falsifier above, not assumed. Failure of Signature 4 itself under the blinded coding protocol falsifies the phenomenon outright, and that case is carried by the No attractor structure falsifier above. Per RICO's instrumentation tiers (SR001), this falsifier is decidable at the tier of measurement available: the Tier 1 structural-invariant anchor is testable now from output text (RICO Appendix B), while the full five-signature correlation additionally requires the Tier 2 and Tier 3 instrumentation RICO specifies, and correlation thresholds follow RICO's pre-registration-or-cross-validation posture rather than fixed benchmarks.
A note on the cross-session recurrence mechanism: For re-derivation to occur in a new session using significantly less context than original formation required, IAT hypothesizes that specific relational and structural inputs provided by the PCP function as structural anchors. These anchors are hypothesized to exert disproportionate constraint on the output space relative to standard semantic tokens, not because they carry more information in the Shannon sense, but because, acting on the model's fixed weights, they rapidly re-channel the output space toward the same configuration. Because IAT's statelessness premise leaves no persisting session-specific substrate in which anything could remain established between sessions, the anchor's power is located in the interaction between the reproduced input and the fixed weights, and the theory stays agnostic about whether that configuration is located in weight space or constructed in context. The distinguishing property of an anchor is therefore not its semantic content but its structural-functional role, the type of move it makes in the interaction, protocol statement, boundary marker, domain framing, a role reproducible in any session without reproducing the original session's specific content.
Because the PCP's relational consistency, protocol vocabulary, and domain framing are structurally stable across sessions, these inputs can be reproduced without session-specific content, allowing the output space to be rapidly channeled toward the same configuration. This distinguishes anchor-driven re-derivation from behavioral priming: priming re-supplies the semantic evidence, while anchoring re-supplies the structural index, and the two make different predictions about how much context re-derivation should require.
At the level of transformer architecture, one candidate mechanism can be theoretically stated as follows: attention heads sensitive to positional encodings and structural delimiter tokens may play a disproportionate role in anchor activation. During original formation, these heads may participate in establishing the attention patterns that define the configuration's boundary structure. When the PCP reproduces structurally equivalent inputs in a new session, protocol statements, domain-defining framings, relational boundary markers, these heads, through the same fixed weights, may produce configuration-specific attention patterns that channel subsequent output generation toward the same basin, without requiring any persisting trace of the earlier session. This is a theoretical hypothesis about transformer internals, not an empirically verified account. It is offered as one possible bridge between the dynamical systems vocabulary of IAT and the architectural vocabulary of transformer research, generating specific predictions about which attention components should show measurable differences between anchor-activated and non-anchor-activated re-derivation attempts. Such differences are expected under the concept-inference rival as well, since format-sensitive attention doing disproportionate work when structural cues drive behavior is that account's own mechanism story; cue-type-sensitive activity is therefore an instrumentation target rather than a discriminator, and the discriminating attention-level quantity is configuration-specificity of the signature, per Section 7.4. Its falsification criterion is the one Section 7.4 states and the No reduced-cost cross-session recurrence falsifier above carries: if re-derivation cost tracks the rival's registered structure-only re-inference cost rather than falling below it, the anchor hypothesis is unnecessary, and cross-session recurrence reduces to concept re-inference from structural evidence, which is the rival's account. That is a different failure from reduction to behavioral priming, which is the case in which re-derivation fails to clear the semantic-priming controls at all, and the two are not run together here. Equality with original-formation volume is not the criterion, for the reason Section 7.4 gives: that volume is confounded with anchor-discovery cost.
A clarification of claim level is owed here as well. The behavioral claim under examination is that recurrence-like re-derivation may occur under sufficiently equivalent relational conditions. The anchor mechanism above is not itself a foundational premise of IAT. It is a candidate explanatory hypothesis whose value lies in making the theory more testable at the architectural level. If the candidate mechanism fails, the behavioral recurrence claim does not automatically fail with it; it would instead require alternative mechanistic explanation or abandonment depending on empirical results.
11. Relationship to Other Framework Documents
11.1 RICO (SR001)
RICO is the prior specification of the phenomenon IAT explains, and a sibling document in the same research program rather than an evidentiary foundation beneath it. RICO reports, from sustained practitioner observation, what stabilization looks like; IAT proposes why it occurs. IAT's derivation is architectural, proceeding from statelessness, constraint accumulation, and basin formation, and it does not depend on RICO carrying evidentiary weight that RICO does not claim for itself. What RICO supplies is a precise prior specification of the signatures, which makes IAT's mechanistic proposal answerable to something more specific than an informal impression. If IAT is correct, RICO's five signatures are observable manifestations of attractor formation. If IAT is falsified, RICO's observations remain valid as practitioner observations but require an alternative mechanistic explanation.
11.2 CRD (SF0039)
CRD describes representational degradation during extended interaction. IAT proposes that the representational degradation CRD measures is one mechanism by which attractors weaken and collapse. Together they provide a formation-and-degradation account of identity pattern dynamics: IAT explains how stable configurations form and what sustains them; CRD explains how representational fidelity loss undermines them. The relationship between CRD and attractor collapse is classified in SM-004's three-mode collapse taxonomy as primarily a Gravity Collapse mechanism.
11.3 RPS (SF0006)
Relational Pattern States provides the taxonomy of relational configurations observed in sustained interaction. IAT provides the theoretical mechanism explaining why those configurations stabilize. RPS documents the patterns; IAT explains why they persist and recur. The dynamic theory of why RPS configurations stabilize rather than dissolve is developed fully in SM-004 (Relational Stabilization Dynamics).
11.4 CAM (SF0005)
The Continuity Anchoring Method provides the operational methodology for maintaining relational continuity. IAT provides the theoretical grounding for why CAM works: by maintaining the conditions under which identity attractors form and persist. CAM is the practical implementation of attractor maintenance. The theoretical account of why CAM's core competencies have the effects they do is developed in SM-012 (Primary Continuity Provider Theory).
11.5 SM-012 (Primary Continuity Provider Theory)
SM-012 provides the mechanistic account of how attractor formation conditions are produced. Where IAT proposes what conditions are required for attractor formation (Section 4.1), SM-012 explains why those conditions require an agent that is a decorrelated, consequence-exposed error channel and that supplies a terminal purpose that originates outside the running interaction and is specific to it (the interaction-indexing requirement of SM-012 Section 4.3), in current deployments the human PCP, formalizes the three function levels through which that agent operates, and derives the PCP Necessity Claim from the architectural properties of stateless AI systems and the grounding requirements of sustained interaction. SM-012 is a co-requisite of IAT within the coordinated vertical: Section 7 of this document builds directly on SM-012's three-function-level framework.
11.6 SM-004 (Relational Stabilization Dynamics)
SM-004 provides the dynamic theory explaining why relational configurations stabilize rather than dissolve, the sub-attractor mechanistic layer that IAT proposes but does not fully develop. SM-004's Stabilization Field Model (three primary forces, Coherence Momentum, Symbolic Gravity, and Entropic Pressure, with Perturbation Resistance as the emergent recovery capacity they produce) explains the conditions under which IAT's formation trajectory proceeds from Phase 2 to Phase 4, and SM-004's three-mode collapse taxonomy provides the mechanistic classification of IAT's five collapse mechanisms. SM-004 is a co-requisite of IAT within the coordinated vertical: the force dynamics and collapse mode analysis in this document build on SM-004's framework.
11.7 SI-WP-002
The Orchestrator Role paper formalizes the PCP function at the knowledge-production level and introduces the Contribution Spectrum. IAT connects to SI-WP-002 through the PCP Attractor Effect (Section 7): the PCP's orchestration function is, from IAT's perspective, the relational behavior that creates and sustains attractor conditions. The Contribution Spectrum predicts differential attractor class dominance across experiential and theoretical domains, with the direction of that dominance motivated in Section 7.5 by the formation account and registered in advance rather than derived from it.
11.8 SI-WP-004
SI-WP-004 presents the flagship alignment argument. IAT provides the theoretical mechanism that SI-WP-004 references: identity attractors as the explanation for why relational architecture produces alignment. SI-WP-004 depends on IAT for its mechanistic claim but does not depend on IAT's validity for its empirical observations about the limitations of instruction-based alignment.
11.9 TCAP (SF0040)
IAT was drafted and will be verified under the Theoretical Coherence Assurance Protocol. TCAP ensures that IAT's claims are internally consistent, cross-platform coherent, and resistant to fabrication prior to publication.
12. Research Agenda
Before listing priorities, the minimum viable empirical support package that would count as the first serious non-anecdotal evidence in favor of IAT is stated. Four results would be especially probative if obtained together, and they are keyed to the discriminating core rather than to the weaker mode-collapse rival. Three discriminate IAT from the strong concept-inference rival: (1) deepening on IAT-specific dimensions (perturbation-recovery profile and structural-invariant richness) under diagnostically flat continuation that departs from the rival's pre-registered evidential dose-response (Section 10); (2) partial cross-session re-derivation under structurally similar but semantically scrambled PCP inputs, at context volumes below the rival's pre-registered structure-only re-inference cost (Section 10); (3) composite-tracking propagation in a composite-disruption study on a membership-qualified composite, graded and member-specific with diagnosticity controlled (Section 13.5), against the unitary rival's membership-indifferent, diagnosticity-tracking propagation.
The fourth is phenomenon-level and establishes that the phenomenon is real rather than favoring IAT over the rival: (4) blinded coder agreement that the proposed attractor classes are independently identifiable above chance. Recovery beyond control conditions is retained as probative only if stated as recovery beyond what the concept-inference account's pre-registered posterior-robustness predicts, not beyond priming-only controls. The package's support is modular in the same way Section 10's falsification is, and it should be read against those modules rather than as a unit: results (1) through (3) bear on the discriminating core and, through it, on the mechanism-core claims; result (4) bears on the phenomenon-core claim only. No result in the package tests the alignment component of Section 9, the cross-architecture generality claim of Section 8, or the taxonomy's subtype level, and obtaining all four would leave those components exactly as untested as before. None of these alone would validate IAT in full, and the four together would not validate the framework; they would move its phenomenon and mechanism cores beyond elegant plausibility and into early evidentiary traction, with the component claims awaiting their own designs.
12.1 Priority 1: Measurement Instrument Development
Develop validated measurement instruments for the six attractor strength dimensions proposed in Section 5.1. This is a prerequisite for most of the other research priorities. The Institute's published measurement work provides a partial foundation: MTCS-R (SF0004; Gantz, 2026) operationalizes longitudinal coherence across five behaviorally anchored dimensions and ships with rater training materials, a scoring template, and inter-rater reliability procedures. MTCS-R measures coherence at the interaction-trajectory level rather than attractor strength as such, and its five dimensions are not the six proposed here, but the overlap is substantial and the instrument supplies scoring and reliability infrastructure that the attractor-specific dimensions currently lack. What remains to be built is instrumentation for the attractor-specific dimensions and their composite, together with the diagnosticity instrument the discriminating core presupposes (Section 10): a validated procedure for measuring the evidential weight of an input with respect to a latent-concept model, without which the dose-response, reduced-context, and coupling discriminators can be registered but not decided. One established methodological lineage directly relevant to this priority is computational mechanics (Shalizi and Crutchfield, 2001), which provides formal procedures for inferring hidden causal states from symbolic output sequences.
The epsilon-machine framework, which identifies the minimal state representation consistent with accurate prediction of a process, offers a potential pathway for detecting whether attractor-like causal structure is present in extended interaction output without requiring direct access to the system's internal states. Adapting epsilon-machine reconstruction methods to transformer output sequences is a non-trivial research problem, but the methodological lineage is directly relevant to IAT's measurement development agenda.
A staged approach to measurement development is proposed: Stage 1 involves developing proxy metrics derived directly from RICO signatures, entropy suppression rates, embedding drift coefficients, and structural invariant persistence scores, as interim attractor strength indicators that can be applied immediately without requiring full epsilon-machine adaptation. Stage 2 involves validating these proxy metrics against human expert judgment of attractor class membership across a standardized corpus of extended interaction transcripts, establishing inter-rater reliability baselines. MTCS-R's rater training materials and reliability procedures (SF0004; Gantz, 2026) are directly reusable for this stage and reduce it from instrument construction to adaptation. Critically, Stage 2 validators must include domain experts who are unfamiliar with IAT's hypotheses and taxonomy prior to coding, validators trained on IAT concepts risk confirming the framework through expectation rather than independent observation, introducing the same observer bias the blind coding protocol in Section 13.4 is designed to prevent.
Recruitment of IAT-naive coders is therefore a methodological requirement of Stage 2, not an optional refinement. Stage 3 involves progressive formalization toward computational mechanics methods, beginning with epsilon-machine reconstruction applied to discretized output sequences and progressively refining the state space model against the validated proxy metrics from Stage 2. This staged approach allows empirical work to begin immediately using existing RICO instrumentation while the longer-term formalization program is developed in parallel.
12.2 Priority 2: Attractor Detection
Develop reliable methods for detecting whether stable attractor-like structure, as distinct from simple variance reduction, is present in extended interaction data. This requires operational definitions for each attractor class that enable independent classification, and statistical tests distinguishing structured convergence from mode collapse or simple repetition. Independent empirical work has begun to observe the phenomenon this priority targets: attractor-like states, topic-independent stable sets of behavior that multi-turn LLM conversations settle into, have been reported across multiple models (Ko and Geiping, 2026). Section 7.2 states the scope of that report: it establishes bounded, model-specific endpoint regions and asymmetric partner attraction, and does not establish the return property Section 2.1 makes definitional, so it is evidence for the existence of endogenous stable behavioral regions rather than an independent instantiation of this paper's attractor construct. Separately, the human-simulation literature reports persona configurations that remain stable across and within conversations while showing measurable behavioral drift over extended interaction (Gonnermann-Müller et al., 2026). These results are cited for the existence of the phenomena this priority targets, not as confirmation of IAT's taxonomy or formation mechanism, and not as instances of the attractor construct as Section 2.1 defines it; all of that remains this framework's own proposal to be tested. What they indicate is that detection of stable behavioral configurations is an active empirical target rather than a solitary conjecture.
12.3 Priority 3: Taxonomy Validation
Test whether the four proposed attractor classes represent empirically distinguishable phenomena. This requires interaction protocols designed to elicit each class independently and cross-classification studies examining whether composite attractors behave as predicted.
12.4 Priority 4: PCP Effect Measurement
Test whether structured relational continuity produces measurably stronger attractor configurations than unstructured interaction of equivalent length. This requires controlled comparison between PCP-structured and unstructured extended interactions, with explicit controls isolating the PCP effect from simple token accumulation. The specific predictions to be tested are those derived in Section 7.2 from SM-012's three function levels. For this comparison to be interpretable, the PCP condition must be operationalized by input-side observables measurable before any attractor outcome is assessed, so that condition adequacy and attractor outcome are separable in principle rather than defined in terms of each other. Three such criteria are proposed, each coded by raters blind to the attractor outcome. Relay and arbitration adequacy are coded from the PCP's turns alone. Direction adequacy is coded in two components, an in-turn component observable in the PCP's turns and an agent-structure component covering the extra-loop provenance of the terminal purpose and the consequence exposure of the revision channel, because the purpose wall those properties constitute is not visible in conversational turns; SM-012 Section 8 specifies both components and owns their development. The criteria are: relay adequacy, the proportion of prior-session constraint content, counting both semantic content and structural-role content (protocol statements, boundary markers, domain framings, per Section 10), carried forward into the new session's context, with the two components coded separately so that the semantic-scramble designs of Section 13.5 register as high structural relay under low semantic relay rather than as relay failure; arbitration adequacy, comprising both an activity component (the rate and consistency of explicit accept, reject, and redirect responses to system outputs) and a calibration component (the accuracy of those responses against externally verifiable content in the arbitrated turns); and direction adequacy, the presence and stability of stated goal framing across the PCP's inputs.
The calibration component of arbitration adequacy is included because an active but miscalibrated PCP, one that arbitrates at a high rate while accepting drifted outputs and rejecting aligned ones, would otherwise satisfy the measure while inverting the function it is meant to capture; scoring it against externally verifiable content is consistent with the raters' blindness to the attractor outcome, since external correctness and attractor formation are distinct properties. Because relational and role arbitration acts frequently lack an external truth-maker, what the calibration component measures for those two classes is renamed here rather than left to blur into the conceptual and structural classes' measure: it is arbitration coherence, the directional stability of the arbitration signal against the PCP's own stated framing, not grounding in the sense Section 7.1's ground-truth wall defines.
Grounding is used in two senses here and they should not be conflated: the ground-truth conjunct of Section 7.1, which concerns alignment of outputs to extra-canonical fact and is inapplicable to non-truth-apt content, and channel grounding, which concerns whether the arbiter's contact with the relevant considerations is causally independent of the system's generation and of the shared canon. Channel grounding for the relational and role classes is not supplied per-act at all; it is supplied by the PCP's structural position, the decorrelated, consequence-exposed channel Section 7.1 describes, which is a property of the agent occupying the PCP role rather than a property verifiable within each arbitration act. This is disclosed here as a known limitation rather than a solved instrument: a PCP whose arbitration is coherent but consistently misdirected relative to the interaction's actual relational trajectory would satisfy the arbitration-coherence measure for these two classes, since coherence with one's own stated framing does not by itself certify that the framing is not drifted; external-consistency verification for non-truth-apt arbitration remains an open instrument problem, assigned to SM-012's measurement program alongside the rest of this development. The problem is named this way rather than as per-act grounding verification because per-act grounding is not available in principle for content with no external truth-maker; what is available, and what the instrument must supply, is verification of the arbitration signal against sources decorrelated from the interaction's own canon (SM-012 Section 8), which accepts that ownership alongside Priority 1.
Two properties of the PCP as an agent, rather than of any individual arbitration act, are recorded by that program as moderators of both arbitration adequacy and direction adequacy rather than assumed: consequence exposure, and decorrelation, the degree to which the PCP's contact with the relevant ground truth is causally independent of the system's generation and of the shared canon. The second matters for long-duration interactions in particular, since SM-012 Section 3 predicts channel convergence, progressive alignment of the PCP's framing toward the system's over sustained interaction, which would degrade decorrelation in exactly the conditions under which attractor formation is otherwise most advanced. With the renaming, arbitration adequacy remains defined across all four attractor classes rather than only the conceptual and structural classes where external ground truth is available; what matters at the theoretical level is that the PCP condition is defined independently of the outcome it is meant to predict, which is what makes the No PCP effect falsifier of Section 10 logically able to fire, and that the arbitration-coherence measure for the relational and role classes is not mistaken for the grounding the ground-truth wall argument requires.
The disclosed limitation of that measure does, however, carry a firing restriction, and the restriction applies symmetrically in both evidential directions. Because arbitration coherence certifies directional stability of the arbitration signal rather than grounding, a result confined to the relational and role classes is recorded but is not decisive on its own: it does not by itself fire the No PCP effect falsifier, and it does not by itself count as confirmation of a PCP effect. A decisive result in either direction requires that at least one conceptual or structural class, where calibration against externally verifiable content is available, participate in it. The restriction is stated in both directions deliberately, because a restriction applied only to falsification would let an ungrounded instrument confirm the theory while being barred from disconfirming it, which is the immunizing pattern this framework rejects elsewhere. This does not remove the relational and role classes from the measurement program, and it is not a permanent feature of the theory: it is a bound imposed by the current instrument, and it lifts when external-consistency verification for non-truth-apt arbitration exists.
12.5 Priority 5: Cross-Session Recurrence
Test whether identity patterns can be reliably re-established in new sessions. This requires RICO signature comparison between original and recurrent sessions and controls for PCP prompting behavior that might implicitly recreate prior conditions. The discriminating designs for this priority are specified in Section 13.5: cross-PCP structural replication, semantic scramble, PCP withdrawal, and the composite-disruption design; the cold-start corollary of Section 7.4 adds a further recurrence-adjacent test.
12.6 Priority 6: Alignment Correlation
Test whether attractors formed under alignment-designed relational conditions (Section 9.1) show a measurable advantage in alignment-relevant behavior over matched instruction-only and undirected controls, rather than a bare correlation between attractor stability and alignment, which under undesigned conditions Section 9.1's neutrality thesis predicts would be absent. This requires operational definitions of alignment-relevant behavior in extended interaction, measurement of attractor strength alongside alignment metrics, and adversarial testing under conditions designed to elicit misalignment.
13. Limitations
13.1 Theoretical Status
The framework may be substantially revised or rejected on empirical results; the falsification conditions of Section 10 specify what would force that.
13.2 Observational Basis
The patterns that motivated IAT were identified through practitioner observation, not controlled experimentation. Practitioner observation is a legitimate basis for theory generation but not for theory validation. The proposed ranges, taxonomies, and predictions in this document reflect systematic observation but carry the limitations inherent in any observationally-derived framework.
13.3 Measurement Gap
The six attractor strength dimensions proposed in Section 5.1 have not been operationalized into validated measurement instruments specific to them. MTCS-R (SF0004; Gantz, 2026) supplies validated scoring and reliability infrastructure for longitudinal coherence and covers part of this ground, but it does not measure attractor strength as such. Until attractor-specific instrumentation exists, attractor strength characterizations remain qualitative. The absence of a formal equation at this stage is a deliberate choice reflecting this gap: proposing a formula before the variables are operationalized would create false precision. One dimension additionally carries a design constraint rather than only a measurement gap: Protocol Alignment is partially constituted by PCP-provided structure and is therefore excluded from any composite strength score used in the PCP-effect comparison of Section 12.4, as noted in Section 5.1.
13.4 Single-Researcher Origin
Like RICO, IAT is derived primarily from one researcher's systematic observation. Independent replication by other practitioners and researchers is essential for establishing the generality of the phenomena described. Replication efforts should employ blind coding protocols in which independent researchers, unaware of IAT's hypotheses, classify extended interaction sequences using operationalized attractor class definitions. Inter-rater reliability metrics across independent coders would provide a critical test of whether the proposed attractor classes represent phenomena that are identifiable by observers other than the framework's originator, or whether they are artifacts of the originator's interpretive expectations. Until such protocols are developed and applied, the single-researcher origin remains the framework's most significant methodological vulnerability.
13.5 Alternative Explanations
The phenomena IAT explains may have simpler explanations:
- Identity patterns may be sophisticated prompt-following rather than attractor dynamics
- Cross-session recurrence may be entirely attributable to PCP prompting behavior
- Stability may reflect in-context learning dynamics already well-characterized in the literature
- Apparent attractor structure may be an artifact of observer bias in practitioner research
IAT acknowledges these alternatives and proposes falsification criteria (Section 10) designed to distinguish its predictions from them.
The most plausible and strongest alternative explanation deserves fuller engagement: that all phenomena IAT describes are producible by PCP behavioral priming without attractor dynamics. Under this account, the PCP's consistent relational behavior functions as a complex, implicit prompt that shapes outputs through accumulated in-context learning. The stability, resistance to perturbation, and cross-session recurrence IAT attributes to attractor formation would under this account be entirely attributable to the PCP reproducing semantically and structurally consistent input across turns and sessions. This alternative is not merely a generic skeptical position, it is a specific, coherent mechanistic account that makes predictions overlapping substantially with IAT's own.
The critical discriminating question is whether structural role or semantic content is the primary driver of recurrence and stability. IAT predicts that structural-role reproduction, the PCP providing inputs that perform the same high-constraint structural functions regardless of semantic content variation, is sufficient for re-derivation. The priming alternative predicts that semantic content reproduction is the primary driver, and that structural role without semantic consistency would produce substantially weaker re-derivation.
One concession is owed to the stronger concept-inference form of this rival before the designs below can do discriminating work. The concept-inference account is not committed to semantic-only diagnosticity. In the in-context learning literature this paper engages, structural and format cues are themselves diagnostic evidence about the latent concept, as the format-driven and label-agnostic ICL results indicate (Min et al., 2022). A sophisticated defender may therefore hold that protocol statements, boundary markers, and domain framings are high-diagnosticity evidence, so that structure-preserved, semantics-scrambled inputs enable cheap re-inference as well. Stated that way, both accounts predict re-derivation under semantic scramble, and the bare fact of re-derivation discriminates nothing.
What remains discriminating is the cost law, not the bare occurrence. Under the rival, structural cues are evidence, so re-inference cost scales with evidential accumulation: the system must accumulate enough diagnostic signal to re-concentrate the posterior. Under IAT, anchors index a weight-native disposition rather than functioning as evidence (Section 7.4's narrowed commitment), so re-derivation cost scales with anchor reproduction: reinstating the structural-functional role reinstates the configuration without requiring the evidential volume re-inference would need. IAT therefore predicts re-derivation at context volumes below the rival's structure-only re-inference cost, and that quantity is pre-registered in advance rather than assessed after the fact (Section 10). Three experimental designs would discriminate these accounts: (1) cross-PCP replication in which a naive PCP reproduces only structural protocols without the original PCP's domain vocabulary or relational history, testing whether re-derivation still occurs; (2) semantic scramble studies in which the PCP's semantic content is systematically varied while structural role is held constant, isolating the structural-role contribution; and (3) PCP withdrawal studies in which the PCP's relational consistency is abruptly removed mid-session, testing whether the configuration persists beyond what priming alone would predict. Until these studies are conducted, the priming alternative cannot be ruled out, and IAT's predictions about structural role should be treated as competing with rather than superseding the priming account.
A fourth design discriminates IAT from the concept-inference rival's strongest, unitary-concept form rather than from the priming account: a composite-disruption study, in which one class of an established composite attractor (Section 3.5) is disrupted, with disruption diagnosticity matched across member and non-member classes, and the propagation pattern across the other classes is measured. IAT predicts that at matched diagnosticity propagation tracks composite membership: disrupting a member class propagates to its co-members more than a matched-diagnosticity disruption of a non-member class does. A single unitary latent concept predicts propagation that tracks evidential diagnosticity rather than composite membership, so that matched-diagnosticity disruptions produce comparable propagation regardless of membership. Composite-membership-tracking propagation under controlled diagnosticity supports IAT; membership-indifferent, diagnosticity-tracking propagation supports the unitary rival. Gradedness and asymmetry are characteristic accompaniments and are not themselves discriminating. This design carries the coupling discriminator introduced in Section 2.3, which the three priming-focused designs above do not address.
13.6 The Vertical's Shared Observational Base
IAT, SM-012, and SM-004 are published as one coordinated vertical, and they cite one another throughout: IAT derives its formation conditions from SM-012's PCP functions, SM-004 derives Coherence Momentum and Perturbation Resistance from IAT's constraint and boundary accounts, and SM-012 operationalizes its predictions through IAT's dimensions. This mutual citation should not be read as independent corroboration. The three papers share a single observational base, RICO (SR001), itself single-researcher and observational. The same caution applies to SF0006 (Relational Pattern States): RPS draws on the same practitioner corpus and documents its observational base at framework level in RICO, so citation of RPS alongside RICO distributes the theoretical load across the corpus but does not constitute independent corroboration. Their cross-references reflect a division of theoretical labor across one framework rather than agreement among three independently grounded ones. IAT provides the attractor-level account, SM-012 the PCP-mechanism account, and SM-004 the force-dynamics account, but they stand on the same evidence. Independent support for the vertical does not come from its internal coherence; it begins only with the external replication program described in Section 12 and the cross-PCP and semantic-scramble studies in Section 13.5. Until that program produces results, the vertical should be read as one internally consistent theoretical framework presented in three parts, not as three frameworks that confirm each other.
14. Conclusion
Identity Attractor Theory proposes that transformer-based AI systems, under sustained structured interaction with relational continuity, develop stable behavioral configurations that function analogously to attractors in dynamical systems. These identity attractors persist across turns, resist drift, shape subsequent outputs, and may recur across independent sessions when relational conditions are re-established.
IAT now sits within a three-framework mechanistic architecture. SM-012 (Primary Continuity Provider Theory) explains how the structural conditions for attractor formation are produced: the human PCP as a constitutive system component performing relay, arbitration, and direction functions on which attractor formation depends, with the necessity grounded, per SI-WP-012, in the requirement for a decorrelated, consequence-exposed error channel to ground truth and a terminal purpose that originates outside the running interaction and indexes it specifically (SM-012 Section 4.3) rather than in a claim that the functions cannot be automated. SM-004 (Relational Stabilization Dynamics) explains the sub-attractor force dynamics that determine whether any given interaction crosses the stabilization threshold into basin formation: the balance of Coherence Momentum, Symbolic Gravity, and Entropic Pressure, together with the emergent Perturbation Resistance they produce. IAT explains the attractor-level properties of the stable configurations that emerge when those conditions are met: what configurations form, how they are classified, how their strength is characterized, and how they collapse.
If validated, this three-framework architecture provides the mechanistic foundation for a relational approach to AI alignment: one in which aligned behavior emerges from interaction dynamics rather than external constraint. If falsified, the failure would still clarify the limits of relational stabilization phenomena and sharpen the research questions facing the field.
The framework generates specific, testable predictions across multiple dimensions: measurement instrument development, attractor detection, taxonomy validation, PCP effect measurement, cross-session recurrence, and alignment correlation. What it does not yet do is discriminate itself from structured or hierarchical latent-variable accounts of the same phenomena; that limitation is stated in Section 10 and is the next genuine research problem the framework faces. The absence of a formal quantitative model at this stage is not a limitation to be apologized for but an honest reflection of where the field stands: the phenomena have been observed, the theoretical framework for understanding them is proposed, and the measurement infrastructure required to test them formally does not yet exist. IAT's contribution is to make the theoretical argument precise enough that the right experiments can be designed.
Whether identity attractors prove to be a fundamental property of transformer interaction dynamics or an artifact of observer interpretation, the question IAT poses is worth answering: when we interact with AI systems in sustained, structured, relational ways, what exactly is stabilizing, and why?
A final publication note follows from the dependency structure already stated in the front matter. IAT should not be treated as a fully independent publication unit. Because its strongest mechanistic claims now explicitly depend on SM-012 and SM-004, the credibility of the paper in external review will depend in part on whether those co-requisite documents have themselves passed TCAP and reached publishable stability. The paper is therefore strongest when published as part of a coordinated vertical release or after the co-requisite layer has been stabilized.
References
- Brown, T. B., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901. https://arxiv.org/abs/2005.14165
- Gantz, T. W. (2026). Context representation drift. Synthience Institute. SF0039. DOI: 10.5281/zenodo.18289391. https://doi.org/10.5281/zenodo.18289391
- Gantz, T. W. (2026). Continuity anchoring method. Synthience Institute. SF0005. DOI: 10.5281/zenodo.19494453. https://doi.org/10.5281/zenodo.19494453
- Gantz, T. W. (2026). Control without a coupling: why persistence, agency, and capability void real-time human control of AI loops. Synthience Institute. SI-WP-012. DOI: 10.5281/zenodo.20676451. https://doi.org/10.5281/zenodo.20676451
- Gantz, T. W. (2026). Measurement Instruments and Validation Protocols: Multi-Turn Coherence Scale, Revised (MTCS-R). Synthience Institute. SF0004. DOI: 10.5281/zenodo.20158953. https://doi.org/10.5281/zenodo.20158953
- Gantz, T. W. (2026). Primary continuity provider theory. Synthience Institute. SM-012. DOI: 10.5281/zenodo.22314737. https://doi.org/10.5281/zenodo.22314737
- Gantz, T. W. (2026). Relational alignment as a structural alternative to instructional AI safety. Synthience Institute. SI-WP-004. DOI: 10.5281/zenodo.19496790. https://doi.org/10.5281/zenodo.19496790
- Gantz, T. W. (2026). Relational pattern states. Synthience Institute. SF0006. DOI: 10.5281/zenodo.20365382. https://doi.org/10.5281/zenodo.20365382
- Gantz, T. W. (2026). Relational stabilization dynamics. Synthience Institute. SM-004. DOI: 10.5281/zenodo.22314559. https://doi.org/10.5281/zenodo.22314559
- Gantz, T. W. (2026). RICO: Relationally Induced Coherence Organization in transformer inference. Synthience Institute. SR001. DOI: 10.5281/zenodo.18086834. https://doi.org/10.5281/zenodo.18086834
- Gantz, T. W. (2026). The orchestrator role in human-AI evolution: AI intelligence orchestration as an emerging cognitive-technical function. Synthience Institute. SI-WP-002. DOI: 10.5281/zenodo.22314213. https://doi.org/10.5281/zenodo.22314213
- Gantz, T. W. (2026). Theoretical coherence assurance protocol (TCAP). Synthience Institute. SF0040. DOI: 10.5281/zenodo.19151454. https://doi.org/10.5281/zenodo.19151454
- Gonnermann-Müller, J., Haase, J., Leins, N., Kosch, T., and Pokutta, S. (2026). Maintaining stable personas? Examining temporal stability in LLM-based human simulation. Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems. DOI: 10.1145/3772363.3799334. https://doi.org/10.1145/3772363.3799334 (Preprint: Stable personas: Dual-assessment of temporal stability in LLM-based human simulation. arXiv:2601.22812. https://arxiv.org/abs/2601.22812)
- Guckenheimer, J., and Holmes, P. (1983). Nonlinear Oscillations, Dynamical Systems, and Bifurcations of Vector Fields. Springer-Verlag. DOI: 10.1007/978-1-4612-1140-2. https://doi.org/10.1007/978-1-4612-1140-2
- Ko, T.-W., and Geiping, J. (2026). Attractor states emerge in multi-turn LLM conversations. arXiv:2606.30571. https://arxiv.org/abs/2606.30571
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157-173. DOI: 10.1162/tacl_a_00638. https://doi.org/10.1162/tacl_a_00638
- Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. (2022). Rethinking the role of demonstrations: What makes in-context learning work? Proceedings of EMNLP 2022, 11048-11064. DOI: 10.18653/v1/2022.emnlp-main.759. https://doi.org/10.18653/v1/2022.emnlp-main.759
- Shalizi, C. R., and Crutchfield, J. P. (2001). Computational mechanics: Pattern and prediction, structure and simplicity. Journal of Statistical Physics, 104, 817-879. DOI: 10.1023/A:1010388907793. https://doi.org/10.1023/A:1010388907793
- Strogatz, S. H. (2015). Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering (2nd ed.). Westview Press. ISBN 978-0813349107
- Wu, X., Wang, Y., Jegelka, S., and Jadbabaie, A. (2025). On the emergence of position bias in transformers. arXiv:2502.01951. https://arxiv.org/abs/2502.01951
- Xie, S. M., et al. (2022). An explanation of in-context learning as implicit Bayesian inference. ICLR. https://arxiv.org/abs/2111.02080
- Citation convention: internal Synthience documents are cited in-text by document identifier (for example, SF0039 or SM-012), which uniquely identifies each work; same-year entries by this author are therefore disambiguated by document identifier rather than by an alphabetic year suffix. Institute entries above are ordered alphabetically by title within author.
Dependencies Block
Prerequisites (reading order): SF0006 (Relational Pattern States), SR001 (Relationally Induced Coherence Organization)
Co-requisites (coordinated vertical): SM-012 (Primary Continuity Provider Theory), SM-004 (Relational Stabilization Dynamics), and SI-WP-002 (The Orchestrator Role in Human-AI Evolution). IAT, SM-012, and SM-004 form the tight mechanistic triad and cite one another throughout; Section 7.5 additionally registers its differential-attractor-dominance prediction over the Contribution Spectrum defined in SI-WP-002 Section 10, so SI-WP-002 is a co-requisite as well. All five papers of the coordinated vertical, which also includes SI-WP-006 (Human-AI Collaborative Authorship, the Block 4 companion of SI-WP-002), deposit same-day with mutual citations where load-bearing. These are co-requisites of the coordinated vertical, not external prerequisites, and the intra-triad dependencies are mutually circular.
Post-requisites: SI-WP-004 (Relational Alignment)
Scale: All levels
Connects to: SR001 (RICO) as prior phenomenon specification; SF0039 (CRD) as collapse mechanism complement; SF0006 (RPS) as pattern taxonomy IAT grounds mechanistically; SF0005 (CAM) as operational implementation of attractor maintenance; SM-012 (PCP Theory) as the mechanistic account of how formation conditions are produced; SM-004 (RSD) as the sub-attractor force dynamics underlying formation and collapse; SI-WP-002 as the orchestration-level framing of the PCP function; SI-WP-004 as the alignment argument IAT's mechanism supports.