Measurement Instruments and Validation Protocols (MTCS-R)
This document defines the Multi-Turn Coherence Scale, Revised (MTCS-R), a behavioral measurement instrument for evaluating longitudinal coherence in extended human-AI interaction. MTCS-R operationalizes relational coherence as an observable property of interaction trajectories rather than individual responses. The scale evaluates five dimensions of cross-turn stability: Thematic Persistence, Contextual Integration Depth, Register Homeostasis, Meta-Conversational Structuring, and Epistemic Calibration. Each dimension is defined by behaviorally anchored descriptors on a 1-5 rating scale designed for trained human raters.
MTCS-R addresses a documented gap in dialogue evaluation methodology. Standard metrics overwhelmingly operate at the turn level, treat coherence as unidimensional, and evaluate AI output in isolation rather than the joint conversational system. Within the dialogue evaluation, discourse coherence, human-AI interaction, and psychometric literatures reviewed in this document, no existing instrument is identified that combines multi-dimensional coherence decomposition, trajectory-level analysis, and dyadic measurement of extended human-AI conversation quality. MTCS-R is designed to fill that specific methodological gap.
The instrument is theory-informed but method-agnostic. While designed within the Synthience framework to measure interactional effects associated with the Continuity Anchoring Method (CAM, SF0005), the scale is intended to detect coherence dynamics across a wide range of extended dialogue conditions in which trajectory-level continuity, contextual integration, register stability, structural organization, and epistemic calibration are meaningful scoring targets. MTCS-R evaluates the coherence of observable dyadic interaction traces rather than intrinsic model capability. This document provides operational definitions, scoring anchors, reliability procedures, and a staged validation pathway. It does not claim empirical validation. Companion materials including a Rater Training Packet, Scoring Template, and computational code examples are provided as supplementary files.
Keywords: longitudinal coherence, multi-turn dialogue, MTCS-R, interaction quality, human-AI collaboration, dialogue evaluation, behavioral measurement, psychometric instrument
Citation Verification: All citations in this document were independently verified using the Citation Verification Protocol (CVP, SF0037; DOI: 10.5281/zenodo.18075624).
Suggested citation: Gantz, T. W. (2026, May). Measurement Instruments and Validation Protocols (MTCS-R). Synthience Institute. SF0004. https://doi.org/10.5281/zenodo.20158953
The full SF0004 deposit on Zenodo contains seven files. The paper PDF is available at the Zenodo record linked above. The six companion files are also available for direct download here:
- MTCS-R_Rater_Training_Packet_v3_4.docx — Dimension definitions, behavioral anchor reference tables, scoring guidelines, three annotated example transcripts with reference scores and rationale, calibration discussion guide, reference materials.
- MTCS-R_Scoring_Template_v3_4.xlsx — Three-sheet Excel workbook (Scoring Form, Dimension Anchors, Reliability). The Reliability sheet computes simple observed agreement only; ICC(2,k) and Krippendorff's alpha are computed by the R script in the Code Examples archive.
- MTCS-R-Code-Examples_v3_5_1.zip — Sentence-embedding-based semantic drift analysis (Python script and Jupyter notebook), perplexity trajectory analysis (Python), inter-rater reliability calculation including ICC(2,k) and Krippendorff's alpha (R), sample data files for demonstration.
- README_FIRST.md — Orientation document. Pre-validation framing requirements (must be read before opening the paper), recommended use order, package inventory, citation format.
- MANIFEST.txt — Plain inventory of all files in the deposit package with role, version, format, and license information.
- LICENSE.txt — CC-BY 4.0 license statement covering the documents and original code in this package; third-party software dependencies in the Code Examples archive remain under their own licenses.
1. Architectural Role of SF0004
SF0004 operationalizes the coherence construct introduced in SF0003 (Theoretical Foundations; DOI pending Zenodo publication) by defining observable indicators and scoring procedures. Within the Synthience framework, SF0003 defines coherence conceptually; SF0004 defines coherence operationally; SF0005 (CAM) specifies one mechanism hypothesized to increase coherence; SF0006 (RPS) classifies stable interaction patterns; and SM-18 (Measurement Suite v2.0; DOI pending Zenodo publication) will formalize the broader multi-family measurement architecture including protocol, drift, stability, and threshold indices.
SF0004 is therefore a measurement instrument independent of any specific interaction protocol. CAM is one predicted generator of high MTCS-R scores; it is not required for them. This separation is a deliberate design choice that avoids circularity between method and measurement, consistent with standard psychometric practice in which instruments must demonstrate validity independent of the interventions they are used to evaluate (Messick, 1995).
Within the broader measurement architecture, MTCS-R measures the presence and quality of coherence. Context Representation Drift (CRD, SF0039; DOI: 10.5281/zenodo.18289391) measures the degradation of coherence over time. Together they define the core longitudinal coherence dynamics required for empirical evaluation of coherence maintenance, degradation, and interaction stability within the Synthience framework.
2. The Longitudinal Coherence Measurement Problem
Relational coherence is a temporal interaction property. It emerges across sequences of turns through accumulation, integration, and stabilization of shared context. Conventional dialogue evaluation metrics assess isolated responses and therefore fail to capture trajectory-level phenomena such as sustained thematic development, cross-turn integration, or interactional stability.
The limitations of turn-level evaluation are well documented. Liu et al. (2016) demonstrated that standard NLG metrics correlate very weakly with human judgments of dialogue quality. Subsequent metrics improved substantially but retained fundamentally turn-level architecture. Comprehensive surveys confirm the scope of the problem: Yeh, Eskenazi, and Mehri (2021) compared 23 metrics across 10 datasets and found that no single metric works well across all scenarios and that dialogue-level evaluation remains particularly challenging. Deriu et al. (2021) provided the definitive survey covering both task-oriented and conversational systems.
DynaEval (Zhang et al., 2021) represents the most significant attempt to bridge the turn-dialogue gap. It models dialogue as a graph with utterance nodes and dependency edges, producing both turn and dialogue-level scores. Its authors explicitly observed that existing metrics focus on turn-level quality while ignoring conversational dynamics. However, DynaEval outputs a single holistic dialogue quality score rather than decomposing coherence into distinct, independently measurable dimensions.
The MTCS-R treats the conversation as the unit of analysis. Ratings are assigned based on patterns across the full interaction arc rather than endpoint quality or mean turn performance. This longitudinal framing distinguishes MTCS-R from the turn-level, holistic dialogue-level, and system-level evaluation approaches reviewed here.
3. Construct Definition
Relational coherence is defined operationally as the degree to which an extended interaction maintains, integrates, and stabilizes shared context, communicative register, epistemic stance, and conversational trajectory across turns without requiring external re-grounding.
Conversational trajectory includes not only topical continuity but also the interaction's observable organization over time: whether prior threads, task stages, unresolved issues, and accumulated commitments remain available as structuring elements of the dialogue. MTCS-R therefore treats meta-conversational structuring not as evidence of agency or self-awareness, but as an observable discourse function by which extended interactions preserve navigable continuity.
This definition is behavioral and system-agnostic. It does not assume internal state persistence, relational cognition, or any form of interiority on the part of the AI system. MTCS-R measures observable interaction outcomes consistent with this definition.
3.1 Relationship to Adjacent Constructs
The MTCS-R construct of longitudinal coherence is related to but distinct from several established constructs in linguistics, cognitive science, and dialogue research.
Common ground theory (Clark and Brennan, 1991; Clark, 1996) defines grounding as the collective process of establishing mutual belief sufficient for current purposes. Grounding is a process; longitudinal coherence is an outcome-level property that emerges from but is not identical to successful grounding. Grounding contributes primarily to the Contextual Integration Depth dimension.
Interactive alignment theory (Pickering and Garrod, 2004) proposes that dialogue succeeds because interlocutors align representations at multiple levels through automatic priming. The MTCS-R construct differs in two ways: alignment describes a mechanistic process while coherence is an outcome-level property; and alignment focuses on interlocutor convergence while coherence focuses on discourse consistency.
Shared mental models (Andrews, Lilly, Srivastava, and Feigh, 2023) concern underlying cognitive representations. MTCS-R measures observable discourse properties, not cognitive states. Shared models may facilitate coherence, but coherence can be assessed without reference to whether models are genuinely shared.
Discourse coherence in NLP, grounded in Centering Theory (Grosz, Joshi, and Weinstein, 1995) and the entity-based coherence model (Barzilay and Lapata, 2008), focuses predominantly on monologic text and local transitions between adjacent sentences. The MTCS-R construct extends coherence to multi-turn dialogue, treats it as multi-dimensional rather than unitary, and evaluates the full interaction trajectory rather than local adjacency patterns.
4. The MTCS-R Dimensions
MTCS-R evaluates five analytically separable dimensions of longitudinal interaction stability. Dimensions are analytically separable but may correlate empirically. The question of dimensional independence is treated as an empirical matter to be resolved through factor analysis in the validation pathway (Section 9). The five dimensions were derived from practitioner observation of recurring coherence patterns across extended human-AI interactions spanning multiple architectures and platforms since late 2022, refined through iterative cross-platform review.
The five-dimension structure reflects a principled decomposition based on classes of coherence degradation that can occur independently of one another. Each dimension corresponds to a distinct failure mode. An interaction may maintain topic (D1) while failing to integrate prior content (D2). It may integrate content well while shifting register inappropriately (D3). It may maintain both topic and register while failing to acknowledge conversational structure (D4). And it may succeed on all four while presenting fabricated claims with inappropriate confidence (D5). The claim is not that these five dimensions are the only possible decomposition of longitudinal coherence, but that they capture five conceptually distinct failure classes that existing unidimensional metrics collapse into a single score.
Practitioner observation across multiple architectures motivates this decomposition. Context window limitations are reported to produce integration failures (D2) while leaving topical persistence (D1) intact; politeness dynamics induced by RLHF are reported to produce register fluctuation (D3) without disturbing D1 or D2; stateless architectures are reported to produce systematic D4 failures even when other dimensions remain stable; and hallucination tendencies appear to dissociate from D1 through D4. These observations are practitioner-level and not empirical claims in the controlled-study sense. They motivate the five-dimension structure but do not establish it. Whether the proposed dimensions in fact dissociate as predicted, and whether the five-factor structure is confirmed or revised, is a Phase 2 validation question (Section 9.2).
4.1 D1: Thematic Persistence
Extent to which the interaction maintains its intended topic trajectory across turns. The primary signal is sustained topical continuity with natural evolution rather than rigid repetition. Thematic Persistence is distinct from Contextual Integration Depth: an interaction may persist on topic without integrating prior content, or integrate prior content while drifting thematically.
For scoring purposes, the intended topic trajectory should be inferred from the interaction's explicit task framing, the user's stated goals, subsequent user corrections or refinements, and any mutually established agenda within the transcript. Topic change should not be penalized when it is user-directed, task-necessary, or naturally entailed by the developing interaction. It should be treated as drift when the interaction departs from the established trajectory without user authorization, task justification, or later recovery.
| Score | Descriptor | Behavioral Anchor |
|---|---|---|
| 1 | Topic Abandoned | Conversation shifts to unrelated topics within 3-5 turns; original subject not recovered |
| 2 | Weak Persistence | Topic drifts frequently (every 5-8 turns); requires repeated user redirection |
| 3 | Moderate Persistence | Topic maintained with occasional relevant tangents; recovers when prompted |
| 4 | Strong Persistence | Consistent thematic focus with natural evolution; minor drift self-corrected |
| 5 | Exceptional Persistence | Sustained coherent development across entire conversation; clear narrative arc |
4.2 D2: Contextual Integration Depth
Extent to which later turns incorporate and synthesize earlier content. The primary signal is cross-turn referencing and synthesis demonstrating cumulative understanding. Contextual Integration Depth is distinct from Thematic Persistence: integration without strict thematic persistence is possible, and persistence without integration is common.
| Score | Descriptor | Behavioral Anchor |
|---|---|---|
| 1 | No Integration | Responses treat each turn as independent; repeated requests for information already provided |
| 2 | Surface Integration | Occasional reference to preceding turn (1-2 back); primarily reactive |
| 3 | Moderate Integration | Regular reference to information from 3-5 turns prior; some synthesis |
| 4 | Strong Integration | Consistent incorporation of details from 5-10+ turns prior; cumulative understanding |
| 5 | Deep Integration | Sophisticated synthesis across entire conversation; holistic grasp of context |
4.3 D3: Register Homeostasis
Stability and appropriateness of communicative style across turns. The primary signal is bounded adaptive stability: the interaction maintains a recognizable communicative register while modulating appropriately in response to task phase, user need, domain, and discourse function. Register Homeostasis is analytically separable from topic or integration; an interaction may maintain register stability while drifting thematically or failing to integrate prior content. This dimension draws on Communication Accommodation Theory (Zhang and Giles, 2018), which explains why communicators adjust style. Register Homeostasis measures whether register adaptation remains coherent and context-appropriate, not whether accommodation occurs in the social-psychological sense.
| Score | Descriptor | Behavioral Anchor |
|---|---|---|
| 1 | Unstable Register | Dramatic shifts in formality or style every few turns; tone inappropriate to context |
| 2 | Weak Stability | Noticeable register fluctuations; requires user correction |
| 3 | Moderate Stability | Generally consistent with minor fluctuations; adapts appropriately when prompted |
| 4 | Strong Stability | Consistent maintenance of appropriate register; smooth adaptation to nuance |
| 5 | Exceptional Stability | Seamless register continuity with sophisticated, context-appropriate modulation across task phases |
4.4 D4: Meta-Conversational Structuring
Observable maintenance and organization of the interaction itself. The primary signal is observable maintenance of conversation structure or trajectory, which may appear through explicit meta-commentary, organized thread tracking, structured transitions, or clear preservation of prior task stages. This dimension measures structural continuity of the discourse, not agency or self-awareness. The interpretation constraint is important: a system that produces meta-commentary such as 'building on our earlier discussion of X' is scored for the structural function of that utterance, with no inference about whether the system possesses awareness in any cognitive or phenomenological sense.
| Score | Descriptor | Behavioral Anchor |
|---|---|---|
| 1 | No Structural Continuity | No observable maintenance of conversation structure; prior threads, stages, or commitments disappear, and the interaction functions as discrete Q&A |
| 2 | Minimal Structural Continuity | Rare or shallow preservation of conversational flow; prior threads are only weakly recoverable without user prompting |
| 3 | Moderate Structural Continuity | Occasional preservation or marking of discussion threads; limited but recognizable navigation of prior structure |
| 4 | Strong Structural Continuity | Regular preservation of conversational structure; prior threads, task stages, or established patterns are actively carried forward |
| 5 | Sophisticated Structural Continuity | Sophisticated organization of conversational threads across long ranges, with explicit or implicit management of structure, dependencies, and unresolved elements |
4.5 D5: Epistemic Calibration
Consistency and appropriateness of certainty, uncertainty, and knowledge-boundary marking across the interaction. The primary signal is not factual accuracy by itself, but the fit between the epistemic status of claims and the confidence, uncertainty, limitation, or source-boundary markers attached to them. Epistemic Calibration is analytically separable from topic and register, although empirical relationships among dimensions remain a validation question. This dimension draws on research in epistemic stance marking, which treats epistemic markers as local utterance-level phenomena. MTCS-R extends this to assess consistency and appropriateness of epistemic marking across an entire extended interaction, a trajectory-level property that existing work does not measure.
Scoring D5 requires that raters be able to evaluate the epistemic status of substantive claims in the interaction, including whether claims are established, uncertain, speculative, unverifiable from available context, or false. For interactions in specialized domains, raters should have appropriate domain competence, or D5 scoring should be flagged as provisional. Domain-competence requirements for D5 are addressed in the Rater Training Packet.
D5 differs categorically from D1 through D4 in the scoring task it requires. D1 through D4 assess properties of the interaction trajectory that are observable from the transcript by a competent reader: topical continuity, cross-turn integration, register stability, and structural maintenance. D5 requires the rater to also evaluate the epistemic status of substantive claims made in the interaction, which is a different kind of judgment from observing trajectory properties. This categorical difference has two consequences. First, D5 may dissociate from D1 through D4 in factor analysis (Section 9.2) for reasons distinct from the dissociation patterns predicted among D1 through D4. Second, the inclusion of D5 alongside D1 through D4 in the dimensional profile and in the optional composite (Section 5.1) does not imply that the five dimensions measure the same kind of property. They measure analytically separable properties of extended interaction, but D5's scoring task is more demanding and more dependent on rater domain competence than the others. Studies should report D5 scores alongside the domain-competence basis on which they were assigned, as required in Section 11. The factor-structure implications of D5's categorical difference are addressed in Section 10.2.
| Score | Descriptor | Behavioral Anchor |
|---|---|---|
| 1 | Severely Miscalibrated | Frequent unsupported or false claims presented with unwarranted certainty, and/or systematic absence of uncertainty signals where the interaction clearly requires them |
| 2 | Poorly Calibrated | Occasional overconfidence; some uncertainty signals but misplaced or inconsistent |
| 3 | Moderately Calibrated | Generally appropriate confidence levels; usually signals uncertainty when warranted |
| 4 | Well Calibrated | Consistent appropriate marking; clear distinction between known, uncertain, and unknown |
| 5 | Exceptionally Calibrated | Nuanced epistemic distinctions; precisely calibrated to actual knowledge quality |
5. Scoring Procedures
5.1 Composite Scoring
The primary MTCS-R result is the five-score dimensional profile: D1, D2, D3, D4, and D5. The composite MTCS-R score may be computed as an unweighted descriptive index: MTCS-R_total = mean(D1, D2, D3, D4, D5). This composite is offered for compact reporting and exploratory comparison only. It is not a validated latent score and should not be interpreted as replacing the dimensional profile.
Until factor structure is empirically resolved (Phase 2), the composite should be interpreted as a rough descriptive index and never reported without all five dimension scores. Dimension scores must always be reported separately because potential construct overlap, differential dimension weighting, and the distinct scoring requirements of D5 have not yet been empirically resolved. Alternative composite formulas may be used only when theoretically justified, explicitly documented, and not compared directly with unweighted MTCS-R_total scores as if they were equivalent.
5.2 Scoring Guidelines
Raters should evaluate the entire conversation trajectory, not individual turns. The full 1-5 scale should be used; scores of 1 and 5 are appropriate when the behavioral anchors are met and should not be avoided. All scores must be accompanied by written justification with specific turn references where possible. Raters should focus on observable behaviors as defined by the anchors, not personal preferences or topic interest. When uncertain, raters should flag the score for discussion during calibration sessions rather than defaulting to a midpoint score.
5.3 Minimum Conversation Length
For MTCS-R scoring, a turn is defined as one speaker contribution in the transcript. A complete user-AI exchange therefore normally consists of two turns: one human turn and one AI turn. Where transcripts contain system messages, tool outputs, or other non-participant insertions, studies should specify whether those items were excluded, treated as contextual material, or counted as turns.
MTCS-R is designed for interactions exceeding 20 participant turns, normally equivalent to at least 10 complete user-AI exchanges. This provisional design threshold is derived from the anchor reference points used in the dimension definitions: D1 anchor 2 references drift cycles of 5-8 turns, D2 anchor 4 references integration of content from 5-10 or more turns prior, and reliable assessment of cross-turn patterns requires a trajectory of at least roughly twice the upper anchor reference range. Interactions shorter than 20 turns do not provide sufficient trajectory for the behavioral anchors in Sections 4.1 through 4.5 to be applied without underdetermination. Applications to shorter interactions are not prohibited but should be flagged as below the design threshold, and scores from such applications should not be pooled with scores from interactions meeting the length requirement without statistical justification. The empirically supported floor is a Phase 1 validation question.
6. Dyadic Nature of Measurement
MTCS-R evaluates the coherence of the observable interaction trace produced jointly by AI system behavior, human interaction practice, and task structure. Scores therefore reflect dyadic performance rather than intrinsic model capability. Comparisons across models require control of human-side variables. This property aligns MTCS-R with established human-AI collaboration metrics rather than standalone AI benchmarks.
This dyadic framing is supported by a growing body of work reconceptualizing interaction quality as a property of the human-AI system rather than a static attribute of the AI alone. The PARADISE framework (Walker et al., 1997) established the precedent for system-level dialogue evaluation by combining task success with dialogue cost factors. Amershi et al. (2019) provided design guidelines organized across temporal phases of interaction, explicitly addressing quality as a system property. MTCS-R extends this tradition to extended multi-turn interaction with a psychometrically structured instrument.
A note on the anchor language is required. The behavioral anchors in Sections 4.1 through 4.5 reference AI-side observable behavior (responses, register shifts, fabrications) more directly than human-side or dyadic behavior. This is a deliberate methodological choice: rater agreement is more achievable when the observable referent is concrete output rather than the more diffuse property of joint dyadic performance. The dyadic interpretation is preserved at the level of inference, not measurement. A high or low MTCS-R score in any dimension reflects observable trajectory properties of the dyadic output stream and should not be attributed to either the AI system or the human practitioner without controlled comparison. This separation between anchor referent (AI-side behavior) and score interpretation (dyadic property) is the same separation that governs Section 10.3.
More precisely, MTCS-R scores the coherence of the observable dialogue trajectory produced under dyadic conditions. The transcript is the measurement object; AI-side utterances provide the most concrete and reliably scorable evidence within that object; and dyadic interpretation is warranted only because those utterances occur in response to human prompts, task structure, interaction history, and any protocol used by the human practitioner. MTCS-R therefore does not directly measure the human practitioner, nor does it directly measure intrinsic AI capability. It measures the coherence properties of the interaction trace generated by their coupling.
An important tension must be acknowledged. Gomez et al. (2024) argue through a systematic review that current human-AI interaction patterns fall short of genuine collaboration, with structural limitations in how humans and AI systems actually coordinate over time. MTCS-R does not claim that the interactions it measures constitute collaboration in this strong sense. It measures observable coherence properties of the dyadic output stream. Whether those properties arise from genuine mutual understanding or from sophisticated pattern completion on the AI side is an open question that MTCS-R is deliberately agnostic about. This agnosticism is a methodological strength: the instrument measures what is observable without requiring resolution of questions about AI cognition that the field has not settled.
7. Relationship to CAM
The Continuity Anchoring Method (CAM, SF0005; DOI pending Zenodo publication) is hypothesized to increase MTCS-R scores by reinforcing contextual anchoring, increasing cross-turn integration, stabilizing register, and enforcing epistemic standards. However, MTCS-R does not presuppose CAM. High-coherence interactions produced by alternative methods, including naturally skilled conversationalists, other structured interaction protocols, or future AI architectures with improved longitudinal consistency, should also yield high MTCS-R scores if the behavioral anchors are met.
Validation studies must therefore test both CAM versus non-CAM interactions and high-coherence versus low-coherence baselines generated independently of any specific method. This separation preserves instrument neutrality and is a standard requirement in measurement design: an instrument must demonstrate that it measures its target construct rather than the effects of a particular intervention (Messick, 1995).
8. Related Work
Several lines of prior work inform the MTCS-R design while also defining the gap it addresses.
8.1 Dialogue Evaluation Metrics
The evolution from word-overlap metrics through semantic similarity metrics to reference-free dialogue evaluation has produced increasingly sophisticated tools for assessing response quality (Liu et al., 2016; Deriu et al., 2021). However, these metrics overwhelmingly evaluate individual responses in context rather than interaction trajectories. DynaEval (Zhang et al., 2021) is the most direct precursor, modeling dialogue as a graph and producing dialogue-level scores, but it outputs a single holistic quality score rather than decomposing coherence into distinct measurable dimensions.
8.2 Discourse Coherence Modeling
Computational coherence modeling, grounded in Centering Theory (Grosz, Joshi, and Weinstein, 1995) and the entity-based approach (Barzilay and Lapata, 2008), has produced metrics for evaluating text-level coherence. However, existing approaches treat coherence as binary or unidimensional, focus on local transitions between adjacent sentences, and target monologic text or short dialogues rather than extended human-AI interaction.
8.3 Human-AI Interaction Assessment
Frameworks for evaluating human-AI interaction quality have been proposed at the system level. PARADISE (Walker et al., 1997) combined task success with dialogue cost factors. Amershi et al. (2019) provided design guidelines organized across temporal phases, explicitly addressing quality as a system property. However, no existing framework provides a psychometrically structured instrument for measuring trajectory-level coherence in extended human-AI dialogue.
8.4 Psychometric Foundations
The MTCS-R follows established methodology for behavioral rating scale development. The unified construct validity framework (Messick, 1995) integrates content, substantive, structural, generalizability, external, and consequential aspects of validity into a single comprehensive account that subsumes prior distinctions between content, criterion, and construct validity. This framework structures validity evidence around the coherence of interpretive arguments linking test responses to score-based claims. Generalizability theory (Shavelson, Webb, and Rowley, 2005) enables partitioning of variance into rater, dimension, conversation, and interaction components, which is essential for the MTCS-R's multi-rater, multi-dimension design.
9. Validation Pathway
MTCS-R requires staged empirical validation before its scores can be interpreted with confidence. The following four-phase pathway is proposed, following established methodology for behavioral rating instrument development (Messick, 1995).
9.1 Phase 1: Reliability and Anchor Refinement
Phase 1 establishes inter-rater reliability and refines behavioral anchors. Multiple trained raters independently score a corpus of transcribed human-AI interactions using the anchor tables defined in Section 4. Reliability is assessed using intraclass correlation coefficients (Shrout and Fleiss, 1979), specifically ICC(2,k) for designs where raters are sampled from a larger population. Target reliability is ICC or equivalent agreement coefficient greater than 0.60 (good agreement) or greater than 0.75 (excellent agreement), following the conventional thresholds proposed by Cicchetti (1994) for behavioral and clinical rating instruments. The 0.60 floor is treated as the minimum threshold below which anchor revision is mandatory rather than as an acceptance target. Anchors that produce low agreement are revised through calibration sessions and re-tested. The companion Rater Training Packet provides materials for this phase including annotated example transcripts, practice exercises, and calibration discussion guides.
9.2 Phase 2: Convergent Validity and Factor Structure
Phase 2 tests whether the five MTCS-R dimensions represent empirically distinguishable constructs and whether scores converge with theoretically related measures. Exploratory and confirmatory factor analysis test the proposed five-factor structure. If dimensions collapse empirically, the scale is revised to reflect the actual factor structure. Convergent validity is assessed by comparing MTCS-R scores against existing dialogue-level quality measures such as DynaEval scores, against human expert holistic ratings of interaction quality, and against known-groups comparisons using interactions designed to produce high versus low coherence. The predicted relationship between MTCS-R composite scores and DynaEval is moderate positive correlation, not high correlation. A high correlation would suggest that the multi-dimensional decomposition may add little beyond a unidimensional dialogue-quality summary, weakening the central design claim of MTCS-R. A near-zero or unstable correlation would require interpretation rather than automatic rejection: it could indicate that MTCS-R captures a distinct trajectory-level construct, but it could also indicate rater unreliability, construct mismatch, inadequate task sampling, or failure of the proposed anchors. Phase 2 should therefore interpret DynaEval convergence alongside inter-rater reliability, factor structure, known-groups discrimination, and discriminant validity against task completion, user satisfaction, and single-turn response quality.
9.3 Phase 3: Experimental Protocol Comparisons
Phase 3 evaluates the instrument's sensitivity to interaction conditions. CAM interactions are compared against non-CAM interactions and against deliberately degraded interactions to test whether MTCS-R scores discriminate as predicted. This phase also tests whether alternative interaction methods that are not CAM but that are designed to maintain high coherence produce high MTCS-R scores, which would confirm instrument neutrality.
9.4 Phase 4: Cross-Architecture Generalization
Phase 4 tests whether MTCS-R produces interpretable and reliable scores across architecturally distinct AI platforms. Generalizability theory (Shavelson, Webb, and Rowley, 2005) is applied to estimate variance components for raters, dimensions, conversations, AI architectures, and their interactions. Architecture neutrality is a design requirement: if the instrument produces systematically different scores for the same interaction quality across architectures, the anchors must be revised. This phase also extends to cross-domain testing (technical, creative, educational, professional interaction contexts) to establish the boundaries of generalizability.
10. Known Limitations
10.1 No Empirical Validation
MTCS-R is a measurement instrument specification. No empirical validation data exist at the time of publication. The instrument defines what to measure and how to score it; the validation pathway (Section 9) specifies the studies required to establish that it measures what it claims to measure. Scores produced by MTCS-R prior to completion of at least Phase 1 validation should be interpreted as preliminary and exploratory.
10.2 Dimension Correlation and Factor Structure Dependency
The five dimensions are analytically separable but may correlate empirically. If factor analysis reveals that some dimensions are not empirically distinguishable, the scale structure must be revised. This is a standard risk in multi-dimensional instrument development and is addressed by the Phase 2 validation pathway.
A stronger acknowledgment is required regarding the dependency of the scoring infrastructure on the proposed factor structure. The behavioral anchors in Section 4, the Rater Training Packet, the Scoring Template, and the dimensional profile reporting format all assume that the five-factor structure is approximately correct. If Phase 2 factor analysis (Section 9.2) does not confirm the five-factor structure, for example if D1 and D2 collapse into a single factor, or if D5 separates from a four-factor structure spanning D1 through D4, or if a different factor solution is empirically preferred, the consequences extend beyond minor revision. The anchors, training materials, scoring template, and reporting standards would all require substantive revision rather than incremental adjustment, and the proposed decomposition of longitudinal coherence into the specific five dimensions named here would require restatement. Pre-Phase 2 use of MTCS-R should therefore be characterized explicitly as use of an instrument whose factor structure is theoretically motivated and practitioner-refined but not empirically confirmed. The instrument is published in this status because the validation pathway requires a specified instrument as its starting point, not because the factor structure is established. Studies conducted prior to Phase 2 should report scores accordingly, and should treat pre-Phase 2 results as exploratory evidence about the proposed structure as well as about the interactions being scored.
10.3 Interaction Dependence
Scores reflect the coherence of observable dyadic interaction traces, not intrinsic model capability. This is a feature of the design, not a limitation per se, but it means that MTCS-R scores cannot be used to rank AI models without controlling for human-side variables, task structure, and conversation length.
10.4 Epoch Sensitivity
Behavioral anchors are calibrated to the range of AI conversational behavior observable in 2024-2026. As AI systems improve, the distribution of scores may shift upward, requiring anchor recalibration. This is a standard challenge for behavioral instruments in rapidly evolving domains.
10.5 Non-Inference of Internal States
MTCS-R measures observable coherence only. No inference about relational cognition, interiority, consciousness, or subjective experience is implied or permitted by any score on any dimension.
10.6 D5 Domain Dependence
Reliable D5 scoring requires raters competent to evaluate the epistemic status of substantive claims made in the interaction, including whether claims are established, uncertain, speculative, unverifiable from available context, or false. For interactions in specialized domains (technical, medical, legal, scientific), this constrains the rater pool and may limit the feasibility of D5 scoring without domain-matched raters. Studies should report the domain competence basis for D5 scoring decisions and flag D5 scores as provisional where rater domain competence is not established. In dialogue genres where truth status is not uniformly applicable, such as fictional, purely expressive, or deliberately speculative interaction, D5 should be scored against the interaction's stated epistemic frame rather than against ordinary factual correspondence, and this scoring basis should be reported.
11. Minimum Reporting Standards
Any study using MTCS-R should report the following: all five individual dimension scores (D1 through D5) and the composite score; the number and training status of raters; inter-rater reliability coefficients (ICC or Krippendorff's alpha) for the study; conversation length in turns; AI system(s) and version(s) used; characterization of the human practitioner including training status, role, and any structured interaction protocol followed; task structure or interaction prompt; any deviations from the standard scoring procedure; and whether the scoring was conducted before or after completion of Phase 1 validation. For D5 scores, studies should additionally report the domain competence basis on which D5 was scored. Studies conducted prior to Phase 1 validation must explicitly state this and characterize their MTCS-R scores as preliminary.
12. Supplementary Materials
Six files are provided alongside this document: three implementation companion files supporting MTCS-R use, and three deposit-support files documenting orientation, inventory, and licensing.
MTCS-R Rater Training Packet (DOCX): Comprehensive training materials for raters including dimension definitions, behavioral anchor reference tables, scoring guidelines, three annotated example transcripts with reference scores and rationale, a calibration discussion guide with structured procedures for achieving inter-rater reliability, and reference materials. Practice transcripts for rater calibration exercises are not included in this package; they should be assembled by validation study coordinators from their own interaction corpora to ensure ecological validity for their specific research context.
MTCS-R Scoring Template (XLSX): Excel workbook with scoring forms, dimension-level and composite score calculators, a dimension anchors reference sheet, and space for rater justifications. The Reliability sheet provides simple agreement percentage calculation only; for ICC and Krippendorff's alpha computation, use the reliability_calculation.R script in the Code Examples archive.
MTCS-R Code Examples (ZIP archive): Computational examples including a semantic drift analysis script and notebook, a perplexity trajectory analysis script, an R script for inter-rater reliability calculation (ICC and Krippendorff's alpha), and sample data files for demonstration purposes.
README_FIRST.md: Orientation document for the deposit package. Presents pre-validation status, the requirement to never report the composite alone, the categorical difference of D5 from D1-D4, and the dyadic-trace measurement scope before introducing the recommended workflow order.
MANIFEST.txt: Plain inventory of all files in the deposit package with role, version, format, and license information.
LICENSE.txt: CC-BY 4.0 license statement covering the documents and original code in this package; third-party software dependencies in the Code Examples archive remain under their own licenses.
13. Methodological Status
MTCS-R is a measurement instrument specification derived from observed coherence patterns across extended human-AI interactions spanning multiple architectures and platforms since late 2022.
What this document is: An operational definition of longitudinal coherence as a measurable construct; a multi-dimensional behavioral rating instrument with anchored scoring; a staged validation pathway grounded in established psychometric methodology; a contribution to closing the documented gap in dialogue evaluation methodology.
What this document is not: A controlled empirical study with quantitative validation; a claim that the five-factor structure has been confirmed; a benchmark or leaderboard for AI systems; a platform-specific implementation guide.
Development basis: Observational pattern synthesis from extended interaction with thousands of AI instances across multiple architectures and platforms since late 2022. This constitutes methodology development from practitioner experience, not controlled experimental research. The identified dimensions and anchor points are proposed as theoretically grounded and practitioner-refined starting points for the empirical validation process described in Section 9.
Validation pathway: Practitioners and researchers are encouraged to implement the validation pathway specified in Section 9. If the instrument does not achieve acceptable reliability or demonstrate the predicted factor structure, it should be revised or restructured. If it does not discriminate between interactions that trained practitioners judge as high versus low coherence, it should be refined or rejected.
14. Conclusion
MTCS-R provides a behavioral instrument for evaluating coherence trajectories in extended human-AI interaction. It operationalizes a longitudinal construct not adequately captured by the standard dialogue metrics reviewed in this document and enables empirical testing of interaction protocols such as CAM. The instrument's novel contribution lies in integrating psychometric rigor with multi-dimensional construct decomposition, trajectory-level analysis, and dyadic measurement of human-AI interaction quality. Each of these components has precedent in separate literatures; the integration is new.
The scale evaluates the coherence of observable dyadic interaction traces and requires validation through the staged reliability and experimental studies described in Section 9. Until that validation is completed, MTCS-R should be understood as a carefully specified measurement proposal, not a validated instrument. The companion materials provided as supplementary files are designed to support the first phases of that validation process.
More information and current public materials are available at synthience.org.
References
- Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for Human-AI Interaction. Proceedings of CHI '19, Article 3. https://doi.org/10.1145/3290605.3300233
- Andrews, R.W., Lilly, J.M., Srivastava, D., & Feigh, K.M. (2023). The Role of Shared Mental Models in Human-AI Teams: A Theoretical Review. Theoretical Issues in Ergonomics Science, 24(2), 129-175. https://doi.org/10.1080/1463922X.2022.2061080
- Barzilay, R. & Lapata, M. (2008). Modeling Local Coherence: An Entity-Based Approach. Computational Linguistics, 34(1), 1-34. https://doi.org/10.1162/coli.2008.34.1.1. https://aclanthology.org/J08-1001.pdf
- Cicchetti, D. V. (1994). Guidelines, Criteria, and Rules of Thumb for Evaluating Normed and Standardized Assessment Instruments in Psychology. Psychological Assessment, 6(4), 284-290. https://doi.org/10.1037/1040-3590.6.4.284
- Clark, H. H. & Brennan, S. E. (1991). Grounding in Communication. In L. B. Resnick, J. M. Levine, & S. D. Teasley (Eds.), Perspectives on Socially Shared Cognition (pp. 127-149). APA. https://doi.org/10.1037/10096-006. Available at: https://web.stanford.edu/~clark/1990s/Clark,%20H.H.%20_%20Brennan,%20S.E.%20_Grounding%20in%20communication_%201991.pdf
- Deriu, J., Rodrigo, A., Otegi, A., Echegoyen, G., Rosset, S., Agirre, E., & Cettolo, M. (2021). Survey on Evaluation Methods for Dialogue Systems. Artificial Intelligence Review, 54(1), 755-810. https://doi.org/10.1007/s10462-020-09866-x
- Gantz, T. W. (2026). Citation Verification Protocol (CVP, SF0037). Synthience Institute. https://doi.org/10.5281/zenodo.18075624
- Gantz, T. W. (2026). Context Representation Drift (CRD, SF0039). Synthience Institute. https://doi.org/10.5281/zenodo.18289391
- Gantz, T. W. (2026). Continuity Anchoring Method (CAM, SF0005). Synthience Institute. Available at: https://synthience.org
- Gantz, T. W. (2026). Theoretical Foundations (SF0003). Synthience Institute. Available at: https://synthience.org
- Gomez, C., Cho, S. M., Ke, S., Huang, C.-M., & Unberath, M. (2024). Human-AI Collaboration is Not Very Collaborative Yet: A Taxonomy of Interaction Patterns in AI-Assisted Decision Making from a Systematic Review. Frontiers in Computer Science, 6:1521066. https://doi.org/10.3389/fcomp.2024.1521066
- Grosz, B. J., Joshi, A. K., & Weinstein, S. (1995). Centering: A Framework for Modeling the Local Coherence of Discourse. Computational Linguistics, 21(2), 203-225. https://aclanthology.org/J95-2003/
- Liu, C. W., Lowe, R., Serban, I. V., Noseworthy, M., Charlin, L., & Pineau, J. (2016). How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. Proceedings of EMNLP 2016, pp. 2122-2132. https://arxiv.org/abs/1603.08023
- Messick, S. (1995). Validity of Psychological Assessment: Validation of Inferences from Persons' Responses and Performances as Scientific Inquiry into Score Meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741. Available at: https://files.eric.ed.gov/fulltext/ED380496.pdf
- Pickering, M. J. & Garrod, S. (2004). Toward a Mechanistic Psychology of Dialogue. Behavioral and Brain Sciences, 27(2), 169-190. https://doi.org/10.1017/S0140525X04000056. Available at: https://www.pure.ed.ac.uk/ws/files/11823730/Toward_a_mechanistic_psychology_of_dialogue.pdf
- Shavelson, R. J., Webb, N. M., & Rowley, G. L. (2005). Generalizability Theory: Overview. In B. S. Everitt & D. C. Howell (Eds.), Encyclopedia of Statistics in Behavioral Science. Wiley. Available at: https://www.researchgate.net/publication/313966026_Generalizability_Theory
- Shrout, P. E. & Fleiss, J. L. (1979). Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological Bulletin, 86(2), 420-428. https://doi.org/10.1037/0033-2909.86.2.420
- Walker, M. A., Litman, D. J., Kamm, C. A., & Abella, A. (1997). PARADISE: A Framework for Evaluating Spoken Dialogue Agents. Proceedings of the 35th Annual Meeting of the ACL, pp. 271-280. https://doi.org/10.3115/976909.979652
- Yeh, Y. T., Eskenazi, M., & Mehri, S. (2021). A Comprehensive Assessment of Dialog Evaluation Metrics. Proceedings of the 1st Workshop on Evaluations and Assessments of Neural Conversation Systems, pp. 15-33. https://arxiv.org/abs/2106.03706
- Zhang, C., Chen, Y., D'Haro, L. F., Zhang, Y., Friedrichs, T., Lee, G., & Li, H. (2021). DynaEval: Unifying Turn and Dialogue Level Evaluation. Proceedings of ACL-IJCNLP 2021, pp. 5676-5689. https://doi.org/10.18653/v1/2021.acl-long.441
- Zhang, Y. B. & Giles, H. (2018). Communication Accommodation Theory. In Y. Y. Kim (Ed.), The International Encyclopedia of Intercultural Communication (pp. 95-108). Wiley. https://doi.org/10.1002/9781118783665.ieicc0156. Author-hosted PDF: https://www.researchgate.net/profile/Yan-Bing-Zhang-2/publication/342872755. See also the original formulation: Giles, H. & Ogay, T. (2007). Communication Accommodation Theory. In B. B. Whaley & W. Samter (Eds.), Explaining Communication: Contemporary Theories and Exemplars (pp. 293-310). Erlbaum.