Verification Report

RICO Citation Verification Report

Report IDSR001-VR1 Versionv1.2 | August 8, 2026 ManuscriptRICO (SR001 v6.3) CVP ProtocolSF0037 v1.4 VerifierClaude Opus 4.6 (claude-opus-4-6), Anthropic Verification DateAugust 8, 2026 Verification TierTier 1 (Single-Platform Verified) with PAR persistence records Prepared byThomas W. Gantz, Synthience Institute LicenseCC-BY 4.0 StatusPublished DOI: 10.5281/zenodo.18082749
Report Version Note

This is a substantive revision, not a reissue. Report versions v1.0 and v1.1 verified RICO v3.8 and v5.7 respectively. Those versions were previously described as permanently locked to the manuscript version they verified. That lock has been lifted by PCP decision of 2026-08-08, on the ground that the lock governs which manuscript version a report verifies, not whether the report’s findings are correct, and v1.1 contained an error that could not be corrected without re-issuing.

The error: v1.1 recorded C2 (Olsson et al., 2022) as verified with no fabrication detected, and dispositioned it KEEP. The RICO reference at that time listed an author, “Rai, A.”, who does not appear anywhere in the source, and gave DasSarma the wrong initial. The v1.1 record shows the verifier encountered the discrepancy and produced an unparseable note about it rather than assigning the AMBIGUOUS rating that CVP Section 5 Part C requires when support cannot be confidently determined. The certification statement issued on that basis was therefore inaccurate as to C2.

This version re-runs the full protocol against RICO v6.3, in which the reference has been corrected. The failure is documented rather than quietly repaired, because a demonstration artifact for a verification protocol is more useful to a practitioner when it shows the protocol’s real failure mode than when it shows only a clean pass.

1. Live Retrieval Access Confirmation

Per SF0037 v1.4 Section 4.1. Live access confirmed prior to verification. Full text was retrieved for every citation; no citation was assessed from abstract or metadata alone.

URL OpenedContent ConfirmedStatus
arxiv.org/abs/2005.14165Brown et al. (2020) landing page; 31-author byline and v4 submission history retrieved.LIVE
arxiv.org/pdf/2005.1416575-page full text extracted; few-shot definition located at p.6.LIVE
transformer-circuits.pub/2022/in-context-learning…Full HTML retrieved (124,410 characters of body text); byline and induction-head definition confirmed.LIVE
aclanthology.org/2024.tacl-1.9.pdf17-page TACL full text; lost-in-the-middle finding located at p.1.LIVE
arxiv.org/pdf/2404.0665427-page RULER full text; context-length degradation results located at p.1.LIVE
arxiv.org/pdf/2309.1745321-page ICLR 2024 full text; attention-sink mechanism located at p.1.LIVE
doi.org/10.5281/zenodo.18289391Concept DOI resolved to record 19155138 (CRD v1.6); PDF openly downloadable.LIVE

2. Citation Inventory

Per SF0037 v1.4 Section 4.2. Total citations in RICO v6.3: N = 6.

IDReferenceCited InAccess Class
C1Brown et al. (2020)5.3FREE-FULLTEXT
C2Olsson et al. (2022)5.3, 7.2FREE-FULLTEXT
C3Liu et al. (2024)1.2FREE-FULLTEXT
C4Hsieh et al. (2024)1.2FREE-FULLTEXT
C5Xiao et al. (2023)7.3FREE-FULLTEXT
C6Gantz (2026) SF0039References onlyFREE-FULLTEXT

3. Individual Citation Verification Records

C1 — Brown et al. (2020) KEEP FULL SUPPORT
ReferenceBrown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., & Amodei, D. (2020). Language models are few-shot learners. arXiv:2005.14165.
ClaimCLAIM-01: “in-context learning, where models adapt to patterns from examples provided within a prompt”
Evidence anchorp.6 (Section 2.1, Figure 2.1 discussion): “few-shot works by giving K examples of context and completion”. Defines few-shot operation as in-prompt example provision at inference time, which is the adaptation the claim attributes to the source.
NotesBibliographic note: the source carries 31 authors. The RICO reference lists the first six followed by the final author, Amodei, with no ellipsis, presenting a truncated list as though complete. All seven named authors are correct and correctly initialled, and no name is fabricated; this is a style defect, not an integrity defect.
C2 — Olsson et al. (2022) KEEP FULL SUPPORT v1.1 DISPOSITION SUPERSEDED
ReferenceOlsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., & Ndousse, K. (2022). In-context learning and induction heads. Transformer Circuits Thread.
ClaimsCLAIM-01 (ICL support) and CLAIM-02: “Olsson et al. (2022) describe induction heads — circuit components that allow transformer models to continue patterns observed earlier in the context window”
Evidence anchorsIntroduction: “induction heads might constitute the mechanism for the actual majority of all in-context learning”. Induction head definition section: “induction heads in our models are implemented by a circuit of two attention heads”. Both FULL SUPPORT.
Author list correctionThe reference verified in v1.1 read “…Mann, B., DasSarma, A., Ndousse, K., & Rai, A.” Two defects were present and both were passed. First: the source byline gives Nova DasSarma; the initial A. was wrong, it is N. Second: no author named Rai appears in the source. An exhaustive search of the retrieved full text returns exactly one occurrence of the string “Rai”, as a substring of “Rebecca Raible” in the contributions list, and zero occurrences of Rai as a standalone surname anywhere in the byline, the BibTeX block, or the acknowledgements. The name was fabricated.
Why v1.1 passed itReport v1.1 recorded that it had encountered this discrepancy and wrote that “Rai A. listed as ‘Kamal Ndousse’ in full list”, a statement that does not resolve to any coherent finding, then concluded no fabrication detected and dispositioned KEEP. Under SF0037 Section 5 Part C, a verifier that cannot confidently determine the position must assign AMBIGUOUS, which is treated as failure. The correct v1.1 disposition was REVISE CLAIM.
ResolutionRICO v6.3 corrects the reference to the source byline order: Olsson, Elhage, Nanda, Joseph, DasSarma (N.), Henighan, Mann, Ndousse. All eight names verified present and correctly initialled against the retrieved byline. Henighan, previously omitted, sits between DasSarma and Mann in the source and is restored.
C3 — Liu et al. (2024) KEEP FULL SUPPORT
ReferenceLiu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173.
ClaimCLAIM-03: models underutilize information positioned in the middle of long contexts, retrieving content near the boundaries more reliably.
Evidence anchorp.1 (Abstract and Introduction): “performance is often highest when relevant information occurs at the beginning or end of the input context”. The source states the boundary-advantage and middle-degradation pattern directly.
NotesBibliographic note cleared: v1.1 recorded that RICO cited the arXiv preprint numbering (vol. 11, 553–569, 2023) rather than the final TACL record. RICO v6.3 uses the final journal citation and the ACL Anthology locator. The v1.1 recommendation was adopted and this note is closed.
C4 — Hsieh et al. (2024) KEEP FULL SUPPORT
ReferenceHsieh, C. Y., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., & Ginsburg, B. (2024). RULER: What’s the real context size of your long-context language models? arXiv:2404.06654.
ClaimCLAIM-04: “the limits of effective context utilization across different architectures”
Evidence anchorp.1 (Abstract): “almost all models exhibit large performance drops as the context length increases”. Across the seventeen long-context models RULER evaluates, only half sustain satisfactory performance at 32K.
C5 — Xiao et al. (2023) KEEP FULL SUPPORT
ReferenceXiao, G., Tian, Y., Chen, B., Han, S., & Lewis, M. (2023). Efficient streaming language models with attention sinks. arXiv:2309.17453.
ClaimCLAIM-05: attention does not distribute evenly across context tokens; certain tokens accumulate disproportionate weight regardless of semantic relevance.
Evidence anchorp.1 (Abstract and Introduction): “strong attention scores towards initial tokens as a sink even if they are not semantically important”. Supports both halves of the claim.
NotesAuthor list verified complete against the source byline: Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis. Five authors, all present, correctly initialled, no truncation.
C6 — Gantz (2026) SF0039 SECTION 4.4 FAILURE REMOVE OR REATTACH
ReferenceGantz, T. W. (2026). Context representation drift. Synthience Institute. SF0039. DOI: 10.5281/zenodo.18289391
ClaimNONE. No inline claim in RICO v6.3 depends on this citation. Part C is not reached, because support verification cannot be performed on a citation that supports no claim.
PVC status resolvedReport v1.1 assigned this citation UNABLE TO VERIFY because the Zenodo record did not resolve in any public index on 2026-03-05. It now resolves: the concept DOI redirects to version record 19155138, CRD v1.6, with the full PDF openly downloadable (178.8 kB, md5 97f250409fedb3557089fb669f49f5dc). The v1.1 flag is cleared. That flag was correctly raised at the time and correctly did not exempt the author’s own work.
New findingFloating citation, SF0037 v1.4 Section 4.4. The reference appears in RICO v6.3’s reference list, and a full-document search finds no inline citation of Gantz, SF0039, CRD, or Context representation drift anywhere in the body text. Section 4.4 states that citations not mapped to claims are invalid and must be removed or reattached with explicit claim linkage. Report v1.1 assessed this under CVP v1.3 and recorded the absence of an inline claim as the reason a PVC failure would not disqualify the manuscript; v1.4’s Section 4.4 makes the same absence itself disqualifying.
DispositionThis is not a fabrication or an accessibility problem. The cited work exists, is the author’s own, is openly accessible, and is architecturally related to RICO. The defect is that the reference list asserts a dependency the manuscript does not have. REMOVE or REATTACH, PCP decision required.

4. Verification Summary

MetricResult
Total citations inventoried (N)6
Citations fully processed6 / 6
Citations passing PVC (FREE-FULLTEXT)6 (C1–C6)
Citations mapped to a manuscript claim5 (C1–C5); C6 unmapped
Citations rated FULL SUPPORT5 (C1–C5), across six claim mappings
Citations failing Section 4.4 claim-map requirement1 (C6)
Citations removed or replaced0 (one pending PCP decision)
Fabricated author names detected and corrected1 (C2, corrected in manuscript v6.3)
Prior-report findings superseded2 (C2 disposition; C6 PVC flag cleared)
Bibliographic style notes recorded1 (C1 truncation without ellipsis)
Verification Tier achievedTier 1 (Single-Platform Verified) with PAR records

Tier note, stated plainly. SF0037 Section 7 sets Tier 2 (dual-platform verification) as the default for public release. This run achieved Tier 1: each citation was verified against one authoritative first-party source with full text opened and evidence anchors captured, and PAR persistence records were recorded for all six. Independent second-repository confirmation was not performed for every citation. The honest tier is therefore Tier 1 with PAR, not Tier 2. Report v1.1 claimed Tier 2 on the basis of dual confirmation for two of six citations, which overstated the tier; that claim is superseded.

Certification Statement

I certify that I executed SF0037 CVP v1.4 using live retrieval against public sources. Full text was opened for every citation; no support rating was assigned from an abstract, from metadata, or from model priors.
Total citations inventoried: N = 6 | Citations fully processed: 6 / 6 | Citations retained: 5 unconditional (C1–C5) + 1 pending PCP disposition (C6) | Citations removed or replaced: 0 | Verification tier achieved: Tier 1 with PAR | Verification completed: August 8, 2026

Verifier: Claude Opus 4.6 (claude-opus-4-6), Anthropic | Prepared by: Thomas W. Gantz, Synthience Institute

5. Recommendations for Publication

1. C6 (Gantz 2026, CRD, SF0039) — action required. The citation satisfies PVC but fails the Section 4.4 claim-map requirement. Two dispositions are available and the choice is a PCP decision. Remove it from the reference list, which costs nothing evidentially since no claim rests on it. Or reattach it by adding an inline citation at a point where RICO’s argument genuinely draws on context representation drift; Section 1.2, which surveys long-context literature, is the natural site. Reattachment is the stronger option if and only if the manuscript actually uses the source; adding an inline citation solely to legitimise a reference-list entry would invert the requirement.

2. C1 (Brown et al. 2020) — optional style correction. The seven-name list is a truncation of a 31-author byline presented without an ellipsis. All named authors are correct. No integrity issue and no urgency.

3. C2 (Olsson et al. 2022) — closed. The fabricated author and the incorrect initial are corrected in RICO v6.3 and verified against the source byline.

4. C3 (Liu et al. 2024) — closed. The final TACL citation replaced the preprint numbering per the v1.1 recommendation.

5. Tier upgrade before deposit. A second independent repository confirmation per citation would achieve the Tier 2 default that SF0037 Section 7 sets for public release.

6. On the value of this report as a demonstration artifact. Report v1.1 was published to show what CVP output looks like. It now also shows what a verification failure looks like from the inside: a verifier that retrieved the correct source, encountered a name it could not place, wrote a fluent sentence that resolved nothing, and passed. Section 5 Part C already prescribes the correct behaviour, which is to rate AMBIGUOUS and treat it as failure. The gap was not in the protocol but in its execution, and the class of error is one an automated author-list comparison cannot commit, since a name either appears in the retrieved byline or it does not. Practitioners running CVP should treat any author-name discrepancy as AMBIGUOUS until resolved by direct string comparison against the retrieved source.

Suggested Citation
Gantz, T. W. (2026). RICO Citation Verification Report (SR001-VR1 v1.2). Synthience Institute. https://doi.org/10.5281/zenodo.18082749

Document: SR001-VR1 Verification Report
Version: v1.2
Author: Thomas W. Gantz
Affiliation: The Synthience Institute
Date: August 8, 2026
License: CC-BY 4.0