Project Interloper

The archive

What this rests on.

The claim running through this investigation is that every statement traces to a specific document you can read yourself. That is only worth anything if the archive behind it is real and handled carefully. This page describes it, and the rules it is kept under.

The holdings

More than twelve thousand documents from the CIA’s declassified STARGATE collection, released through the agency’s records search tool, which holds CIA, DIA, Army and contractor records rather than one agency’s files. That part alone runs to roughly 104,000 pages.

Around it sit declassified records from other releases, documents obtained through records requests, archival collections, the press of the period, and published books. In total the working surface is about 159,000 transcribed pages drawn from several dozen separate sources, every page carrying a header that records where it came from, how it was read, and when.

That figure counts pages transcribed, not unique pages. The same document often survives in more than one release, sometimes with different redactions. Those copies are mapped against each other rather than merged, because the differences between them are frequently the most interesting thing on the page.

More than thirty records requests have been filed with federal, state and local agencies.

How a scan becomes a source

A declassified page arrives as an image of a photocopy of a typescript, often third generation, often stamped and annotated by hand. Machine reading handles the bulk of it: roughly 157,000 pages through commercial optical character recognition, with about 2,200 difficult pages given a second, independent reading by a vision model, most of them handwritten, plus a few thousand pages from older optical character recognition passes.

Those machine readings are searchable, and they are what makes a corpus this size usable at all. They are not what gets quoted.

A machine transcription is a pointer, not a source. Before any quoted word is relied on, it is checked against the scan or photograph of the original page, and the check is logged.

A third lane sits above both: pages I have read off the original image myself. There are 34 so far, and each outranks every machine reading of the same page. Separately, every quoted passage is checked against its page image before use, and that ledger runs to a few hundred entries.

What is deliberately not trusted

Two earlier transcription passes are still on disk and are no longer usable as sources.

The first was an early attempt at reading photographs of handwritten pages. It did not merely misread words, it produced words that are not on the page. The second is a pre-2026 text extraction, superseded by the current per-page process; it is less accurate and carries no per-page provenance at all.

Both are marked as superseded, and that marking exists because of a failure rather than foresight: two separate audits scored the dead set by mistake and reported wrong numbers before anyone noticed. Keeping obsolete data where tooling can still reach it is its own hazard, and the correction is part of the record.

How a source is rated

Every source is weighed before it is used, and the rating stays attached to the claim wherever that claim travels.

Tier 1: the record itself
Records made in the normal course of someone's duties, before anyone expected them to be public. Provenance decides this, not content and not age. Only Tier 1 can establish a fact.
Tier 2: accounts with access
People with first-hand knowledge, or researchers with document access, writing afterwards and for an audience. Tier 2 can corroborate a Tier 1 finding. It can never found one.
Tier 3: everything of unknown provenance
Uncited web pages, anonymous claims, documented disinformation, and answers from AI chatbots. A lead, never evidence.

Where a later account conflicts with the record, the conflict is stated and the record is preferred, unless something better than the record contradicts it. That has mattered more than once.

The discipline around it

Since October 2026, every reading is logged with its scope
So a claim that something was read in full says exactly what was read, and a later pass can tell the difference between a document that was searched and a document that was read.
A register of what turned out to be false
Claims that failed are recorded along with where the error came from, so the same mistake is not made twice and a reader can see what has been withdrawn.
Retractions are swept, not banners
When something is withdrawn, every file that repeated it is corrected, or the exception is logged with a reason. A notice at the top of one page is not a retraction.
A negative search is not a negative finding
Optical character recognition misreads names. Nothing is reported as absent until the plausible misspellings and variant renderings have been searched too.
Every change is committed
The version history is the provenance log. What was claimed, when it changed, and why, is recoverable rather than remembered.

Limits I hold myself to

Living private individuals appear through their published or professional record and nowhere else. Some archival material is held under personal-use terms and is read but not reproduced. Where a document cannot be shown, that is said rather than worked around.

Seeing it work

The clearest public example is the training app built from the remote viewing manuals. Every document behind its method is listed, with what each one contributed and a link to read it.

The documents behind Scanate

The method itself, and how a claim is rated before it is published, is set out in How I work.

Follow along

New documents and dispatches as they publish. The newsletter is the one place to keep up.