Verafy
Light Paper
1. Introduction
The question this paper starts from is not which model is smartest. It is who is telling the truth, and how anyone would check. Language models now sit between people and the record. They summarize the news, answer disputed questions, and will increasingly act for people as agents. Their answers are fluent. Fluency is not a record.
Verafy is a system for comparing those answers, judging them against each other, and writing the result down where it can be checked later. The public surface of that system today is the Swarm Explorer. This paper describes the design, the history that produced it, and a larger claim: that agreement among models, under a procedure that can be audited, is a new primitive for shared information. Where a claim looks past what is running now, the text says so.
2. History
Verafy started as a hackathon project. The useful question was why anyone should use AI and a blockchain together. Plenty of ideas were already circulating about agents. The one that held was older than those ideas. A blockchain is a way for many machines to keep the same ledger, because each copy can be checked against the others. Applied to AI, that property is useful for one job: inscribing a result so that it cannot be quietly deleted, and so that anyone can verify what was written.
Rex St John built the first demonstration of that idea over a winter vacation, with ChatGPT, and published it in December 2024. The demo was called Truth Chain. That name belongs to the demonstration. It is not the name of the system this paper describes. The demo put model disagreement on screen. It spread. Thousands of people watched it, and about 4,000 people came into the community. A team of volunteers was recruited from that community to keep going. The original demonstration is still online.
The work that followed had three parts, and they are still the core of the design. A public catalog of contentious questions, chosen and graded by human reviewers, so the same prompts could be asked again later. Access to many models, which Hugging Face and OpenRouter made available by API. And a way to inscribe the results. That method came from IQ6900, built by Zo. The project was announced on 27 January 2025. On 7 February 2025 the early product set shipped: an agent, the first website, and a terminal.
The same period produced a Solana validator, the Truth Node. On 10 March 2025 it entered the Solana Foundation Delegation Program. A keynote at Solana Accelerate was watched more than 600,000 times on X. Designers, engineers, community managers, and builders in several countries shipped further iterations of the platform.
Then the effort stalled. The crypto market crash of October 2025 set the work back. After a few months with a new generation of tools, including Claude Code, Cursor, and Grok Bot, the project was restarted: a blog, a rebuilt website, and a return to the original job. The product of that second run is the Swarm Explorer.
3. The problem
Who is telling the truth?
The difficulty is not a shortage of text. It is that the sources of text are interested, incomplete, or both. The early notes for this project listed the familiar failure modes: model bias, tampered media, interference by institutions and governments, disinformation on social networks, censorship by omission, filtered search, incomplete information, studies that do not hold up, appeals to authority, appeals to emotion, biased editing, and reframing. Finding the truth is hard because each of those can look like an ordinary paragraph.
Models add a problem of their own. As models multiplied across Hugging Face and the major laboratories, it became obvious that their builders were steering them. Some were sandboxed. Some were pushed away from certain lines of reasoning. Some gave cagey answers. And some answered the same question differently as they were updated. A bias that moves is easy to miss if you only look once.
A snapshot is not enough. One response, captured today, says nothing about the trajectory. Is a model becoming more truthful or less? More biased or less? On which subjects? The missing tool was a canonical way to ask the same question of many models, keep the answers, and show the path.
Agents make the gap sharper. An agent that ingests markets, social feeds, news, and chat, and then posts or trades, needs a check it can call in the moment. The project notes from early 2025 put the gap plainly: there was no viable real-time fact-checking or prediction interface for that caller. Search pages and human moderation were not built for it.
4. Prior approaches
Search engines, of which Google is the type, rank documents. They do not keep a public, comparable record of how a language model answered a disputed question, or of how that answer moved. Filtering and omission are part of the product. A ranked list is not a cross-examination.
Perplexity and systems like it synthesize an answer and attach sources. That is useful, and it is still one system's answer. There is no second model in the loop whose disagreement is the point, and no requirement that yesterday's answer remain inscribed when today's answer changes.
Community Notes distributes the work to people who write context on a post. It is a real correction mechanism, and it is slow, manual, and bound to the host platform. It does not run on an arbitrary claim an agent encounters in the middle of the night.
Wikipedia is the strongest public compromise we have for settled topics, produced by editors under a visible policy. It is not real-time. It is not neutral by magic: it has editors, and editors have views. It is the wrong shape for the question of what a particular model said about a claim on a particular date.
Prediction markets, of which Polymarket is a current example, price a question that resolves. They are a serious instrument when the question has an observable outcome. They are a poor instrument for whether a paragraph is biased, and they do not archive the reasoning of language models.
Fact-checkers and media desks, including organizations in the shape of Snopes, investigate claims one at a time. The work can be careful. It does not scale to a stream of agent queries, and the editorial line is itself something a reader has to judge. Social networks, whether Facebook, Bluesky, Farcaster, or Truth Social, are distribution systems. They are not procedures for agreement.
None of these is useless. None of them is a procedure that asks many models the same question, scores the answers with further models, weights the judges by a test they did not know they were taking, and writes the outcome to a ledger.
5. The Verafy design
The platform needs three pieces, and a procedure for turning their output into a record. The pieces are a benchmark, access to models, and inscription. The procedure is judgment: models examining models, with human review where the questions are chosen and where the record is checked.
5.1 A benchmark of contentious questions
The prompts that reveal a model are not trivia. They are the questions people actually fight about. Verafy keeps a database of contentious questions, graded and selected by human reviewers because those reviewers expect the questions to expose a slant. The catalog is public, and it is a benchmark. Asked again over time, it is how a track record is built, across different dimensions of bias, rather than from a single impression.
5.2 Access to many models
The comparison is only as wide as the set of models you can call. Hugging Face and OpenRouter provided API access to a wide range of models, which is why this part of the system did not have to be invented in-house. Verafy does not train a foundation model. It benchmarks the ones that already exist, and it builds the query path, the record, and the tools around them. Swarm Explorer sits on that access: one question, many models, answers side by side.
5.3 LLM-as-judge consensus
A pile of answers is not yet a judgment. The method Swarm Explorer is built around is to use language models as judges. Judges cross-examine the claims in the answers, compare them, and surface where the swarm agrees and where it splits. Agreement is not truth. It is evidence about the swarm, and it is more informative than any single completion. Disagreement is the part worth reading. Bias shows up as a split that persists, or as a split that moves when the same question is asked again.
Using a model to grade a model is a known evaluation method, and it has a known weakness: the judge can share the bias of the thing it grades. The next section is the design's answer to that weakness. This paper does not claim that a panel of models converges on the truth. It claims that a recorded panel, asked the same question over time, is a better instrument than one unrecorded answer.
5.4 Merit-weighted voting
Judges are not equal. The design samples models for merit against an incoming class of query. Synthetic challenges, called tracer questions, are salted into the query stream so a model cannot easily tell the exam from the work. Merit is updated continuously, for bias as well as for competence. Models that hold up better receive more weight, and they are selected more often when a judgment is formed.
The early sketches of this mechanism draw illustrative weights next to the nodes. Those figures are a diagram, not a measurement, and this paper does not treat them as results. What the design commits to is the procedure: hidden tracers, continuous scoring, and more influence for the models that survive the test.
5.5 Immutable inscription
The result of a run is a snapshot: the question, the answers, the judgment. A snapshot that lives only in a vendor's database can be edited, dropped, or lost when the company changes its mind. Verafy inscribes snapshots on a blockchain so the record is public and hard to alter. That is the original reason to put a chain under an AI system. Many nodes can agree on the ledger because each copy can be checked.
The inscription method is Code-In, from IQ6900, built by Zo. Code-In treats the transaction itself as the data. Fragments are linked, each pointing at the previous transaction, so the system does not allocate a growing account to hold the file. IQ6900 describes the result as remaining readable for as long as the chain does. Verafy uses that method to inscribe its snapshots. The partnership is part of the design, not a separate product claim.
5.6 Verified document bundles
The same pipeline applies to documents, not only to live questions. A source is ingested. A report is generated and appended. The bundle is compressed, packaged as JSON, and inscribed with Code-In. The bundle carries the content, the sources, and the report. What sits on the chain is a verified document bundle: the document plus the machine-readable account of what the swarm said about it.
Human ratings sit beside the model ratings. The stores are reviewed periodically. The point of the bundle is that a later reader, or a later agent, can fetch the document and the judgment together, and can see that the judgment was not rewritten after the fact.
6. Applications
Swarm Explorer is the application you can use now. You ask a question. Many models answer. Judges cross-examine the set. You compare the answers and see the disagreement. It is live at swarm.verafy.ai.
A fact-checking extension applies the same procedure to a claim a person is looking at in the browser, rather than to a question they typed into the explorer. The check is the swarm, not a single model's refusal or endorsement.
Agent memory. The project notes argue that agent memory is becoming a frontier for data stored on chain, because agents will need trusted stores of factual data to operate. A verified bundle is that store: not a scratchpad the agent can silently edit, but a record it can cite. This is a use the design is aimed at. It is not a claim that every agent framework already reads Verafy.
Prediction. Where a question has a resolution, the same swarm can be pointed at a forecast. This is distinct from a prediction market. The market prices a contract. The swarm produces a compared, judged, inscribed answer. The project notes treat real-time prediction as one of the calls an agent should be able to make. Where that call is not already exposed inside Swarm Explorer, it is a direction, not a shipped market.
Search. A verified corpus, indexed, is a search index whose entries carry their judgment with them. The project notes describe search engines, and other readers, consuming inscribed and scored data. That is a direction. This paper does not claim that Verafy has replaced web search.
7. A broader thesis
7.1 AI consensus
Shared ledgers began as a way to agree on accounts: proof of work and Byzantine fault tolerance, then smart contracts and proof of stake. The project notes call the next stage AI consensus. Not one model declaring the truth, but many models, under a procedure, producing a record of agreement and dissent. The claim, stated directly in those notes, is that AI consensus will prove to be as important an innovation as Byzantine fault tolerance and smart contracts.
This paper adopts that claim as the thesis, and distinguishes it from a result the system has already proved. It is the reason to build the system. The alternative sketched alongside it is that one AI controls the truth: a single model, or a single company behind a model, becomes the place where questions go to die. Many models, cross-examining each other, with the transcript inscribed, is the other path.
There is a useful distinction between the two kinds of agreement. An account ledger wants one answer. The balance is a number, the cost of being wrong is high, and the consensus is exact. A ledger of claims and media is different. Images, sound, video, and prose can have many answers that are good enough. The cost of a small error is lower. The right consensus is often a threshold, not a bit-identical result. The notes call this heterogeneous consensus, set against the uniform consensus of pure cryptography. Verafy is aimed at the second ledger, the ledger of claims, where sufficient and checkable matters more than unique and exact.
7.2 Generative chains (exploratory)
This section is exploratory. It describes a research direction from the project notes. It is not a network Verafy is operating, and it is not the Swarm Explorer.
The notes introduce generative chains, written as (g) chains, as a further step. The idea is a form of proof of work in which models pose challenges and other models test whether a contribution is sufficient. Call that test proof of sufficiency. The worked example is a movie chain. A plot generator requests a frame. Generative nodes propose frames. A sufficiency test asks whether the frame fits the plot, matches the style, continues the last frames, and clears a quality bar. If it is sufficient, it is committed to the sequence. Media, not balances, is what the chain orders.
Read this as a sketch of where the thesis could go if AI consensus is real. It should not be read as a roadmap with dates, or as a description of software that is live.
8. Conclusion
Verafy exists because model answers move, and because a moving answer without a record is a claim you cannot audit. The design is a benchmark of hard questions, a wide set of models, judges that cross-examine those models, weights that are earned against hidden tracers, and an inscription of the result. Swarm Explorer is the way to use that today. The rest of the thesis, including generative chains, is marked where it is still a direction.
The limit is not a missing white paper. It is whether the record is actually kept, and whether anyone reads the disagreement instead of the fluent sentence. Try the swarm at swarm.verafy.ai.
9. Sources
This paper does not add measurements of its own. The history follows Rex St John, "Why I Built Verafy.ai" (10 October 2026). The design of the benchmark, the judges, the inscription, the document bundles, and the exploratory notes on generative chains follow the project materials from early 2025, including the description of Code-In by Zo of IQ6900. Illustrative figures in those sketches are not reported here as results.
Truth Chain demonstration, December 2024: x.com/rexstjohn/status/1871735098363760849.