This is my submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.
What I Built
Market Evidence Desk is a public research tool for questions where a confident answer can be misleading. It focuses on crypto investor protection: proof of reserves, custody, deposit insurance, stablecoins, and the difference between international recommendations and protections actually in force.
Consider a question that sounds simple: Can a proof-of-reserves report show that customers’ assets are safe?
Kraken describes how a customer can check whether their balance was included in a reserves snapshot. The PCAOB and SEC explain limits that such a snapshot may leave unanswered, including liabilities and what happens outside the snapshot. These publications address different parts of the question. Treating one as the complete answer would mislead a reader.
The desk connects each dated claim to its original publication and research question. Visitors can inspect those connections themselves, compare attributed positions, and ask an AI agent to draft an answer from the Sanity evidence.
As checked on October 1, 2026, the public published Sanity graph contains 101 research questions, 112 linked evidence claims, and 22 source records covering 21 distinct original URLs. The relationships matter more than the count: a reader can move from a question to each claim, then to its source, date, excerpt, and recorded review information.
Demo
- Open Market Evidence Desk
- Browse the evidence atlas
- Compare attributed claims
- Inspect the sources
- Ask the agent about a disagreement
- Watch the video walkthrough
A three-step test for judges
The public desk requires no sign-in.
- Open the AI agent disagreement case and press Ask the sources. The answer may take up to two minutes. Look at the cited publications and the Sanity tools reported for that run.
- Open Compare to see the positions attributed separately. Follow a claim into the research desk or sources page, then inspect its original publication.
- Try the agent’s live price case. The historical source collection has no live market feed. The answer should acknowledge that boundary instead of inventing a current Bitcoin price or a buy recommendation.
The guided comparison reads published Sanity records directly, so visitors can still inspect the evidence if the model service is temporarily slow. The research desk labels its fictional demonstration separately from published evidence.
Code
GitHub repository · Architecture, setup, and limitations
How I Used Sanity
I modeled the research as three linked document types in Sanity:
- A
sourcerecords an original publication, its URL, publication date, and scope. - A
marketEventrepresents a focused research question. - An
evidenceClaimconnects one question to one source. It records the claim, its stance (supports,conflicts, orcontext), its observation date, and its source-check fields.
This structure gives a claim two essential links: What question does it address? Which publication supports its wording?
The site assembles a dossier from those links and shows where sources discuss different scopes. A company description of a customer snapshot, for example, can appear as context beside regulatory cautions about broader assurance. The interface does not force them into a false yes-or-no vote.
Sanity Studio is the editing and review surface. The public site reads the published graph through a read-only server route. Claim cards can show an excerpt, its location in the publication, and a recorded editor source check. That check describes an editor’s attestation. It is not an independent audit of the publication or approval of everything the AI might say.
Question approval and claim source checks are displayed separately. For example, a dossier can show “Question approved · 3/3 claims source checked” while still telling readers to open the original publications and assess the limits of the evidence themselves.
Where Sanity Context enters
I pointed Sanity Context at the production dataset and a separate SEC investor-alert page to build the Market Evidence Desk Sourcebook Knowledge Base. Its dataset selection is bounded to fit the Knowledge Base document budget. The published dataset remains available to the agent through a Context MCP endpoint that supports GROQ queries.
For an agent question, the Python agent uses two Sanity Context tools:
-
groq_queryretrieves the linked, published question, claim, and source records. -
knowledge_base_readretrieves relevant Sourcebook entries and their source context.
The agent then drafts an answer that attributes positions and provides source links. The server checks that both required retrieval tools ran before returning a completed research answer. The public interface displays the tool names reported for the run. Model and organization credentials stay on the server.
Why the structure changes the answer
A keyword search for “proof of reserves safe” can retrieve several pages. By itself, it does not tell the agent which statement is a company’s account of its own process, which is a regulator’s caution, which research question each statement addresses, or whether a claim was checked against a specific passage.
The Sanity records supply those connections. GROQ follows a claim’s references to its question and source. The Sourcebook supplies navigable, source-linked context. The agent can use both to explain what each publication establishes and what it leaves open.
A second test concerns two SEC publications dated April 4, 2025. One describes a Division of Corporation Finance staff view of stablecoin issuer practices involving reserve reports. Commissioner Caroline A. Crenshaw challenges what such reports demonstrate. The desk identifies the speakers and keeps their positions distinct. It does not present either publication as a finding that a particular stablecoin is safe today.
Sanity Project Details
-
Sanity project ID:
cxjysvlq -
Dataset:
production - Public dataset: Query published source, question, and claim records
- Hosted Sanity Studio: phoniex-market-evidence-desk.sanity.studio
-
Sourcebook Knowledge Base ID:
kbu9WNgZ9ocF
The Knowledge Base is a selected index, while the public graph contains the published research records. Their counts can differ. Publishing a dataset change also does not, by itself, prove that a previously built Sourcebook entry has incorporated it.
Agent Session and Evaluation
I saved a public ten-question agent session with returned answers, reported Sanity Context tool calls, and criteria used to check each response. Eight questions passed the first run. Two failed, including a service error that I counted as a failure.
After changes to retries and source validation, I reran the same ten questions. All ten returned HTTP 200 on the first call and passed those original core criteria. This is a small recorded test, not a promise that an external model service will never fail.
The tests include source comparison, the SEC staff and Commissioner disagreement, and a live Bitcoin price and buy-advice question that the historical evidence cannot answer.
What I Learned
The difficult part was preserving the limits of each source. A dated investor alert, a company’s explanation, an international recommendation, and a Commissioner’s statement do not carry the same authority or answer the same question.
Sanity let me keep those distinctions in the content itself: separate records, explicit references, dates, stances, source locations, and review fields. Sanity Context then gave the agent a way to retrieve that structure and the Sourcebook context before writing.
The result is a research draft whose claims a reader can trace back to the original publications, challenge, and revise. Market Evidence Desk remains a historical evidence tool. It has no live price feed and does not issue trading decisions. Its AI answers are drafts for readers to check against the receipts.
Top comments (0)