journal

The research needed infrastructure


We started by trying to answer one question at a time.

The question that started it was ordinary enough: what does someone actually need in order to be onboarded as a Cargon agent in New York? Not the marketing version — the real version. Which licence classes count. What "active" means, and whether there is a grace period. Who has to sponsor whom. Which disclosures have to be handed to a consumer, and when.

Each of those is answerable. You read the regulator's page, you read the statute it points at, you write down what you found, and you move on.

Then you need the same answers for New Jersey. Then for the next state. Then someone asks whether the answer you wrote down four months ago is still true, and you discover that you cannot tell — because what you kept was a conclusion, not the evidence for it.

That is the moment the problem changes shape. It stops being a research problem and becomes an infrastructure problem.

What was actually wrong

The failure was not that the research was bad. Nationwide licensing and brokerage-onboarding research across all 51 jurisdictions produced 561 distinct findings, carrying 2,100 citations to 1,010 unique sources. As research, that is real work.

The failure was that the research had nowhere to live that could tell the truth about itself.

Three things kept collapsing into each other:

What we found. A regulator's page said something on a particular day.

What we think it means. Someone — or something — read that page and drew a conclusion.

What Cargon has actually decided to do. A policy the company will operate by, that a named human approved, that has an effective date and can be superseded.

Those are three different kinds of object with three different lifecycles, and in a document, a spreadsheet, or a model's context window they all look identical. They are all just sentences. The distance between them — which is the only thing that tells you whether you may act on something — is invisible by default, and it is the first thing to go.

So the second version of the question was not "what does New York require?" It was: where does an answer like that live so that, two years from now, anyone can see where it came from, how confident it was, who reviewed it, and whether Cargon ever approved acting on it?

That is Cargon Knowledge.

The four ideas it is built on

Research confidence is not counsel approval. A finding marked CONFIRMED means it was well-sourced. It says nothing about whether Cargon may act on it. Those are different claims and they are stored as different claims.

AI interpretation is not approved policy. An interpretation is an attributed opinion, carrying its author, its inputs and its date — never a fact the system speaks in its own voice.

Policy is one class of knowledge, not the spine. This one took a correction to get right. The first design routed everything through a legal-approval lifecycle, which is correct for regulatory requirements and actively wrong for everything else. Facts, public records, time series, entities and derived values are first-class and never forced through counsel review. A building's square footage does not need a lawyer.

Unknown, missing, expired, superseded or unreviewed policy fails closed. The honest answer is a structured refusal with the research attached. Never a guess.

There is a fifth idea underneath those, and it is the one that made the rest worth building: the model is not the memory. If every model Cargon uses were swapped out tomorrow, nothing the company knows may be lost and no approved policy may change. A model that has read the documents is not a company that remembers.

What exists today

I want to be precise here, because "we built a knowledge system" is the kind of sentence that can mean almost anything.

What runs today is a canonical store: 29 tables, standard PostgreSQL, no proprietary extensions. Historical records are append-only, enforced by a database trigger rather than by convention. Dataset 0001 — the nationwide licensing research — is ingested with zero byte drift across all 561 statements, and re-running the ingestion changes nothing.

On top of that there is a read-only, deterministic query layer with three distinct absence states, and an answer contract that assembles results with element-level basis and first-class citations.

And here is the part that matters most, which reads at first like a failure:

Zero of 561 claims have been reviewed by a human. Zero approved Cargon policies exist, in any jurisdiction.

Which means every operational question the system is asked today correctly refuses to answer.

That is the system working. A query that asks "may this person be onboarded in New York?" returns NO_APPROVED_POLICY and hands back the underlying research with its confidence and its sources attached. It does not helpfully summarise 11 New York findings into something that sounds like an answer. The refusal is the feature. The whole point of separating evidence from interpretation from policy is that the system is able to refuse — and a system that cannot refuse cannot be trusted when it doesn't.

Two things we learned by building it

The evidence was thinner than the citations suggested.

Dataset 0001 had strong source URLs and 2,100 citations. What it did not have was a copy of what those sources actually said on the day they were read. A citation to a regulator's page is a pointer, and pointers rot: pages change, get reorganised, and disappear. A knowledge system whose provenance is a list of URLs is a knowledge system that quietly becomes fiction.

So capture came before retrieval. 1,010 sources attempted, 595 captured, 115 megabytes, and — this part was not optional — zero access restrictions circumvented. Every refusal was recorded as a refusal.

There were a lot of refusals. Forty-one official government hosts actively restrict automated access, including state real-estate commissions, legislatures, and licence-lookup systems. Another 26 simply failed to respond. That is not a complaint about anyone's website. It is a fact about the domain, and it is the kind of fact you only learn by trying: a meaningful fraction of American regulatory source material is not durably archivable by a machine, which means the human capture queue is a permanent feature of this system rather than a temporary gap. Triage reduced that queue from 311 items to 122 by classifying every uncaptured source into an action class — but it does not go to zero, and pretending otherwise would be the same mistake as pretending a citation is evidence.

One discipline that came out of this and I think is right: when a blocked source had an alternate official endpoint — a state open-data CSV serving the same dataset as a blocked portal — we captured the alternate and left the original marked uncaptured. An alternate endpoint is a related source with its own identity. It is never a substitute for the citation. The temptation to mark the original green is exactly the kind of small dishonesty that makes a provenance system worthless.

We planned to add a vector database, and the evidence said don't.

Semantic retrieval was on the roadmap. Before building it, we measured what the existing structured and full-text queries actually did against the real corpus: 561 claims, 2,100 citations. Absence queries came back at 1.2 milliseconds. A topic across all 51 jurisdictions, 20.1 milliseconds. A deterministic alias resolver handled the synonym cases it was given — "moving to a new brokerage" resolving to affiliation_transfer — with no embeddings involved.

So semantic retrieval is deferred by evidence, not cancelled. No vector database, no embeddings, no model-vendor dependency added on the strength of having once been on a roadmap. The seam stays in place, and there are three written triggers that reopen it — a materially larger corpus, a measured latency budget being missed, or a benchmark of real questions showing insufficient recall.

I find this the most useful thing the build produced, and it has nothing to do with real estate. The roadmap was wrong in the direction roadmaps are usually wrong: it specified a solution to a problem we had not yet demonstrated we had.

Where it goes next

Two directions, and they are at very different stages.

The near one is review. The system holds unreviewed research and zero approved policy, and the next build is the graduation path: unreviewed research → reviewed interpretation → approved Cargon policy, with reviewer identity, effective dates, version history and supersession all preserved. The test is deliberately narrow — six claims out of 561, enough to answer one real onboarding question in one state — and the bar is that the same query flips from NO_APPROVED_POLICY to an approved answer without the answer contract being redesigned. If the shape has to change to accommodate its first real policy, the shape was wrong.

The further one is that this stops being about real estate.

Cargon Knowledge was deliberately built domain-neutral. Real-estate licensing is its first dataset, not its schema. The second domain is already visible: housing vouchers and housing programs are modelled — a voucher program as a first-class entity with identity, provenance and history; the authority hierarchy from federal down through state, county, city and public housing authority as layered scope, because "what is the payment standard?" is answered at different layers by different authorities and the layer is part of the answer.

Modelled, and not yet ingested. There is no voucher dataset in Cargon Knowledge today, and no claim about voucher rules may be made from it.

What makes that direction credible rather than aspirational is that the same pattern has already appeared independently. vouchers.cargon.io is a shipped product, and its jurisdiction and program rules are already configuration rather than code — payment standards, bedroom rules, required documents, expiration and extension rules, all data underneath a core that does not know New York from anywhere else. It arrived at that shape for its own reasons, before any of this existed.

Which is the actual argument. Two products, in two domains, both discovering that the hard part is not the software but the knowledge underneath it, and both needing that knowledge to carry its own provenance. That is what makes it infrastructure rather than a database.

The honest state

Nothing here is deployed. Nothing here is legal advice. Nothing in it may drive product behaviour, and nothing in it may be published. Zero of 561 findings have been reviewed; 271 open questions are waiting on counsel; 122 sources still need a human to capture them.

The version number is 0.1.0, not 1.0.0, and that is not modesty. It is queryable knowledge, not a validated release.

I am writing about it now because the interesting part is not the finished thing. It is the moment where you notice that the work you have been doing by hand, one question at a time, has quietly become a system you were building without admitting it — and you either go back and build it properly or you keep paying for it forever.

We went back.