Notes

Owning the hardware

Documents

Why sensitive documents are processed on GPUs I control, and what that makes possible.

The question every prospective client asks in the first conversation, usually within ten minutes, is some form of “where does our data go?” For most AI vendors the honest answer is “to us, and then to a model provider.” My answer is different, and it starts with a rack.

What runs locally

I own GPU hardware, and the parts of the work that touch sensitive material run on it: optical character recognition of scanned documents, extraction of payroll and plan figures into tables, and any pass over records that might contain protected health information. The machines sit inside the client’s network boundary or on a private link to it. Nothing in that pipeline leaves for a hosted service.

That is not a preference for tinkering. Workers’ compensation and benefits administration produce records that HIPAA governs, and payroll detail is sensitive in its own right. A design that sends those records to a third party, however reputable, has a question at its center that never quite goes away. A design that keeps them on hardware I control has a plain answer.

What hosted models are for

This is not a rejection of hosted models. The strongest reasoning models are hosted, and for work on data that permits it, they are the right tool. The agent that connects a client’s stores and writes the final answer is usually one of them. The rule is about classification, not ideology: each kind of data has a boundary, and the model that touches it is chosen by which side of that boundary it sits on.

Getting that right depends on knowing what a corpus contains before it is moved. So the first thing that runs over a new set of documents is a survey that checks for the identifiers HIPAA lists and for health context around them. It is built so that no document text ever enters a language model’s context during the survey; it reports counts and categories, not passages. Anything it flags stays local.

What owning the hardware makes possible

Beyond the boundary question, local hardware changes what is affordable. Running OCR over ten years of scanned agreements is a one-time cost of electricity rather than a per-page bill. Re-extracting a payroll archive because the schema improved is an afternoon, not a budget line. Experiments that would be too expensive to justify against a metered API get run, and some of them turn out to matter.

It also removes a dependency. A model provider can change terms, prices, or availability. The open models running on my hardware do not, and the documents they process never leave.

The cost

The honest cost is that I maintain machines, and that open models trail the best hosted ones for hard reasoning. I accept both. The maintenance is my job, not the client’s. The reasoning gap is closed by routing: local models where the data demands it, hosted models where it allows it, and an agent that knows which is which.

“Where does our data go?” deserves a short answer. Owning the hardware is what makes a short answer possible.