Private AI · Installed on site

AI that never leaves the building.

We install private AI inside businesses that cannot send client data to the cloud. The models run on hardware you own, the data stays behind your firewall, and the system keeps working when the internet does not.

Start a conversation See how it works

01Nothing confidential leaves.

Patient charts, case files, tax returns, closing packages. The models that read them sit in your server closet, on your network. There is no upload, so there is no vendor to trust, no retention policy to parse, and nothing for a third party to hold.

Confidential records sent off site: zero

02Routine work costs electricity.

Most office AI work is routine: reading, sorting, drafting, reconciling. Dedicated local models absorb roughly 60% of a typical workload, and once the hardware is yours, each task costs power instead of tokens. At enterprise document volumes, the build typically pays for itself in token savings inside three months, and we model that with your numbers before anything is bought.

Typical enterprise payback: under 3 months  ·  Marginal cost after: electricity

03An outage elsewhere changes nothing.

When a major AI provider goes down, every business built on it stops and falls back to manual. Your models and your data live on your machine. A bad day in someone else's data center is not a bad day in your office.

Dependency on cloud AI uptime: none

Thousands of firms resell cloud AI. Almost no one will put hardware in your building, sit beside your staff, and leave you owning the result. This work does not scale like software. That is why it holds up.

What this is, exactly.

No mystique. This is a specific, buildable system with six parts, and you should know what each one is before anyone quotes you a number.

The node

A dedicated inference machine, GPU workstation to small-server class with 64 to 128GB of memory depending on workload, installed on your network in your own rack or closet. It serves models over your LAN only. You own it outright.

The models

Open-weight language models: the class released with downloadable weights that can legally run on hardware you control. We select and tune per workflow, quantize to fit the node, and serve them behind a local API. No calls leave the building for this tier.

The agents

Small programs, one per workflow, that connect the models to the systems you already run: practice management, document management, accounting, email. They read from and write to your systems through their existing interfaces. Nothing about how your office works gets replaced.

The gates

Each workflow is paired with a deterministic acceptance test that runs outside the model: the reconciliation balances to the cent, the cited figure traces to a source page, the code matches the chart. Plain code, not another AI's opinion. Output either passes or comes back flagged, and the model has no write access to the files that grade it. That is enforced in the tooling, not in a policy document.

The escalation path

A small, defined set of task categories may call a frontier model when local capability is not enough. A sanitization step strips client identifiers before anything leaves. Which categories, and what gets stripped, is a signed document you approve before the system runs.

The record

Every run is logged on your own systems under your own retention policy: which records were touched, by which process, what passed, what failed. When a stronger open model ships, it is benchmarked against your gates before it replaces anything.

Some operations are not allowed to stop.

The moment AI enters the operational path, its availability becomes your availability. If a provider outage can halt the work, the risk has been imported. These are the places where that is unacceptable.

Hospitals and clinics

Clinical documentation, coding support, and discharge summaries cannot pause mid-shift. Downtime procedures exist because downtime hurts patients.

Government and public safety

Dispatch support, benefits processing, records requests. Public services carry continuity requirements that a third-party dependency quietly violates.

Banks and credit unions

Fraud screening and compliance monitoring run continuously or not at all. An outage during business hours is direct exposure.

Investment and advisory firms

Client books, review prep, and compliance files live under SEC and FINRA duties. Confidentiality is not optional, and quarter-end does not move because a vendor had an incident.

Pharmacies and labs

Interaction checks and result processing sit inside patient safety. They cannot be waiting on a status page.

Legal, title, and closings

Wire deadlines and filing dates are hard-edged. A closing does not reschedule because the cloud did.

On-premises systems remove the imported risk. Your uptime is decided in your building, by hardware you own.

Boutique, by design.

Every deployment is built for one business: your systems, your documents, your definition of correct. Nothing here is a template, and nothing is built without your sign-off.

STEP 1

We watch

We sit with your team and map how the work actually moves: what gets retyped, what waits on one person, where the hours disappear.

STEP 2

We scope

You approve which workflows move over and what correct output means for each one. That definition becomes the standard every build is tested against.

STEP 3

We build

Hardware goes into your building. Each workflow is built against your real systems and proven on your real documents before it goes live.

STEP 4

We hand over, and stay

Your team runs it. We stay on call while they settle in, and when your staff wants new agents for new workflows, we help them build those too.

The system is yours to grow. Because it lives on your hardware and is shaped around your processes, extending it is natural. Many teams start with two workflows and are running six within the year, most of them proposed by their own staff.

Deep engineering, for the hard cases.

Behind every deployment is a research-grade engineering bench that ships local AI in constrained environments most vendors never touch. Given the time, there is no workflow too tangled to build: multi-system, multi-step, decades of accumulated process included.

Based in Austin. Deployed anywhere.

Kept On Site works from Austin, Texas. Central Texas is home ground, and when the work justifies the trip, we get on the plane. The deployment happens in your building either way, because that is the entire point.

What it costs.

Published ranges, not quotes. These are estimates, calibrated to a specific shape of business described below, and the exact number comes out of the assessment in writing before anything is built.

Assessment

Begins after a complimentary intro call, once payment is received. On site, mapping how the work actually moves, with findings presented to your team within 14 to 30 days: workflow map, hours, gate specifications, and a build plan with fixed pricing. Yours to keep whether or not you continue. Complex regulatory environments, law in particular, run $7,500.

$6,500Fixed

Build

Scoped per workflow from the assessment, so the number is known before anything starts. Includes the agents, the gates, integration with your systems, staff training, and an acceptance run on your real documents.

$28k to $65kPer scope

Hardware

The inference node, at cost plus 20%. Cloud AI is a rental: pay forever, own nothing at the end. This one is yours: a depreciable asset on your books that stays if we ever part ways.

$6k to $14kOnce, you own it

Ongoing

Monitoring, gate maintenance, model updates re-tested against your gates before promotion, and a monthly report of runs, passes, failures, and hours recovered.

$2,400 to $6,000Per month

Who these numbers fit

An 8 to 25 person professional office with steady document volume: a CPA firm through filing season, a title company closing weekly, a law firm with active matters, an advisory firm with a full client book. Typically three to five workflows on a single node at one location.

It runs lower when

Fewer workflows, modern systems with clean interfaces, a single location, lighter regulatory overhead. A two-workflow build for a small insurance agency can land under the published floor.

It runs higher when

More workflows, legacy or fragmented systems that need custom integration, air-gap requirements, multiple entities or locations. Hospital systems and government agencies are a different scale entirely and are quoted separately.

Every number here is an estimate until the assessment turns it into a fixed quote in writing. At enterprise document volumes, token savings typically cover the build inside three months, modeled with your numbers before you commit. If your volume is light, the math may not work, and the assessment will say so plainly.

Trust is earned in writing.

You found us on the internet and we are asking to put hardware inside your firm. Here is how that gets earned without asking you to take anyone's word.

01

The assessment is yours either way

Fixed fee, delivered as a written document: the workflow map, the numbers, the plan. If you take it to another vendor or do nothing at all, it was still worth having.

02

Nothing signs off until it passes your test

Every workflow ships with an acceptance check you approved, run on your real documents. The system does not grade its own work, and neither do we.

03

Our lab publishes its failures

The engineering bench behind our deployments logs its runs publicly, failed runs included, and open-sources the evaluation harness. Ask for the log and we send it.

04

The boundary is a document, not a promise

Exactly which categories of work may escalate off site, and what gets stripped first, written down and signed before anything runs.

05

We say no when it is not a fit

A small office with light volume may never pay back the hardware. When the math does not work, we tell you, because a bad-fit deployment costs us more than the invoice earns.

Start with a conversation.

Curious what this would look like in your building? Reach out. Every engagement begins with a complimentary 30-minute call where we answer your questions, no obligation and no pitch.

  1. 01Complimentary 30-minute call. You bring the questions, we bring straight answers. If it is not a fit, we say so and that is the end of it.
  2. 02The assessment. If you decide to move forward, the audit begins once payment is received, and findings come back within 14 to 30 days.
  3. 03The presentation. We present the findings to your team. Move ahead with a build, or keep the insights and run with them yourself. Both are fine outcomes.

No sales sequence. A person replies within two business days to set up the call.