Back to blog
AI Architecture

Private AI Infrastructure, Built Four Ways

Sympatric AI · Jul 22, 2026 · 14 min read

Four private AI deployment patterns — local workstation, owned server, rented regional GPU, and managed open-model endpoint — followed by a practical architecture guide covering control, context, inference, operations, model sizing, and cost.

The awe around chatbots and coding agents does wear off. A few uncanny answers become screenshots, then habits, then ordinary buttons in ordinary software. People stop asking whether a model can write a plausible memo or patch a function. They ask whether it can help finish Tuesday's work before Wednesday arrives.

That quieter phase is where company data enters the work. A useful request now carries material the company cannot casually replace or disclose. Its answer becomes input to the next decision, so the job continues after the chat window closes.

AI has moved inside the company.

Privacy has to travel with the work. A private model call means little if the source leaks during search or the answer is saved somewhere the company cannot control.

The best place to start is a blocked job. Find useful work that cannot move because its data must stay private or in one region. Then choose the simplest setup that lets people finish it safely.

Prepare the filing while the research is still unpublished

At 4:20 on a Thursday, a research lead is preparing evidence for outside patent counsel. One assay in the invention disclosure shows a response that is stronger than the value recorded in the protocol summary. The difference might be a harmless normalization choice. It might also change which result supports the proposed claim. Counsel's review starts Friday morning, and the underlying compound series has not been published. His job is to reconcile one result against the protocol and the lab notebook, work out what happened, and give counsel an account tied to the source files.

The assay export, signed protocol, and notebook pages already reside on an encrypted workstation approved for the research program. A local open-weight model and a small document index run on that same machine. Network rules prevent the workflow from reaching an external inference service, and the application writes its working files only to the encrypted project volume.

Private AI on a laptop or workstation — a person connected to a local model runtime on an encrypted workstation with local files inside a protected boundary.
Private AI on a laptop or workstation

He starts with the number, but the number does not settle anything. The local tool finds the reported value in the assay export and the lower value in the protocol summary, then cites both locations. It also pulls a notebook entry made two days after the protocol was signed. That entry records a plate-reader calibration failure and a repeated run at a different dilution. The mismatch came from the gap between the original plan and the experiment as performed.

Now the task has somewhere to go.

The research lead asks for an experiment timeline rather than another summary. The first entry is the approved protocol, with its planned concentration and analysis method. The second is the failed calibration check. The third is the notebook instruction to repeat the plate. The final entry is the assay export produced by that repeat. Each event includes a file name, page or row reference, recorded time, and a short note about whether it describes intent or execution.

He opens every citation. One timestamp came from the file's modification date rather than the notebook entry, so he corrects it in the template. Patent counsel needs the order of the research, not the order in which someone copied PDFs into a folder.

The corrected timeline becomes the input to an evidence table. Each proposed statement for the filing sits beside the experiment that supports it and the exact source. The stronger response is tied to the repeated run. The original protocol value remains visible as planned work. One open issue says the filing must explain the dilution change before comparing the measurements.

The model's role ends with a traceable packet built from scattered research records. The research lead and counsel make the claim decision with fewer hidden assumptions. When counsel asks why the filing uses the repeated result, the answer starts at the evidence-table row, follows it into the timeline, and ends at the notebook instruction and assay export.

The boundary never moves. The authorized workstation is the whole system, and updates are reviewed before installation on that machine.

The workflow fits because one specialist owns a bounded record and checks every important statement beside the source files. Careful review sets the pace. Concurrent demand never enters the problem.

By Friday morning, counsel receives an evidence table instead of a confident narrative assembled from memory. The filing discussion can focus on what the experiment supports. The unpublished record never had to leave the machine already trusted to hold it.

Restore the line before the next shift inherits the fault

The packaging line stops six minutes after a routine restart. Its controller reports an intermittent drive fault, but the alarm clears before the maintenance lead reaches the cabinet. Operators have seen brief versions of the same alarm twice during the week. Neither event lasted long enough to trigger the usual escalation, and the previous shift left its observations in a free-text handover note.

Production cannot resume on a guess. The maintenance lead needs to join the controller's alarm history with the equipment service manual and the previous shift's notes, identify a safe diagnostic sequence, and leave a record that keeps the next crew from repeating the investigation.

The plant is air-gapped. Inside its isolated network, one company-owned GPU server holds the approved open-weight model and a retrieval index for controlled maintenance material. The line application sends the three sources to that server over the plant network. Inference, document parsing, retrieval, and generated artifacts stay there.

An owned private AI server inside a company-controlled plant network, connected to internal tools, encrypted storage, and monitoring behind a private gateway.
An owned private AI server inside the plant network

The alarm history alone looks random. Once aligned with the handover note, a pattern appears: each fault followed a warm restart after the washdown cycle. The service manual says that this alarm can reflect a marginal encoder connection, but it requires the technician to isolate power and verify connector seating before testing resistance. The model produces a diagnosis packet with the relevant alarm times, the operator's observation about restart conditions, and the exact manual section governing the check. A missing fact stops the packet from becoming an instruction. The handover says that a connector was "checked," without naming the connector or stating whether the drive was isolated. The maintenance lead treats the note as an observation, not proof that the manual procedure was completed. He inspects the cabinet under the plant's lockout process and finds slight movement at the encoder connector.

That finding narrows the repair.

Using the cited manual sequence, he drafts a repair instruction in the maintenance application: isolate the drive, reseat and secure the encoder connector, inspect the cable strain relief, restore power, then run the prescribed low-speed verification cycle. The server checks the draft against the manual and points out that the verification cycle has a specified minimum duration. He adds it before releasing the instruction to the technician.

The technician completes the steps and records the measured result. The line runs through the low-speed cycle without another alarm, then returns to production under observation. The diagnosis packet has fed the repair instruction, and the completed instruction now supplies facts rather than prose for the handover.

The next shift receives a compact account: when the fault appeared, which sources supported the suspected cause, what physical condition was found, what procedure was followed, and how the line behaved afterward. The shift lead can open the alarm entries and manual passage directly. If the fault returns, the crew starts with a known repair state instead of another vague note that something was checked.

The accepted handover becomes a new maintenance record for the asset. It links the intermittent alarm signature to the secured connector and stores the verification result with the work order. Future retrieval can distinguish this confirmed event from earlier operator observations. The line's history has gained one trustworthy episode.

The owned server follows directly from plant operations. The site has no live path to an outside service, and maintenance work must continue under that constraint. Predictable plant use justifies hardware the company operates itself. Access follows plant roles, records stay inside the isolated network, and the server's model revision is pinned so a troubleshooting session can later be reconstructed.

Model and index updates arrive as signed offline bundles reviewed at the transfer station. A failed signature leaves the current approved revision untouched.

No fallback is hiding in that process. If the server is unavailable, the maintenance team uses the manual procedure without AI assistance. Sending plant records somewhere else would defeat the operating boundary that made the server appropriate.

When the next shift arrives, the line is running and the handover is specific enough to challenge. More important, the repair is no longer an isolated success in one technician's head. The alarm history led to a supported diagnosis, the diagnosis led to a controlled instruction, and the completed work became the record that improves the next response.

Decide the claim inside the approved region

A commercial property claim crosses two national operations. The insured is based in one country, the damaged equipment sits in another, and the policy restricts processing of the file to an approved region. The dispute turns on when the interruption became covered. A reviewer must eventually reconstruct that decision without gaining access to unrelated claims.

The claims application reaches one dedicated rented GPU through a private connection. Its storage and retrieval index remain in the contractually approved region, and claim requests cannot route to another endpoint.

A dedicated rented GPU server in an approved region, reached from a company gateway through an encrypted private connection.
A dedicated rented GPU in an approved region

The analyst begins by turning the file into a timeline. The application extracts dated events and keeps their sources visible. When two documents disagree about the first outage, the timeline shows both statements instead of quietly choosing one.

He asks an investigator to check the disputed time against the operating log. The investigator finds a time-zone notation, confirms the conversion with the site record, and sends the corrected event back with its source. The analyst adds it to the timeline. They can now answer the coverage question: did the interruption begin before or after the waiting period expired?

The timeline drives a search for earlier decisions about the same waiting-period clause. The analyst's access is checked before ranking. A close ruling from a restricted file stays out, while an allowed decision involving the same clause appears.

Permission checks belong at retrieval because a citation cannot be made safe after its text has already entered the prompt.

The analyst reads the permitted decision and sees an important difference. In the precedent, the interruption began when the equipment stopped. In the current claim, the equipment kept running at reduced output until the later shutdown. He adds the passage to the case record and writes the difference beside it.

The reviewer receives the decision through the claims application and works backward from the proposed outcome to the original file. The request record keeps the source map and the analyst's edits, giving him enough evidence to check the work without putting claim documents in general debugging logs.

The location is easy to prove as well. The request ID maps to the dedicated endpoint in the approved region, and the provider record shows that the attached storage stayed there for the run. If the GPU is unavailable, the AI-assisted step waits or the analyst works manually inside the claims system.

The insurer rents the regional GPU to obtain provable location without turning the claims team into a hardware operator. The request trace must still show that the real claim used the promised resource. A region name on an invoice proves little.

Once the reviewer accepts the decision, the explanation and its source map join the claim record. The insurer can later show how the file became a timeline, how the timeline found an allowed earlier decision, and how both supported the outcome. The same regional GPU handled the work from start to finish.

Resolve the seller dispute when demand refuses to be smooth

A flash sale produces a familiar marketplace mess. Orders surge, carrier scans arrive late, and sellers challenge automatic refunds issued under the delayed shipment policy. One seller submits a dispute with tracking evidence, a warehouse closure notice, and three affected orders. Thousands of similar disputes may arrive during the promotion, then volume may fall back by Monday.

The platform team is small, so the entire AI path runs through one managed private open-model endpoint in the approved region. The service isolates company traffic, retains no request content under the contract, and does not train on it. There is no second inference route.

A managed private open-model endpoint in an approved region, isolating company traffic from shared infrastructure.
Managed private open-model inference

The dispute first becomes a constrained case type: delayed shipment with a claimed carrier exception. That label is useful only because it selects the correct policy family and required evidence. The application validates the output against its schema; an invalid label stops for manual handling rather than wandering into a nearby refund policy.

Policy retrieval uses the seller's market, order dates, program tier, and the classified exception type. It returns the version that governed when the orders were shipped, plus the section describing acceptable proof of a carrier disruption. The warehouse notice is relevant context, but the policy requires a carrier event tied to each order. The difference would be easy to blur in a generic apology. The retrieved passage and case evidence then feed a draft resolution. For two orders, carrier scans show acceptance before the cutoff and a qualifying network delay, so the draft recommends reversing the automatic refunds. For the third, the first scan occurred after the deadline and the warehouse notice does not satisfy the carrier-evidence rule. The draft recommends keeping that refund in place and explains what evidence was missing.

A marketplace specialist reviews the recommendation. He accepts the first two outcomes and corrects the third after finding a carrier incident code in a structured field that the retrieval query missed. The application reruns policy retrieval with that evidence and produces a narrower revision. He accepts it, and all three orders are resolved in the seller's favor with citations to the policy version that applied.

The correction matters more than the first draft.

The company measures cost after the specialist accepts the dispute. Review time and corrections count alongside the managed inference bill. A cheap first answer can still be expensive when it sends the specialist into the wrong policy.

Elsewhere in the company, the same endpoint is doing completely different work.

A catalog operator receives a seller feed with product attributes shifted into the wrong columns. He asks the endpoint to map one broken row to the marketplace schema, checks the proposed changes against the seller's source file, and approves the repair. The corrected mapping then processes the rest of the feed. No dispute policy is involved, and no new AI service is created for the catalog team.

Later that afternoon, a finance analyst investigates a payout that does not match the seller's expected total. The endpoint pulls the fee rule that applied on the order date and explains the difference against the ledger entries. He checks the calculation, sends the explanation to seller support, and saves the reviewed case as a test for the next payout query.

The jobs are unrelated. Their infrastructure is shared.

During the sale, demand spikes sharply. The managed endpoint absorbs the burst under its reserved limits, while the application queues work when those limits are reached. A delayed dispute keeps the policy date it arrived with, so waiting cannot switch it to a newer rule. The platform team watches how long accepted work takes and how often people correct it.

One burst exposes a delay after an emergency carrier notice. Several drafts use the old version and specialists reject them. The team fixes how approved notices enter the index, reloads the notice, and saves those disputes as tests for the same endpoint.

The managed endpoint fits because demand jumps and falls, while the platform team stays small. The contract keeps request data private. The marketplace teams still own the evidence and final sign-off for their work.

By the end of the promotion, each team has a reviewed case it can reuse as a test. The endpoint is shared, but every team still defines what good work looks like.

What a private AI system actually contains

The four deployments differ in ownership and scale, but their basic architecture is similar. Each has a boundary around the data, a way to decide who may ask for what, a context layer that selects permitted source material, an inference runtime, and a record of what happened. The GPU sits near the middle of that chain. It is rarely the most difficult part.

A practical private system has four layers.

The control layer receives requests from people and applications. It authenticates the caller, applies role and application permissions, limits usage, chooses an approved route, and records which model version handled the request. In a shared deployment, this is usually an internal gateway rather than a collection of direct connections to the model server. The gateway also gives the company somewhere to disable a route, rotate credentials, or replace a model without rewriting every application.

The context layer prepares the material the model is allowed to see. It parses documents, retrieves relevant passages, checks source permissions before prompt assembly, and preserves citations or source IDs. This layer often takes more engineering than inference. Long context does not solve stale documents, missing provenance, or a search result that the user was never allowed to retrieve.

The inference layer holds the model and its runtime. Local tools such as llama.cpp, MLX, Ollama, or LM Studio are practical for one person and early tests. Shared production systems usually move to vLLM or SGLang because they expose standard APIs and handle batching, memory, and concurrent requests.

The vLLM scaling guidance is deliberately conservative: keep a model on one GPU when it fits, use tensor or pipeline parallelism only when memory or throughput forces the split. Every additional GPU and node adds coordination and recovery work.

The evidence and operations layer stores the minimum logs needed to reconstruct important work. Useful records include the caller, model revision, retrieved source IDs, route, latency, output status, and human acceptance or correction. Monitoring also needs queue depth, GPU memory, context use, and error rates. Backups, deletion, model updates, incident response, and an explicit manual fallback belong here. A private endpoint without these controls is only a model running somewhere the company pays for.

The same architecture can be read as one controlled path from request to evidence:

A unified private AI architecture showing people, applications, agents, gateway, context, audit, and model runtime layers.
One controlled path from request to evidence

Model size changes the engineering envelope

The smallest model that clears the company's own evaluation should win. Public benchmarks can narrow the field, but the acceptance test has to use real documents, real tool calls, and known failure cases from the workflow.

An 8B-class model is the practical local tier. A model such as Qwen3-8B, usually quantized for local use, can handle extraction, classification, first-pass document work, simple tool calls, and offline assistance on a capable laptop or workstation. It gives up quality on difficult reasoning and long multi-step work, but it is cheap to replicate and easy to keep close to sensitive files.

The 27B-to-32B tier is a sensible first shared model for many companies. Models in this range can support stronger document analysis, coding, structured generation, and internal agents. A quantized 32B model can fit on a 48GB professional GPU with careful context limits; FP8 or higher-precision deployments are more comfortable on an 80GB accelerator. Qwen's release of Qwen3.6-27B model is widely considered to be a great balance between compact size and intellectual power.

Large 70B+ models and frontier mixture-of-experts systems move the project into multi-GPU or managed-inference territory. Sparse models may activate only part of their weights for each token, yet the full weight set still has to live in memory. They make sense when an evaluation proves that the smaller tier cannot complete valuable work reliably. Starting there turns model ambition into cluster operations before the company has proved demand.

Context can change these sizing assumptions. A model that fits comfortably for 8,000-token requests may lose most of its concurrency at 128,000 tokens because the key-value cache consumes GPU memory. Retrieval, summaries, structured memory, and smaller task-specific packets are usually cheaper than sending an entire repository or policy library with every request.

The cost starts with compute and ends with ownership

Private AI can begin inexpensively when the company already owns a suitable machine. A new local system with enough memory for serious experimentation generally costs several thousand dollars. NVIDIA's DGX Spark, for example, provides 128GB of unified memory in a desktop system; it is useful as a local development box, though the ability to load a large model says little about interactive speed or multi-user capacity.

An owned production server moves into the low five figures once it includes a professional 48GB GPU, server chassis, memory, storage, warranty, and power protection. Higher-memory datacenter accelerators raise that figure quickly. Purchase price still omits installation, monitoring, security work, spare capacity, electricity, and the engineer who gets called when a driver update breaks serving.

Renting makes the first production estimate easier. Hugging Face's published dedicated endpoint prices list one 48GB L40S at $1.80 per hour and one 80GB A100 at $2.50 per hour on AWS. At continuous use, that is roughly $1,314 and $1,825 per month before storage, data services, support, and application engineering. An H100 endpoint at $4.50 per hour reaches about $3,285 per month. Scale-to-zero or scheduled operation can cut the bill sharply when the workload is intermittent.

Those figures explain why ownership follows utilization. A rented or managed GPU is often cheaper during discovery because the company can stop it, resize it, or abandon the model. Owned hardware becomes attractive when demand is steady, the data boundary rules out external infrastructure, or the organization needs a fixed platform for several workflows. Managed inference stays attractive when traffic is volatile and the internal team is small, even if its hourly compute price is higher.

The larger budget sits around the GPU. Identity integration, document ingestion, permission-aware retrieval, application changes, evaluation sets, logging, security review, and ongoing operations can cost more than inference. A small pilot may need one engineer for a few weeks. A shared production service touching regulated records can become a platform project involving infrastructure, security, data, application owners, and legal review. Hardware quotes are precise; organizational effort is where estimates usually fail.

Four practical build shapes

  • Build shape: Local encrypted workstation; Practical starting envelope: One expert, low concurrency, bounded files; Model fit: Quantized 8B; larger quantized models where memory allows; Cost shape: Existing machine or several thousand dollars once; Main operational burden: Device security, reviewed updates, weak central oversight
  • Build shape: Owned server on a private network; Practical starting envelope: Stable shared demand or a hard offline boundary; Model fit: 8B replicas or one quantized 27B-32B model; larger models need more GPUs; Cost shape: Low-five-figure capex upward, plus power and staff; Main operational burden: Patching, monitoring, capacity, hardware failure, offline update process
  • Build shape: Dedicated rented GPU in an approved region; Practical starting envelope: Pilot or steady regional workload without hardware ownership; Model fit: 27B-32B on 48-80GB; larger tiers on multi-GPU instances; Cost shape: Roughly hundreds per active month or $1,300-$3,300 for one continuously running GPU at the cited rates; Main operational burden: Provider evidence, storage policy, network isolation, availability
  • Build shape: Managed private endpoint; Practical starting envelope: Volatile demand and a small platform team; Model fit: Small replicated models through frontier open-weight models; Cost shape: Usage or reserved-capacity bill, usually with a management premium; Main operational burden: Contract terms, route control, retention proof, dependency on provider capacity

The safest rollout is smaller than the final diagram. Start with one workflow and a manually reviewed evaluation set. Test models on rented hardware before buying anything. Put the first shared model behind a gateway, connect only the sources the workflow needs, and record corrections. Add replicas when queues prove a throughput problem. Add larger models when failures prove a quality problem. Buy hardware when measured use and the required data boundary make ownership more practical than rental.

Private AI becomes a serious capability when the surrounding system can answer five questions: who used it, which information entered the request, where inference ran, which model produced the answer, and whether the work was accepted. The architecture should grow only when one of those answers, or the economics of providing it, requires another layer.

Want this in your company?

Book a 30-minute AI Readiness Call to see where to start.

Book an AI Readiness Call