Uniqcli

Solutions

On-Premises Generative AI

Agencies that cannot send a prompt outside their boundary need the model to come to them. Uniqcli sizes, quotes and delivers the in-boundary hardware and licensing that makes that possible — the models, the data and the guardrails stay yours.

Category
In-boundary generative AI infrastructure
We quote
GPU nodes, host platforms, memory, storage and OS subscriptions
Sized against
Model footprint, concurrency and the boundary it must stay inside
Boundary
We deliver infrastructure — models, prompts, data and guardrails stay yours
Overview

When the prompt cannot leave the building

A great deal of public-sector work involves information that cannot be sent to a service outside the accreditation boundary — not because a vendor is untrustworthy, but because the authorization, the data agreement or the classification does not permit it. The practical answer is to run the model where the data already is. That turns an abstract policy question into a concrete procurement: accelerators with enough memory to hold the weights, hosts and system memory to feed them, storage for models and retrieved content, and a rack that can take the power. Uniqcli sizes and supplies that infrastructure. Which model you run on it, how you retrieve context, and what you allow it to answer are decisions that belong inside your program, and we do not take them.

What in-boundary really requires

Three constraints that shape the build

Memory capacity comes first. The size of the model you intend to serve, at the quantization you intend to use, sets a floor on accelerator memory before anything else is negotiable — and the concurrent-session cache sits on top of that floor. Getting this wrong is the difference between a system that serves a department and one that serves a demo.

Isolation comes second. An in-boundary deployment usually means no vendor telemetry path, no cloud licensing callback and no automatic update channel, so the operating system, the driver stack and the entitlement model all have to work in an environment that cannot phone home. We quote subscriptions and support terms with that constraint stated rather than discovered at activation.

Growth comes third. Generative deployments almost never stay at their pilot size, and the second wave is much cheaper if the first one left PCIe slots, power headroom, rack space and network ports available. We quote the pilot with the expansion path written down, including what it would cost, so the follow-on is a purchase order rather than a redesign.

Where the line is

What we do not do, stated plainly

We do not train or fine-tune models, we do not label or prepare data, and we do not build or operate the guardrail, retrieval or evaluation layer that decides what a system is permitted to say. We do not host anything on your behalf, and we do not operate a service you would consume — we source and integrate the equipment, and your team or your integration partner operates it.

We also do not make compliance claims on the strength of a purchase. FedRAMP authorization belongs to a cloud service offering, not to a server or to a reseller. An accreditation boundary is defined and defended by the organization that owns it. What in-boundary hardware does is remove one specific problem — data leaving the boundary to reach a model — and that is the honest description of the benefit.

Model and generative-platform software vendors do not carry priced rows in our catalog. Where a design calls for one, it is sourced on request through authorized US distribution and quoted as such, rather than described as something we shelve.

Questions

In-boundary AI questions

Can you tell us which model to run?

No, and any supplier who answers that question from a catalog should be treated with suspicion. Model selection depends on your task, your data, your evaluation criteria and your risk posture. Tell us the footprint of what you have chosen and we will size hardware that runs it with headroom.

What does a realistic pilot look like?

Commonly a single well-specified node with enough accelerator memory for the target model and a modest concurrency, plus local NVMe for weights and retrieved content. We quote it with the expansion path written down, so growing it later does not mean replacing it.

Does buying this hardware make us FedRAMP or CMMC compliant?

No. FedRAMP authorization applies to a cloud service offering, and CMMC certifies organizations rather than products. No purchase — hardware, license or service — confers either. What the equipment does is support an architecture your program has designed and will be assessed on.

Can this run in a disconnected environment?

That is exactly the case it is usually bought for, and it changes the licensing and update model rather than the hardware. We quote operating system and support entitlements with the disconnected constraint stated up front, so activation and renewal are worked out before delivery instead of during it.

How does this differ from UniQ AI?

This page describes vendor-agnostic infrastructure we quote for any in-boundary deployment. UniQ AI is our own in-enclave offering with its own scope and terms. If you are comparing the two, start with the model footprint and the boundary you have to stay inside, and the right path becomes fairly obvious.

Ask AI about Uniqcli

On-Premises Generative AI

Bring us the footprint, not the hype

Model size, quantization, concurrency and the boundary it has to stay inside. We will quote the in-boundary node and the licensing that runs on it.