Solutions
On-Premises Generative AI
Agencies that cannot send a prompt outside their boundary need the model to come to them. Uniqcli sizes, quotes and delivers the in-boundary hardware and licensing that makes that possible — the models, the data and the guardrails stay yours.

- Category
- In-boundary generative AI infrastructure
- We quote
- GPU nodes, host platforms, memory, storage and OS subscriptions
- Sized against
- Model footprint, concurrency and the boundary it must stay inside
- Boundary
- We deliver infrastructure — models, prompts, data and guardrails stay yours
When the prompt cannot leave the building
A great deal of public-sector work involves information that cannot be sent to a service outside the accreditation boundary — not because a vendor is untrustworthy, but because the authorization, the data agreement or the classification does not permit it. The practical answer is to run the model where the data already is. That turns an abstract policy question into a concrete procurement: accelerators with enough memory to hold the weights, hosts and system memory to feed them, storage for models and retrieved content, and a rack that can take the power. Uniqcli sizes and supplies that infrastructure. Which model you run on it, how you retrieve context, and what you allow it to answer are decisions that belong inside your program, and we do not take them.
Three constraints that shape the build
Memory capacity comes first. The size of the model you intend to serve, at the quantization you intend to use, sets a floor on accelerator memory before anything else is negotiable — and the concurrent-session cache sits on top of that floor. Getting this wrong is the difference between a system that serves a department and one that serves a demo.
Isolation comes second. An in-boundary deployment usually means no vendor telemetry path, no cloud licensing callback and no automatic update channel, so the operating system, the driver stack and the entitlement model all have to work in an environment that cannot phone home. We quote subscriptions and support terms with that constraint stated rather than discovered at activation.
Growth comes third. Generative deployments almost never stay at their pilot size, and the second wave is much cheaper if the first one left PCIe slots, power headroom, rack space and network ports available. We quote the pilot with the expansion path written down, including what it would cost, so the follow-on is a purchase order rather than a redesign.
The in-boundary stack we can put on a quote
Hardware and licensing sourced through authorized US distribution, screened per line for TAA and NDAA §889. Naming a manufacturer describes the market, not a Uniqcli partnership or endorsement.
Accelerated nodes
NVIDIA data-center accelerators sized by memory capacity first, quoted against a chassis and power envelope that actually exists in your room. Single-node pilots and multi-node builds are quoted the same way — to the model footprint, not to a marketing tier.
Hosts, memory and storage
Lenovo host platforms with Micron registered ECC memory and the NVMe that holds model weights and retrieved content close to the accelerators, so a restart is a minute rather than an afternoon.
Operating system and support
Red Hat and SUSE enterprise Linux subscriptions, virtualization entitlements and vendor support terms quoted for environments without an outbound update path. Software, subscriptions and support entitlements are quoted and sourced through authorized US distribution rather than held on a shelf, so stock language does not apply to them.
What we do not do, stated plainly
We do not train or fine-tune models, we do not label or prepare data, and we do not build or operate the guardrail, retrieval or evaluation layer that decides what a system is permitted to say. We do not host anything on your behalf, and we do not operate a service you would consume — we source and integrate the equipment, and your team or your integration partner operates it.
We also do not make compliance claims on the strength of a purchase. FedRAMP authorization belongs to a cloud service offering, not to a server or to a reseller. An accreditation boundary is defined and defended by the organization that owns it. What in-boundary hardware does is remove one specific problem — data leaving the boundary to reach a model — and that is the honest description of the benefit.
Model and generative-platform software vendors do not carry priced rows in our catalog. Where a design calls for one, it is sourced on request through authorized US distribution and quoted as such, rather than described as something we shelve.
In-boundary AI questions
Can you tell us which model to run?
No, and any supplier who answers that question from a catalog should be treated with suspicion. Model selection depends on your task, your data, your evaluation criteria and your risk posture. Tell us the footprint of what you have chosen and we will size hardware that runs it with headroom.
What does a realistic pilot look like?
Commonly a single well-specified node with enough accelerator memory for the target model and a modest concurrency, plus local NVMe for weights and retrieved content. We quote it with the expansion path written down, so growing it later does not mean replacing it.
Does buying this hardware make us FedRAMP or CMMC compliant?
No. FedRAMP authorization applies to a cloud service offering, and CMMC certifies organizations rather than products. No purchase — hardware, license or service — confers either. What the equipment does is support an architecture your program has designed and will be assessed on.
Can this run in a disconnected environment?
That is exactly the case it is usually bought for, and it changes the licensing and update model rather than the hardware. We quote operating system and support entitlements with the disconnected constraint stated up front, so activation and renewal are worked out before delivery instead of during it.
How does this differ from UniQ AI?
This page describes vendor-agnostic infrastructure we quote for any in-boundary deployment. UniQ AI is our own in-enclave offering with its own scope and terms. If you are comparing the two, start with the model footprint and the boundary you have to stay inside, and the right path becomes fairly obvious.
Related AI and mission infrastructure
The solutions atlas
Every solution, one accountable partner.
UniQ platforms
By technology
By customer
- TAA & NDAA-889 Compliance Screening
- CMMC & CUI Solutions
- Federal & DoD
- State, Local & Education
- Healthcare
- Enterprise
- Rapid Procurement & GPC Buys
- Multi-Vendor Integration Projects
- eProcurement & Custom Catalogs
- FISMA Modernization
- CJIS-Compliant Justice Cloud & Local AI
- Federal Storage Modernization
- Government ERP & Business Systems Infrastructure
- Managed Procurement
- Secure AV & Conferencing
- Fiber Network Infrastructure
- Satellite & Resilient Connectivity
- Wavelength & Optical Transport
- Decentralized Data Centers
- Data Center Design & Build
Bring us the footprint, not the hype
Model size, quantization, concurrency and the boundary it has to stay inside. We will quote the in-boundary node and the licensing that runs on it.