Uniqcli

Solutions

AI Inference Servers

Inference is a throughput-per-watt purchase, not a training one. Uniqcli sizes the accelerator and the host against the model and the concurrency you actually serve, quotes the licensing that rides with it, and delivers the node burned in rather than boxed.

Sized against
Model footprint, batch size, concurrency and latency target
We quote
Accelerators, hosts, memory, NVMe and platform subscriptions
Boundary
We size and build the server; the model and its outputs stay yours
Overview

Serving a model is a different purchase from training one

Training rewards raw floating-point throughput and enormous interconnect bandwidth. Serving rewards something much less glamorous: enough accelerator memory to hold the weights and the KV cache for the concurrency you expect, enough memory bandwidth to keep tokens moving, and a power and thermal envelope you can actually pay for in a rack that already exists. A training-class node bought for an inference workload is usually two-thirds idle and entirely over budget; an undersized one falls off a cliff the moment a second user arrives. Uniqcli sizes the node to the workload you describe — model, quantization, batch and latency target — and quotes the accelerator, host, memory and storage together.

The economics

Throughput per watt is the number that survives contact with a budget

An inference deployment is judged over years, not benchmarks. Two nodes can serve the same model at the same latency while drawing very different power, and the difference compounds into the facility bill, the UPS runtime, the cooling load and eventually the rack count. That is why we ask for the concurrency and latency target up front: it turns a vague request for a GPU server into an arithmetic problem with a defensible answer.

The second economic lever is consolidation. Smaller quantized models often serve acceptably on far less silicon than the original request assumed, and a single well-specified node with headroom beats three underspecified ones on every axis a facilities team cares about. We will say so when the sizing points that way, even though it is a smaller quote.

The third is what the node runs. Operating-system subscriptions, virtualization entitlements and vendor support terms are real recurring lines, and we quote them alongside the hardware so the first renewal is not a surprise. Software, subscriptions and support entitlements are quoted and sourced through authorized US distribution rather than held on a shelf, so stock language does not apply to them.

Before it ships

Burned in, imaged and documented

A node that arrives as a pallet of parts costs a week nobody budgeted. We build the server, level the firmware across the fleet so every node in a group is identical, install the base operating system and driver stack against your build sheet, and run it under load long enough for infant-mortality failures to show up in our facility instead of yours.

What ships is a documented unit: asset tags and serials recorded, configuration captured, and the accelerator, driver and firmware versions written down. If you are adding to an existing fleet, we match the build rather than introducing a second variant — the quiet cause of most support tickets six months later.

  • Firmware leveled across the group before shipment
  • Base OS and driver stack installed to your build sheet
  • Loaded burn-in run before release, not a power-on test
  • Serials, configuration and version record delivered with the unit
  • Matched to an existing fleet build where one exists
Accelerated compute hardware staged in a build and integration environment
Accelerated compute hardware staged in a build and integration environment
Questions

Inference sizing questions

What do you need from us to size a node?

The model or model family, roughly how it is quantized, the number of concurrent sessions you expect at peak, and the latency you consider acceptable. If you also tell us the rack power available, we can rule out configurations that would not fit before we quote them.

Do you recommend a specific model or serving stack?

No. Model selection, serving framework and prompt handling are engineering decisions inside your program, and they change faster than any hardware quote. We size the infrastructure to what you tell us you are running and re-size it when that changes.

Can inference run on hardware we already own?

Often, yes — adding accelerators, memory or NVMe to existing hosts is a common and much cheaper path. We will quote the upgrade honestly, including the cases where the existing chassis power supply or PCIe topology makes it a false economy.

Is software licensing included?

It is quoted as its own lines — operating system subscriptions, virtualization entitlements and vendor support terms, sourced through authorized US distribution. We list them with the hardware so the recurring cost is visible before you commit rather than at the first renewal.

Ask AI about Uniqcli

AI Inference Servers

Tell us what you are serving

Model, concurrency, latency target and the power you have. We will size the accelerator and host against it and quote the node as one screened bill of materials.