Buying ten thousand GPUs is the easy part. Rizzo Infrastructure designs, builds and operates large-scale GPU clusters for the companies that own them. Four NVIDIA data-center generations in production, configuration automation proven at 4,500 nodes, and an operations desk staffed every hour of the year.
Rizzo has operated owned bare-metal GPU compute since 2021, under economics where every minute of downtime is paid for directly.
Most operators pick one. We hold all three, because the handoffs between them are where large clusters go wrong: a topology nobody validated, a build nobody inspected, a runbook written by people who never touched the floor. One accountable owner from the first drawing to the last night shift.
Full cluster topology: intra-rack NVLink domains, external fabric to NVIDIA reference architecture, parallel storage layout, out-of-band management, four separated network planes, and a power and cooling strategy sized to the hall. Decisions live in version control, not in slideware, so the second site inherits the first.
Specification and vendor coordination, bill-of-materials discipline, and supervision of rack, stack, cabling and labeling against a written standard. Deviations raised the day they happen, not at handover. Every rack photographed, inventory reconciled, as-builts delivered, acceptance checklist signed jointly.
Continuous monitoring across GPU, interconnect, storage and fabric. Proactive fault detection, patching through image rebuilds rather than in-place drift, spares managed against observed failure rates, and an escalation path that ends with named senior engineers rather than a queue.
A service you cannot see is a service you cannot enforce. For every area below we state what happens, how often, who owns it, and the artifact you hold afterwards as evidence that it happened.
Topology, fabric design, storage layout, addressing, power and cooling strategy. Design freeze before hardware ships.
Specification and vendor coordination. Hardware invoiced to you directly at cost. No markup, no rebate, no undisclosed vendor relationship.
High-density power, interconnect and fabric cabling, labeling to standard, as-built documentation, and a signed acceptance checklist.
Bare-metal Linux from an image factory, firmware and BMC to a defined baseline, fabric bring-up, storage mount, identity and access.
Fleet-wide GPU health sweep, interconnect and collective benchmarks, storage throughput validation, measured against criteria agreed in advance.
Continuous telemetry, proactive fault detection, patch and CVE campaigns on an agreed cadence, monthly written reporting.
A dedicated always-on channel. The person who answers the page is the person who can fix it, with a senior bench on a short page behind them.
Spares depth modeled per component class by observed failure rate. RMA and hot-swap discipline, OEM warranty coordination, live pool reporting.
Node-level authentication and certificates, segmented network planes, access logging, and change made through code so it is reviewable and reversible.
You see the same instruments we do, not a curated extract. Dashboards, incident logs, utilization feeds, weekly digests at two levels.
Tooling and operational data remain the client's property. Runbooks, as-builts, inventory records and the infrastructure-as-code that provisions and rebuilds the fleet sit in repositories the client owns from day one, so continuity never depends on us.
Stated plainly so it can be checked. What carries from one generation to the next is operating discipline, automation and the depth of the bench. What does not carry is platform-specific, and we say which is which rather than let a logo do the arguing.
A twenty-node DGX P100 Kubernetes cluster on Flannel and Canal.
A thirty-node DGX A100 Kubernetes cluster. Separately, a DGX BasePOD backed by InfiniBand spine-and-leaf with dedicated redundant fabric managers.
OpenShift hosted control planes with MIG, vGPU and GPU passthrough, alongside a hundred-node NVIDIA GRID enterprise fleet.
HGX B200 nodes brought into live revenue-bearing inference service, run alongside a Hopper fleet of roughly 1,200 H200 GPUs across approximately 149 nodes.
That fleet was scored publicly and penalized automatically at the protocol level. We have been paid for paging discipline for years, under measurement we did not control. These are our own operating records, offered for diligence rather than as audited third-party statistics.
Coverage is a follow-the-sun roster across three desks in three time zones, with a one-hour overlap at every handoff so the outgoing pair walks the incoming pair through open tickets and anything still burning. No cold handoffs. Leave, sickness and training are covered from the wider bench rather than by leaving a desk thin.
The desk triages and resolves at Tier 2. Behind it, senior specialists sit on a short page, a dedicated automation engineer holds the toil down, and both founders are named on the first page of the runbook for any escalation that needs a decision above the desk.
Rack-scale GPU systems punish a slow response. Coolant leak, CDU failure, flow imbalance, delta-T drift: none of these are software problems, and none of them are learned on a live cluster. The mechanical depth on our bench is direct and on the record.
Rizzo's own production fleet is owned bare metal, operated by our engineers in a New Mexico colocation facility. Leak detection, CDU redundancy and failover get proven at commissioning, before a single GPU carries load, and the result goes into the acceptance package.
Not a team, not a function. Every area of an engagement has a single accountable seat. The engineers who fill them have worked together for decades, most of them out of the national-laboratory circuit, where response against a written service level is the daily condition of the job.
Single accountable owner: milestones, service-level performance, root-cause reviews, monthly reporting and the vendor chain end to end. Holds acceptance authority at handover.
The facility interface: mechanical loop, power chain, commissioning acceptance and build specification. Grew a two-person facility into a regional operator and served as CTO of the acquirer.
Physical build oversight, OEM and integrator coordination, cabling and labeling standards, as-built documentation and the acceptance package. Tier 3 on hardware integration.
Six to eight petabytes of production file, block and object storage, full data lifecycle. A multi-petabyte migration across the country. 4,500 nodes under automated configuration management. FedRAMP, NIST and DISA STIG hands-on.
Provisioning and platform architecture across four NVIDIA generations, with Base Command Manager and Warewulf. Standing Tier 3 authority on the platform layer.
Metrics, logs and traces off ten thousand GPUs. Kubernetes, infrastructure-as-code, bare metal, VMware, OpenShift and OpenStack at scale. Security+, Network+, A+.
The image factory and provisioning pipeline. STIG hardening automation to roughly 90 percent compliance in production, and ~700 systems deployed from a single standardized build.
Automates the switching layer across large Arista and Juniper estates: configuration generation, provisioning, validation and lifecycle. Ansible at scale on H100 clusters.
Access control and identity, segmentation, access logging, CVE and patch posture, secrets management, and audit support. FedRAMP, NIST and STIG environments.
Delivered performance per GPU, and the escalation point when nodes underperform the fleet rather than fail outright. Ran ~1,200 H200 GPUs under revenue-bearing inference load.
Model-serving and agentic deployments, orchestration troubleshooting, and bringing workloads into reliable production use. High-consequence secure and air-gapped systems.
Technical escalation of last resort and commercial accountability. Named on the first page of the runbook, reachable directly, and standing behind every priority-one event without being asked.
Physical hands at each site come from a small pre-badged on-call crew with call-out, backed by a contracted technical services partner holding Tier 1 to Tier 3 depth. That is deliberately not a fixed crew sized for the worst night of the year: a hot-spare pool turns most physical faults into scheduled next-day work, and the contracted bench absorbs the days when the work arrives all at once.
A capability statement that only says yes is not a capability statement. Here is where we are new, stated before anyone has to ask.
Rizzo, Inc. is a young entity and has not operated under a third-party service-level agreement. We therefore have no priority-one resolution statistics of our own, and we will not present estimates as measurements. What we do have is a bench whose engineers have spent careers running infrastructure to three and four nines. From day one of any engagement, response and resolution are timed automatically on instrumentation the client can see, and we publish the first ninety days as the baseline we are then measured against.
The liquid-cooling depth on this page belongs to our facilities lead and is real, but it is not something Rizzo's own production fleet exercises daily. That is why the mechanical interface and commissioning acceptance sit with one named owner alongside the facility team and the OEM, and why failover gets proven before load rather than discovered under it.
Every hardware generation brings a surface that is genuinely new: a fabric generation, a filesystem, a failure-domain arithmetic that changes spares depth and scheduling. We name those items at the start of an engagement, close each one with vendor professional services, a recruited specialist, or both, and say which one we think is most likely to bite. Guessing quietly is how large clusters stall.
Hardware, colocation agreements and client paper stay with the owner. We operate as authorized implementation agent and operator, with no title to equipment and no direct relationship with your end client.
One fee covering the full scope. No revenue sharing, no mobilization charge, no non-recurring charges. Any procurement assistance passes through at actual cost with no markup.
Operational records and infrastructure-as-code live in repositories you own, with step-in and successor rights, so continuity never depends on us being there.
Operational activities, cost invoices and incident records are open to audit on reasonable notice. We expect to be checked.
The first cluster stands up the operation. Every one after arrives into a running NOC, tooling stack, image factory and program office, and is priced on that basis by change order once the facility is known.
Where a counterparty is taking a risk on a new operator, we are willing to fund a performance security out of our own fee and have it released against measured performance rather than judgement.
Design, build and operations for large-scale GPU infrastructure. If you own the metal and need an operator who will be measured on it, start here.