InfrastructureA division of Rizzo, Inc.

Someone has to run it.

Buying ten thousand GPUs is the easy part. Rizzo Infrastructure designs, builds and operates large-scale GPU clusters for the companies that own them. Four NVIDIA data-center generations in production, configuration automation proven at 4,500 nodes, and an operations desk staffed every hour of the year.

The operator

We have run this kind of infrastructure with our own money on the line.

Rizzo has operated owned bare-metal GPU compute since 2021, under economics where every minute of downtime is paid for directly.

4
NVIDIA generations in production
4,500
Nodes under automated management
24/7
Follow-the-sun operations desk
2021
Operating owned bare metal since
01 · The engagement

Design. Build. Operate.

Most operators pick one. We hold all three, because the handoffs between them are where large clusters go wrong: a topology nobody validated, a build nobody inspected, a runbook written by people who never touched the floor. One accountable owner from the first drawing to the last night shift.

Design

Full cluster topology: intra-rack NVLink domains, external fabric to NVIDIA reference architecture, parallel storage layout, out-of-band management, four separated network planes, and a power and cooling strategy sized to the hall. Decisions live in version control, not in slideware, so the second site inherits the first.

Build

Specification and vendor coordination, bill-of-materials discipline, and supervision of rack, stack, cabling and labeling against a written standard. Deviations raised the day they happen, not at handover. Every rack photographed, inventory reconciled, as-builts delivered, acceptance checklist signed jointly.

Operate

Continuous monitoring across GPU, interconnect, storage and fabric. Proactive fault detection, patching through image rebuilds rather than in-place drift, spares managed against observed failure rates, and an escalation path that ends with named senior engineers rather than a queue.

02 · Scope

Ten service areas, one accountable operator.

A service you cannot see is a service you cannot enforce. For every area below we state what happens, how often, who owns it, and the artifact you hold afterwards as evidence that it happened.

01

Solution architecture

Topology, fabric design, storage layout, addressing, power and cooling strategy. Design freeze before hardware ships.

02

Procurement support

Specification and vendor coordination. Hardware invoiced to you directly at cost. No markup, no rebate, no undisclosed vendor relationship.

03

Build & integration

High-density power, interconnect and fabric cabling, labeling to standard, as-built documentation, and a signed acceptance checklist.

04

Provisioning & bring-up

Bare-metal Linux from an image factory, firmware and BMC to a defined baseline, fabric bring-up, storage mount, identity and access.

05

Acceptance testing

Fleet-wide GPU health sweep, interconnect and collective benchmarks, storage throughput validation, measured against criteria agreed in advance.

06

Operations, 24/7/365

Continuous telemetry, proactive fault detection, patch and CVE campaigns on an agreed cadence, monthly written reporting.

07

Tier 2 and Tier 3 support

A dedicated always-on channel. The person who answers the page is the person who can fix it, with a senior bench on a short page behind them.

08

Hardware maintenance

Spares depth modeled per component class by observed failure rate. RMA and hot-swap discipline, OEM warranty coordination, live pool reporting.

09

Security & compliance

Node-level authentication and certificates, segmented network planes, access logging, and change made through code so it is reviewable and reversible.

10

Reporting & transparency

You see the same instruments we do, not a curated extract. Dashboards, incident logs, utilization feeds, weekly digests at two levels.

Tooling and operational data remain the client's property. Runbooks, as-builts, inventory records and the infrastructure-as-code that provisions and rebuilds the fleet sit in repositories the client owns from day one, so continuity never depends on us.

03 · The record

Four NVIDIA generations, in production.

Stated plainly so it can be checked. What carries from one generation to the next is operating discipline, automation and the depth of the bench. What does not carry is platform-specific, and we say which is which rather than let a logo do the arguing.

P100

A twenty-node DGX P100 Kubernetes cluster on Flannel and Canal.

A100

A thirty-node DGX A100 Kubernetes cluster. Separately, a DGX BasePOD backed by InfiniBand spine-and-leaf with dedicated redundant fabric managers.

H100 · H200

OpenShift hosted control planes with MIG, vGPU and GPU passthrough, alongside a hundred-node NVIDIA GRID enterprise fleet.

Blackwell

HGX B200 nodes brought into live revenue-bearing inference service, run alongside a Hopper fleet of roughly 1,200 H200 GPUs across approximately 149 nodes.

~1T
Tokens per week at peak, in and out
#1
Operator slot held against larger fleets
4,500
Nodes under automated deployment, national-laboratory scale
Live
GPU confidential computing and attestation in production

That fleet was scored publicly and penalized automatically at the protocol level. We have been paid for paging discipline for years, under measurement we did not control. These are our own operating records, offered for diligence rather than as audited third-party statistics.

04 · The desk

Two engineers live on every hour of the year.

Coverage is a follow-the-sun roster across three desks in three time zones, with a one-hour overlap at every handoff so the outgoing pair walks the incoming pair through open tickets and anything still burning. No cold handoffs. Leave, sickness and training are covered from the wider bench rather than by leaving a desk thin.

The desk triages and resolves at Tier 2. Behind it, senior specialists sit on a short page, a dedicated automation engineer holds the toil down, and both founders are named on the first page of the runbook for any escalation that needs a decision above the desk.

The tooling

  • Telemetry. DCGM and Fluent Bit on every node for deep hardware metrics and logs.
  • Pipeline. A buffering aggregation layer into ClickHouse, because scraping ten thousand GPUs through a single time-series database saturates the network.
  • Observability. SigNoz for dashboards, alerting and traces, exposed to the client rather than summarized for them.
  • Health. An agent that learns the normal signature of each GPU, NIC and optic and surfaces pre-failure patterns before they become outages.
  • Drift control. An automated auditor compares every live node against the gold configuration and reverts or raises unauthorized change. This is what prevents snowflake nodes.
  • Provisioning. Terraform, Ansible and Packer, with an image factory and a local package mirror so environments are reproducible and insulated from the public internet.
05 · Metal and mechanical

Liquid cooling is a mechanical system with its own failure modes.

Rack-scale GPU systems punish a slow response. Coolant leak, CDU failure, flow imbalance, delta-T drift: none of these are software problems, and none of them are learned on a live cluster. The mechanical depth on our bench is direct and on the record.

5 MW
Full design and engineering scope for a 4,000 sq ft hall at 110 kW per rack, covering liquid to the rack, CDUs, chillers and towers, A and B utility, switchgear, UPS, generator and ATS, transformers and 480 V distribution
2 MW
Vertiv installation running row-based CDUs into in-rack heat exchangers
500 kW
Direct-to-chip liquid cooling across twelve racks, tied into the building chiller plant
20 racks
Rear-door heat exchanger deployment, built and operated

Rizzo's own production fleet is owned bare metal, operated by our engineers in a New Mexico colocation facility. Leak detection, CDU redundancy and failover get proven at commissioning, before a single GPU carries load, and the result goes into the acceptance package.

06 · The bench

One name against every ask.

Not a team, not a function. Every area of an engagement has a single accountable seat. The engineers who fill them have worked together for decades, most of them out of the national-laboratory circuit, where response against a written service level is the daily condition of the job.

Program & delivery lead

26 years

Single accountable owner: milestones, service-level performance, root-cause reviews, monthly reporting and the vendor chain end to end. Holds acceptance authority at handover.

Data-center & facilities lead

25+ years

The facility interface: mechanical loop, power chain, commissioning acceptance and build specification. Grew a two-person facility into a regional operator and served as CTO of the acquirer.

Solutions & integration architect

12 years

Physical build oversight, OEM and integrator coordination, cabling and labeling standards, as-built documentation and the acceptance package. Tier 3 on hardware integration.

Principal DevSecOps & cloud architect

27 years

Six to eight petabytes of production file, block and object storage, full data lifecycle. A multi-petabyte migration across the country. 4,500 nodes under automated configuration management. FedRAMP, NIST and DISA STIG hands-on.

Principal platform engineer

30 years

Provisioning and platform architecture across four NVIDIA generations, with Base Command Manager and Warewulf. Standing Tier 3 authority on the platform layer.

Principal infrastructure engineer

19 years

Metrics, logs and traces off ten thousand GPUs. Kubernetes, infrastructure-as-code, bare metal, VMware, OpenShift and OpenStack at scale. Security+, Network+, A+.

Linux & automation engineer

10 years

The image factory and provisioning pipeline. STIG hardening automation to roughly 90 percent compliance in production, and ~700 systems deployed from a single standardized build.

Network automation engineer

5+ years

Automates the switching layer across large Arista and Juniper estates: configuration generation, provisioning, validation and lifecycle. Ansible at scale on H100 clusters.

Security, identity & compliance

12 years

Access control and identity, segmentation, access logging, CVE and patch posture, secrets management, and audit support. FedRAMP, NIST and STIG environments.

Senior AI & performance architect

10+ years

Delivered performance per GPU, and the escalation point when nodes underperform the fleet rather than fail outright. Ran ~1,200 H200 GPUs under revenue-bearing inference load.

AI workload & integration architect

10+ years

Model-serving and agentic deployments, orchestration troubleshooting, and bringing workloads into reliable production use. High-consequence secure and air-gapped systems.

Founders · CEO and CTO

25+ years each

Technical escalation of last resort and commercial accountability. Named on the first page of the runbook, reachable directly, and standing behind every priority-one event without being asked.

Physical hands at each site come from a small pre-badged on-call crew with call-out, backed by a contracted technical services partner holding Tier 1 to Tier 3 depth. That is deliberately not a fixed crew sized for the worst night of the year: a hot-spare pool turns most physical faults into scheduled next-day work, and the contracted bench absorbs the days when the work arrives all at once.

07 · Straight answers

What we don't claim.

A capability statement that only says yes is not a capability statement. Here is where we are new, stated before anyone has to ask.

We have no customer SLA history.

Rizzo, Inc. is a young entity and has not operated under a third-party service-level agreement. We therefore have no priority-one resolution statistics of our own, and we will not present estimates as measurements. What we do have is a bench whose engineers have spent careers running infrastructure to three and four nines. From day one of any engagement, response and resolution are timed automatically on instrumentation the client can see, and we publish the first ninety days as the baseline we are then measured against.

Our own fleet is air cooled.

The liquid-cooling depth on this page belongs to our facilities lead and is real, but it is not something Rizzo's own production fleet exercises daily. That is why the mechanical interface and commissioning acceptance sit with one named owner alongside the facility team and the OEM, and why failover gets proven before load rather than discovered under it.

Some things we plan for rather than assume.

Every hardware generation brings a surface that is genuinely new: a fabric generation, a filesystem, a failure-domain arithmetic that changes spares depth and scheduling. We name those items at the start of an engagement, close each one with vendor professional services, a recruited specialist, or both, and say which one we think is most likely to bite. Guessing quietly is how large clusters stall.

08 · How we engage

You own it. We run it.

You hold title

Hardware, colocation agreements and client paper stay with the owner. We operate as authorized implementation agent and operator, with no title to equipment and no direct relationship with your end client.

Flat monthly fee

One fee covering the full scope. No revenue sharing, no mobilization charge, no non-recurring charges. Any procurement assistance passes through at actual cost with no markup.

Your repositories from day one

Operational records and infrastructure-as-code live in repositories you own, with step-in and successor rights, so continuity never depends on us being there.

Audit rights retained

Operational activities, cost invoices and incident records are open to audit on reasonable notice. We expect to be checked.

Priced per additional site

The first cluster stands up the operation. Every one after arrives into a running NOC, tooling stack, image factory and program office, and is priced on that basis by change order once the facility is known.

We put money behind it

Where a counterparty is taking a risk on a new operator, we are willing to fund a performance security out of our own fee and have it released against measured performance rather than judgement.

Tell us what you're building.

Design, build and operations for large-scale GPU infrastructure. If you own the metal and need an operator who will be measured on it, start here.