Open server rack with GPU servers mid-installation and structured cabling

02 Build · Infrastructure You Own

GPU infrastructure built around your workload.

Peregrine designs, specifies, sources, integrates and commissions GPU infrastructure that your organisation owns — then hands it over, or operates it for you.

Lifecycle

From discovery to handover.

A fixed-price project with defined stages. Each stage produces documentation your team keeps.

  1. DISCOVERY

    Workloads, constraints, site

  2. ARCHITECTURE

    Compute, network, storage

  3. PROCUREMENT

    Specify and source

  4. COMMISSIONING

    Rack, power, cabling

  5. PLATFORM INSTALLATION

    Drivers, CUDA, K8s / Slurm

  6. TESTING

    Burn-in and validation

  7. HANDOVER

    Documentation or Operate

Deployment options

Where the infrastructure lives is your decision.

YOUR PREMISES

Infrastructure installed in your own facility where power, cooling and physical security requirements can be met.

AUSTRALIAN DATA CENTRE

Peregrine can arrange hosting in Australian data-centre facilities suited to GPU power densities.

COLOCATION

Your hardware in a colocation facility you select or we recommend, with Peregrine handling integration.

HYBRID

Owned infrastructure for steady workloads, complemented by Peregrine capacity or public cloud for peaks.

Network core rack with dense fibre bundles fanning into overhead trays

Potential components

Specified per workload, not from a catalogue.

Architecture is determined per workload. The components below are potential elements of a build — not a mandatory list — and are selected once requirements are understood.

Peregrine designs, specifies, sources, integrates and commissions infrastructure from established hardware vendors.

  • NVIDIA GPUs
  • CPU
  • RAM
  • NVMe storage
  • Shared storage
  • High-speed Ethernet
  • InfiniBand where required
  • Rack infrastructure
  • Power
  • Cooling requirements
  • Kubernetes
  • Slurm
  • Drivers
  • CUDA
  • Monitoring

Example reference architectures

Illustrative, not inventory.

These examples show how components can combine for common objectives. Actual node counts, accelerators and interconnect are determined by your Compute Blueprint.

Research GPU cluster and storage array in a university data room

RESEARCH CLUSTER

Example reference architecture for universities and research groups.

  • Multi-user scheduling with Slurm or Kubernetes
  • Shared high-performance storage
  • Quota and project-based allocation
  • Mixed training and experimentation workloads
Compact inference server racks in a modern colocation hall

INFERENCE CLUSTER

Example reference architecture for serving production models.

  • GPU nodes sized for model memory footprint
  • Model serving with autoscaling endpoints
  • Low-latency networking to applications
  • Observability for latency and utilisation
Liquid-cooled simulation cluster with coolant distribution unit

TRAINING CLUSTER

Example reference architecture for large training runs.

  • Tightly coupled GPU nodes with high-speed interconnect
  • Parallel file system or fast shared storage
  • Checkpointing and job orchestration
  • Power and cooling planned for sustained load

How it works

Five steps, one accountable partner.

  1. ASSESS

    Understand workloads and constraints

  2. DESIGN

    Architecture and specification

  3. DEPLOY

    Source, install, commission

  4. CONNECT

    Networking, storage, platform

  5. SCALE

    Expand or add capacity

Commercial model
Fixed-price project
At completion
HANDOVER or PEREGRINE OPERATE
You own
Hardware, software configuration, documentation
Optional
Managed AI Platform (Operate)

Discuss your GPU build.

Share your workload, site and timeframe. We will recommend a deployment option and scope a fixed-price project.