I will set up, validate and troubleshoot your gpu cluster and ai infrastructure
Senior HPC and AI Infrastructure Engineer, GPU SLURM Linux
About this Gig
You have GPUs. You need them to behave like a reliable compute platform.
I build and run GPU and HPC infrastructure in production: provisioning, driver and CUDA stack validation, scheduling, high-speed networking and monitoring. I work at the infrastructure layer. I am not an ML engineer and will not pretend to be one.
WHAT I DO
GPU node bring-up: NVIDIA driver, CUDA toolkit and runtime alignment, persistence mode, MIG
Health validation: DCGM diagnostics, Xid and ECC error triage, thermal and power checks
Multi-node GPU communication: nccl-tests all-reduce benchmarking, NCCL tuning, network path validation
Scheduling: SLURM GPU partitions and GRES, cgroups and binding, or Kubernetes with the NVIDIA device plugin and GPU Operator
Monitoring: Prometheus, DCGM exporter, Grafana dashboards
Automation: Ansible roles so the build is repeatable
BEFORE YOU ORDER
Message me with node count, GPU model, interconnect and what you are seeing. If it is not something I can genuinely fix, I will say so.
You get working configuration, benchmark output showing what changed, and written documentation.
Server:
Database server
Operating system:
Linux
Other Support & IT Services I Offer
FAQ
Do you write CUDA kernels or train models?
No. I work on the infrastructure underneath: provisioning, drivers, scheduling, networking and monitoring. Model code and training scripts are outside my scope, and I would rather say that than take on work I cannot do well.
What access do you need to get started?
SSH with sudo on the nodes, and access to BMC or IPMI if hardware-level checks are needed. For the Basic diagnostic, read-only access is usually enough to start.
Do you work with cloud GPUs or only on-premise?
Both. Bare-metal on-premise clusters and AWS GPU instances.
Our GPUs are barely utilised. Can you find out why?
Usually yes. Low utilisation is normally scheduling, data loading, NUMA and affinity, or the interconnect. The Basic diagnostic is built for exactly this question, and you get the measurements behind the answer.
Do you support InfiniBand fabrics?
My high-speed networking experience is Ethernet and RoCE, not InfiniBand. If your fabric is InfiniBand I will tell you before you order rather than learn on your cluster.
2 reviews for this Gig
| (2) | ||
| (0) | ||
| (0) | ||
| (0) | ||
| (0) |
Rating Breakdown
- Seller communication level
- Quality of delivery
- Value of delivery
Sort By
A apreda1

United States
I hired him again he assisted me with my project very effectively. I definitely recommend him
$50-$100
Price
1 day
Duration
Helpful?A aimen39

Algeria
Best seller + communication + skills + knowledge
$50-$100
Price
1 day
Duration
H 
Seller's Response
Helpful?
2 reviews for this Gig
| (2) | ||
| (0) | ||
| (0) | ||
| (0) | ||
| (0) |
Rating Breakdown
- Seller communication level
- Quality of delivery
- Value of delivery
Sort By
A apreda1

United States
I hired him again he assisted me with my project very effectively. I definitely recommend him
$50-$100
Price
1 day
Duration
Helpful?A aimen39

Algeria
Best seller + communication + skills + knowledge
$50-$100
Price
1 day
Duration
H 
Seller's Response
Helpful?
