I will deploy your llm for production with vllm on your cloud GPU

I
ike_kolawole
I
ike_kolawole
Joshua

About this gig

A model that works in a notebook and one that survives 1,000 concurrent users are two different projects. I do the second one.


I'm a senior ML-infrastructure engineer, 7 years of production reliability underneath. Most recently: serving infrastructure handling 50,000+ concurrent tasks at sub-100ms for 1,000+ users, built on vLLM and SGLang with GPU Kubernetes node pools, at 35% lower infra cost than when I started.


What you get:

  • Your open-weight model (Llama, Mistral, Qwen, DeepSeek) served with vLLM
  • Deployed in YOUR cloud account: you keep the keys, data, and endpoint
  • GPU autoscaling and batching tuned for inference, not web traffic
  • Load-test results, so you know real capacity before your users find it
  • Monitoring dashboards; runbook on Standard/Premium

If your GPU budget and latency target don't match, I'll say so before you spend a cent on compute.


Basic: one model, one GPU node, load-tested.

Standard: adds autoscaling, tuning, dashboards.

Premium: multi-model, cost pass, handover.


Send your model and traffic details; the requirements form covers the rest.

Get to know Joshua

Joshua

Senior Platform Engineer

  • FromNigeria
  • Member sinceJul 2026
  • Languages

    English
Deployments on my platforms are boring, and that's the point. Seven years of production infrastructure: multi-cloud Kubernetes (EKS/AKS), Terraform at 20+ module scale, CI/CD with automatic rollback at 99.9% deploy success. One team went from 40% failed deploys to 5% after I rebuilt their pipelines; a fintech's PCI-DSS audits returned zero critical findings. I also build ML serving infrastructure: 50,000+ concurrent tasks at sub-100ms using vLLM on GPU node pools. Hand me a breaking pipeline, an untrusted cluster, or a cloud bill that grew quietly, and get it fixed properly.

My Portfolio