I will deploy and manage mlops, llm and gpu infrastructure on aws and kubernetes
Senior AI DevOps Engineer: AWS, Kubernetes, ComfyUI APIs, DevSecOps
About this Gig
I put AI models into production and keep them running. A training script is not a product. A monitored, versioned, autoscaling service is.
Tools:
Kubernetes
•
Docker
•
Amazon EKS
Frameworks:
Terraform
•
Ansible
Programming language:
Bash
•
Go
•
Python
Expertise:
Installation
•
Migration
•
Debugging
Other DevOps Engineering Services I Offer
FAQ
Which models can you deploy?
Any open model on Hugging Face plus Llama, Mistral, Qwen, DeepSeek and your own fine tunes. For image and video I also handle Flux, SDXL and Wan. Tell me the model and I will confirm the GPU it needs.
Do I need my own cloud account?
Yes, and that is deliberate. Everything runs in your AWS, GCP or RunPod account so you own the infrastructure and control the billing. I set it up and hand it over.
What will it cost to run per month?
Depends on model size and traffic. Before you order I give you a sizing estimate: GPU type, expected utilisation and a monthly range. Spot instances and scale to zero usually cut it a lot.
Do I get the Terraform and Helm source?
Always. You get the full repository, Dockerfiles, Terraform or Helm charts and API docs, so you or another engineer can maintain it later.
Can you fix an existing setup instead of building new?
Yes. Broken deploys, OOM crashes, GPU nodes that will not schedule, runaway cloud bills and undocumented infrastructure are a large part of what I do.
Do you do self hosted LLMs on premise?
Yes. Ollama or vLLM with Open WebUI on your own server or GPU workstation, fully offline, with no data leaving your network.
What monitoring do I get?
Prometheus and Grafana with dashboards for latency, throughput, GPU utilisation, error rate and cost per request, plus alerts that reach you before your users notice.
What are your working hours?
I am available 24/7 and I work US hours, so we can talk in your timezone. I usually reply within an hour.
