I will deploy self hosted llm, rag, ollama, vllm, ai agents on vps or gpu server
AIOps And Linux DevOps Engineer
Level 1
Has met certain performance criteria and shows strong potential in the marketplace.
About this Gig
671 five-star reviews · 600+ orders · 13+ years in production infrastructure
I deploy open-source LLMs and AI systems on your own servers, so you stop paying per token and keep your data in-house.
What I deploy
- Ollama, vLLM, llama.cpp, OpenWebUI
- Llama, Mistral, Qwen, DeepSeek, Gemma
- RAG with Qdrant, Chroma, Weaviate, pgvector
- AI agents via n8n, LangChain, LlamaIndex
Where I deploy it
- Your VPS: Hetzner, DigitalOcean, Contabo, OVH
- GPU servers and bare metal: A100, L40S, RTX 4090
- AWS, GCP, Azure, RunPod
- On-premise and air-gapped setups
How it works
- You send your server specs and the model you want
- I confirm what your hardware can realistically run
- I deploy, tune and document it
What you get
- Docker or systemd deployment, reproducible
- Nginx reverse proxy with SSL and API authentication
- GPU and VRAM tuning so the model runs fast
- Monitoring, logging and auto-restart on failure
- Written handover documentation
Track record
- Nobel Prize Organization, Starbucks, T-Mobile, Beasley Media
- Platforms serving 45,000 concurrent users at 40 Gbps
Message me before ordering and I will tell you honestly if your hardware can handle it.
Tools:
Kubernetes
•
Docker
•
Amazon EKS
•
Google Kubernetes Engine
Frameworks:
Terraform
•
Ansible
•
Puppet
Programming language:
Bash
•
Python
Expertise:
Installation
•
Migration
•
Debugging
My Portfolio
FAQ
What hardware do I need to run an LLM?
It depends on the model. A 7B or 8B model in quantised form runs on a 16GB GPU, or even CPU-only with enough RAM. A 70B model needs 48GB+ of VRAM or multiple GPUs. Send me your server specs before ordering and I will tell you exactly what will run well on it.
Can you deploy on my existing VPS, or do I need a GPU server?
Both work. Smaller models and RAG setups run fine on a standard CPU VPS. If you need larger models or fast response times, I can help you choose a GPU host such as RunPod, Hetzner or Lambda, or deploy on hardware you already own.
Will my data stay private?
Yes. Everything runs on your own infrastructure. No prompts, documents or embeddings leave your server, which is the main reason most clients move off hosted APIs. Air-gapped and on-premise setups are supported.
2 reviews for this Gig
| (2) | ||
| (0) | ||
| (0) | ||
| (0) | ||
| (0) |
Rating Breakdown
- Seller communication level
- Quality of delivery
- Value of delivery
Sort By
L lara_chad2

United Kingdom
Excellent service! The deployment was completed quickly and professionally. The seller understood the issue, handled the Dockerized Python AI application deployment smoothly, and kept me updated throughout the process. Everything is working as expected. Highly recommended!
$50
Price
1 day
Duration
Helpful?D drezin

United States
Ongoing collaborationShazeb was great in understanding the problem and working to fix It. would recommend.
$100-$200
Price
2 days
Duration
Helpful?
2 reviews for this Gig
| (2) | ||
| (0) | ||
| (0) | ||
| (0) | ||
| (0) |
Rating Breakdown
- Seller communication level
- Quality of delivery
- Value of delivery
Sort By
L lara_chad2

United Kingdom
Excellent service! The deployment was completed quickly and professionally. The seller understood the issue, handled the Dockerized Python AI application deployment smoothly, and kept me updated throughout the process. Everything is working as expected. Highly recommended!
$50
Price
1 day
Duration
Helpful?D drezin

United States
Ongoing collaborationShazeb was great in understanding the problem and working to fix It. would recommend.
$100-$200
Price
2 days
Duration
Helpful?
