I will deploy and host your llm or ai app on the cloud

Sri Lanka

I speak English

Senior Software Engineer POS, ERP and AI Powered Web Mobile Solutions

Senior Software Engineer building scalable POS, ERP, MVP, web & mobile apps with a focus on performance, security & maintainability. Skilled in microservices, API development, database optimization, s...
About this Gig

Need your LLM deployed to production, fast, optimized, and cost-efficient?


You're in the right place!


I deploy and configure open-source LLMs on GPU cloud servers and Kubernetes, with tuned inference engines for fast responses and lower GPU costs.


What I Offer:


  • LLM deployment (Llama, Mistral, Qwen, DeepSeek, Gemma)
  • Inference engine configuration (vLLM, TGI, Ollama, Triton)
  • Model quantization (AWQ, GPTQ, GGUF)
  • GPU optimization & cost reduction
  • OpenAI-compatible LLM API setup
  • Private & self-hosted LLM deployment
  • RAG & AI app backend deployment
  • Kubernetes LLM cluster with auto-scaling
  • Docker & CI/CD for LLM apps
  • Monitoring & performance tuning


Tech Stack:


Inference: vLLM || TGI || Ollama || NVIDIA Triton || TensorRT-LLM || llama.cpp || SGLang

Models: Hugging Face || Llama || Mistral || Qwen || DeepSeek

Cloud & GPU: AWS || Google Cloud || Azure || RunPod || Lambda Labs

Containers: Docker || Kubernetes || Helm

Monitoring: Prometheus || Grafana


Why Choose Me?


  • Fast delivery
  • Free consultation
  • Optimized for speed & cost
  • Full documentation
  • Post-deployment support


Let's get your LLM live today!


Tools:

Kubernetes

•

Docker

•

Amazon EKS

•

Google Kubernetes Engine

Frameworks:

Npm

•

Terraform

•

Ansible

•

Puppet

•

Crossplane

Cloud Provider:

Amazon Web Services

•

Microsoft Azure

Programming language:

Bash

•

C

•

Go

•

Java

•

JavaScript

•

Lua

•

PHP

•

Python

•

Ruby

Expertise:

Installation

•

Debugging

•

Configuration