I will quantize your ai model

K
kalebcadenhead
K
kalebcadenhead
Kaleb C

About this gig

Are your cloud GPU bills getting out of hand, or is your fine-tuned model too heavy to run on edge hardware?


I am a machine learning systems engineer specializing in low-bit quantization, inference optimization, and edge deployment. I take large language, vision, and mixture-of-experts (MoE) models and shrink their memory footprint by 50% to 80% without destroying accuracy.


Whether you need a 70B model squeezed onto consumer GPUs, a 7B running on an Apple Silicon Mac/iPhone, or a calibrated GGUF file for local corporate compliance, I will build the exact runtime pipeline you need.


What I Specialize In:

* Formats & Runtimes: GGUF (llama.cpp), MLX (Apple Silicon Mac/iOS), EXL2, AWQ, GPTQ, bitsandbytes (NF4/FP8).

* Model Architectures: Llama 3, Qwen 2.5 / 3.8, Mistral, Gemma, MoE models (OLMoE, Mixtral), and Vision-Language Models (VLMs).

* Target Hardware: Single NVIDIA GPUs (L4, A10, RTX 4090), Apple Silicon (M-series unified memory, iOS), and CPU inference.

* Precision & Quality:Advanced calibration (imatrix), activation-aware quantization, and deterministic quality verification (Perplexity & downstream evals).


Get to know Kaleb C

Kaleb C

Senior Data AI Leader

  • FromUnited States
  • Member sinceMay 2024
  • Languages

    English
I am a Senior Manager of Data & AI and Founding IT Leader with experience building data and IT functions from the ground up. I specialize in architecting scalable infrastructure, Snowflake data warehouses, and production AI agents that drive operational leverage. I thrive at the intersection of data engineering and applied AI to deliver reliable, HIPAA-compliant systems.

Other AI Development Services I Offer