I will optimize and accelerate your ai model with onnx and tensorrt

R
riticpathania
R
riticpathania
Ricky

About this gig

Is your AI model accurate but too slow for real-world use?


I optimize machine-learning and computer-vision inference pipelines to reduce latency, improve throughput, and simplify model deployment on NVIDIA GPUs, edge devices, APIs, and production systems.


I can help with:


PyTorch model profiling and inference bottleneck analysis

PyTorch to ONNX export

ONNX Runtime optimization

NVIDIA TensorRT engine conversion

FP16 and INT8 optimization where supported

Latency, throughput, and FPS benchmarking

Output validation against the original model

Preprocessing and postprocessing optimization

Dynamic input and batch-size handling

Dockerized inference environments

FastAPI model-serving integration

NVIDIA Jetson / GPU deployment workflows

YOLO and computer-vision inference optimization


My focus is not simply converting a file format. I verify that the optimized model still produces correct outputs and measure the actual performance improvement.


Typical deliverables may include:

Optimized model / engine files

Python inference scripts

Before/after benchmark results

Output validation results

Setup instructions


Please message me before ordering.

Get to know Ricky

Ricky

High Performance Python and AI Engineer, APIs, Vision, Optimization

4.9(6)
  • FromIndia
  • Member sinceJul 2023
  • Last delivery2 weeks
  • Languages

    Hindi, English
I'm a Python and AI engineer focused on performance and production machine learning. I help clients solve technical problems such as slow Python/Pandas code, broken or incomplete FastAPI backends, custom YOLO computer-vision pipelines, and slow AI inference workflows. My work includes: - Python, Pandas, NumPy, Cython and C/C++ performance optimization - FastAPI and Flask backend development and debugging - REST APIs, databases, Docker and production deployment - YOLO object detection, tracking and computer-vision pipelines - PyTorch, ONNX and TensorRT inference optimization

My Portfolio