I will optimize and accelerate your ai model with onnx and tensorrt


About this gig
Is your AI model accurate but too slow for real-world use?
I optimize machine-learning and computer-vision inference pipelines to reduce latency, improve throughput, and simplify model deployment on NVIDIA GPUs, edge devices, APIs, and production systems.
I can help with:
PyTorch model profiling and inference bottleneck analysis
PyTorch to ONNX export
ONNX Runtime optimization
NVIDIA TensorRT engine conversion
FP16 and INT8 optimization where supported
Latency, throughput, and FPS benchmarking
Output validation against the original model
Preprocessing and postprocessing optimization
Dynamic input and batch-size handling
Dockerized inference environments
FastAPI model-serving integration
NVIDIA Jetson / GPU deployment workflows
YOLO and computer-vision inference optimization
My focus is not simply converting a file format. I verify that the optimized model still produces correct outputs and measure the actual performance improvement.
Typical deliverables may include:
Optimized model / engine files
Python inference scripts
Before/after benchmark results
Output validation results
Setup instructions
Please message me before ordering.
Get to know Ricky
High Performance Python and AI Engineer, APIs, Vision, Optimization
- FromIndia
- Member sinceJul 2023
- Last delivery2 weeks
Languages
Hindi, English
My Portfolio
FAQ
What models can you optimize?
I primarily work with PyTorch models, including computer-vision and deep-learning models that can be exported to ONNX. Compatibility depends on the model architecture and operators used.
Will ONNX or TensorRT change my model’s predictions?
The goal is to preserve model behavior. I validate optimized outputs against the original model and report any meaningful numerical differences introduced by conversion or reduced precision.
What kind of speedup can I expect?
It depends on the model, hardware, batch size, preprocessing, and current bottleneck. I benchmark before and after optimization rather than promising a fixed speedup.
Do you support FP16 and INT8 quantization?
Yes, where the model and target hardware support them. INT8 may require calibration data and can introduce accuracy trade-offs, so I validate results before delivery.
Can you optimize YOLO models?
Yes. I can export and optimize YOLO models for ONNX/TensorRT workflows and help improve real-time image or video inference.
Can you deploy the optimized model?
Yes. Premium or custom projects can include FastAPI, Docker, NVIDIA Jetson, GPU server, or other deployment workflows.
What if my model cannot be exported to ONNX?
Some models contain unsupported or custom operators. I can investigate the export failure and, where practical, modify or replace incompatible components. Complex custom-operator work may require a custom offer.

