I will turn your messy data into clean insights
About this Gig
Turn messy data into decisions you can actually act on.
I'm a software engineer and data scientist with a dual degree in Software Engineering and Business Analytics. I've spent my whole carrer building real data pipelines and machine learning models forecasting revenue for restaurant clients, building credit scoring systems, and diagnosing data quality problems that are hidden.
What I do:
- Clean and structure messy datasets (duplicates, missing values, inconsistent formats)
- Exploratory analysis to surface trends, patterns, and outliers
- Clear visualizations that make findings obvious, not confusing
- Predictive models (classification, regression, forecasting) with honest performance metrics
What makes this different:
I don't just hand you a chart and disappear. You get a plain-language explanation of what the data actually says, what it doesn't say, and what I'd be cautious about. If a model looks too good to be true, I'll tell you why.
Tools: Python, pandas, scikit-learn, XGBoost, SQL, Power BI, Jupyter
Message me before ordering with a description of your data and what you're trying to learn from it !
FAQ
What file formats do you accept?
CSV, Excel (.xlsx, .xls), JSON, XML, TSV, Parquet, SQL database exports, Google Sheets, and plain text files. If you have something else, just message me. I can usually work with it!
Is my data kept confidential?
Yes. I don't share, reuse, or store your data beyond the project. Happy to sign an NDA if needed.
What if my dataset is larger than the package limit?
Message me first and I'll send a custom offer based on the actual size and complexity.
Do I get the code, or just the results?
Results are always included. The full commented Jupyter notebook is available as an extra if you want to rerun or modify the analysis yourself.
Can you guarantee a specific model accuracy?
Accuracy depends on your data, and be cautious of anyone who promises a number before seeing it. That said, I've worked with genuinely difficult datasets (fragmented sources, missing values, imbalanced classes) and consistently pulled strong results out of them. I'll push for the best performance

