I will do molecular docking, qsar modeling, and ml drug discovery
About this Gig
Accelerate your drug discovery pipeline with rigorous machine learning and structure-based computational analysis!
Are you working on lead optimization, target profiling, or bioactivity prediction? Standard machine learning models often fail to generalize because they rely on random splits that overfit to chemical series. I deliver end-to-end, publication-grade computational drug discovery workflows tailored to your biological target. By combining scaffold-aware machine learning (Random Forest, Extra Trees, SVR), explainable AI (SHAP), and molecular docking (AutoDock Vina), I help you identify, interpret, and prioritize high-potential candidate molecules while accounting for resistance mutations.
FAQ
What input data do I need to provide?
You can provide a CSV, Excel, or SDF file containing your compounds with SMILES strings and bioactivity values (like IC50 or pIC50), or simply give me your target protein name and I can curate data from ChEMBL for you.
Why use scaffold-aware splitting instead of a random train/test split?
Random splits often place chemically similar analogs in both sets, leading to falsely inflated accuracy. Scaffold splitting tests the model's true ability to generalize to entirely novel chemical structures.
Are docking scores treated as absolute experimental binding affinities?
No. AutoDock Vina scores and residue-contact analyses are used as robust structural ranking and prioritization signals to guide experimental lead selection.
What programming languages and tools do you use?
The pipeline is built using Python (scikit-learn, RDKit, SHAP) for machine learning and cheminformatics, alongside AutoDock Vina and Meeko for molecular docking.
What will I receive upon order completion?
You will receive well-documented Python scripts or Jupyter Notebooks, high-resolution figures (such as SHAP plots and docking poses), and a comprehensive summary report of the results.

