Amir Saleem

Amir Saleem

Data Scientist / ML Engineer | 8 Years Data Engineering, ML & LLM Fine-Tuning

8 yrs experience · Islamabad, Pakistan · <10 hrs/week

About

Data Scientist / ML Engineer with 8 years of experience building scalable ETL pipelines and production ML systems. Specialized in LLM fine-tuning, cloud (AWS/GCP), and large-scale data processing, delivering solutions that improve model performance, reduce manual effort, and enable data-driven decision making.

Skills

A/B TestingA/B testingAzure FunctionsBASHBashCI/CDData pipelinesDatabricksDjangoDockerFastAPIGitGraphQLJupyterKerasLinuxMatplotlibMySQLPandasPyTorchPythonRSQLScikit-learnSeleniumStatistical AnalysisTableaupandasscikit-learn

Experience

  • Python Data Scientist · Turing11-01-2023 – 03-31-2026

    Improved model accuracy by 22% using advanced feature engineering. Enhanced, evaluated, and reviewed large language model (LLM) through prompt engineering techniques, reducing LLM error rate by 35% Designed and maintained high-throughput ETL pipelines and data workflows using Python and SQL, processing ~ 5M+ records daily with improved reliability. Eliminated repetitive data processing tasks using Python scripting, which reduced manual effort by 40%. Partnered with cross-functional teams (engineering, product, and business) to convert data insights into actionable product and strategy improvements. Optimized existing data pipelines for scalability and performance, enabling faster iteration cycles for the data science team.

  • Data Engineer (Freelance) · Medusa06-01-2023 – 02-29-2024

    Built and managed ETL processes on AWS using Redshift, Athena, and MySQL within containerized AWS Lambda functions running on Docker. Integrated with existing CI/CD pipelines via Bitbucket to enable automated, reliable code deployment workflows. Deployed Docker images to AWS infrastructure and monitored system performance and health using AWS CloudWatch dashboards and alerts. Implemented reverse ETL pipelines for Monday.com using GraphQL API integration, enabling seamless bi-directional data synchronization.

  • Data Scientist · LoveForData01-01-2023 – 11-30-2023

    Designed and implemented End-to-End ETL workflows to extract, transform, and load data from diverse sources into a centralized data warehouse. Performed exploratory data analysis (EDA) using Python and SQL, identifying key business patterns and trends that guided strategic decisions. Built and deployed predictive machine learning models using scikit-learn and PyTorch, improved recall from 55% to 73% while maintaining constant precision for a default prediction model

  • Data Analyst · LoveForData01-01-2021 – 12-31-2022

    Developed Django-based web applications for interactive data visualization and self-service analytics, increasing stakeholder data accessibility by 50% Built and deployed a Python-based ETL pipeline on Azure Functions, enabling serverless, scalable data processing on the cloud. Applied market basket analysis (association rules) to a restaurant client's transaction data, generating actionable insights that improved menu design, marketing strategy, and customer experience. Developed multiple production-grade web scraping solutions in Python to automate data collection from external websites. Shipped a forecasting model, supporting data-driven planning and operational decisions.

  • Junior Data Analyst · LoveForData04-01-2019 – 12-31-2020

    Performed comprehensive data audits to identify and correct data quality issues, reducing data entry errors by 25% across 5 client databases. Built an interactive, real-time dashboard using R-Shiny for data visualization and exploratory analysis, enabling stakeholder self-service reporting. Developed data workflows to extract, transform, and load data from multiple heterogeneous sources into a structured data warehouse.

  • Data Science Trainee · LoveForData10-01-2018 – 03-31-2019

    Automated data extraction from 20+ external websites daily. Developed an NLP pipeline supporting sentiment analysis, text summarization, and question answering using Python-based NLP libraries.

  • Data Science Intern · LoveForData08-01-2018 – 09-30-2018

    Collected, cleaned, and structured datasets.

  • Data Science Intern · LoveForData08-01-2018 – 09-01-2018

    - Collected, cleaned, and structured datasets.

Similar talent on Pangea

Hire Amir through Pangea

Describe your project to the Pangea agent — see if Amir is a fit, with transparent pricing and interviews booked straight onto your calendar. No contact details change hands until you hire.

See if Amir is a fit