Yash Rupani
Data Engineer & ML Infrastructure Builder | RAG · Vector Search · Real-Time Pipelines | AWS, GCP, Spark, Snowflake | OSU M.Eng ’25
1 yrs experience · Corvallis, OR · <10 hrs/week
About
I've built data pipelines that process 102 billion+ YouTube views. I've deployed RAG systems that improved grant retrieval accuracy by 40%. And I've cut manual data work at Infosys by 10 hours a week through automation. That's what I do, I turn messy, complex data problems into clean, scalable, production-ready infrastructure. I'm Yash Rupani, a Data Engineer and ML Infrastructure specialist with a Master of Engineering in Computer Science from Oregon State University (GPA: 3.87 | Dec 2025). Before grad school, I spent 2+ years at Infosys as a Senior Systems Engineer, where I built and automated ETL pipelines, reduced data validation time by 35%, and generated dashboards from chatbot performance logs at scale. My technical focus lives at the intersection of: Data Engineering — ETL/ELT pipelines, Data Lakes, Lakehouse architecture (Kafka, Spark, Airflow, PySpark) AI/ML Infrastructure — Graph RAG, Vector Search (FAISS), NLP, real-time AI systems Cloud Platforms — AWS (S3, Glue, Lambda, Athena, EMR, QuickSight) & GCP Data Warehousing — Snowflake, Redshift, BigQuery Recent highlights: → Built a real-time crypto monitoring platform processing 500+ events/second using FastAPI + Kafka + WebSockets → Architected a Lakehouse pipeline (Kafka + Spark Streaming) achieving sub-second latency → Designed a serverless AWS Data Lake converting JSON to Parquet, cutting storage costs by 65% and Athena scanning costs by 50% → Achieved 90% NLP classification accuracy on a financial news sentiment analyzer (NLTK + TensorFlow) Beyond the code, I've mentored 15+ students on industry-sponsored analytics capstone projects for clients like the Portland Trailblazers and Port of Portland, because I genuinely believe the best engineers can also communicate their work clearly. I hold certifications from Snowflake (Data Warehousing) and Databricks (AI Agent Fundamentals), and I'm actively growing my expertise in RAG architectures and real-time AI product infrastructure. I'm actively seeking full-time Data Engineering or ML Engineering roles, especially in teams building intelligent, cost-efficient, real-time data systems. If that sounds like your team, let's connect. Reach me here on LinkedIn or visit my portfolio: https://rupaniyash.github.io/portfolio/
Skills
Experience
- Research Assistant for the School of Marketing/Design/Analytics · Oregon State University10-01-2025 – 12-01-2025
Engineered an end-to-end autonomous ETL pipeline for unstructured multimedia data using Python-based preprocessing, eliminating 40% of manual data preparation overhead and accelerating Agentic AI workflows for active research initiatives. Refined backend data structures and indexing strategies for AI Agent systems, reducing retrieval latency by 25% and improving real-time inference performance for production-ready AI pipelines. Developed robust data validation frameworks ensuring 100% data integrity across terabyte-scale dataset ingestion, enabling reliable and scalable processing for downstream AI and analytics research.
- Intern AI Agent Graph RAG/ Machine Learning Scientist · GrantAide07-01-2025 – 09-01-2025
Architected a Retrieval-Augmented Generation (RAG) system using FAISS vector databases and Python to power intelligent grant discovery, increasing search relevance and retrieval accuracy by 40% over keyword-based approaches, directly improving match quality for nonprofit clients. Built and optimized Flask + Google Firestore backend APIs with CORS handling and data consistency controls, improving overall application response performance by 30% under concurrent user load. Deployed and monitored production AI applications across GCP and AWS environments, maintaining 99.9% uptime and reducing cross-region API latency by 25% through architecture optimization.
- Sr. Systems Engineer (Data Engineering) · Infosys06-01-2021 – 06-01-2023
Automated end-to-end data quality pipelines in Python across chatbot performance and user behavior datasets, eliminating 10+ hours of manual validation work weekly and reducing data error rates by 35%. Designed and deployed ETL workflows using SQL and automation scripting to process structured and semi-structured log data at scale, boosting pipeline throughput efficiency by 40% and enabling near-real-time reporting for analytics teams. Built chatbot KPI dashboards from raw IVR/IVA interaction logs, standardizing 25% more accurate reporting outputs consumed by product and operations stakeholders across multiple business units.
Education
- Oregon State UniversityMasters of Engineering, Computer Science
- Pandit Deendayal Energy UniversityBachelors of Technology, Electrical Engineering
Similar talent on Pangea
Hire Yash through Pangea
Describe your project to the Pangea agent — see if Yash is a fit, with transparent pricing and interviews booked straight onto your calendar. No contact details change hands until you hire.
See if Yash is a fit