I clean, structure, and pipeline messy datasets so they're actually usable — whether that's a CRM export full of duplicates, a spreadsheet with inconsistent formats, or raw web data that needs to be scraped and organized from scratch.
Recent work: cleaned and reconciled a contact list against a CRM export for a client (5-star review, delivered fast). Built a full web scraping pipeline (Python, Scrapy, Docker) that pulls and filters e-commerce data, bypassing anti-bot protections, running automatically on a schedule. Cleaned and imputed a 100,000+ row dataset, fixing structural errors and reconstructing missing data using spatial/statistical methods.
I also had a bug fix merged into Apache Wayang, a real Apache Software Foundation open-source project — so I'm comfortable with code that has to actually work correctly, not just look right.
Skills: Python, Pandas, NumPy, SQL, ETL, Web Scraping (Scrapy), Data Cleaning & Validation, Docker, MySQL/PostgreSQL.
I respond fast, explain every fix I make so nothing is a black box, and I'm available to start immediately.