H
Haseeb Khan
From United States 12:47 PM (GMT-04:00)
$45/hr or $110,000/yr

Active 21 hours ago


Member since Aug 2026

Share this profile:

Sr Data Engineer

Data Engineer
Available for hire
Years of experience
8+ years
Experience level
Senior
Available for
Full-time, Part-time, Contract, Freelance
Available from
20 Sep 2026
Download Resume / CV

I’ve spent the last several years owning data pipelines end to end — building medallion lakehouses, keeping freshness SLAs predictable, and supporting teams that rely on clean, stable datasets. Most of my work has been in Fabric and Databricks, with hands‑on experience in Python, PySpark, SQL, incremental processing, and practical debugging when something breaks. What makes me different is that I’m comfortable taking responsibility for a pipeline from ingestion through on‑call, and I focus on keeping things simple, reliable, and easy for others to use.

I’m looking for a role where I can continue building and maintaining production pipelines, improve data quality, and reduce manual work through better automation. A place that values straightforward engineering and clear ownership is the right fit for me.

Skills

No skills.

Languages

Employment History

Senior Data Engineer at Strive Health 2021 - 2026
- Built Microsoft Fabric medallion lakehouses (bronze, silver, gold) using OneLake, PySpark notebooks, Data Factory pipelines, and Warehouse T-SQL; owned silver-to-gold transforms for 20-40 EHR, claims, and lab feeds supporting 100,000+ CKD/ESKD members. - Automated Fabric workspace, lakehouse, and warehouse promotion with Git and Azure DevOps CI/CD, cutting partner environment setup from about 2 days of manual tickets to under 30 minutes. - Built HIPAA and HITRUST PHI controls with security and compliance teams, including Microsoft Purview sensitivity labels, workspace RBAC, column-level security on 100+ sensitive fields, audit logging, and isolated dev/staging/production environments. - Built 15+ gold-layer marts in Spark SQL and Warehouse T-SQL with incremental loads, schema contracts, and pipeline CI, turning partner feeds into risk stratification, hospitalization, and total-cost-of-care datasets. - Monitored data quality rules and T+8 hour freshness SLAs (about 99% hit rate) for owned pipelines, catching late or broken clinical files before they reached care team workflows. - Built Power BI Direct Lake semantic models and Fabric Copilot reporting with analytics and clinical operations stakeholders (25-40 people), reducing recurring ad-hoc SQL requests by about 30-40%. - Tuned Fabric Spark jobs and warehouse capacity, partitioning, and incremental refresh with the platform team, contributing to a 20-30% reduction in compute spend, and supported lakehouse on-call with typical resolution time under 45 minutes. - Partnered with product, clinical, and care team stakeholders to turn data needs into data contracts and SLAs, shortening new-partner dataset onboarding from weeks to days. - Built Warehouse and Power BI semantic models with row-level security for 10+ payor and health-system partners, so each partner could see only their own members, supporting value-based care reporting without cross-tenant data exposure.
Junior Data Engineer at Cornerstone OnDemand 2018 - 2021
- Built Databricks medallion pipelines (bronze, silver, gold) using PySpark, Delta Lake MERGE/CDC, Auto Loader, and Databricks Workflows; owned silver-to-gold transforms for 10-20 learning, skills, performance, and HRIS/ATS feeds. - Implemented Unity Catalog schemas and access grants with the platform team through CI/CD using Databricks Asset Bundles, cutting catalog and setup tickets from about 1-2 days to under 30 minutes. - Built 25-40 production dbt models in SQL and Python with incremental logic, automated tests, and CI on pull requests, turning LMS, skills, and recruiting events into gold marts for course completion, skills-gap, and compliance-training reporting. - Monitored owned jobs with dbt tests, daily freshness checks, and pipeline alerts, maintaining 98%+ on-time delivery and catching late or broken HR and learning files before they reached People Analytics dashboards. - Built Databricks Genie reporting on Unity Catalog governed gold tables with analytics and HR partners (15-25 stakeholders), reducing recurring ad-hoc SQL requests by about 25-35%. - Built Algolia indexing jobs on curated course, skills, and people records with product partners, cutting employee search response time to under 200 milliseconds. - Tuned Databricks job clusters and Delta tables using Photon and partitioning/liquid clustering with senior engineers, contributing to a 15-25% reduction in DBU spend. - Shared lakehouse on-call rotation, recovering failed learning and HR jobs with typical resolution time under 60 minutes, and documented runbooks for the team. - Built bronze-layer ingestion from Cornerstone REST/OData and HRIS/ATS APIs, landing incremental employee and course data so L&D reporting stayed current across 50+ customer tenants.

Education

Bachelors of Software Engineering at Foundation University 2014 - 2017