H
hamza_khawar's photo
Hamza Khawar
From United States
$35/hr or $100,000/yr

Active 19 hours ago


Member since Aug 2026

Share this profile:

Data Engineer

Database Engineer
Available for hire
Years of experience
7+ years
Experience level
Senior
Available from
22 Sep 2026
Download Resume / CV

Data Engineer with 7+ years of experience building scalable cloud data platforms and production data pipelines using Azure Databricks, PySpark, Python, SQL, Delta Lake Azure Data Factory Azure Synapse and Snowflake with strong expertise in healthcare data engineering and regulated environments, including processing over 50 million healthcare records annually while ensuring data quality, governance, security, reliability, and analytics-ready data solutions

Languages

No languages.

Employment History

Data Engineer at UBC (United BioSource) Current 2022 - Now
Designed and implemented a Medallion Architecture (Bronze/Silver/Gold) on Azure Databricks and Delta Lake, enabling structured data promotion of raw clinical and pharmacovigilance data through validated, analysis-ready layers across ADLS Gen2. Built large-scale PySpark transformations on Azure Databricks for adverse event aggregation, drug safety signal detection, and real-world evidence processing, handling 50M+ healthcare records annually from 15+ pharmaceutical partners. Architected end-to-end ELT pipelines on Azure Data Factory, orchestrating data ingestion from clinical study feeds, EHR systems, claims databases, and patient access platforms into the Delta Lake data lake on ADLS Gen2. Engineered the Delta Lake storage layer with partition optimization, schema evolution, and data versioning to support audit-ready immutable Bronze-layer logs and full transformation lineage for FDA-adjacent REMS regulatory reporting. Developed HIPAA-compliant data pipelines with PHI de-identification and masking logic in Python and PySpark, ensuring all patient-level data processed through the Databricks platform met regulatory and privacy standards. Built dbt transformation models on Azure Synapse Analytics for REMS program compliance metrics and adverse event reporting, delivering clean analytical datasets to medical affairs and biostatistics teams. Engineered pharmacovigilance data consolidation pipelines aggregating adverse event reports from multiple safety databases into a unified data warehouse, reducing safety report generation time by 50%. Built real-world evidence (RWE) data pipelines integrating EHR, claims, and registry data from multiple external sources via Airbyte and Azure Data Factory, supporting late-stage clinical research and post-market safety studies. Implemented pipeline monitoring and alerting using Azure Monitor and Azure Data Factory diagnostics, reducing mean time to detect (MTTD) for pipeline failures by 65%. Deployed Azure Purview to catalogue all clinical and pharmacovigilance data assets across ADLS Gen2 and Synapse, establishing end-to-end data lineage for compliance and audit teams. Established CI/CD pipelines via Azure DevOps for Azure Data Factory, Databricks notebooks, and dbt deployments, automating testing, environment promotion from development to staging to production, and rollback, reducing deployment errors by 70%. Collaborated with biostatistics and medical affairs teams to model patient access and therapy adherence data across 10+ specialty therapy programs. Defined data contracts, SLAs, and lineage documentation across all production pipelines in partnership with data governance and compliance teams. Mentored junior data engineers on Azure Databricks best practices, Delta Lake patterns, dbt modeling, and HIPAA data handling standards.
Junior Data Engineer at SERHANT 2019 - 2022
Built and maintained Azure Data Factory pipelines to ingest MLS property listing data, CRM records, and agent transaction data from 20+ regional markets into ADLS Gen2, enabling unified reporting across the brokerage network. Developed dbt models on Snowflake for agent performance analytics, deal pipeline tracking, and market trend reporting, covering average days on market, close rate, and revenue per agent across 500+ agents and 10,000+ annual transactions. Designed and implemented a star schema data warehouse on Snowflake, consolidating listing, transaction, lead, and media performance data to support self-service BI dashboards for sales leadership and marketing teams. Built Python-based ingestion scripts to extract and normalize property data from third-party listing APIs and MLS feeds, reducing data onboarding time for new markets from weeks to days. Orchestrated data workflows using Apache Airflow, managing DAGs for daily MLS synchronizations, weekly agent performance aggregations, and monthly revenue attribution reports. Implemented dbt tests and schema validation to enforce data quality across property listing and transaction pipelines, catching upstream issues before they impacted executive dashboards. Built pipelines for digital content performance data, including video views, lead conversions, and property page traffic, integrating advertising platform APIs into the analytics warehouse. Led the migration from legacy PostgreSQL reporting to modular dbt models on Snowflake.

Education

Bachelor of Science in Computer Science at National University of Science and Technology 2014 - 2018