Woman wearing a hardhat outside, taking notes while looking at the sunset

Senior Databricks Engineer - 100% Remote

Richland, WA, US

Job description

Senior Databricks Engineer

Objective & Purpose

The Pacific Northwest National Laboratory (PNNL) requires an experienced Senior Databricks Engineer contractor to augment the enterprise data engineering team during the implementation of a new ERP platform. The contractor will directly develop code, build and optimize data pipelines, and work agilely across a variety of day-to-day engineering tasks to support technical execution within the team. The contractor will work alongside an existing technical contractor team and internal staff to develop, refine, and operationalize five inaugural domain lakehouse pipelines.

As a collaborative and adaptable technical contributor, the contractor will lead development across the Silver through Platinum medallion layers, implement dimensional data models, establish automated testing and quality assurance frameworks, maintain technical documentation (including Architecture Decision Records (ADRs) in GitHub), and mentor incoming staff on daily operations.

Reporting Structure

The contractor will work under the technical and operational direction of the Senior Manager, Data & Analytics and the Data Architecture Workstream Lead to maintain alignment with critical ERP project milestones and enterprise architecture standards.

Work Location & Property

100% remote
All work must be conducted during standard business hours aligned with the Pacific Time Zone (PST)
Government-furnished laptop provided
Contractor must possess and maintain a reliable internet connection at their own cost

Responsibilities

  • Hands-on development, tune, and operationalize high-throughput batch and streaming ETL/ELT pipelines in PySpark and SQL across the Medallion Architecture (Bronze → Silver → Gold/Platinum).
  • Ingest complex, enterprise transactional data from the ERP platform and a legacy data warehouse into Azure Data Lake Storage (ADLS Gen2) and Delta Lake.
  • Implement governed data distribution patterns optimized for downstream applications, analytics, reporting layers, and Power BI semantic models.
  • Lead the implementation of enterprise dimensional models, star schemas, slowly changing dimensions (SCDs), and curated analytical aggregates.
  • Collaborate with data architects and business analysts to translate functional data mappings into performant, query-optimized physical data structures.
  • Design and implement automated data quality validations, schema enforcement, and end-to-end integration tests across development, test, and production environments.
  • Establish automated monitoring, alerting, error-handling routines, and pipeline telemetry to guarantee data accuracy, operational resilience, and SLA compliance.
  • Operationalize and optimize Databricks Workflows and Jobs for cost efficiency, scalability, and enterprise fault tolerance.
  • Standardize deployment workflows using Git/GitHub CI/CD patterns such as Databricks Asset Bundles (DAB).
  • Apply fine-grained access controls, object governance, and lineage tracking leveraging Databricks Unity Catalog.
  • Document pipeline designs, runbooks, and Architecture Decision Records (ADRs) directly within GitHub.
  • Mentor and train incoming Databricks Engineers to transfer platform domain knowledge and ensure smooth handover of day-to-day lakehouse operations.

Required Qualifications

  • 7+ years of data engineering/platform engineering experience, with 3-5+ years focused on production cloud data architectures.
  • 5+ years of production experience with Azure Databricks, Delta Lake, PySpark, Spark SQL, Workflows, and Unity Catalog.
  • Proven ability to write clean, modular, and maintainable production code in Python/PySpark and SQL while agilely adapting to diverse technical tasks.
  • Demonstrated expertise in dimensional data modeling (Kimball methodology), star schemas, fact/dimension structures, and Platinum layer curation.
  • Proven background ingesting, transforming, and modeling large-scale, complex transactional or ERP data structures.
  • Proficient in CI/CD pipelines (e.g., Databricks Asset Bundles, GitHub Actions), automated testing frameworks, and managing ADR documentation.
  • Exceptional interpersonal skills, an amenable and collaborative team-first mindset, and a proven ability to mentor and upskill fellow engineers.

Preferred Qualifications

  • Demonstrated familiarity or hands-on experience with legacy data warehouses, SQL Server databases, data marts, and migrating legacy SQL/ETL workloads to Databricks.
  • Prior experience delivering data pipelines in highly regulated or security-conscious environments.
  • Experience leveraging GenAI/LLM developer tools (e.g., GitHub Copilot, Databricks Assistant) to accelerate development and test creation.
  • Databricks Certified Data Engineer Associate/Professional or Microsoft Certified: Azure Data Engineer Associate.

Key Deliverables

  • Five Operational ERP Domain Pipelines: Fully tested, production-grade Silver-to-Platinum pipelines supporting ERP integration.
  • Curated Platinum Data Models: Validated, performant dimensional models serving downstream business applications and analytical consumers.
  • Automated QA & Monitoring Framework: Automated test suites, pipeline observability metrics, alerting rules, and operational runbooks.
  • Architecture & Deployment Documentation: Approved ADRs, workflow architecture diagrams, and CI/CD deployment configurations managed in GitHub.
  • Knowledge Transfer & Training Artifacts: Onboarding sessions, technical walkthroughs, and operational handover documentation for internal engineering staff.
  • Weekly Status Summaries: A weekly progress report submitted to leadership detailing key activities, completed tasks, upcoming priorities, and any identified blockers or challenges.

Performance Expectations & Success Criteria

  • Milestone Delivery & Reporting: Timely delivery and execution of agreed-upon sprint goals, milestones, and project deliverables, supported by consistent and transparent weekly progress reporting.
  • Collaboration & Agility: Demonstrated ability to jump right in, work agilely on diverse technical tasks, and collaborate amenably with fellow contractor and internal engineering peers.
  • Data Accuracy & System Integration: Robust ERP integration with the Databricks Lakehouse, maintaining complete data fidelity, consistency, and reliability across all domains.
  • Architecture Alignment: Implementation of resilient, scalable solutions that adhere to enterprise data principles and security standards.
  • Documentation Quality: Complete, clear, and version-controlled documentation for all models, mappings, workflows, and testing suites.
  • Staff Enablement: Measurable progress in mentoring internal staff, evidenced by their increasing independence in managing and operating the Databricks platform.

Compliance Requirements

Citizenship & Security Requirements

  • U.S. Citizenship: The contractor must be strictly a United States Citizen.
  • Information Security: The contractor must execute agreements to safeguard sensitive data in accordance with laboratory requirements and federal stipulations.

Mandatory Onboarding Training

  • Prior to being granted access to PNNL networks and computing environments, the contractor must successfully complete all prerequisite onboarding courses, including Cyber Security, Human Resources, and System Access Training, administered online via the PNNL Web Portal.

Period of Performance

  • Estimated from Date of Award through September 30, 2027, with the possibility of extension based on project milestones and organizational need.

For persons with disabilities: If you require assistance completing, or are unable to complete the online resume/application process, please contact:

VAS Human Resources @ 803-644-0070
Mon. - Fri. 9am - 4pm

237 High Gate Loop 
Aiken, SC 29803

Value Added Solutions, Inc. is an Equal Opportunity Employer and supports a drug-free work environment.