Primary record

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Shield AI Indexed employerRemote
Source-hosted applyChecked 1h ago$180K–$270K/yrFull-Time
Apply at Shield AI

Shield AI receives this application through Lever. Babu Careers does not claim delivery.

Workplace

remote

Employment

Full-Time

Published

Aug 12, 2026

Closes

No date supplied

The role

Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide. For more information, visit www.shield.ai. Follow Shield AI on LinkedIn, X, Instagram, and YouTube.

Requirements

What you'll do:: Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbook patterns for data jobs and platform services. Define and maintain CI/CD and promotion standards for Databricks assets, including workflows, jobs, notebooks, code packages, infrastructure configuration, and environment promotion from dev to prod. Design and maintain platform standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and production execution reliability. Establish reusable operational templates and enablement patterns for new domains onboarding to Databricks, including logging conventions, job tagging, metadata capture, and support handoff expectations. Partner with the Senior Data Engineer to ensure ingestion and medallion patterns are implemented in a way that is observable, recoverable, cost-aware, and secure in production. Work with the cloud/infrastructure team to align Databricks configuration and usage patterns with broader enterprise cloud standards, especially where commercial and future government-hosted environments are involved. Help enforce technical controls for data segregation, access boundaries, and operational compliance in a highly regulated environment. Track and improve platform health metrics such as job success rates, incident trends, data pipeline reliability, cost efficiency, and environment drift. Document platform standards, operational expectations, and support models so the Databricks platform can scale beyond a small founding team. Mentor internal engineers who are growing into platform responsibilities, helping expand Databricks operational knowledge within the team. Required qualifications:: 12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure. Hands-on experience with Databricks or a closely related cloud