Current role

Staff Engineer, Data Platform (R5659)

Shield AI Indexed employerSan Diego, California
Source-hosted applyChecked 1h ago$150K–$230K/yrFull-Time
Apply at Shield AI

Shield AI receives this application through Lever. Babu Careers does not claim delivery.

Workplace

On-site

Employment

Full-Time

Published

Aug 26, 2026

Closes

No date supplied

The role

Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide. For more information, visit www.shield.ai. Follow Shield AI on LinkedIn, X, Instagram, and YouTube. Job Description: We are looking for a Staff Data Platform Engineer to help define and build the data foundation of the AI Factory. The Data Platform provides a unifying, knowledge-graph-centered API layer for human and agentic workflows. It connects configurations, requirements, software versions, test executions, files, signals, training data, and results through stable identities and typed relationships. It also provides consistent access to the storage and compute systems behind those data products. This is a hands-on technical leadership role. You will design platform architecture, implement production software, evaluate storage and compute technologies, establish data-modeling patterns, and work directly with teams collecting and consuming mission-critical data. Success requires balancing developer productivity, semantic clarity, operational reliability, system performance, portability, and long-term maintainability.

Requirements

What you'll do:: Develop a unifying Graph API: Lead the architecture and implementation of the knowledge graph and multi-modal API layer that serves as the backbone for human, service, and agentic workflows. Own DataOps infrastructure: Research, optimize, and maintain the storage, indexing, query, ingestion, and compute infrastructure used throughout the data lifecycle. Establish best-practices: Establish durable, best-practice patterns for schema modeling, relationships, lineage, and schema evolution. Turbocharge agentic data access: Build APIs that enable agents to retrieve structured, connected, and explainable context rather than relying only on keyword or vector similarity. Develop reference architectures: Establish recommended storage and compute profiles, deployment patterns, benchmarks, and operational guidance for both internal and customer-managed infrastructure. Advise downstream teams: Partner directly with autonomy, ML, test, infrastructure, product, and customer-facing teams to turn real workflows into reusable platform capabilities from modeling to integrations. Build first-party integrations: Deliver integrations that make important data easy to collect and aggregate, including data produced by simulations, test infrastructure, training systems, and edge devices. Improve developer experience: Create self-service APIs, SDKs, tools, examples, and diagnostics that make correct data modeling and ingestion the easiest path. Drive technical direction: Evaluate emerging data and AI infrastructure technologies, make principled build-versus-buy decisions, and guide implementation across team boundaries. Raise operational quality: Establish expectations for observability, performance, reliability, security, data integrity, disaster recovery, and lifecycle management. Key outcomes:: Human and agentic workflows use one coherent API for discovering data, traversing relationships, and accessing specialized payloads. Teams spend their time deciding how to model and