Workplace
hybrid
Employment
Full-Time
Published
May 29, 2025
Closes
No date supplied
The role
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, primarily across on-prem environments for the US Government. Forward Deployed Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll travel to various locations where you will be the expert for Palantir’s infrastructure, helping partner teams build & configure their hardware and network for software to operate reliably within. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
Requirements
Core Responsibilities: Maintaining availability of physical Linux servers that power the Palantir platform in air-gapped production environments Design, deploy, and operate infrastructure to support customer & product requirements via modern orchestration & monitoring platforms Collaborate closely with product teams on requirements & SLOs for deploying software into air-gapped environments Identifying, troubleshooting, and solving network & systems issues Scripting to automate away routine operational tasks Provide technical troubleshooting support for production issues, ensuring timely resolution and minimal impact on operations. Participate in a support on-call schedule What We Value: Confidence in troubleshooting complex systems issues independently using stack traces and observability & systems tools Comfort with configuration management, load balancing, monitoring & alerting infrastructure, and container orchestration on small hardware form factors. Demonstrated ability to continuously learn and work independently, making decisions with minimal supervision while working in secure facilities Experience with containers (Docker/Podman) and orchestration (OpenShift/Kubernetes) at scale is a plus Preferred Certifications: DOD 8570 IAT Level II or greater (CISSP, Sec+), Unix/Linux Computing Environment (e.g Linux+, RHCE) What We Require: Available for 50% travel (domestic and international) 4+ years of experience with Linux system administration (RHEL or equivalent preferred) Experience with hardware environments, including setup, configuration, and management of physical servers and networking equipment Familiarity with monitoring systems using tools like Prometheus and writing health checks Proficiency with at least one programming or scripting language, such as Java, Go, Python, JavaScript, Bash, or similar languages. Strong engineering background, preferred in fields such as Computer Science, Mathematics, Software Engineering, Physics, and Data Science. Active US