Primary record

Engineering Manager, Kernel Reliability

Cerebras Indexed employerUS and Canada Offices
Source-hosted applyChecked 1h agoFull-Time
Apply at Cerebras

Cerebras receives this application through Ashby. Babu Careers does not claim delivery.

Workplace

On-site

Employment

Full-Time

Published

Jan 8, 2026

Closes

No date supplied

The role

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership https://openai.com/index/cerebras-partnership/ with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About The Role We're looking for a deeply technical, hands-on engineering leader for our on-field Kernel Reliability team. You will lead a high performing team to tackle a critical challenge: improving the reliability of our advanced compute clusters and the underlying inference, training, and internal production services. In this role, you'll set the technical vision while staying close to the code and designing solutions that will scale to our exponentially growing system production and software service offerings. If you have proven expertise in software or hardware reliability, diagnostic tool building, or failure analysis and debugging, we want to hear from you. Responsibilities - Provide hands-on technical leadership, owning the technical vision and roadmap for the kernel-centric reliability of our internal and customer-facing systems - Assist System and Cluster Operations teams on reducing system and service downtime after failure by providing tooling and manual intervention for failure analysis and diagnostic - Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis - Collaborate with SW teams to improve the software stack, including Kernels, to improve on-field debugging and failure analys

Requirements

Department: Software; Team: Software