Primary record

Customer Success Engineer (CSE), GPU Cluster

Together AI Indexed employerSan Francisco · San Francisco, California, United States
Source-hosted applyChecked 3h agoInternship
Apply at Together AI

Together AI receives this application through Greenhouse. Babu Careers does not claim delivery.

Workplace

hybrid

Employment

Internship

Published

Aug 11, 2026

Closes

No date supplied

The role

About the role

As a Customer Success Engineer at Together AI, you will serve as the named technical owner for one of our most strategic customer relationships. You will be the primary technical point of contact across all infrastructure domains — compute, networking, storage, and facilities — ensuring flawless delivery and operational health of large-scale GPU deployments. This role sits at the intersection of deep infrastructure expertise and high-stakes customer partnership, making you a critical driver of both customer success and company growth.

Responsibilities

• Serve as the named technical point of contact for a dedicated strategic customer, owning the end-to-end technical relationship across compute, networking, storage, and facilities

• Drive structured engagement through regular cadences — status reporting, technical steering meetings, quarterly business reviews (QBRs), and executive business reviews (EBRs) — spanning both operational and strategic levels

• Translate customer operational feedback into actionable input for Engineering, Product, and Infrastructure roadmaps

• Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains in partnership with Support, SRE, DC Ops, and Engineering teams

• Own end-to-end RMA coordination and hardware lifecycle management, including acceptance testing, spare inventory management, and hardware health reporting for large-scale GPU deployments

• Maintain deep technical expertise across the customer's infrastructure stack — GPU compute, high-speed fabric, and large-scale storage systems — advising on configuration, operational best practices, and incident resolution

• Own the observability strategy for the customer estate, including alert policy definition, dashboard development, and proactive health management across all infrastructure layers

• Coordinate DC

Requirements

Department: Customer Success