Workplace
remote
Employment
Internship
Published
Aug 11, 2026
Closes
No date supplied
The role
About the Role
Together AI is building the AI Native Cloud, an end-to-end platform for the full
generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art
AI cloud infrastructure. The Together Cloud team builds the Together GPU Clusters
product, which provides high-performance, AI-ready GPU clusters through a self-serve
cloud console and is the virtualized infrastructure layer powering Together’s inference,
RL, and fine-tuning products.
As a Senior Software Engineer in Together Cloud Infrastructure, you will play a key role
in building the next generation AI cloud platform – a highly available, global, blazing-fast
cloud infrastructure that virtualizes cutting-edge ML hardware (GB200s/GB300s,
BlueField DPUs). You'll enable state-of-the-art ML practitioners with self-serve AI cloud
services, such as on-demand + managed Kubernetes and Slurm clusters, for both our
internal SaaS products (inference, fine-tuning, RL) and our external cloud customers,
spanning dozens of data centers across the world.
Responsibilities
• Design, build, and maintain performant, secure, and highly-available backend services/operators that run in our data centers and automate hardware management, such as Infiniband partitioning, in. DC parallel storage provisioning, and VM provisioning.
• Design and build out the IaaS software layer for a new GB200 data center with thousands of GPUs.
• Design and build distributed GPU scheduling and the global management plane that power on-demand and managed clusters across dozens of data centers
• Develop infrastructure that powers our internal inference, RL, and fine-tuning products in addition to external cloud customers
• Design and build systems that scale per-cluster capacity limits and automate the onboarding of new capacity
• Work on a global multi-exabyte high
Requirements
Department: Engineering