Workplace
On-site
Employment
Internship
Published
Aug 20, 2026
Closes
No date supplied
The role
Role Overview
As a Senior Software Engineer, AI Ops at Scale AI, you will own the long-term technical health, performance, and stability of AI solutions deployed across our strategic public sector partners.
While our Delivery Teams build and launch new use cases, you are the technical steward ensuring these deployments operate with Operational Excellence. You will bridge software engineering, MLOps, and client governance, managing tiered SLAs, tracking model drift, and executing maintenance protocols that protect both system integrity and operational margins.
Key Responsibilities
• Handover Gate & Onboarding: Act as the technical gatekeeper during the formal transition from Delivery to Maintenance. Conduct deep-dive reviews to ensure baseline code, prompts, and architecture meet strict maintainability and documentation standards before sign-off.
• Tiered SLA & Incident Management: Own technical response and resolution targets across multi-tiered service models (from Business-Hours Essential to 24/7 Mission-Critical). Lead Incident Governance, Root Cause Analysis (RCA), and P1/P2 mitigations within strict active support windows.
• AI Lifecycle Governance: Monitor production model performance, latency, and data drift. Manage prompt configuration repositories to maintain behavioral consistency and perform regression testing when LLM providers update underlying endpoints.
• Request Classification & Technical Scope: Operationalize the boundary between Routine Maintenance (In-Scope) and System Evolution (Out-of-Scope). Assess incoming client requests and run comparative benchmarking on new AI models.
• Automation & Reliability Engineering: Eliminate operational toil by engineering self-healing data pipelines, automated RAG indexing syncs, and telemetry tooling. Influence upstream "Delivery" teams to adopt architectural patterns that simplify ongoing maintenance.
• Client Techni
Requirements
Department: GPS Engineering