Data & AI ProductOps practice established, adopted and scaled.
Clear and reusable standards established across DataOps, MLOps, LLMOps and AgentOps.
ProductOps capabilities successfully implemented within clients’ environments.
Clear production ownership, operational readiness and L0/L1/L2/L3 support models established.
Improved product health and SLO compliance.
Reduced incident detection and restoration time.
Reduced recurring production issues and unnecessary engineering escalations.
Increased automation, proactive monitoring and self-service.
Strong integration between Data & AI ProductOps and Enterprise Service Management.
Increased reuse of standards, patterns, playbooks and delivery accelerators.
Strong internal capability development and effective client knowledge transfer.
Responsibilities
Practice Development
Establish and lead the Data & AI ProductOps practice, including the operating model, standards, methods, governance, reusable patterns, playbooks and implementation approach.
Define the practice across DataOps, MLOps, LLMOps and AgentOps, with clear standards for operating Data, BI, ML, GenAI and Agent-based products in production.
Develop reusable ProductOps assets, including operational readiness standards, support models, SLO frameworks, monitoring patterns, runbooks, implementation templates and delivery accelerators.
Product Operations
Define and implement operational readiness requirements covering product ownership, criticality, support levels, SLAs/SLOs, monitoring, alerting, runbooks, escalation, recovery, dependencies and rollback.
Establish clear L0/L1/L2/L3 support models, with L0 focused on automation and self-service, L1 on Service Desk support and initial triage, L2 on ProductOps-led operational support and product-level troubleshooting, and L3 on complex issues requiring Engineering expertise.
Establish product health and observability across availability, performance, data freshness and quality, pipeline and integration health, model and AI performance, usage, cost and other product-specific operational measures.
Lead product-level incident investigation, restoration and root-cause analysis, coordinating across Engineering, Platform, Service Management, Security and vendors.
Integrate ProductOps with enterprise Incident, Problem, Request, Change, Release, Knowledge, Service Level and Major Incident Management processes.
Ensure recurring incidents and operational issues are converted into permanent corrective actions, automation opportunities and product improvement backlog items.
Establish effective change and release practices supported by automated testing, CI/CD, versioning, controlled deployment, and rollback.
Assess clients Data & AI ProductOps maturity and define target operating models, support structures, service levels, observability, automation and implementation roadmaps.
Design and implement ProductOps capabilities for Data, BI, ML, GenAI and Agent-based products within clients’ environments.
Lead clients solutioning, architecture workshops and implementation engagements, defining roles, responsibilities, support models, operational processes, tooling and Service Management integration.
Provide architecture and implementation oversight for complex Data and AI production environments, engaging senior stakeholders across Product, Engineering, Platform, Governance, Security and Service Management.
Support proposals, technical solutioning and advisory engagements across Data & AI ProductOps, DataOps, MLOps, LLMOps and AgentOps.
Leadership Expectations
Build and scale a new Data & AI ProductOps practice.
Provide strong technical leadership while remaining hands-on where required.
Lead multidisciplinary teams across Data, AI, Engineering, Platform, Product and Service Management.
Provide architecture and implementation oversight across complex enterprise environments.
Engage confidently with senior clients and internal stakeholders.
Translate emerging Data and AI technologies into practical enterprise operating practices.
Develop internal talent and build sustainable technical capability.
Qualifications
Bachelor’s degree in Computer Science, Engineering, Data/AI, Information Systems, or a related field.
Relevant certifications across cloud, Data & AI platforms, DevOps/MLOps, Service Management, architecture, or data management are an advantage.
10+ years of experience across Data Engineering, AI/ML, DataOps, DevOps, SRE, Product Operations, Platform Engineering or related disciplines.
Strong hands-on experience with modern enterprise data platforms and production Data/AI solutions.
Demonstrated experience establishing or leading production operating capabilities for Data and AI products.
Strong practical experience across DataOps, MLOps, LLMOps and AgentOps.
Strong experience with operational readiness, observability, SLAs/SLOs, monitoring, support models, incident and problem management, CI/CD, automated testing, release management, recovery and rollback.
Experience operating BI and data products, data pipelines, APIs, integrations, ML models, GenAI applications and AI Agents in production.
Practical experience with model lifecycle management, model serving, drift monitoring, RAG, vector search, AI evaluation, AI observability, Agent tracing, tool execution and guardrails.
Strong understanding of enterprise ITSM and Service Management, with experience integrating Data and AI product teams into established support, incident, problem, change and release processes.
Strong incident leadership, troubleshooting, root-cause analysis and production problem-solving capabilities.