Permhunt
See how well you match this role
Upload your résumé — we'll score your fit and check off the skills you already have. Free.
Our client is a technology consulting company providing operational and engineering services to the high-tech sector. They support platform and infrastructure teams across cloud and distributed environments, helping operate scalable, reliable, and production-ready technology platforms. The Role Our client is looking for a Senior AI Platform Engineer to operate, maintain, and continuously improve production AI platforms running on Kubernetes across on-premise, AWS, and GCP environments . The role combines AI platform engineering, MLOps, Kubernetes, Python, observability, and production operations , working with environments similar to AI on EKS and Kubeflow-based machine learning platforms. You will also play a senior role in improving engineering practices, mentoring team members, and helping shape the platform roadmap. Key Responsibilities • Deploy platform releases and configuration changes using GitOps and DevOps practices. • Monitor AI platform and service health through logs, metrics, monitoring, and observability tools. • Improve platform reliability through automation, operational tooling, observability, and self-service capabilities. • Participate in incident response, root cause analysis, and 24/7 operational rotations. • Investigate and resolve user, platform, integration, and configuration-related issues. • Promote strong standards across platform security, reliability, and operational engineering. • Mentor junior engineers in Python fundamentals and help develop their MLOps capabilities. • Drive the adoption of MLOps best practices across the engineering team. • Identify gaps in tooling, technical capabilities, and processes required to support production-grade AI systems. • Contribute to the technical direction and ongoing development of the AI platform. Qualifications • 3+ years of experience supporting production AI, ML, or data platforms using technologies such as Ray, Jupyter, AWS SageMaker, Kubeflow, or similar platforms . • 5+ years of experience across the AI/ML lifecycle, including development, deployment, DevOps, or MLOps. • 5+ years of hands-on Python experience supporting AI/ML workflows, applications, or data engineering pipelines. • Strong practical experience with Kubernetes , including managed platforms such as AWS EKS or Google GKE . • Good understanding of microservices architectures and service communication patterns . • Strong troubleshooting skills across application crashes, resource contention, service latency, performance, and scaling issues. • Experience analysing logs, metrics, monitoring systems, and service-level KPIs within production environments. Nice to Have • Exposure to additional AI and data platforms such as Flyte, Hugging Face, Vertex AI, LangChain, Claude Code, or other AI agent platforms . • Hands-on automation or scripting experience using Bash or Python . • Relevant Kubernetes or cloud certifications such as CKAD or AWS certifications . Originally posted on Himalayas
Sourced from Himalayas. Confirm the details and apply on the employer's site.