Vingroup

Senior DevOps / MLOps Engineer (5+ Years of Experience)

Hanoi, Vietnam · On site · Senior

OV members work here

Job description

We are looking for Senior MLOps & Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast’s ADAS and autonomous driving programme — on-premise GPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.

This is a fullstack infrastructure role. You will work across compute, storage, networking, automation and security. We do not divide the team into narrow specialists: everyone owns a part of the platform end to end — provisioning, deployment, monitoring and incident response — and is expected to grow across the whole stack over time. Work is assigned according to what the programme needs at that moment, not by a fixed job boundary.

Much of the platform runs on-premise, including air-gapped segments. Experience operating without a public-cloud safety net is valued here

Key Responsibilities

1. AI Compute & Serving

- Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling, utilisation and capacity

- Deploy and scale high-throughput model serving; tune GPU memory and runtime performance

- Manage multi-GPU distributed training and reproducible environment sandboxing

- Scale pipeline orchestration on Kubernetes for large data processing jobs

2. Platform Automation & Delivery

- Drive GitOps-based deployment and maintain infrastructure-as-code across the platform

- Build and maintain CI pipelines, secure container builds and release automation

- Build monitoring, logging and alerting so that failures are detected and actionable, never silent

- Lead incident response and drive follow-up actions to closure

3. Data Infrastructure & Security

- Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery

- Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction

- Implement access control and single sign-on across platform services, with audit logging

- Apply data protection measures to sensitive content before it reaches external users

Requirements

- 4+ years in MLOps, DevOps, SRE or HPC platform engineering, with production ownership of AI/ML infrastructure

- Kubernetes administration at production scale: Helm, ingress (Traefik or Envoy), CNI, and GitOps (Flux or ArgoCD)

- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging

- Strong Linux systems skills and infrastructure-as-code (SaltStack or Ansible)

- Python and Bash for automation, including Airflow DAGs and custom operators

- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)

- Willingness to work across the full stack — compute, storage, networking, automation and security — rather than within a single specialty

- Good communication in English — technical documentation and working with international partners

Get access to all 2,144 jobs.

Free, with your email. New roles that fit you, every Monday.

OV members work at Vingroup. Members see who they are and can ask them for a short chat.

Apply to join

OV member? Use the email you use for OV.

Hiring? Reach Overseas Vietnamese.Post roles