Job description
We are looking for Senior MLOps & Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast’s ADAS and autonomous driving programme — on-premise GPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.
This is a fullstack infrastructure role. You will work across compute, storage, networking, automation and security. We do not divide the team into narrow specialists: everyone owns a part of the platform end to end — provisioning, deployment, monitoring and incident response — and is expected to grow across the whole stack over time. Work is assigned according to what the programme needs at that moment, not by a fixed job boundary.
Much of the platform runs on-premise, including air-gapped segments. Experience operating without a public-cloud safety net is valued here
Key Responsibilities
1. AI Compute & Serving
- Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling, utilisation and capacity
- Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
- Manage multi-GPU distributed training and reproducible environment sandboxing
- Scale pipeline orchestration on Kubernetes for large data processing jobs
2. Platform Automation & Delivery
- Drive GitOps-based deployment and maintain infrastructure-as-code across the platform
- Build and maintain CI pipelines, secure container builds and release automation
- Build monitoring, logging and alerting so that failures are detected and actionable, never silent
- Lead incident response and drive follow-up actions to closure
3. Data Infrastructure & Security
- Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
- Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
- Implement access control and single sign-on across platform services, with audit logging
- Apply data protection measures to sensitive content before it reaches external users
Requirements
- 4+ years in MLOps, DevOps, SRE or HPC platform engineering, with production ownership of AI/ML infrastructure
- Kubernetes administration at production scale: Helm, ingress (Traefik or Envoy), CNI, and GitOps (Flux or ArgoCD)
- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
- Strong Linux systems skills and infrastructure-as-code (SaltStack or Ansible)
- Python and Bash for automation, including Airflow DAGs and custom operators
- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
- Willingness to work across the full stack — compute, storage, networking, automation and security — rather than within a single specialty
- Good communication in English — technical documentation and working with international partners
