Overview
Site Reliability Engineer Jobs in Ho Chi Minh City, Vietnam at Talentnet Careers
Title: Site Reliability Engineer
Company: Talentnet Careers
Location: Ho Chi Minh City, Vietnam
Purpose of the position:
Talentnet is building a modern HR technology platform with hybrid infrastructure across on-premise systems and cloud-native services. We are looking for an experienced SRE to design, standardize, automate, and operate scalable infrastructure and platform operations across data, application, and AI workloads.
The role requires strong operational discipline, infrastructure engineering capability, and the ability to establish reliable engineering processes in a fragmented enterprise environment.
Job Accountabilities:
1. Infrastructure Design & Platform Operations
- Design, deploy, and operate hybrid infrastructure environments across on-premise systems and Microsoft Azure cloud.
- Maintain reliability, availability, scalability, and security of production systems.
- Build standardized infrastructure provisioning and operational processes.
- Manage Kubernetes/container platforms and CI/CD environments.
- Monitor platform health, capacity, incidents, and performance.
2. Reliability Engineering & Incident Management
- Define and maintain SLO/SLA/SLI, incident response procedures, RCA, and disaster recovery strategies.
- Implement observability stack including logging, metrics, tracing, and alerting.
3. Security & Access Governance
- Implement operational security controls, secrets management, and access governance.
4. Cross-Team Collaboration
- Collaborate closely with engineering, AI, and data platform teams.
Requirements:
- Bachelor’s degree in Information Technology, Computer Science, Software Engineering, or a related field
- 5–7 years of experience in DevOps, Infrastructure, SRE, Platform Engineering, or related roles
- Strong expertise in Linux, Docker, Kubernetes, and Microsoft Azure (especially AKS)
- Hands-on experience with CI/CD, GitOps, Helm Charts, and infrastructure automation
- Experience with monitoring and observability tools (Prometheus, Grafana, ELK/OpenSearch, Azure Monitor)
- Scripting skills in Bash, Python, or PowerShell
- Strong ownership mindset with a focus on automation, reliability, and operational excellence
- Experience with Data Platforms or AI Infrastructure is a plus