Lead Site Reliability Engineer

ABOUT CLIENT

Our client is a global technology company that specializes in providing innovative IT solutions for the financial services industry

JOB DESCRIPTION

Lead a team of SREs, offering technical guidance and coaching while promoting a culture of reliability and continuous improvement.

Define and advance SRE practices such as SLIs/SLOs, error budgets, and incident response processes across production systems.

Take ownership of the design and evolution of automated cloud operations, driving the adoption of Infrastructure-as-Code (Terraform, CloudFormation) and CI/CD pipelines.

Oversee major incident responses, prioritize rapid resolution, conduct root cause analysis, and implement preventive measures.

Collaborate closely with Development, DevOps, and Cloud Engineering teams to incorporate reliability and resilience at every delivery stage.

Establish and track key reliability metrics (availability, latency, error rates) and drive initiatives for continuous improvement.

Assess and implement AWS-native and third-party tools to enhance monitoring, alerting, and automation.

Serve as the main contact for Service Reliability topics with clients, ensuring transparency and alignment on reliability goals.

Ensure compliance with industry standards and internal policies related to security, audit, and operational risk.

JOB REQUIREMENT

Minimum 7 years of experience working as an SRE Engineer, with exposure to data platform solutions being beneficial.

Extensive experience with cloud platform, including IAM, ECS, EKS, Lambda, and CloudWatch.

Proficiency in deploying and managing containerized services, particularly on Kubernetes.

Hands-on experience with Infrastructure-as-Code and automation tools like Terraform or Scalr.

Strong knowledge of cloud architecture, with an emphasis on maintaining service SLAs and ensuring high availability.

Experience with cloud security practices, SSO solutions, and authentication protocols (e.g., Auth0, SAML/OIDC, OAuth).

Familiarity with deploying and maintaining data processing frameworks and ML platforms such as Airflow, Airbyte, Superset, Metabase, Databricks, Snowflake, MLflow, etc., is advantageous.

Certifications such as AWS Certified DevOps Engineer - Professional or AWS Solutions Architect - Professional.

Experience in highly regulated industries.

Knowledge of advanced security practices and compliance frameworks (PCI-DSS, ISO 27001, SOC2).

Multi-region/multi-AZ architecture design for high availability and disaster recovery.

WHAT'S ON OFFER

We offer a professional and enjoyable working atmosphere.

We prioritize your long-term development.

We are dedicated to creating a future-ready digital bank platform.

Competitive salary

13th-month salary guarantee

Performance bonus

Access to professional English courses

Premium health insurance

Generous annual leave allowance

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666

We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

System, Devops

Location:

Ho Chi Minh, Ha Noi - Viet Nam

Working Policy:

Job ID:

J00771

Status:

Related Job:

Senior Full-Stack Engineer (AI Service Desk)

Ho Chi Minh - Viet Nam

Outsource

.NET
ReactJS
Azure

Create and maintain backend services and APIs using C# and .NET technologies Construct contemporary frontend applications and dashboard interfaces using React Establish scalable integrations between operational systems, analytics services, and reporting layers Possess full-stack development experience across frontend and backend systems Collaborate with engineering and product teams to translate business requirements into technical solutions Enhance application performance, maintainability, and scalability Engage in technical discussions, code reviews, and architecture decisions Contribute to engineering standards and development best practices

Negotiation

View details

Senior AI DevSecOps Engineer

Ho Chi Minh - Viet Nam

Product

Devops
AWS
Azure
Security

Management of CI/CD Pipeline: Ensure automation, security, and scalability across all stages of the development lifecycle. Infrastructure & Security: Design and implement secure multi-cloud infrastructure solutions leveraging cloud services, containerization, and orchestration tools. Policy as Code: Define and enforce security and compliance policies across Kubernetes clusters using OPA or Kyverno, ensuring guardrails are automated and auditable. AI & Platform Automation: Drive the adoption of AI-powered tools and workflows to automate infrastructure operations, optimize CI/CD pipelines, accelerate root cause analysis, improve security posture, and enhance engineering productivity. Observability & Alerting: Build and maintain a comprehensive observability stack with proactive alerting, dashboards, and runbooks for critical business flows and security events. Secret & Credential Management: Design and enforce secrets management practices across all environments ensuring zero hardcoded credentials in codebases and pipelines. Incident Response & On-Call: Own and continuously improve incident response processes, define runbooks, lead post-mortems, track MTTR, and participate in on-call rotation to maintain platform reliability and SLO adherence. Threat Modelling & Penetration Testing: Conduct regular threat modelling sessions with engineering teams and coordinate or perform penetration testing activities to proactively identify attack surfaces before they reach production. Code Security: Conduct regular code reviews and static/dynamic analysis to identify and remediate security vulnerabilities. Compliance and Best Practices: Ensure compliance with industry standards and best practices. Collaboration: Collaborate with development, operations, and security teams to foster a culture of automation and security-first thinking. Mentorship: Mentor junior engineers and other team members on security best practices. Documentation: Maintain thorough and up-to-date documentation of security policies, procedures, and incident reports. Trend Scouting: Stay updated with the latest trends in technology and AI to integrate innovative solutions into our processes.

Negotiation

View details

Senior Data Engineer (C++, Python, AI/LLM)

Ho Chi Minh - Viet Nam

Outsource

Data Engineering
C/C++
Python

Refine a wide array of structured and unstructured data to produce high-quality datasets for quantitative analysis and financial engineering. Improve data integrity and quality by creating validation tools and frameworks to assess the effectiveness of data enrichment pipelines. Gain expertise in machine learning, deep learning, and emerging AI/LLM applications, and analyze the underlying dynamics and behaviors in the data. Derive insights from large-scale datasets and collaborate with research teams to pinpoint opportunities for tradable signals. Create utility tools to automate software development, testing, deployment, and monitoring workflows. Offer technical support to global researchers, including diagnosing technical issues, troubleshooting Python and C++ code, and suggesting scalable fixes and improvements. Troubleshoot and resolve issues in C++ applications and data pipelines with a strong focus on performance, stability, correctness, and maintainability. Investigate and implement AI/LLM-based solutions to enhance data processing, workflow efficiency, troubleshooting, documentation, and research support processes.

Negotiation

View details