Lead Site Reliability Engineer

ABOUT CLIENT

Our client is a global technology company that specializes in providing innovative IT solutions for the financial services industry

JOB DESCRIPTION

Lead a team of SREs, offering technical guidance and coaching while promoting a culture of reliability and continuous improvement.
Define and advance SRE practices such as SLIs/SLOs, error budgets, and incident response processes across production systems.
Take ownership of the design and evolution of automated cloud operations, driving the adoption of Infrastructure-as-Code (Terraform, CloudFormation) and CI/CD pipelines.
Oversee major incident responses, prioritize rapid resolution, conduct root cause analysis, and implement preventive measures.
Collaborate closely with Development, DevOps, and Cloud Engineering teams to incorporate reliability and resilience at every delivery stage.
Establish and track key reliability metrics (availability, latency, error rates) and drive initiatives for continuous improvement.
Assess and implement AWS-native and third-party tools to enhance monitoring, alerting, and automation.
Serve as the main contact for Service Reliability topics with clients, ensuring transparency and alignment on reliability goals.
Ensure compliance with industry standards and internal policies related to security, audit, and operational risk.

JOB REQUIREMENT

Minimum 7 years of experience working as an SRE Engineer, with exposure to data platform solutions being beneficial.
Extensive experience with cloud platform, including IAM, ECS, EKS, Lambda, and CloudWatch.
Proficiency in deploying and managing containerized services, particularly on Kubernetes.
Hands-on experience with Infrastructure-as-Code and automation tools like Terraform or Scalr.
Strong knowledge of cloud architecture, with an emphasis on maintaining service SLAs and ensuring high availability.
Experience with cloud security practices, SSO solutions, and authentication protocols (e.g., Auth0, SAML/OIDC, OAuth).
Familiarity with deploying and maintaining data processing frameworks and ML platforms such as Airflow, Airbyte, Superset, Metabase, Databricks, Snowflake, MLflow, etc., is advantageous.
Certifications such as AWS Certified DevOps Engineer - Professional or AWS Solutions Architect - Professional.
Experience in highly regulated industries.
Knowledge of advanced security practices and compliance frameworks (PCI-DSS, ISO 27001, SOC2).
Multi-region/multi-AZ architecture design for high availability and disaster recovery.

WHAT'S ON OFFER

We offer a professional and enjoyable working atmosphere.
We prioritize your long-term development.
We are dedicated to creating a future-ready digital bank platform.
Competitive salary
13th-month salary guarantee
Performance bonus
Access to professional English courses
Premium health insurance
Generous annual leave allowance

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

System, Devops

Location:

Ho Chi Minh, Ha Noi - Viet Nam

Working Policy:

Job ID:

J00771

Status:

Close

Related Job:

Data Platforms Team Leader

Ho Chi Minh - Viet Nam


Product

  • Data Engineering
  • Devops
  • AI

#Team Leadership and Management: Lead and mentor agile team of data platform and support engineers towards self-organization and empowerment Conduct performance reviews and provide professional development opportunities Present team achievements in internal and external events#Data Platform Administration & Support​: Install, configure, maintain and monitor data platforms, including databases, filesystems, and other related technologies​ Regularly update and patch systems to maintain security and stability​ Identify, diagnose, and resolve data platform problems in a timely manner​ Implement and maintain backup and recovery strategies for data platforms, ensuring data integrity and availability​ Inform, support and educate users of historical and new platform features and related processes#Continuous Improvement, Technical Excellence and Accountability: Ensure proper data platform design and platform features enabling AI/ML model training/serving, data management, data quality tracking Oversee the technical design aligned with company strategy and policies Follow, select and apply the latest technologies and best practices in platform code, management, and AI/ML area#Delivery Tracking: Ensure projects are delivered on time, within scope, in alignment with product owner and users' expectations, including non-functional requirements (performance, security, stability, observability, continuity) Manage expectations, risks and resolve issues that arise during platform lifecycle#Collaboration with Product Owner (s) and Development Team (s): Work closely with the Product Owners and Development Teams depending on platform features and performance to align platform roadmaps with their goals Provide platform-level feasibility assessments for new features and enhancements Ensure clear communication and coordination between the team, Product Owner (s) and Development Team (s)#Automation, code quality and documentation: Implement and maintain quality standards for platform code Ensure regular testing and validation are implemented as vital part of automated CI/CD#Data Governance and Compliance: Ensure data platform features comply with relevant data privacy and security regulations and policies Manage data access and usage to protect personal data and other sensitive information

Negotiation

View details

Senior AI Data Platform Engineer

Ho Chi Minh - Viet Nam


Product

  • Data Engineering
  • Devops
  • AI

#ML Data Platform Engineering & Administration Install, configure, and maintain ML data platforms on top of Kubernetes, Object Storage, Cassandra, Postgres and related technologies Monitor platform performance and optimize as needed for reliability and efficiency#Platform Configuration and Maintenance Implement and manage platform configurations, ensuring adherence to best practices and security standards Regularly update and patch systems to maintain security and stability#Collaborate with Cross-Functional Teams Work closely with ML and data engineers and other roles to align IT needs and strategies Provide expert guidance on ML data platform best practices and optimization#Troubleshoot and Resolve Technical Issues Identify, diagnose, and resolve data platform problems in a timely manner Escalate complex issues to upper-level support when necessary#Backup and Recovery Management Implement and maintain backup and recovery strategies for data platforms, ensuring data integrity and availability#Maintain and Update Documentation Create and maintain documentation related to data platform administration, configuration, and maintenance Share knowledge with team members and contribute to a culture of continuous learning#Enhance Data Security and Compliance Ensure data platforms adhere to security best practices and comply with relevant regulations Stay up-to-date on industry trends and evolving security standards#Drive Continuous Improvement Evaluate and implement new technologies and techniques to enhance data platform performance and administration Proactively identify areas for improvement, prepare plan for implementation and get support from management and development teams Always prefer automation and code-first approach over hard to reproduce manual tasks

Negotiation

View details

Senior IT Infrastructure Engineer

Ho Chi Minh - Viet Nam


Product

  • System
  • IT inhouse

Plan, design, and build enterprise hardware systems such as PDC, DRC, evaluating and selecting the best options to meet the organization's needs. Operate, maintain, and upgrade VMWare, OpenStack environments, backup software and libraries. Maintain, plan, deploy and protect the infrastructure environment server hardware and firmware, consolidating and improving system components such as virtualization, backup servers, SAN, Hardware security module, rack, and storage systems mainly in PDC, DRC to meet the organization's current and future requirements. Participate in IT projects, providing technical expertise and guidance for infrastructurerelated components. Ensure information security by understanding key security objectives such as confidentiality, integrity, and availability, and implementing appropriate measures in the infrastructure environment, fulfill the requirements of internal and external audits. Design converged SAN architecture and maintain SAN devices such as switches and storage, upgrading firmware when needed. Perform capacity planning for free space, performance, and usage, and proactively monitor systems to identify and address performance issues. Manage and monitor data center systems such as fire alarm and air conditioning, and plan server room layout and cabling. Test all changes to SAN networks, hardware, and software to ensure smooth implementation and operation. Document issues and resolutions and prepare status reports and reviews for management. Foster a strong collaborative relationship with Vendor ( Hardware & Storage Team, Virtualization team, Backup Team), seamlessly merging the efforts into a harmonious and unified team. Hold primary accountability for maintaining the smooth operation and availability of the datacenter, in strict adherence to SLA requirements. Prioritize swift resolution of any emerging issues to minimize downtime.

Negotiation

View details