Site Reliability Engineer (Shift-working)

ABOUT CLIENT

Our client is a global technology company that specializes in providing innovative IT solutions for the financial services industry

JOB DESCRIPTION

The Senior SRE plays a vital role in overseeing the everyday operations of the organization. It is crucial for this position to have a solid understanding of various technical aspects such as production system access and control, production deployment, Amazon Web Services, Kubernetes, continuous deployment, and systems observability.
 
Key Responsibilities
Take part in on-call rotations to provide round-the-clock support for critical systems.
Address system incidents promptly and effectively
Implement changes in staging and production environments
Collaborate with Platform Engineers to comprehend the changes
Establish deployment pipeline for changes
Comprehend the changes and build observability (monitoring and alert) as per the changes
Design and execute resiliency testing solutions
Continuously improve monitoring solutions
Create and update operational runbooks
Automate operational runbooks

JOB REQUIREMENT

Technical Skills
Proficient in Amazon Web Services
Proficient in Kubernetes system
Proficient in Python or Bash scripting
Familiarity with continuous deployment tools
Familiarity with Harness is a plus
Familiarity with infrastructure as code (IaC) tools, particularly Terraform
Experience with observability solutions like Prometheus and Grafana
Familiarity with SumoLogic is a plus
 
Soft Skills
Effective communication skills, fluent in English
Strong problem-solving abilities
Self-motivated and quick learner

WHAT'S ON OFFER

Attractive salary
13th-month salary and performance bonus
Professional English course available for all employees
Comprehensive health insurance package

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Information Technology & Services

Technical Skills:

Devops, AWS, Google Cloud

Location:

Ho Chi Minh, Ha Noi - Viet Nam

Working Policy:

Hybrid

Salary:

Negotiation

Job ID:

J01150

Status:

Close

Related Job:

Senior Deep Learning Engineer - AI for Wireless Systems

Ho Chi Minh, Ha Noi - Viet Nam


Computer Hardware

  • Machine Learning

Develop and test deep learning models for various wireless signal processing tasks including channel estimation, beam alignment, link adaptation, and scheduling. Utilize simulation tools and real-world datasets to create models that can be applied across different wireless scenarios. Build, train, and assess neural networks (e.g., CNNs, Transformers, GNNs) with PyTorch or TensorFlow. Engage in teamwork with researchers and system engineers to incorporate models into complete RAN systems. Enhance model efficiency for real-time processing and hardware acceleration. Participate in model assessment, performance comparison, and deployment preparation on GPU platforms.

Negotiation

View details

Director Engineering – Software Engineering and AI Inferencing Platforms

Ho Chi Minh, Ha Noi - Viet Nam


Computer Hardware

  • Management
  • Backend
  • Cloud
  • Data Engineering
  • AI

Lead and expand engineering teams in Vietnam across system software, data science, and AI platforms. Drive the creation, structure, and delivery of high-performance system software platforms that support AI products and services. Collaborate with global teams across Machine Learning, Inference Services, and Hardware/Software integration to guarantee performance, reliability, and scalability. Oversee the development and optimization of AI delivery platforms in Vietnam, including NIMs, Blueprints, and other flagship services. Collaborate with open-source and enterprise data and workflow ecosystems to advance accelerated AI factory, data science, and data engineering workloads. Promote continuous integration, continuous delivery, and engineering best practices across multi-site R&D Centers. Work with product management and other stakeholders to ensure enterprise readiness and customer impact. Establish and implement standard processes for large-scale, distributed system testing including stress, scale, failover, and resiliency testing. Ensure security and compliance testing aligns with industry standards for cloud and data center products. Mentor and develop talent within the organization, fostering a culture of quality and continuous improvement.

Negotiation

View details

Senior Natural Language Processing Engineer

Ho Chi Minh - Viet Nam


Computer Hardware

  • Machine Learning
  • NLP

Create and enhance AI language models for different NLP tasks such as translation and sentiment analysis. Deploy and integrate models for various language pairs to ensure smooth integration into projects. Work with cross-functional teams to support the effective deployment of NLP models. Constantly assess and enhance model performance to upkeep high standards.

Negotiation

View details