Platform Reliability Engineer

ABOUT CLIENT

Our client is a reputable company specializing in software development and IT consulting services

JOB DESCRIPTION

Maintain production reliability of the Linux-based research and trading platform within a globally distributed engineering team.
Respond quickly to production infrastructure issues.
Comprehend internal client needs and effectively communicate them to regional and global leadership.
Identify risks, develop contingency plans, and implement solutions to mitigate them.
Enhance the observability platform to monitor the performance and health of critical computing environments.
Take part in occasional on-call rotations and support on-call staff during their shifts.
Contribute to organizational knowledge through documentation, education, and writing maintainable code.

JOB REQUIREMENT

At least 2 years of experience in SRE, DevOps, or similar infrastructure engineering roles, with a preference for experience in the financial industry.
Knowledge of Linux system internals, including kernel operations, memory management, and performance optimization.
Familiarity with storage technologies, especially those used in high-performance computing (experience with GPFS is a bonus).
Broad understanding of IT infrastructure components such as networking, DNS, NTP/PTP, and NIS.
Proficiency in system automation, monitoring, and self-healing, with experience in Salt seen as a positive attribute.
Experience with container orchestration and virtualization technologies like Kubernetes, Nomad, and VMware.
Understanding of on-premises and cloud-based HPC infrastructure, with operational knowledge of Slurm and GPU considered a bonus.
Awareness of AI technologies and their applications in infrastructure automation and management.
Experience or strong interest in implementing AI/ML solutions for infrastructure optimization, anomaly detection, or predictive analytics.
Passion for technology and automation, with a deep sense of curiosity and ownership.
Hands-on problem-solving approach and enthusiasm for technology.
Excellent verbal and written English communication skills.

WHAT'S ON OFFER

Be part of a dynamic and passionate team working on cutting-edge projects using the latest technology.
Collaborate with experts from around the globe to enhance your skills and knowledge.
Embrace a culture of transparency and support, valuing individual growth and potential.
Additional month's salary and performance bonuses.
Comprehensive healthcare and accident insurance coverage.
Yearly health checkup package.
Various allowances such as lunch, marriage, newborn baby, bereavement, and more.
Well-equipped pantry for a comfortable lunch break.
Diverse sports and social activities like yoga, football, badminton, and tech clubs.
Annual company retreats and team-building events.
Recognition awards for outstanding individual and team performance and long-term service.
Professional development opportunities including advanced English and soft skills training.
Regular social events such as gatherings, games, birthday celebrations, and year-end parties.

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Devops

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Onsite

Salary:

Negotiation

Job ID:

J01977

Status:

Active

Related Job:

Android Engineer - Hanoi

Ha Noi - Viet Nam


Product

  • Android

Creating and managing Android applications using Kotlin Constructing Android services for production use and contributing to live service operations Utilizing Jetpack Compose to design modern Android UI Incorporating asynchronous programming through Coroutines and Flow Developing scalable Android app architecture with modularization and dependency injection Collaborating with cross-functional teams through effective communication

Negotiation

View details

iOS Engineer - Hanoi

Ha Noi - Viet Nam


Product

  • iOS

Create and update iOS applications with Swift Utilize UIKit and SwiftUI for building user interfaces Integrate and design APIs for effective data processing Employ reactive and asynchronous programming for strong app architecture Communicate effectively with cross-functional teams Enhance code quality, performance, and maintainability of iOS applications

Negotiation

View details

Engineering Manager (Data Platform)

Ho Chi Minh - Viet Nam


Offshore

  • Data Engineering
  • Management

Agile Team Leadership: Guide and coach Agile teams to uphold engineering standards, manage sprint backlogs, clarify responsibilities, ensure code quality, enforce development guardrails, and drive rigorous testing practices. Agile Data Delivery: Oversee Agile execution across data platforms, maintaining excellence in data quality, testing, code review practices, CI/CD pipelines, documentation, and operational readiness. Cross-Functional Collaboration: Partner with data architects, product managers, analytics teams, platform engineers, and governance stakeholders to deliver data capabilities aligned with business priorities. Roadmap Ownership: Lead the execution of the data engineering roadmap, balancing immediate delivery needs with long-term platform sustainability. Architecture & Design: Contribute to the design of data platform architecture across ingestion, transformation, storage, and consumption layers. Engineer Development: Coach engineers to become T-shaped professionals, capable of working across batch processing, streaming, analytics engineering, and platform operations. Technical Debt Remediation: Own and prioritize the resolution of technical and data debt, including legacy pipelines, performance bottlenecks, and data quality issues. Modern Practices: Stay current with evolving data engineering tools, methodologies, and patterns-particularly within the Databricks ecosystem. Lifecycle Accountability: Ensure end-to-end ownership of data solutions, from design and build through deployment, monitoring, and ongoing support. Team Empowerment: Foster self-sufficient, disciplined teams accountable for the reliability and resilience of data products. Process Excellence: Lead initiatives to enhance data delivery through automation, observability, and operational best practices. Continuous Improvement: Inspire teams to innovate, experiment, and embrace continuous delivery as part of their culture. Career Growth: Drive career development for data engineers, partnering with HR to manage performance and define growth pathways.

Negotiation

View details