Platform Reliability Engineer

ABOUT CLIENT

Our client is a reputable company specializing in software development and IT consulting services

JOB DESCRIPTION

Maintain production reliability of the Linux-based research and trading platform within a globally distributed engineering team.
Respond quickly to production infrastructure issues.
Comprehend internal client needs and effectively communicate them to regional and global leadership.
Identify risks, develop contingency plans, and implement solutions to mitigate them.
Enhance the observability platform to monitor the performance and health of critical computing environments.
Take part in occasional on-call rotations and support on-call staff during their shifts.
Contribute to organizational knowledge through documentation, education, and writing maintainable code.

JOB REQUIREMENT

At least 2 years of experience in SRE, DevOps, or similar infrastructure engineering roles, with a preference for experience in the financial industry.
Knowledge of Linux system internals, including kernel operations, memory management, and performance optimization.
Familiarity with storage technologies, especially those used in high-performance computing (experience with GPFS is a bonus).
Broad understanding of IT infrastructure components such as networking, DNS, NTP/PTP, and NIS.
Proficiency in system automation, monitoring, and self-healing, with experience in Salt seen as a positive attribute.
Experience with container orchestration and virtualization technologies like Kubernetes, Nomad, and VMware.
Understanding of on-premises and cloud-based HPC infrastructure, with operational knowledge of Slurm and GPU considered a bonus.
Awareness of AI technologies and their applications in infrastructure automation and management.
Experience or strong interest in implementing AI/ML solutions for infrastructure optimization, anomaly detection, or predictive analytics.
Passion for technology and automation, with a deep sense of curiosity and ownership.
Hands-on problem-solving approach and enthusiasm for technology.
Excellent verbal and written English communication skills.

WHAT'S ON OFFER

Be part of a dynamic and passionate team working on cutting-edge projects using the latest technology.
Collaborate with experts from around the globe to enhance your skills and knowledge.
Embrace a culture of transparency and support, valuing individual growth and potential.
Additional month's salary and performance bonuses.
Comprehensive healthcare and accident insurance coverage.
Yearly health checkup package.
Various allowances such as lunch, marriage, newborn baby, bereavement, and more.
Well-equipped pantry for a comfortable lunch break.
Diverse sports and social activities like yoga, football, badminton, and tech clubs.
Annual company retreats and team-building events.
Recognition awards for outstanding individual and team performance and long-term service.
Professional development opportunities including advanced English and soft skills training.
Regular social events such as gatherings, games, birthday celebrations, and year-end parties.

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Devops

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Onsite

Salary:

Negotiation

Job ID:

J01977

Status:

Close

Related Job:

Android Engineer (Java/Kotlin)

Ho Chi Minh - Viet Nam


Product

  • Android

Develop Android App part of various Services Develop new services and improve structures Analyze and apply new technologies to services

Negotiation

View details

C++ Engineer - Market Data

Ho Chi Minh, Ha Noi - Viet Nam


Product

  • C/C++

Maintain/enhance our legacy C++-based tick data processing platform as needed, demonstrating product ownership, and helping migrate datasets to our new tick data processing platform. Contribute to our new, modern C++-based tick data processing platform - enhancing the platform to support additional tick data feeds across asset classes, developing/reviewing the implementation of tick data-based interval features/statistics, and adding new functionality to the platform all while maintaining high software standards and best practices. Collaborate with the Research and Portfolio Management organizations to facilitate the transition from our legacy platform to our new platform, supporting their price volume data needs for signal generation.

Negotiation

View details

Platform Lead

Others - Singapore


Product

  • Backend
  • Devops
  • Data Engineering

Develop and expand distributed systems to handle large volumes of sensory, telemetry, and control data across cloud and edge environments, facilitating real-time connections for fleets of robots. Create the API Platform with a focus on high reliability, exceptional developer experience, and robust multimodal AI capabilities accessible through user-friendly APIs and SDKs. Establish extensive training and inference platforms for foundation models used in robot autonomy, teleoperation, and developer integrations. Devise data ingestion and streaming pipelines for real-time connectivity of robot fleets to the cloud, covering various data inputs such as video, LiDAR, joint states, and audio. Oversee and advance a modern cloud native infrastructure stack employing Kubernetes, Docker, and infrastructure as code tools. Ensure platform reliability through telemetry, monitoring, alerting, autoscaling, failover, and disaster recovery measures. Make infrastructure decisions pertaining to distributed storage, consensus protocols, GPU orchestration, network reliability, and API security. Foster collaboration across ML, robotics, and product teams to facilitate hardware in the loop simulation, policy rollout, continuous learning, and CI/CD workflows. Implement secure APIs featuring fine-grained access control, usage metering, rate limiting, and billing integration to accommodate a growing user base.

Negotiation

View details