Platform Reliability Engineer

ABOUT CLIENT

Our client is a reputable company specializing in software development and IT consulting services

JOB DESCRIPTION

Maintain production reliability of the Linux-based research and trading platform within a globally distributed engineering team.

Respond quickly to production infrastructure issues.

Comprehend internal client needs and effectively communicate them to regional and global leadership.

Identify risks, develop contingency plans, and implement solutions to mitigate them.

Enhance the observability platform to monitor the performance and health of critical computing environments.

Take part in occasional on-call rotations and support on-call staff during their shifts.

Contribute to organizational knowledge through documentation, education, and writing maintainable code.

JOB REQUIREMENT

At least 2 years of experience in SRE, DevOps, or similar infrastructure engineering roles, with a preference for experience in the financial industry.

Knowledge of Linux system internals, including kernel operations, memory management, and performance optimization.

Familiarity with storage technologies, especially those used in high-performance computing (experience with GPFS is a bonus).

Broad understanding of IT infrastructure components such as networking, DNS, NTP/PTP, and NIS.

Proficiency in system automation, monitoring, and self-healing, with experience in Salt seen as a positive attribute.

Experience with container orchestration and virtualization technologies like Kubernetes, Nomad, and VMware.

Understanding of on-premises and cloud-based HPC infrastructure, with operational knowledge of Slurm and GPU considered a bonus.

Awareness of AI technologies and their applications in infrastructure automation and management.

Experience or strong interest in implementing AI/ML solutions for infrastructure optimization, anomaly detection, or predictive analytics.

Passion for technology and automation, with a deep sense of curiosity and ownership.

Hands-on problem-solving approach and enthusiasm for technology.

Excellent verbal and written English communication skills.

WHAT'S ON OFFER

Be part of a dynamic and passionate team working on cutting-edge projects using the latest technology.

Collaborate with experts from around the globe to enhance your skills and knowledge.

Embrace a culture of transparency and support, valuing individual growth and potential.

Additional month's salary and performance bonuses.

Comprehensive healthcare and accident insurance coverage.

Yearly health checkup package.

Various allowances such as lunch, marriage, newborn baby, bereavement, and more.

Well-equipped pantry for a comfortable lunch break.

Diverse sports and social activities like yoga, football, badminton, and tech clubs.

Annual company retreats and team-building events.

Recognition awards for outstanding individual and team performance and long-term service.

Professional development opportunities including advanced English and soft skills training.

Regular social events such as gatherings, games, birthday celebrations, and year-end parties.

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666

We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Devops

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Onsite

Job ID:

J01977

Status:

Related Job:

Senior Full-Stack Engineer (AI Service Desk)

Ho Chi Minh - Viet Nam

Outsource

Design and develop backend services and APIs using C# and .NET technologies Build modern frontend applications and dashboard interfaces using React Develop scalable integrations between operational systems, analytics services, and reporting layers Full-stack development experience across frontend and backend systems Work closely with engineering and product teams to translate business requirements into technical solutions Improve application performance, maintainability, and scalability Participate in technical discussions, code reviews, and architecture decisions Contribute to engineering standards and development best practices

Negotiation

View details

Senior AI DevSecOps Engineer

Ho Chi Minh - Viet Nam

Product

Devops
AWS
Azure
Security

Management of CI/CD Pipeline: Ensure automation, security, and scalability across all stages of the development lifecycle. Infrastructure & Security: Design and implement secure multi-cloud infrastructure solutions leveraging cloud services, containerization, and orchestration tools. Policy as Code: Define and enforce security and compliance policies across Kubernetes clusters using OPA or Kyverno, ensuring guardrails are automated and auditable. AI & Platform Automation: Drive the adoption of AI-powered tools and workflows to automate infrastructure operations, optimize CI/CD pipelines, accelerate root cause analysis, improve security posture, and enhance engineering productivity. Observability & Alerting: Build and maintain a comprehensive observability stack with proactive alerting, dashboards, and runbooks for critical business flows and security events. Secret & Credential Management: Design and enforce secrets management practices across all environments ensuring zero hardcoded credentials in codebases and pipelines. Incident Response & On-Call: Own and continuously improve incident response processes, define runbooks, lead post-mortems, track MTTR, and participate in on-call rotation to maintain platform reliability and SLO adherence. Threat Modelling & Penetration Testing: Conduct regular threat modelling sessions with engineering teams and coordinate or perform penetration testing activities to proactively identify attack surfaces before they reach production. Code Security: Conduct regular code reviews and static/dynamic analysis to identify and remediate security vulnerabilities. Compliance and Best Practices: Ensure compliance with industry standards and best practices. Collaboration: Collaborate with development, operations, and security teams to foster a culture of automation and security-first thinking. Mentorship: Mentor junior engineers and other team members on security best practices. Documentation: Maintain thorough and up-to-date documentation of security policies, procedures, and incident reports. Trend Scouting: Stay updated with the latest trends in technology and AI to integrate innovative solutions into our processes.

Negotiation

View details

Senior Data Engineer (C++, Python, AI/LLM)

Ho Chi Minh - Viet Nam

Outsource

Data Engineering
C/C++
Python

Refine a wide array of structured and unstructured data to produce high-quality datasets for quantitative analysis and financial engineering. Improve data integrity and quality by creating validation tools and frameworks to assess the effectiveness of data enrichment pipelines. Gain expertise in machine learning, deep learning, and emerging AI/LLM applications, and analyze the underlying dynamics and behaviors in the data. Derive insights from large-scale datasets and collaborate with research teams to pinpoint opportunities for tradable signals. Create utility tools to automate software development, testing, deployment, and monitoring workflows. Offer technical support to global researchers, including diagnosing technical issues, troubleshooting Python and C++ code, and suggesting scalable fixes and improvements. Troubleshoot and resolve issues in C++ applications and data pipelines with a strong focus on performance, stability, correctness, and maintainability. Investigate and implement AI/LLM-based solutions to enhance data processing, workflow efficiency, troubleshooting, documentation, and research support processes.

Negotiation

View details