Platform Reliability Engineer

ABOUT CLIENT

Our client is a reputable company specializing in software development and IT consulting services

JOB DESCRIPTION

Maintain production reliability of the Linux-based research and trading platform within a globally distributed engineering team.
Respond quickly to production infrastructure issues.
Comprehend internal client needs and effectively communicate them to regional and global leadership.
Identify risks, develop contingency plans, and implement solutions to mitigate them.
Enhance the observability platform to monitor the performance and health of critical computing environments.
Take part in occasional on-call rotations and support on-call staff during their shifts.
Contribute to organizational knowledge through documentation, education, and writing maintainable code.

JOB REQUIREMENT

At least 2 years of experience in SRE, DevOps, or similar infrastructure engineering roles, with a preference for experience in the financial industry.
Knowledge of Linux system internals, including kernel operations, memory management, and performance optimization.
Familiarity with storage technologies, especially those used in high-performance computing (experience with GPFS is a bonus).
Broad understanding of IT infrastructure components such as networking, DNS, NTP/PTP, and NIS.
Proficiency in system automation, monitoring, and self-healing, with experience in Salt seen as a positive attribute.
Experience with container orchestration and virtualization technologies like Kubernetes, Nomad, and VMware.
Understanding of on-premises and cloud-based HPC infrastructure, with operational knowledge of Slurm and GPU considered a bonus.
Awareness of AI technologies and their applications in infrastructure automation and management.
Experience or strong interest in implementing AI/ML solutions for infrastructure optimization, anomaly detection, or predictive analytics.
Passion for technology and automation, with a deep sense of curiosity and ownership.
Hands-on problem-solving approach and enthusiasm for technology.
Excellent verbal and written English communication skills.

WHAT'S ON OFFER

Be part of a dynamic and passionate team working on cutting-edge projects using the latest technology.
Collaborate with experts from around the globe to enhance your skills and knowledge.
Embrace a culture of transparency and support, valuing individual growth and potential.
Additional month's salary and performance bonuses.
Comprehensive healthcare and accident insurance coverage.
Yearly health checkup package.
Various allowances such as lunch, marriage, newborn baby, bereavement, and more.
Well-equipped pantry for a comfortable lunch break.
Diverse sports and social activities like yoga, football, badminton, and tech clubs.
Annual company retreats and team-building events.
Recognition awards for outstanding individual and team performance and long-term service.
Professional development opportunities including advanced English and soft skills training.
Regular social events such as gatherings, games, birthday celebrations, and year-end parties.

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Devops

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Onsite

Salary:

Negotiation

Job ID:

J01977

Status:

Close

Related Job:

DevOps Engineer

Others - Viet Nam


Product

  • Devops
  • Kubernetes
  • Network

Operate and evolve our Kubernetes platform across multiple clusters and environments (Prod, Dev, hybrid on-prem and public cloud), covering control plane operations, node lifecycle, upgrades, and autoscaling at every layer (Cluster Autoscaler, HPA, KEDA). Architect and manage hybrid cloud infrastructure spanning on-premises and public clouds (GCP, AWS), including workload placement, cross-cloud networking, and unified resource management. Own the CI/CD and GitOps experience end-to-end: container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD. Own the observability stack as a single pane of glass across all clusters: Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus -- and help push toward agent-assisted SRE workflows. Manage and improve our inference platform: vLLM serving and AIBrix for multi-model orchestration and autoscaling across a fleet of NVIDIA GPUs. Operate platform services: Kafka, Redis, PostgreSQL, OpenSearch. Manage identity and access via Keycloak integrated with Google Workspace; harden SSO, RBAC, and secrets management across the platform. Harden network security across private load balancers, firewalls, and VPC segmentation; design and maintain hub-and-spoke / multi-AZ topologies. Support training infrastructure: self-service VM provisioning, RunPod burst capacity, Weights and Biases integration. Drive infrastructure reliability, cost efficiency, and capacity planning as the platform scales.

Negotiation

View details

Platform Engineer

Ho Chi Minh - Viet Nam


Product

  • Backend
  • Devops
  • Data Engineering

Build and maintain distributed infrastructure handling telemetry, sensory, and control data across cloud and edge environments Design and operate data ingestion and streaming pipelines connecting robot fleets to the cloud in real time, covering video, joint states, audio, and LiDAR Develop and maintain backend services and APIs that power the Company's developer-facing platform, with a focus on reliability and developer experience Manage and evolve cloud native infrastructure using Kubernetes, Docker, and infrastructure as code tooling Ensure platform reliability through monitoring, alerting, autoscaling, failover, and incident response Support ML and robotics teams with data infrastructure for training pipelines, policy rollout, and hardware-in-the-loop simulation Implement secure APIs with access control, rate limiting, and usage metering as we scale

Negotiation

View details

Software Engineer (Digital Twin)

Ho Chi Minh - Viet Nam


Product

  • Python
  • C/C++

Build and maintain high-fidelity digital twin environments for Asimov across MuJoCo, Isaac Sim, and Unreal Engine, calibrated to real hardware behavior. Design and own the systems -- not just the environments -- that let locomotion, autonomy, and perception teams generate, validate, and iterate on simulation scenarios at scale. Build pipelines for asset import, USD and MJCF workflows, sensor modeling, and real-to-sim calibration to keep digital twins synchronized with evolving hardware. Develop photorealistic rendering pipelines in Unreal Engine for synthetic data generation and perception model training. Work with hardware and mechatronics teams to model actuator dynamics, contact physics, and structural behavior, ensuring simulation parameters reflect physical ground truth. Integrate digital twin environments with the Company's locomotion training pipeline (Cyclotron) and autonomy stack, enabling teams to run experiments and close the sim-to-real gap. Contribute to the open-source Asimov simulation stack, including tooling, documentation, and reproducible environment workflows.

Negotiation

View details