Senior DevOps Engineer (Cloud Solutions)

JOB DESCRIPTION

Planning, design, management, maintenance and support of cloud infrastructure for high-traffic workloads that operate at an enterprise scale
Drive automation of tasks and implementing of infrastructure services
Identify and rectify potential risks within the infrastructure, network or security
Research, setup, testing and implementation of technologies and solutions to improve the performance, reliability, availability, security and efficiency of infrastructure on AWS
Troubleshoot, perform root cause analysis and working closely with the developers to implement corrective/preventive actions during and after an incident

JOB REQUIREMENT

8+ years of experience with using a broad range of AWS technologies (e.g. EC2, IAM, VPC, CloudWatch, EKS, ECS, Security Hub, DynamoDB, SecretManager, GuardDuty, etc)
Solid experience in Terraform as Infrastructure as Code
Knowledge of Golang and Python is a plus
Experience with containerized workloads
Experienced in a 24x7x365 uptime Amazon AWS environment leveraging git repositories and CI/CD tools like Jenkins
Ability to analyze and resolve complex infrastructure resource and application deployment issues (e.g. by using APM tools like NewRelic, Dynatrace, etc)
Knows the best practice and cloud security (AWS Well-Architected Framework)
 
Nice to have
AWS Data Engineering expertise (e.g. AWS Certified Big Data - Specialty)
Experienced with Data Engineering services like Lake Formation, Glue, Athena, Redshift, Sagemaker, Kineses, Kafka, etc.
Experienced with SQL and NoSQL Databases like DynamoDB, RDS Aurora, MySQL, ElasticSearch, Solr, etc.

WHAT'S ON OFFER

We are an equal opportunity employer and do not discriminate based on gender, race, age, religion, disability, or other local protected class. We are committed to cultivating an inclusive environment for all employees, and we welcome the diversity that you will bring!
If you are looking for a rapid-growth environment and great teams to work with, you should apply now.
We are sorry to inform you that only shortlisted candidates will be notified as we may be overwhelmed by the number of applicants coming into our system; hence if you do not get a reply from us - don’t give up on us just yet!
18 days Annual leaves
Quarter bonus based on employee performance and company's business
Health insurance, private insurance provided that covers yourself and your immediate dependents (spouse and children if any)
Laptop provided
Work from home allowances
Wellness benefit (cover for gym membership etc; can also be used to top up personal insurance too) up to 60usd/quarter
Any government-regulated perks
We have an L&D budget that supports the following:
Online or classroom courses held by an external provider
Conferences, workshops, or seminars
Coursework for a relevant diploma, degree, or professional certification

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Internet, Payment, Product

Technical Skills:

Devops, AWS, Security

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Salary:

Negotiation

Job ID:

J00564

Status:

Close

Related Job:

DevOps Engineer

Others - Viet Nam


Product

  • Devops
  • Kubernetes
  • Network

Operate and evolve our Kubernetes platform across multiple clusters and environments (Prod, Dev, hybrid on-prem and public cloud), covering control plane operations, node lifecycle, upgrades, and autoscaling at every layer (Cluster Autoscaler, HPA, KEDA). Architect and manage hybrid cloud infrastructure spanning on-premises and public clouds (GCP, AWS), including workload placement, cross-cloud networking, and unified resource management. Own the CI/CD and GitOps experience end-to-end: container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD. Own the observability stack as a single pane of glass across all clusters: Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus -- and help push toward agent-assisted SRE workflows. Manage and improve our inference platform: vLLM serving and AIBrix for multi-model orchestration and autoscaling across a fleet of NVIDIA GPUs. Operate platform services: Kafka, Redis, PostgreSQL, OpenSearch. Manage identity and access via Keycloak integrated with Google Workspace; harden SSO, RBAC, and secrets management across the platform. Harden network security across private load balancers, firewalls, and VPC segmentation; design and maintain hub-and-spoke / multi-AZ topologies. Support training infrastructure: self-service VM provisioning, RunPod burst capacity, Weights and Biases integration. Drive infrastructure reliability, cost efficiency, and capacity planning as the platform scales.

Negotiation

View details

Platform Engineer

Ho Chi Minh - Viet Nam


Product

  • Backend
  • Devops
  • Data Engineering

Build and maintain distributed infrastructure handling telemetry, sensory, and control data across cloud and edge environments Design and operate data ingestion and streaming pipelines connecting robot fleets to the cloud in real time, covering video, joint states, audio, and LiDAR Develop and maintain backend services and APIs that power the Company's developer-facing platform, with a focus on reliability and developer experience Manage and evolve cloud native infrastructure using Kubernetes, Docker, and infrastructure as code tooling Ensure platform reliability through monitoring, alerting, autoscaling, failover, and incident response Support ML and robotics teams with data infrastructure for training pipelines, policy rollout, and hardware-in-the-loop simulation Implement secure APIs with access control, rate limiting, and usage metering as we scale

Negotiation

View details

Software Engineer (Digital Twin)

Ho Chi Minh - Viet Nam


Product

  • Python
  • C/C++

Build and maintain high-fidelity digital twin environments for Asimov across MuJoCo, Isaac Sim, and Unreal Engine, calibrated to real hardware behavior. Design and own the systems -- not just the environments -- that let locomotion, autonomy, and perception teams generate, validate, and iterate on simulation scenarios at scale. Build pipelines for asset import, USD and MJCF workflows, sensor modeling, and real-to-sim calibration to keep digital twins synchronized with evolving hardware. Develop photorealistic rendering pipelines in Unreal Engine for synthetic data generation and perception model training. Work with hardware and mechatronics teams to model actuator dynamics, contact physics, and structural behavior, ensuring simulation parameters reflect physical ground truth. Integrate digital twin environments with the Company's locomotion training pipeline (Cyclotron) and autonomy stack, enabling teams to run experiments and close the sim-to-real gap. Contribute to the open-source Asimov simulation stack, including tooling, documentation, and reproducible environment workflows.

Negotiation

View details