Senior DevOps Engineer

JOB DESCRIPTION

Work within the team on the various products that the team supports.
Develop tools to improve our ability to rapidly deploy and effectively monitor services in a large-scale distributed environment.
Work with teams to design, develop, and implement innovative software solutions related to the DevOps and Agile transformation of the enterprise.
Ensure Cloud environments are compliant with security policies.
Find out the solutions to achieve highly available, highly scalable systems and reliability.
Maintain, support, and enhance CI/CD environment.
All are monitored and measured.
Evaluate infrastructure cost and find out the solution to optimize cost.
Troubleshoot and performs root cause analysis as well as implement corrective/preventive actions when needed.
Define and document best practices and operational procedures regarding solution deployment and infrastructure maintenance to ensure a smooth handover to other teams.

JOB REQUIREMENT

Must have:
3+ years of professional DevOps or Site Reliability Engineering experience in a fast-paced work environment.
Solid understanding of Linux and Container technology.
Hands-on knowledge of Docker.
Experience hands-on Infrastructure and Configuration as Code abilities e.g. Terraform(strongly preferred), Ansible, Packer.
Experience in building CI/CD pipeline automation, tooling (Github Action, Jenkins(strongly preferred)), and Compliance as code.
Experience with cloud services is essential, in particular, our core AWS Technologies (Organizations, Account Design, VPC, Subnet and Network segmentation, EC2, ASG, Lambda, S3, SQS, SNS, ECS, EKS, RDS, Lambda, Cloudwatch, etc).
Ability to create scripts using Bash, Python, or Golang. Must have the habit of cleaning code, reusing code, and implementing the unit test.
Excellence in analytical and problem-solving skills.
English communication, focus on writing.
Experience with Kubernetes (K8S).
Agile development experience is a plus.
Nice to have:
Strongly preferred: AWS Certificates (SysOps, DevOps, SAA, SAP)
Hashicorp Certificate or Experience with Harshicorp stacks such as Terraform Cloud/Terraform Enterprise, Packer, and Vault.
Kubernetes certifications (CKA/CKD).

WHAT'S ON OFFER

Great salary package and Semiannual performance-salary review.
100% official salary during the probation period.
13th-month salary & bonus (Token Bonus/ Investment Allocation)
Full-paid compulsory insurance according to Vietnam Labor Law
Premium Healthcare insurance (support for spouse and children)
12 days annual leave & other leaves, as below:
Birthday Leave: Company encourages employees to take time off and spend it with their loved ones on their birthdays with 1 day of paid leave.
Charity Leave: In an effort to encourage all of our employees to get involved in various charity organizations, we provide 2 days of paid leave per year to yours perform volunteer work
Self-development Leave: Company values employee self-development, employees not only are given a budget to be used solely to pursue personal development but also 2 days of paid leave per year for these activities.
Transportation & Lunch Allowances
Macs/ Laptop and other work-equipments will be provided
Yearly company trips and many outing trips/ team-bonding activities
Unlimited potential for the career path
Budget for your training & self-development
Fantastic yet professional working environment
Lovely, friendly, and talented colleagues
Working hours: 8 hours x 5 days/week (Monday to Friday) with flexible working hours
A pantry is full of tasty food & beverage.
Other benefits will surprise you!

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product

Technical Skills:

Devops, AWS, Google Cloud

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Job ID:

J01286

Status:

Close

Related Job:

Software Engineer (Node.js) - Database

Ho Chi Minh - Viet Nam


Product

  • NodeJS

Design system architectures, establish coding standards, and construct cohesive, cloud-native solutions. Develop high-quality Node.js code, optimize system performance, and tackle complex software integration challenges. Oversee the testing, deployment, and comprehensive documentation of integrated systems. Mentor less-experienced engineers, engage in cross-functional teamwork, and ensure solutions meet business requirements and international standards. Participate actively in all Agile software development phases, including creating user stories and executing sprint planning Engage with multinational companies, demonstrating flexibility to occasionally adapt to US and EU time zones.

Negotiation

View details

Software Engineer (Node.js) - Platform Security

Ho Chi Minh - Viet Nam


Product

  • NodeJS

Design system architectures, establish coding standards, and construct cohesive, cloud-native solutions. Develop high-quality Node.js code, strengthen system security and reliability, and tackle complex software integration challenges. Design and implement platform security controls across web applications, APIs, and cloud services, including authentication, authorization, session management, secrets management, encryption, and audit logging. Identify and remediate security risks through threat modeling, secure code reviews, automated security testing, dependency scanning, and investigation of security-related issues. Oversee the testing, deployment, and comprehensive documentation of integrated systems. Mentor less-experienced engineers, engage in cross-functional teamwork, and ensure solutions meet business requirements and international standards. Participate actively in all Agile software development phases, including creating user stories and executing sprint planning Engage with multinational companies, demonstrating flexibility to occasionally adapt to US and EU time zones.

Negotiation

View details

AI Agent Ops Engineer

Ho Chi Minh - Viet Nam


Product

  • AI

#Agent Engineering & operation Design, build, and maintain production-grade AI agent systems, including: context engineering and instruction architecture, prompt hardening and safe execution boundaries, tool integrations and multi-step orchestration, memory strategies and reliability patterns. Own the full agent lifecycle: prototype → evaluate → deploy → monitor → iterate. Build and maintain an evaluation pipeline to measure agent quality, catch regressions, and enforce deployment gates (golden datasets, scenario suites, automated checks). Instrument agents and agent platforms for production observability: structured logging, tracing, and metrics; latency and cost monitoring; tool-call success rates and failure analysis. Define operational readiness standards including: rollback criteria, incident response playbooks, recovery paths for common failure modes.#Team Enablement & Coaching Embed with product engineering teams to identify high-value use cases ready for agent automation. We will be operating in a Central Agent Ops role enabling Ai product builders through AI enablers. Translate business workflows into agent-executable tasks with clear: contact boundaries/interfaces, assumptions and inputs/outputs, failure modes and safe fallbacks. Deliver targeted coaching to engineers on: context engineering best practices, harness design and regression testing patterns, agent skill design and tool-contract discipline. Reduce onboarding time for teams adopting AI capabilities-from first conversation to a production-ready agent. Train product engineers to extend and maintain agent skills independently.#Standards & Knowledge operations Author and maintain org-level standards for agents, including: naming conventions, context file structures and ownership rules, skill interface contracts (inputs/outputs, invariants, error handling), evaluation criteria and release quality bars. Establish and enforce "repo-as-discipline" practices so agent knowledge is: versioned, reviewable, discoverable, reusable; not trapped in prompt snippets or individual heads. Build and grow a shared agent skills library that teams can reuse and extend. Track and aggregate AI tooling/framework updates and external best practices, serving as a central intake so product teams don't each have to follow the entire AI landscape. Run internal knowledge-sharing sessions, showcases, and retrospectives to propagate learnings efficiently.

Negotiation

View details