Senior Cloud Operations Engineer (Azure, AWS)

JOB DESCRIPTION

Proactive monitoring of Alerts through and take action as per the statement of work defined
Perform operations and administration support to the Azure and AWS VM’s and PaaS components
Perform patching, monitoring and other maintenance support activities to the servers support by Cloud services team
Linux experience for administration of VM’s
Change management - Change Coordination during maintenance/Releases (Admin Access for Application team, Change Ticket creation & Closure, Coordination, SPOI Communication, CAB Meeting)
Proactive Event Handling and Incident Management based on the Alerts from Monitoring Systems
Provide incident management support for the issues reported by customer and alerts generated by system
Backup and Restoration (testing and implementation)
Azure IMG recommendations and Cloud Optimization
Review and implement ASC/AWS Trusted Advisory recommendations
Prepare weekly and monthly reports on incidents, alerts, changes which can be presented to the customer
Prepare documents on the operations activities – Process documents, technical documents
Certificate Management (Tracking Certificate expire and Reporting, renewal of certificates)
Collaboration - Operational meetings with project and management Team. Provide clarification about Incidents, Requests, Problems, Changes and other Issues 

JOB REQUIREMENT

ITIL knowledge
Has 4+ years experience working on AWS and Azure
Having strong skills on following services:
AWS, Azure IAAS, PAAS services including App Proxy; CGW; VNET; NSGs; VPN; DNS; API Management; Stream Analytics
Azure PowerShell, Storage Account; ADLS; Key Vault & BYOK; IAM; Policies; Application Gateway
Azure Automaton, Azure Kubernetes, IOT, Linux, Azure DevOps, Azure backup
Cloud monitoring, load balancer, NACL, Cloud security, Application Insights, Log Analytics
Terraform, Server less , Docker , Container
Collaborate efficiently
Clear & crisp email writing
Should be good in English communications both written & verbal 

WHAT'S ON OFFER

Working in one of the Best Places to Work in Vietnam
Join a dynamic and fast growing global company (English-speaking environment)
13th-month salary bonus + attractive performance bonus (you'll love it!) + annual performance appraisal
100% monthly basic salary and mandatory social insurances in 2-month probation
Onsite opportunities: short-term and long-term assignments
15++ days of annual leave + 1 day of birthday leave
Premium health insurance for employee and 02 family members
Flexible working time
Lunch and parking allowance
Various training on hot-trend technologies/ foreign language (English/Chinese/Japanese) and soft-skills
Fitness & sport activities: football, badminton, yoga, Aerobic
Free in-house entertainment facilities and snack
Join in various team building, company trip, year-end party, tech talks and a lot of charity events  

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Cloud, Azure, AWS

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Job ID:

J00620

Status:

Close

Related Job:

Software Engineer (Node.js) - Database

Ho Chi Minh - Viet Nam


Product

  • NodeJS

Design system architectures, establish coding standards, and construct cohesive, cloud-native solutions. Develop high-quality Node.js code, optimize system performance, and tackle complex software integration challenges. Oversee the testing, deployment, and comprehensive documentation of integrated systems. Mentor less-experienced engineers, engage in cross-functional teamwork, and ensure solutions meet business requirements and international standards. Participate actively in all Agile software development phases, including creating user stories and executing sprint planning Engage with multinational companies, demonstrating flexibility to occasionally adapt to US and EU time zones.

Negotiation

View details

Software Engineer (Node.js) - Platform Security

Ho Chi Minh - Viet Nam


Product

  • NodeJS

Design system architectures, establish coding standards, and construct cohesive, cloud-native solutions. Develop high-quality Node.js code, strengthen system security and reliability, and tackle complex software integration challenges. Design and implement platform security controls across web applications, APIs, and cloud services, including authentication, authorization, session management, secrets management, encryption, and audit logging. Identify and remediate security risks through threat modeling, secure code reviews, automated security testing, dependency scanning, and investigation of security-related issues. Oversee the testing, deployment, and comprehensive documentation of integrated systems. Mentor less-experienced engineers, engage in cross-functional teamwork, and ensure solutions meet business requirements and international standards. Participate actively in all Agile software development phases, including creating user stories and executing sprint planning Engage with multinational companies, demonstrating flexibility to occasionally adapt to US and EU time zones.

Negotiation

View details

AI Agent Ops Engineer

Ho Chi Minh - Viet Nam


Product

  • AI

#Agent Engineering & operation Design, build, and maintain production-grade AI agent systems, including: context engineering and instruction architecture, prompt hardening and safe execution boundaries, tool integrations and multi-step orchestration, memory strategies and reliability patterns. Own the full agent lifecycle: prototype → evaluate → deploy → monitor → iterate. Build and maintain an evaluation pipeline to measure agent quality, catch regressions, and enforce deployment gates (golden datasets, scenario suites, automated checks). Instrument agents and agent platforms for production observability: structured logging, tracing, and metrics; latency and cost monitoring; tool-call success rates and failure analysis. Define operational readiness standards including: rollback criteria, incident response playbooks, recovery paths for common failure modes.#Team Enablement & Coaching Embed with product engineering teams to identify high-value use cases ready for agent automation. We will be operating in a Central Agent Ops role enabling Ai product builders through AI enablers. Translate business workflows into agent-executable tasks with clear: contact boundaries/interfaces, assumptions and inputs/outputs, failure modes and safe fallbacks. Deliver targeted coaching to engineers on: context engineering best practices, harness design and regression testing patterns, agent skill design and tool-contract discipline. Reduce onboarding time for teams adopting AI capabilities-from first conversation to a production-ready agent. Train product engineers to extend and maintain agent skills independently.#Standards & Knowledge operations Author and maintain org-level standards for agents, including: naming conventions, context file structures and ownership rules, skill interface contracts (inputs/outputs, invariants, error handling), evaluation criteria and release quality bars. Establish and enforce "repo-as-discipline" practices so agent knowledge is: versioned, reviewable, discoverable, reusable; not trapped in prompt snippets or individual heads. Build and grow a shared agent skills library that teams can reuse and extend. Track and aggregate AI tooling/framework updates and external best practices, serving as a central intake so product teams don't each have to follow the entire AI landscape. Run internal knowledge-sharing sessions, showcases, and retrospectives to propagate learnings efficiently.

Negotiation

View details