Site Reliability Engineer

JOB DESCRIPTION

Maintain systems and troubleshoot system issues.
Identifying bottleneck in various Java applications and implement performance improvements.
Identify and analyze user requirements.
Prioritize, assign, and execute tasks throughout the software development life cycle.
Develop, configure, and deploy tools for cloud-based systems and services.
Containerize new and legacy applications.
Maintain awareness of new and emerging technologies.
Support development and operations teams.
Enhance, modify or debug developer code as needed.

JOB REQUIREMENT

Must-have
Understanding of an object-orientated language, preferably the latest version of Java (with experience in Hibernate, Multi-thread, Spring Boot)
Experience in configuration, in Jenkins for CI/CD pipeline creation, automation scripts and Kubernetes implementation with Google.
Proficiency in supporting a 24×7 critical operation.
Experience in a cloud computing platform and associated automation patterns it provides, preferably GCP.
Proficient in production systems design including High Availability, Disaster Recovery, Performance, Efficiency, and Security user, application performance, system, log, time-series, and dashboarding.
Familiarity with Open-Source concepts and tools like Prometheus, Grafana, ELK etc. 
Proficient in a modern infrastructure automation toolkit such as Terraform/Helm
Proficient in a Linux or Unix based environment.
Experience in destructive testing methodologies and tools such as chaos monkey
Experience in defensive coding practices and patterns for high availability
Nice-to-have
Experience in a cloud computing platform and associated automation patterns it provides, preferably GCP
Proficient in a modern scripting language like GO or Python
Knowledge of APM fundamentals or experience in tools like New Relic or AppDynamics.

WHAT'S ON OFFER

Open to deal base salary with additional project allowances.
Full salary during probation & Full coverage of social insurance.
Performance & salary review: twice a year
Monthly childcare support.
Premium Healthcare insurance and Health check-up services for employee and family ones.
15 Annual Leaves plus 10 days for Bereavement leave and 1.5 months for Paternity leave.
Premium package at top Gym service provider.
Diverse internal activities: Football, Billiards, Badminton, E-sport clubs & other regular company events.
Frequent opportunities to travel to US headquarter from 3-6 months.
Free parking for motorbike and car

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Devops, Java

Location:

Ho Chi Minh, Da Nang - Viet Nam

Working Policy:

Job ID:

J01196

Status:

Close

Related Job:

AI Engineer

Ho Chi Minh - Viet Nam


Product

  • AI
  • Python

#AI-Powered Insurance Solutions Develop and deploy machine learning models to support insurance recommendation, pricing, risk assessment, and claims automation Build and maintain AI-powered features integrated into our embedded insurance platform - from data pipelines to inference APIs Design and implement RAG (Retrieval-Augmented Generation) pipelines and LLMbased assistants for internal and customer-facing use cases Collaborate with product and engineering teams to translate business requirements into practical AI solutions Participate in the full ML/AI lifecycle: data collection, feature engineering, model training, evaluation, deployment, and monitoring Contribute to our internal AI agent framework - building, testing, and improving multiagent workflows using LangChain, LangGraph, or similar tools Ensure AI models and services meet performance, reliability, and compliance standards in a regulated (insurance) environment Participate in developing CI/CD workflows for AI/ML/Data pipelines. Contribute to Agent Development Toolkits (ADK) and the automated Software Development Life Cycle (SDLC) that accelerate AI-native feature delivery#R&D & Engineering Tasks Research and experiment with state-of-the-art AI/ML techniques and evaluate their applicability to insurance use cases Prototype new AI capabilities quickly, iterate based on feedback, and graduate successful experiments into production Document experiments, model architectures, and engineering decisions clearly for team knowledge sharing Stay current with the AI ecosystem (LLM advances, agent frameworks, tooling) and bring relevant insights back to the team Support Full Stack Engineers in using the AI Development Toolkit (ADK) to develop applications - guiding and enabling their use of the toolkit rather than building the applications yourself - and maintain a basic understanding of the SDLC to help ensure consistent engineering standards across the platform Contribute to and continuously enhance the system knowledge base, ensuring documentation, patterns, and learnings are accessible to the agents

Negotiation

View details

Backend Engineer

Ho Chi Minh - Viet Nam


Product

  • Typescript
  • NodeJS
  • Python
  • AI

Build the core gateway: a unified, OpenAI-compatible API in front of multiple providers (OpenAI, Anthropic, Google, plus self-hosted and OSS models). Own provider routing and reliability: load balancing, automatic failover, and cost and latency-aware routing. Build billing and metering that is correct, not approximate: per-request token accounting, usage ledgers, cost attribution per team, user, and key, budgets, and spend limits. Ship org controls: API key management, per-team and per-user quotas, rate limiting, and RBAC. Handle streaming and performance: low-overhead proxying, streaming responses, connection handling, and caching where it helps. Contribute to the Jan Agent and connect it to the router: route its model and tool calls through the gateway, and make agent traffic first-class in metering, controls, and observability. Make deliberate speed-versus-correctness calls: move fast where iteration is cheap, refuse to cut corners where a bug means a bad charge or a leaked key, and pay down debt on your own initiative.

Negotiation

View details

Senior Backend Engineer (Shop 6.0)

Ho Chi Minh - Viet Nam


Outsource

  • NodeJS
  • Azure

Design and develop the backend services for the core areas of Shop 6.0 (catalog, orders, payments, ERP Cloud integration) Define API boundaries and ensure consistent, scalable communication between services Own reliable data synchronization between the ERP Cloud and the platform, including error handling and recovery mechanisms Optimize systems for performance and scalability (caching, asynchronous processing, read/write optimization) Establish observability standards (logging, monitoring, alerting) across all backend services Conduct code reviews, promote best practices, and support less experienced team members

Negotiation

View details