Site Reliability Engineer

JOB DESCRIPTION

Maintain systems and troubleshoot system issues.
Identifying bottleneck in various Java applications and implement performance improvements.
Identify and analyze user requirements.
Prioritize, assign, and execute tasks throughout the software development life cycle.
Develop, configure, and deploy tools for cloud-based systems and services.
Containerize new and legacy applications.
Maintain awareness of new and emerging technologies.
Support development and operations teams.
Enhance, modify or debug developer code as needed.

JOB REQUIREMENT

Must-have
Understanding of an object-orientated language, preferably the latest version of Java (with experience in Hibernate, Multi-thread, Spring Boot)
Experience in configuration, in Jenkins for CI/CD pipeline creation, automation scripts and Kubernetes implementation with Google.
Proficiency in supporting a 24×7 critical operation.
Experience in a cloud computing platform and associated automation patterns it provides, preferably GCP.
Proficient in production systems design including High Availability, Disaster Recovery, Performance, Efficiency, and Security user, application performance, system, log, time-series, and dashboarding.
Familiarity with Open-Source concepts and tools like Prometheus, Grafana, ELK etc. 
Proficient in a modern infrastructure automation toolkit such as Terraform/Helm
Proficient in a Linux or Unix based environment.
Experience in destructive testing methodologies and tools such as chaos monkey
Experience in defensive coding practices and patterns for high availability
Nice-to-have
Experience in a cloud computing platform and associated automation patterns it provides, preferably GCP
Proficient in a modern scripting language like GO or Python
Knowledge of APM fundamentals or experience in tools like New Relic or AppDynamics.

WHAT'S ON OFFER

Open to deal base salary with additional project allowances.
Full salary during probation & Full coverage of social insurance.
Performance & salary review: twice a year
Monthly childcare support.
Premium Healthcare insurance and Health check-up services for employee and family ones.
15 Annual Leaves plus 10 days for Bereavement leave and 1.5 months for Paternity leave.
Premium package at top Gym service provider.
Diverse internal activities: Football, Billiards, Badminton, E-sport clubs & other regular company events.
Frequent opportunities to travel to US headquarter from 3-6 months.
Free parking for motorbike and car

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Devops, Java

Location:

Ho Chi Minh, Da Nang - Viet Nam

Working Policy:

Job ID:

J01196

Status:

Close

Related Job:

Senior Bioinformatics Engineer

Ho Chi Minh - Viet Nam


Product

  • Data Science

Serve as the primary bioinformatics subject matter expert for engineering teams developing cloud-native bioinformatics software. Collaborate with software architects to translate scientific workflows into scalable distributed computing architectures. Help engineers understand the computational characteristics, assumptions, and limitations of existing bioinformatics tools. Validate that modernized applications preserve scientific correctness and produce reproducible results. Define biological data models, metadata standards, controlled vocabularies, and best practices for data harmonization. Guide engineering teams in designing scalable approaches for processing large genomic and multi-omics datasets. Evaluate open-source bioinformatics software (e.g. PLINK, Regenie, BCFtools, GATK, Nextflow workflows, etc.) and identify opportunities for cloud-native modernization. Work with distributed computing specialists to determine how algorithms can be parallelized using Spark and other large-scale execution frameworks. Develop validation datasets, benchmarking methodologies, and acceptance criteria for transformed applications. Review engineering designs to ensure biological accuracy and scientific integrity. Collaborate with AI engineering teams on using AI-assisted software transformation while ensuring scientific correctness. Stay current with advances in bioinformatics, computational biology, distributed computing, and cloud-based scientific software.

Negotiation

View details

Senior Software Engineer (Distributed Computing)

Ho Chi Minh - Viet Nam


Product

  • Python
  • AWS
  • Spark

Design and provide input on system architectures, contribute to coding standards, and mentor junior engineers. Work with various teams to ensure software solutions meet business requirements, international standards, and objectives. Tackle complex software development and integration challenges, optimizing system performance. Create high-quality Python code and integrate diverse software components into cohesive solutions, with a focus on cloud computing and life sciences applications. Keep up-to-date with new technologies and Python frameworks, particularly in cloud computing. Oversee testing, deployment, and comprehensive documentation of integrated systems. Actively participate in all phases of the software development lifecycle. This includes creating user stories and engaging in sprint planning to align development efforts with business objectives. Engage with multinational companies, demonstrating flexibility to occasionally adapt to US and EU time zones.

Negotiation

View details

Senior Software Engineer - DevOps (ERP)

Ho Chi Minh - Viet Nam


Outsource

  • Devops
  • Azure
  • Terraform
  • Kubernetes

Lead the migration of our ERP from a single virtual-machine setup to containerized workloads on Linux and Docker. Design, implement, and maintain agent-based deployment for both virtual machines and containers to ensure reliable provisioning, deployment, and updates. Manage virtual machines, containers, deployment targets, and delivery infrastructure in Azure. Ensure the cloud onboarding wizard (React) and surrounding cloud services run reliably for uninterrupted customer work. Instrument the platform with telemetry, monitoring, and observability for early issue detection and transparent operations. Design, implement, and maintain CI/CD pipelines to automate build, packaging, testing, and deployment processes. Support the migration of existing build and release processes towards GitHub Actions and modern delivery workflows. Manage and improve self-hosted runners and build infrastructure. Collaborate with Quality Engineering to integrate automated tests into delivery pipelines. Work with Release Management to enhance release reliability, transparency, and automation. Identify and address bottlenecks in the software delivery process to enhance developer productivity. Help establish the technical foundation for future Continuous Delivery and Continuous Deployment scenarios. Shared responsibility for platform operations ensuring disruption-free operations primarily through automation, telemetry, and clear escalation paths. Ownership of reliable operation of machines and containers to support uninterrupted customer operations. Develop tools and processes to maintain smooth operations and prevent incidents. Ensure correct functioning of deployments and the cloud wizard (React) in production. Share responsibility for operational availability in a rotation with the team, aiming for a defined reaction time to incidents, including weekends. Help build alerting, telemetry, and on-call processes for fast and reliable incident response.

Negotiation

View details