Head Of System Infrastructure

JOB DESCRIPTION

1. Storage: vận hành hệ thống high availability master-slave multipath redundant storage server với filesystem ZFS.
Đảm bảo thời gian gián đoạn dịch vụ ngoài mong muốn ở mức thấp nhất do cơ chế auto failover.
Định kì kiểm tra khả năng auto failover.
Định kì vá các lỗi phát sinh và cập nhật các cải tiến tốc độ.
Đảm bảo vận hành tốc độ và ổn định như giới hạn phần cứng (băng thông 1.5-2 GByte/s và >200.000 IOPS) 
Tư vấn khi cần mở rộng hoặc thay thế thiết bị
Giải quyết các sự cố storage server: hardware failure (L1), system crash (L1), performance (L2)
Chạy incremental instant backup và full backup định kỳ tuỳ mức độ quan trọng của các volume dữ liệu.
 
2. Server và docker swarm: đa phần các app của fahasa chạy trên môi trường docker swarm.
Quản lý lượng tải vả tài nguyên server
Hỗ trợ cài đặt khi có phát sinh thay thế, thêm mới server mới.
Tư vấn cấu hình, phối hợp nhà cung cấp
Giải quyết các sự cố server và docker
Tư vấn lộ trình nâng cấp OS khi hết vòng đời OS.
 
3. Xây Dựng và Bảo trì hệ thống: Linux OS và các phần mềm quan trọng sau: nginx (web server), php, mariadb, magento và salt. 
Cần có kiến thức chuyên sau về việc setup các hệ thống có chịu tải lớn sử dụng: Nginx, Redis và Docker Swarm, Mariadb, Php và phpfpm …
Có kiến thức về xây dựng hệ thống sử dụng Kubernetes.
Một số phần mềm fahasa sử dụng đã hết được hỗ trợ chính thức từ nhà phát triển phần mềm. Các lỗi security cần được tự sửa hoặc lấy từ các bản vá lỗi ở các phiên bản mới hơn. Đây là những phần mềm trọng yếu, lỗi security sẽ gây ra tổn thất rất lớn. (L2)
Đảm bảo tương thích giữa hệ thống phần mềm hiện tại với các phần cứng POS server mới.
 
4. Troubleshoot các vấn đề gây gián đoạn dịch vụ hệ thống: hệ thống hoặc lỗi performance và security của web, idempiere và POS server
Nền tảng TMĐT chịu 1 lượng traffic rất lớn tại các kỳ Flashsale, cần các kiến thức chuyên sâu về performance, load balancing và scalability để hỗ trợ, troubleshoot và đưa ra hướng giải quyết cho vấn đề.
Phản ứng nhanh, xử lý các lỗi xảy ra bất ngờ này. Hỗ trợ xác định nguyên nhân và tư vấn giải pháp. L1 cho web và idempiere. L2 cho POS server nhà sách. Riêng pos server nhà sách chỉ xử lý các vấn đề mà phòng IT chưa được hướng dẫn xử lý.
Cung cấp giải pháp load balancing và chống DOS

JOB REQUIREMENT

Tốt nghiệp Đại học hoặc sau Đại học chuyên ngành Công nghệ thông tin
Có kỹ năng quản lý đội nhóm, phản ứng nhanh với sự cố của hệ thống
Tư duy tốt trong làm việc độc lập lẫn làm việc nhóm
Có kinh nghiệm trong lĩnh vực Thương Mại Điện Tử
Kinh nghiệm làm việc ở vị trí tương tự: 3 – 7 năm
Ngoại ngữ: tiếng Anh
Năng động, nhạy bén, có tinh thần trách nhiệm cao

WHAT'S ON OFFER

Chế độ bảo hiểm y tế, bảo hiểm xã hội
Lương thưởng theo quy định nhà nước
Chăm sóc sức khỏe hàng năm
Du lịch mỗi năm 1 lần
Môi trường làm việc trẻ trung, thân thiện

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product, Book

Technical Skills:

Devops, System

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Job ID:

J00820

Status:

Close

Related Job:

Senior Software Engineer - DevOps (ERP)

Ho Chi Minh - Viet Nam


Outsource

You drive the step-by-step migration of our ERP from a single virtual-machine setup towards containerized workloads on Linux and Docker. You design, implement, and maintain agent-based deployment for both virtual machines and - increasingly - containers, making sure they are provisioned, deployed, and updated reliably. You provision and operate virtual machines, containers, deployment targets, and delivery infrastructure in Azure. You ensure the cloud onboarding wizard (React) and the surrounding cloud services run reliably so that customers can work without interruption. You instrument the platform with telemetry, monitoring, and observability so that issues are detected early and operations stay transparent. You design, implement, and maintain CI/CD pipelines and automate build, packaging, testing, and deployment processes. You support the migration of existing build and release processes towards GitHub Actions and modern delivery workflows. You manage and evolve self-hosted runners and build infrastructure. You collaborate with Quality Engineering to integrate automated tests into delivery pipelines. You work with Release Management to improve release reliability, transparency, and automation. You identify bottlenecks in the software delivery process and continuously improve developer productivity. You help establish the technical foundation for future Continuous Delivery and Continuous Deployment scenarios#OPERATIONS & ON-CALL RESPONSIBILITYTogether with your future colleague and the wider team, you take shared responsibility for keeping the platform running for our customers. We ensure disruption-free operations primarily through the right tooling and welldefined processes - automation, telemetry, and clear escalation paths - not through manual firefighting. This is explicitly not about working fixed shifts: it is about ownership and reachability, being there within a defined reaction time when problems occur, including on weekends. You take ownership of the reliable operation of our machines and - increasingly - containers, so that customers can keep working. You build the tools and processes that keep operations running smoothly and prevent incidents before they happen. You make sure that deployments and the cloud wizard (React) keep functioning correctly in production. You share responsibility for operational availability in a rotation with the team - the goal is to respond within a defined reaction time when incidents arise, not to staff fixed shifts. You help build the alerting, telemetry, and on-call processes that make fast and reliable incident response possible.

Negotiation

View details

Engineering Manager – Shop 6.0

Ho Chi Minh - Viet Nam


Outsource

  • Management
  • Backend
  • Frontend
  • Devops
  • Azure

Take on overall technical and organizational responsibility for delivering Shop 6.0 across frontend, backend, and infrastructure Plan delivery scope, milestones, and releases, and coordinate the work of parallel workstreams (frontend, backend, DevOps, QA) Make and facilitate key architectural decisions together with the senior engineers, particularly around microservices, APIs, and cloud-native implementation on Azure Identify and manage risks related to integration (especially the ERP Cloud connection), scalability, and technical dependencies Serve as the central point of contact for stakeholders on scope, prioritization, and timelines Ensure code quality through reviews, mentoring, and clear standards, including a pull request process with mandatory checks and senior/lead review Ensure transparency on progress, risks, and decisions towards the team and management Hire, lead, develop and retain a team of 6 engineers (4 backend and 2 frontend)

Negotiation

View details

Senior AI DevSecOps Engineer

Ho Chi Minh - Viet Nam


Product

  • Devops
  • AWS
  • Azure
  • Security

Management of CI/CD Pipeline: Ensure automation, security, and scalability across all stages of the development lifecycle. Infrastructure & Security: Design and implement secure multi-cloud infrastructure solutions leveraging cloud services, containerization, and orchestration tools. Policy as Code: Define and enforce security and compliance policies across Kubernetes clusters using OPA or Kyverno, ensuring guardrails are automated and auditable. AI & Platform Automation: Drive the adoption of AI-powered tools and workflows to automate infrastructure operations, optimize CI/CD pipelines, accelerate root cause analysis, improve security posture, and enhance engineering productivity. Observability & Alerting: Build and maintain a comprehensive observability stack with proactive alerting, dashboards, and runbooks for critical business flows and security events. Secret & Credential Management: Design and enforce secrets management practices across all environments ensuring zero hardcoded credentials in codebases and pipelines. Incident Response & On-Call: Own and continuously improve incident response processes, define runbooks, lead post-mortems, track MTTR, and participate in on-call rotation to maintain platform reliability and SLO adherence. Threat Modelling & Penetration Testing: Conduct regular threat modelling sessions with engineering teams and coordinate or perform penetration testing activities to proactively identify attack surfaces before they reach production. Code Security: Conduct regular code reviews and static/dynamic analysis to identify and remediate security vulnerabilities. Compliance and Best Practices: Ensure compliance with industry standards and best practices. Collaboration: Collaborate with development, operations, and security teams to foster a culture of automation and security-first thinking. Mentorship: Mentor junior engineers and other team members on security best practices. Documentation: Maintain thorough and up-to-date documentation of security policies, procedures, and incident reports. Trend Scouting: Stay updated with the latest trends in technology and AI to integrate innovative solutions into our processes.

Negotiation

View details