Principal Engineer, System Software Platform Engineering

ABOUT CLIENT

Our client is a leading technology company specializing in graphics processing units (GPUs) and artificial intelligence (AI).

JOB DESCRIPTION

Create and manage a platform for AI that provides services for multiple users, handles identity and policy management, configures quotas, and controls costs. Additionally, this platform should offer easy paths for teams to work on AI projects.
Oversee the deployment of AI models at scale, including routing, autoscaling, and implementing safety measures to ensure reliability and observability.
Manage GPU resources in a Kubernetes environment, including device plugins, feature discovery, and scheduling strategies, among other responsibilities.
Take charge of the entire lifecycle of GPUs, ensuring that driver, firmware, and runtime updates are implemented safely and consistently.
Implement virtualization strategies for GPU resources, such as vGPU and PCIe passthrough, while defining policies for resource placement, isolation, and preemptive actions.
Establish secure traffic and networking protocols, including gateways, service mesh, and authentication/authorization measures.
Enhance observability and operational efficiency through monitoring tools for GPUs, response protocols for incidents, and optimization of costs.
Develop reusable templates, integrate SDKs and CLIs, and implement infrastructure-as-code standards for the platform.
Influence the platform's direction by creating design documents, mentoring engineers, and aligning platform development with the needs of AI products.

JOB REQUIREMENT

Requires a minimum of 15 years' experience in building and operating large-scale distributed systems or platform infrastructure, with a strong track record of shipping production services.
Proficiency in one or more of Python, Go, Java, C++, with a deep understanding of concurrency, networking, and systems design.
Expertise in containers, orchestration, Kubernetes, cloud networking, storage, IAM, and infrastructure-as-code.
Practical experience with GPU platforms, including Kubernetes GPU operations, scheduling, isolation, preemption, and utilization tuning.
Background in virtualization, including deploying and operating vGPU, PCIe pass-through, and mediated devices in production.
Experience in Site Reliability Engineering (SRE) or equivalent, including SLOs, incident management, performance tuning, resource management, and financial oversight.
Strong security mindset, with experience in TLS/mTLS, RBAC, secrets, policy-as-code, and secure multi-tenant architectures.
In-depth experience in GPU operations, including MIG partitioning, MPS sharing, NUMA/topology awareness, DCGM telemetry, GPUDirect RDMA/Storage.
Exposure to inference platforms, including serving runtimes, caching/batching, autoscaling patterns, and continuous delivery.
Exposure to agentic platforms, including workflow engines, tool orchestration, policy/guardrails for tool access, and data boundaries.
Experience in traffic/data plane technologies, such as gRPC, HTTP, Protobuf, service mesh, API gateways, CDN/caching, and global traffic management.
Proficiency with tools such as Terraform, Helm, GitOps, Prometheus, Grafana, OpenTelemetry, and policy engines; bare-metal provisioning experience is a plus.

WHAT'S ON OFFER

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product

Technical Skills:

Devops, Backend, AI

Location:

Ho Chi Minh, Ha Noi - Viet Nam

Working Policy:

Onsite

Salary:

Negotiation

Job ID:

J01969

Status:

Close

Related Job:

Presales Consultant

Ho Chi Minh - Viet Nam


Outsource

  • Presale
  • Network
  • Security
  • System

Support Sales &Product Managers: Deliver presales technical consultancy and solution design. Lead Customer Engagements: Conduct technical presentations, product demos, and Proof of Concepts (PoCs). Develop Technical Documents: Analyze customer requirements and prepare solution proposals, architecture designs, and BOMs. Assist with RFP/RFI Responses: Support partner enablement, opportunity qualification, and tender compliance. Conduct Training Workshops: Provide training for partners/resellers and internal teams. Maintain Technical Certifications: Obtain and maintain certifications for relevant vendor technologies. Build Product Expertise: Develop deep knowledge of company solutions to support sales and marketing initiatives. Provide Post-Sales Support: Ensure smooth handover and contribute to customer satisfaction. Offer Market Insights: Support solution positioning and provide content recommendations. Support Business Growth: Perform other duties as required.

Negotiation

View details

DevOps Engineer

Ha Noi - Viet Nam


Outsource

  • Devops

Client Consulting: Directly engage with clients to design solutions and implement DevSecOps practices. Tool Deployment & Configuration: Deploy, install, and configure DevSecOps and CI/CD tools, including: Container Orchestration: Kubernetes, OpenShift Source Code Management: GitLab, GitHub Automation Tools: Jenkins, GitLab CI Artifact Management: Nexus, JFrog Code Scanning: SonarQube, Semgrep, BlackDuck, Coverity Observability Solutions: Deploy, install, and configure logging, monitoring, and tracing systems. CI/CD Pipeline Development: Build and optimize CI/CD pipelines for application delivery. Operational Support: Provide ongoing operational and administrative support for DevSecOps tools and solutions. Research & Innovation: Conduct R&D on emerging technologies in DevOps, DevSecOps, Cloud-Native, and AI.

Negotiation

View details

Software Architect

Ho Chi Minh - Viet Nam


Outsource

  • Azure
  • .NET

Responsible for creating and overseeing integration architectures on Azure Converting business requirements into integration patterns, data flows, error handling, monitoring, and resiliency models Defining and directing the usage of Azure Integration Services, covering Logic Apps, Functions, API Management, Service Bus, and Event Hubs Leading integration platform design following Infrastructure as Code principles and cloud landing zone considerations Ensuring secure integration architectures utilizing OAuth, OIDC, and API security best practices Guiding development teams through architecture reviews, best practices, and reference implementations using C# and .NET Providing support for SAP integrations, including SAP S/4HANA, SAP PI/PO, and SAP BTP Integration Suite Contributing to integration platform modernization and legacy transformation initiatives, such as BizTalk migrations Collaborating with stakeholders, vendors, and delivery teams to ensure alignment and drive technical decisions

Negotiation

View details