Principal Engineer, System Software Platform Engineering

ABOUT CLIENT

Our client is a leading technology company specializing in graphics processing units (GPUs) and artificial intelligence (AI).

JOB DESCRIPTION

Create and manage a platform for AI that provides services for multiple users, handles identity and policy management, configures quotas, and controls costs. Additionally, this platform should offer easy paths for teams to work on AI projects.
Oversee the deployment of AI models at scale, including routing, autoscaling, and implementing safety measures to ensure reliability and observability.
Manage GPU resources in a Kubernetes environment, including device plugins, feature discovery, and scheduling strategies, among other responsibilities.
Take charge of the entire lifecycle of GPUs, ensuring that driver, firmware, and runtime updates are implemented safely and consistently.
Implement virtualization strategies for GPU resources, such as vGPU and PCIe passthrough, while defining policies for resource placement, isolation, and preemptive actions.
Establish secure traffic and networking protocols, including gateways, service mesh, and authentication/authorization measures.
Enhance observability and operational efficiency through monitoring tools for GPUs, response protocols for incidents, and optimization of costs.
Develop reusable templates, integrate SDKs and CLIs, and implement infrastructure-as-code standards for the platform.
Influence the platform's direction by creating design documents, mentoring engineers, and aligning platform development with the needs of AI products.

JOB REQUIREMENT

Requires a minimum of 15 years' experience in building and operating large-scale distributed systems or platform infrastructure, with a strong track record of shipping production services.
Proficiency in one or more of Python, Go, Java, C++, with a deep understanding of concurrency, networking, and systems design.
Expertise in containers, orchestration, Kubernetes, cloud networking, storage, IAM, and infrastructure-as-code.
Practical experience with GPU platforms, including Kubernetes GPU operations, scheduling, isolation, preemption, and utilization tuning.
Background in virtualization, including deploying and operating vGPU, PCIe pass-through, and mediated devices in production.
Experience in Site Reliability Engineering (SRE) or equivalent, including SLOs, incident management, performance tuning, resource management, and financial oversight.
Strong security mindset, with experience in TLS/mTLS, RBAC, secrets, policy-as-code, and secure multi-tenant architectures.
In-depth experience in GPU operations, including MIG partitioning, MPS sharing, NUMA/topology awareness, DCGM telemetry, GPUDirect RDMA/Storage.
Exposure to inference platforms, including serving runtimes, caching/batching, autoscaling patterns, and continuous delivery.
Exposure to agentic platforms, including workflow engines, tool orchestration, policy/guardrails for tool access, and data boundaries.
Experience in traffic/data plane technologies, such as gRPC, HTTP, Protobuf, service mesh, API gateways, CDN/caching, and global traffic management.
Proficiency with tools such as Terraform, Helm, GitOps, Prometheus, Grafana, OpenTelemetry, and policy engines; bare-metal provisioning experience is a plus.

WHAT'S ON OFFER

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product

Technical Skills:

Devops, Backend, AI

Location:

Ho Chi Minh, Ha Noi - Viet Nam

Working Policy:

Onsite

Job ID:

J01969

Status:

Close

Related Job:

Backend Engineer

Ho Chi Minh - Viet Nam


Product

  • Typescript
  • NodeJS
  • Python
  • AI

Build the core gateway: a unified, OpenAI-compatible API in front of multiple providers (OpenAI, Anthropic, Google, plus self-hosted and OSS models). Own provider routing and reliability: load balancing, automatic failover, and cost and latency-aware routing. Build billing and metering that is correct, not approximate: per-request token accounting, usage ledgers, cost attribution per team, user, and key, budgets, and spend limits. Ship org controls: API key management, per-team and per-user quotas, rate limiting, and RBAC. Handle streaming and performance: low-overhead proxying, streaming responses, connection handling, and caching where it helps. Contribute to the Jan Agent and connect it to the router: route its model and tool calls through the gateway, and make agent traffic first-class in metering, controls, and observability. Make deliberate speed-versus-correctness calls: move fast where iteration is cheap, refuse to cut corners where a bug means a bad charge or a leaked key, and pay down debt on your own initiative.

Negotiation

View details

AI Transformation Lead

Ho Chi Minh - Viet Nam


Outsource

  • AI

Develop a 3-year AI transformation roadmap aligned with business goals in IT outsourcing, product development, and ODC services. Prioritize highest-impact AI use cases in the organization using a build-vs-buy-vs-partner framework. Establish AI governance for model selection, cost management, data privacy, IP protection, and ethical AI guidelines. Implement AI-assisted development workflows for 1,000+ engineers, including AI code generation, review, automated testing, and AI-powered debugging. Drive adoption of AI orchestration platforms to automate repetitive engineering tasks. Create internal AI skills/training programs and a culture of continuous AI experimentation. Measure and report on productivity gains, quality improvements, and time-to-market acceleration from AI adoption. Collaborate with product teams to define AI features for various products. Lead the development of AI Copilots, intelligent assistants, and autonomous agents embedded within products. Guide the architecture of an Ontology-Based AI ERP/MES Platform, including Knowledge Graphs, GraphRAG, and Multi-Agent Systems. Identify new AI-powered product opportunities in logistics, manufacturing, and supply chain to create new revenue streams. Advocate for AI internally, communicate the vision, celebrate wins, and address concerns across all levels. Partner with HR to define new AI-focused roles and refine hiring criteria. Collaborate with ODC/Client Delivery teams to package and sell AI capabilities to existing and new clients. Represent the company externally in conferences, thought leadership, and talent branding to position the company as an AI leader in the IT services industry.

Negotiation

View details

Head of Engineering - Marketing Technology

Ho Chi Minh - Viet Nam


Product

  • Management

Develop an integrated roadmap for strategic execution based on the organization's strategic aspirations and lead the implementation process from planning to delivery. Manage multiple engineering teams throughout the organization to achieve desired outcomes, hence having knowledge of the organization's specific areas of operation is an advantage. Collaborate closely with business teams and product owners to verify requirements before and after delivery through showcases and post-production monitoring. Take responsibility for both the development and operation of applications in production, providing active operational support and establishing a clear support model with a focus on site reliability engineering. Direct and oversee the implementation of cybersecurity updates, including keeping software versions current and patching infrastructure regularly. Oversee investment allocation across the organization to maintain alignment, ensure effective spending, and offer insights on the effectiveness and prioritization of investments. Integrate engineering excellence with domain expertise in marketing while pursuing strategic technology objectives.

Negotiation

View details