Platform Engineer

JOB DESCRIPTION

Build and maintain distributed infrastructure handling telemetry, sensory, and control data across cloud and edge environments
Design and operate data ingestion and streaming pipelines connecting robot fleets to the cloud in real time, covering video, joint states, audio, and LiDAR
Develop and maintain backend services and APIs that power the Company's developer-facing platform, with a focus on reliability and developer experience
Manage and evolve cloud native infrastructure using Kubernetes, Docker, and infrastructure as code tooling
Ensure platform reliability through monitoring, alerting, autoscaling, failover, and incident response
Support ML and robotics teams with data infrastructure for training pipelines, policy rollout, and hardware-in-the-loop simulation
Implement secure APIs with access control, rate limiting, and usage metering as we scale

JOB REQUIREMENT

4 or more years of professional software engineering experience in platform, infrastructure, or data engineering
Proficiency in one or more of Go, Rust, Python, or TypeScript, with strong fundamentals in concurrency and systems performance
Hands-on experience with cloud native tooling: Kubernetes, Docker, Helm, and gRPC
Experience building and operating data pipelines and streaming systems -- Kafka, Flink, or similar
Solid understanding of API design patterns including REST, gRPC, and WebSockets
Experience with databases spanning PostgreSQL, Redis, and modern vector databases
Familiarity with observability tooling: Prometheus, Grafana, Datadog, or OpenTelemetry
Bonus Points
Experience with real-time data streams from physical sensors or robotics systems
Familiarity with MLOps workflows including model versioning, inference pipelines, and model registries
Background in distributed training or large-scale simulation infrastructure
Contributions to open-source infrastructure, robotics middleware, or AI frameworks
Experience on developer platforms or API products

WHAT'S ON OFFER

Collaborate with a world-class research team on meaningful, high-impact projects
Own and shape the core training code infrastructure used daily by the team
Work on real models, real data, and real scale - not toy problems
Help bridge the gap between research velocity and engineering quality
Flexible work environment with a culture that values depth, clarity, and curiosity

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product

Technical Skills:

Backend, Devops, Data Engineering

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Onsite

Salary:

Negotiation

Job ID:

J02106

Status:

Active

Related Job:

DevOps Engineer

Others - Viet Nam


Product

  • Devops
  • Kubernetes
  • Network

Operate and evolve our Kubernetes platform across multiple clusters and environments (Prod, Dev, hybrid on-prem and public cloud), covering control plane operations, node lifecycle, upgrades, and autoscaling at every layer (Cluster Autoscaler, HPA, KEDA). Architect and manage hybrid cloud infrastructure spanning on-premises and public clouds (GCP, AWS), including workload placement, cross-cloud networking, and unified resource management. Own the CI/CD and GitOps experience end-to-end: container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD. Own the observability stack as a single pane of glass across all clusters: Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus -- and help push toward agent-assisted SRE workflows. Manage and improve our inference platform: vLLM serving and AIBrix for multi-model orchestration and autoscaling across a fleet of NVIDIA GPUs. Operate platform services: Kafka, Redis, PostgreSQL, OpenSearch. Manage identity and access via Keycloak integrated with Google Workspace; harden SSO, RBAC, and secrets management across the platform. Harden network security across private load balancers, firewalls, and VPC segmentation; design and maintain hub-and-spoke / multi-AZ topologies. Support training infrastructure: self-service VM provisioning, RunPod burst capacity, Weights and Biases integration. Drive infrastructure reliability, cost efficiency, and capacity planning as the platform scales.

Negotiation

View details

Software Engineer (Digital Twin)

Ho Chi Minh - Viet Nam


Product

  • Python
  • C/C++

Build and maintain high-fidelity digital twin environments for Asimov across MuJoCo, Isaac Sim, and Unreal Engine, calibrated to real hardware behavior. Design and own the systems -- not just the environments -- that let locomotion, autonomy, and perception teams generate, validate, and iterate on simulation scenarios at scale. Build pipelines for asset import, USD and MJCF workflows, sensor modeling, and real-to-sim calibration to keep digital twins synchronized with evolving hardware. Develop photorealistic rendering pipelines in Unreal Engine for synthetic data generation and perception model training. Work with hardware and mechatronics teams to model actuator dynamics, contact physics, and structural behavior, ensuring simulation parameters reflect physical ground truth. Integrate digital twin environments with the Company's locomotion training pipeline (Cyclotron) and autonomy stack, enabling teams to run experiments and close the sim-to-real gap. Contribute to the open-source Asimov simulation stack, including tooling, documentation, and reproducible environment workflows.

Negotiation

View details

Product Manager (Platform)

Ho Chi Minh - Viet Nam


Product

  • Product Management

Own the product roadmap across the Company Platform and related software infrastructure, working with engineering leads to define priorities, sequencing, and success criteria. Partner with autonomy, locomotion, and simulation teams to understand their tooling requirements and translate those into platform product decisions. Define and track developer experience metrics for internal teams building on top of the Company Platform, and drive improvements based on usage and feedback. Work with research engineers and infrastructure teams to scope, prioritise, and deliver software capabilities that enable faster iteration on robot behaviour. Build the product documentation, release processes, and internal communication structures needed as the platform matures and the team grows. Engage with external partners and customers deploying on the Company Platform to gather requirements and validate product direction. Identify and resolve cross-functional dependencies between software, hardware, and operations as Asimov moves from development to deployment at scale.

Negotiation

View details