Platform Lead

ABOUT CLIENT

Our client is a leading research company specializing in technology innovation

JOB DESCRIPTION

Develop and expand distributed systems to handle large volumes of sensory, telemetry, and control data across cloud and edge environments, facilitating real-time connections for fleets of robots.
Create the API Platform with a focus on high reliability, exceptional developer experience, and robust multimodal AI capabilities accessible through user-friendly APIs and SDKs.
Establish extensive training and inference platforms for foundation models used in robot autonomy, teleoperation, and developer integrations.
Devise data ingestion and streaming pipelines for real-time connectivity of robot fleets to the cloud, covering various data inputs such as video, LiDAR, joint states, and audio.
Oversee and advance a modern cloud native infrastructure stack employing Kubernetes, Docker, and infrastructure as code tools.
Ensure platform reliability through telemetry, monitoring, alerting, autoscaling, failover, and disaster recovery measures.
Make infrastructure decisions pertaining to distributed storage, consensus protocols, GPU orchestration, network reliability, and API security.
Foster collaboration across ML, robotics, and product teams to facilitate hardware in the loop simulation, policy rollout, continuous learning, and CI/CD workflows.
Implement secure APIs featuring fine-grained access control, usage metering, rate limiting, and billing integration to accommodate a growing user base.

JOB REQUIREMENT

At least 7 years of professional experience in software engineering focusing on distributed systems, backend infrastructure, or data platforms.
Proven track record of developing and managing high-scale systems for critical workloads.
Proficiency in Go, Rust, C++, Python, or TypeScript, with a solid understanding of concurrency, networking, and systems performance.
Deep knowledge of cloud native architectures, including Kubernetes, Docker, Helm, gRPC, Kafka, Ray, and service mesh technologies like Istio or Linkerd.
Solid understanding of API architecture and design patterns such as REST, gRPC, WebSockets, OAuth2, and OpenAPI.
Experience with various databases including PostgreSQL, Redis, and modern vector databases such as Pinecone, Weaviate, or FAISS.
Strong understanding of data consistency, replication, and fault tolerance in diverse environments.
Familiarity with observability tools like Prometheus, Grafana, Datadog, or OpenTelemetry in the context of large-scale production systems.
Experience in developing distributed training, large-scale simulation, or fleet-scale telemetry systems.
Familiarity with real-time robotics workloads, including streaming from physical sensors and actuators.
Experience with MLOps tools and AI workflows, encompassing model versioning, inference pipelines, and model registries.
Knowledge of billing systems, quota enforcement, chargeback models, and multi-tenant security and isolation.
Previous involvement in developer platforms, API products, and a strong focus on developer UX and documentation.
Contributions to open source infrastructure, AI frameworks, or robotics middleware such as ROS, gRPC, or Mediasoup are a plus.

WHAT'S ON OFFER

Work remotely in an environment that promotes open-source collaboration
Enjoy 14 days of leave and unlimited sick days
Access to GPUs, AI credits, opportunities for fast career progression, and other perks.

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product

Technical Skills:

Backend, Devops, Data Engineering

Location:

Others - Singapore

Working Policy:

Onsite

Salary:

Negotiation

Job ID:

J02064

Status:

Active

Related Job:

AI & DATA Engineer/Databricks

Ho Chi Minh - Viet Nam


Outsource

  • Data Engineering
  • AI

Lead the design and implementation of cloud-native data pipelines for large-scale analytics and AI applications using Databricks Develop and maintain high-quality backend APIs and AI Agents supporting internal tools and customer-facing products Execute and manage data migration projects with a focus on performance, reliability, and maintainability Access Databricks environments directly or via CLI to develop, orchestrate, test, and deploy jobs and pipelines Promote and implement CI/CD best practices, Git workflows, and engineering standards across the data team Collaborate closely with AI engineers, consultants, and external stakeholders to translate requirements into scalable solutions

Negotiation

View details

Associate Manager – Software Engineer

Ho Chi Minh - Viet Nam


Product

  • Java
  • ReactJS

Lead the decisions around scalable full-stack and cloud-native systems architecture. Advocate for best practices in system design, reliability, and observability. Take charge of delivering essential platform capabilities, including crew training and assessment (OCL), compliance systems (Track & Trace), restaurant monitoring and reporting (MRD), virtual restaurant assessments, and intelligent operational action systems for RGMs. Collaborate with product and stakeholders, translating business requirements into scalable solutions. Build and design scalable applications using React/React Native, Spring Boot, and NestJS. Develop robust APIs and microservices that support restaurant operational systems. Ensure high code quality through testing, code reviews, and performance enhancement. Manage and create cloud infrastructure using AWS (EKS, Lambda). Set up CI/CD pipelines using GitLab CI. Guarantee strong monitoring and system reliability through the use of Datadog. Collaborate closely with engineering and product teams across global locations to deliver platform capabilities. Address complex technical challenges and provide scalable solutions to enhance platform reliability and operational efficiency.

Negotiation

View details

PreSales Solutions Engineer

Ho Chi Minh - Viet Nam


Product

  • Presale
  • System
  • Google Cloud

PreSales Support: Collaborating with the Sales team to understand client needs and develop tailored solutions using Google Maps and Google Cloud services. This involves conducting technical presentations, product demonstrations, and creating proof of concepts (POCs) for prospective clients, as well as contributing to proposals and RFP responses with detailed technical information. Post-Sales Support: Leading the technical implementation of Google Maps and Google Cloud services, ensuring smooth deployment and integration. Providing ongoing technical support and troubleshooting for clients after implementation, working closely with cross-functional teams to ensure client satisfaction and build long-term relationships. Technical Expertise: Staying up-to-date with the latest Google Maps and Google Cloud technologies, serving as a subject matter expert (SME) for both internal teams and clients. Integrating new features and services into client solutions and providing guidance on best practices. Collaboration: Working closely with Sales, Product, Infrastructure, Data, and Engineering teams to align solutions with client needs and company goals. Mentoring junior team members and contributing to training initiatives.

Negotiation

View details