MLOps Engineer

ABOUT CLIENT

Our client is a leading research company specializing in technology innovation

JOB DESCRIPTION

Develop and maintain training and inference pipelines using PyTorch, which includes DDP support, mixed precision, checkpointing, experiment versioning, and reproducible evaluation workflows.
Take ownership of and advance inference serving infrastructure using vLLM and SGLang, with a focus on debugging issues in inference stacks like tool call parsers and reasoning parsers, and optimizing for throughput and latency.
Create and sustain robust tooling in Python and C++ to aid the complete training lifecycle, from data ingestion to model release.
Optimize compute workloads for bare-metal environments, encompassing CPU/GPU utilization, memory bandwidth, and I/O throughput.
Address low-level networking issues, distributed training errors, and hardware bottlenecks across NCCL, MPI, and high-speed interconnects like InfiniBand and RoCE.
Set up and manage ML environments, covering containers, package management, GPU drivers, and runtime configurations.
Establish CI/CD patterns for AI workloads, encompassing training, evaluation, quantization, and model release workflows.
Integrate monitoring, alerting, anomaly detection, and incident response for both training jobs and inference services.
Contribute to shared platform capabilities across reliability, observability, and cost management.
Develop and maintain scalable runtime infrastructure for model-backed services and APIs, including support for LLM-backed APIs, MCP servers, and agentic systems.

JOB REQUIREMENT

Proficiency in PyTorch internals, including DDP, FSDP, mixed precision training, TorchScript, and torch.compile.
Strong programming skills in Python and C++, with the ability to understand and modify unfamiliar codebases.
Solid understanding of computer science basics including data structures, concurrency, operating systems, and memory management.
Practical experience with vLLM and SGLang for production inference serving, serving quantized models such as FP8, INT8, and NVFP4.
Experience with RLHF and PPO training pipelines, including frameworks like veRL and TRL, and integration of reward models.
Solid understanding of distributed training setups, networking, and interconnects including NCCL, MPI, InfiniBand, and RoCE.
Experience in debugging and optimizing bare-metal Linux servers, including kernel parameters, NUMA topology, and GPU driver configuration.
Familiarity with job schedulers such as Airflow and experience in operating production-grade distributed infrastructure.
Strong understanding of containerized and cloud-native environments using Docker and Kubernetes.
Familiarity with ML compiler stacks such as LLVM, MLIR, TensorRT, or XLA.
Knowledge of model quantization techniques and deployment optimization, including GPTQ, AWQ, and bitsandbytes.
Contributions to open source ML projects, including PyTorch, vLLM, SGLang, or related inference and training tooling.
Experience with infrastructure-as-code tools such as Ansible, Terraform, or Nix for reproducible cluster setup.
Experience with custom or on-premise deployments, local clusters, or edge inference.
Familiarity with observability stacks like Prometheus, Grafana, or OpenTelemetry applied to training and inference workloads.
Experience building infrastructure for agentic systems including secure tool access, orchestration, and isolation boundaries.
Passion for clean, well-documented code and detail-oriented engineering.

WHAT'S ON OFFER

Work remotely in an environment that promotes open-source collaboration
Enjoy 14 days of leave and unlimited sick days
Access to GPUs, AI credits, opportunities for fast career progression, and other perks.

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Product

Technical Skills:

Machine Learning, Devops

Location:

Ho Chi Minh - Viet Nam

Working Policy:

Onsite, Remote

Job ID:

J01855

Status:

Close

Related Job:

Salesforce Engineer (IC2)

Ho Chi Minh - Viet Nam


Outsource

  • Salesforce

#Salesforce Engineering (Quote-to-Cash) Build and maintain Salesforce processes from Quote through to Cash Work full-stack on the platform: declarative (data model, Flows, validation, security model) and programmatic (Apex, Lightning Web Components) Implement changes safely on a production system used by Go-to-Market every day#Integrations & System Interfaces Maintain and further develop the integrations between Salesforce and internal/external systems (customer portal, HubSpot, license management, billing/ERP, and others) Design and operate integrations using Salesforce REST/SOAP APIs and standard integration patterns Own the health of these interfaces: monitoring, error handling, and data consistency#Collaboration with Go-to-Market / RevOps Partner with RevOps and GTM to translate business needs into robust technical implementations Challenge requirements where they create risk or technical debt Coordinate handoffs at the Lead-to-Quote / Quote-to-Cash boundary#Reliability, Security & Change Safety Protect a revenue-critical system: change control, testing, backup/rollback awareness Apply JTL security and compliance standards to Salesforce work#Optional / stretch With spare capacity, support tech-heavy Lead-to-Quote topics in Salesforce (optional; primary ownership of Lead-to-Quote stays with RevOps)

Negotiation

View details

Senior Bioinformatics Engineer

Ho Chi Minh - Viet Nam


Product

  • Data Science

Serve as the primary bioinformatics subject matter expert for engineering teams developing cloud-native bioinformatics software. Collaborate with software architects to translate scientific workflows into scalable distributed computing architectures. Help engineers understand the computational characteristics, assumptions, and limitations of existing bioinformatics tools. Validate that modernized applications preserve scientific correctness and produce reproducible results. Define biological data models, metadata standards, controlled vocabularies, and best practices for data harmonization. Guide engineering teams in designing scalable approaches for processing large genomic and multi-omics datasets. Evaluate open-source bioinformatics software (e.g. PLINK, Regenie, BCFtools, GATK, Nextflow workflows, etc.) and identify opportunities for cloud-native modernization. Work with distributed computing specialists to determine how algorithms can be parallelized using Spark and other large-scale execution frameworks. Develop validation datasets, benchmarking methodologies, and acceptance criteria for transformed applications. Review engineering designs to ensure biological accuracy and scientific integrity. Collaborate with AI engineering teams on using AI-assisted software transformation while ensuring scientific correctness. Stay current with advances in bioinformatics, computational biology, distributed computing, and cloud-based scientific software.

Negotiation

View details

Senior Software Engineer (Distributed Computing)

Ho Chi Minh - Viet Nam


Product

  • Python
  • AWS
  • Spark

Design and provide input on system architectures, contribute to coding standards, and mentor junior engineers. Work with various teams to ensure software solutions meet business requirements, international standards, and objectives. Tackle complex software development and integration challenges, optimizing system performance. Create high-quality Python code and integrate diverse software components into cohesive solutions, with a focus on cloud computing and life sciences applications. Keep up-to-date with new technologies and Python frameworks, particularly in cloud computing. Oversee testing, deployment, and comprehensive documentation of integrated systems. Actively participate in all phases of the software development lifecycle. This includes creating user stories and engaging in sprint planning to align development efforts with business objectives. Engage with multinational companies, demonstrating flexibility to occasionally adapt to US and EU time zones.

Negotiation

View details