Lead Data Engineer

ABOUT CLIENT

Our client is a global technology company that specializes in providing innovative IT solutions for the financial services industry

JOB DESCRIPTION

Design, build and maintain scalable data infrastructure, including data lakes, pipelines, and metadata repositories to ensure timely and accurate data delivery.
Collaborate with data scientists to develop and support data models, integrate data sources, and assist in machine learning workflows and experimentation environments.
Create and enhance large-scale batch and real-time data processing systems to improve operational efficiency and meet business objectives.
Utilize Python, Apache Airflow, and AWS services to automate data workflows and processes, ensuring efficient scheduling and monitoring.
Leverage AWS services to manage data storage and compute resources, aiming for high performance, scalability, and cost-efficiency.
Implement rigorous testing and validation procedures to guarantee the reliability, accuracy, and security of data processing workflows.
Keep abreast of industry best practices and emerging technologies in data engineering and data science to suggest optimizations and innovative solutions.

JOB REQUIREMENT

Proficient in Python for data processing (pandas, pyspark), workflow automation (Apache Airflow), and experience with AWS services (Glue, S3, EC2, Lambda).
Experience working with Kubernetes and Docker for managing containerized environments in the cloud.
Hands-on experience with columnar and big data databases (Athena, Redshift, Vertica, Hive/Hadoop), along with version control systems like Git.
Strong familiarity with AWS services for cloud-based data processing and management.
Experience with CI/CD tools such as Jenkins, CircleCI, or AWS CodePipeline for continuous integration and deployment.
Expertise in building and managing robust data architectures and pipelines for large-scale data operations.
Ability to support data science workflows, including collaboration on data preparation, feature engineering, and enabling experimentation environments.
Familiarity with Langchain for building data applications involving natural language processing or conversational AI frameworks.
Experience with AWS Sagemaker or Databricks for enabling machine learning environments.
Familiarity with both RDBMS (MySQL, PostgreSQL) and NoSQL (DynamoDB, Redis) databases.
Experience with enterprise BI tools like Tableau, Looker, or PowerBI.
Familiarity with distributed messaging systems like Kafka or RabbitMQ for event streaming.
Experience with monitoring and log management tools such as the ELK stack or Datadog.
Knowledge of best practices for ensuring data privacy and security, particularly in large data infrastructures.

WHAT'S ON OFFER

Attractive salary
Bonus equivalent to a month's salary
Performance-based incentives
Access to professional English training
Comprehensive health coverage
Generous annual leave allowance

CONTACT

PEGASI – IT Recruitment Consultancy | Email: recruit@pegasi.com.vn | Tel: +84 28 3622 8666
We are PEGASI – IT Recruitment Consultancy in Vietnam. If you are looking for new opportunity for your career path, kindly visit our website www.pegasi.com.vn for your reference. Thank you!

Job Summary

Company Type:

Outsource

Technical Skills:

Data Engineering

Location:

Ho Chi Minh, Ha Noi - Viet Nam

Working Policy:

Hybrid

Job ID:

J01942

Status:

Active

Related Job:

Senior Backend Engineer

Ho Chi Minh, Ha Noi - Viet Nam


Product

Design, build, and operate backend services that power a leading B2B SaaS platform for the construction industry. Design clean, maintainable, and extensible software architecture. Identify and resolve performance bottlenecks (query optimization, caching, async processing). Drive software quality through automated testing, code reviews, observability, and CI/CD practices. Review and validate AI-generated code to ensure maintainability, correctness, and security. Mentor engineers and raise the team's engineering bar.#Development environment: Backend: Ruby (on Rails), Golang, Amazon Aurora (MySQL), DynamoDB FrontEnd: Nuxt.js, Vue.js, Next.js, ReactJS Mobile App: Kotlin, Swift, Flutter Deploy/ Build: AWS Amplify, CodePipeline, CodeBuild, CircleCI, GitHub Actions Others: Swagger, Docker, Figma, Confluence, JIRA, esa, gRPC Infrastructure: Helm, Terraform Automation test: Autify, Magicpod

Negotiation

View details

Application Engineer

Ho Chi Minh - Viet Nam


Product

  • Frontend
  • ReactJS
  • NodeJS

Build and maintain the Jan desktop app across macOS, Windows, and Linux using Tauri. Work full stack across the app: React and Node.js on the application layer, Rust on the native side. Own packaging, builds, and distribution: installers, auto-update, release pipelines, and code signing. Give Linux first-class support: packaging and distribution across common distros. Integrate local model runtimes and app features end to end. Work agent-natively as your default, running agents in parallel with your own review and guardrail discipline.

Negotiation

View details

AI Engineer

Ho Chi Minh - Viet Nam


Product

  • AI
  • Python

#AI-Powered Insurance Solutions Develop and deploy machine learning models to support insurance recommendation, pricing, risk assessment, and claims automation Build and maintain AI-powered features integrated into our embedded insurance platform - from data pipelines to inference APIs Design and implement RAG (Retrieval-Augmented Generation) pipelines and LLMbased assistants for internal and customer-facing use cases Collaborate with product and engineering teams to translate business requirements into practical AI solutions Participate in the full ML/AI lifecycle: data collection, feature engineering, model training, evaluation, deployment, and monitoring Contribute to our internal AI agent framework - building, testing, and improving multiagent workflows using LangChain, LangGraph, or similar tools Ensure AI models and services meet performance, reliability, and compliance standards in a regulated (insurance) environment Participate in developing CI/CD workflows for AI/ML/Data pipelines. Contribute to Agent Development Toolkits (ADK) and the automated Software Development Life Cycle (SDLC) that accelerate AI-native feature delivery#R&D & Engineering Tasks Research and experiment with state-of-the-art AI/ML techniques and evaluate their applicability to insurance use cases Prototype new AI capabilities quickly, iterate based on feedback, and graduate successful experiments into production Document experiments, model architectures, and engineering decisions clearly for team knowledge sharing Stay current with the AI ecosystem (LLM advances, agent frameworks, tooling) and bring relevant insights back to the team Support Full Stack Engineers in using the AI Development Toolkit (ADK) to develop applications - guiding and enabling their use of the toolkit rather than building the applications yourself - and maintain a basic understanding of the SDLC to help ensure consistent engineering standards across the platform Contribute to and continuously enhance the system knowledge base, ensuring documentation, patterns, and learnings are accessible to the agents

Negotiation

View details