Data Engineer (PySpark)

Job Overview

We are seeking a Regular Data Engineer (PySpark) to join a team working with production ETL pipelines on AWS. The person in this role will apply strong PySpark skills within an existing architecture, independently deliver tasks, and diagnose common production issues. This position focuses on practical implementation and operational excellence rather than end-to-end architecture design.

Responsibilities

  • Implement data transformations and ETL logic using PySpark.
  • Maintain and develop existing AWS Glue jobs and related pipelines.
  • Investigate and resolve pipeline failures using logs, error messages, and basic metrics.
  • Diagnose and mitigate Spark performance issues (partitioning, shuffle, skew, joins).
  • Choose appropriate partitioning and join strategies for given workloads.
  • Handle error scenarios in ETL processes, including retries, reprocessing, and partial failures to avoid duplicate processing.
  • Work with large volumes of data and many files without creating uncontrolled parallelism.
  • Follow established architectural patterns and implement solutions that integrate with the current system design.
  • Support data migration or batch processing tasks as needed.

Qualifications

  • 2–4 years of experience as a Data Engineer or in a similar role.
  • Practical, hands-on experience with PySpark / Apache Spark.
  • Solid understanding of key Spark concepts, including partitioning, shuffle, repartition vs coalesce, broadcast joins, and data skew impact.
  • Practical experience with AWS services, especially AWS Glue, Amazon S3, and Amazon CloudWatch.
  • Experience building, maintaining, or enhancing ETL/data processing pipelines.
  • Ability to diagnose common Glue and PySpark issues from logs and basic metrics.
  • Experience handling ETL error scenarios: retries, reprocessing, partial failures, and preventing duplicate processing.
  • Basic understanding of idempotency and safe restart strategies after failures.
  • Experience processing large data volumes or many files and familiarity with batch migration processes.
  • Strong SQL skills and familiarity with joins, aggregations, filtering, grouping, and ranking.

Benefits

  • Remote working.
  • Unique TEAL culture, relationship- and respect-driven community, non-corporate atmosphere.
  • Agile approach and no bureaucracy.
  • Outstanding integration trips to various places in Europe.
  • Activities to support your well-being and health.
  • Luxmed Gold Extended medical care and Multisport Plus benefit.
ID: 680 job_post.published_on: 27/08/2026
announcement.apply