Data Engineer
5 views
28.07.2026
Data Engineer
Senior
Data Science and Analytics
7-10 years

Schedule

Full-time

Location

Kyiv

Type of Work

Remote

Sphere

IT Consulting
About Job
Summary

We are looking for a Senior Data Engineer with strong hands-on experience in Apache Spark, PySpark and GCP Dataproc to join Implex and work on a long-term data platform project for a public-sector in the Middle East.

The engineer will contribute to the development and support of data-processing pipelines that integrate multiple enterprise data sources, including PostgreSQL, SAP HANA and S3-compatible object storage. The role involves not only writing PySpark jobs but also troubleshooting distributed workloads, configuring Dataproc environments, securing external connections and ensuring reliable delivery of data into the target data warehouse.

Responsibilities

  • Develop, maintain, and optimize ETL pipelines using Apache Spark and PySpark.
  • Configure, run, and troubleshoot Spark workloads on GCP Dataproc.
  • Package and submit Spark jobs, manage dependencies, and analyze driver and executor logs.
  • Identify and resolve performance issues related to memory usage, shuffling, partitioning, data skew, and distributed data processing.
  • Integrate Spark workloads with S3-compatible object storage using the Hadoop S3A connector.
  • Configure and troubleshoot Spark connectivity with MinIO, including bucket policies, custom endpoints, path-style access, TLS, signature compatibility, and redirect handling.
  • Configure custom CA certificates and Java truststores for Spark drivers and executors.
  • Implement and optimize scalable Spark JDBC reads and writes for PostgreSQL and SAP HANA, including partitioning, batching, pushdown, and data type mapping.
  • Build reliable incremental data loads, retries, backfills, idempotent processing, and data reconciliation mechanisms.
  • Design and support staging-to-publish data flows, bulk data loads, and data warehouse integration patterns.
  • Contribute to data modeling, schema evolution, Slowly Changing Dimension patterns, data lineage, technical documentation, and data quality practices.
  • Apply secure secrets management and least-privilege access principles across data integrations.


Results / KPI

Must-Have Technical Skills

  • Apache Spark / PySpark development (Dataproc): driver/executor behavior, job packaging/submission, performance tuning
  • GCP Dataproc operations: cluster configuration, init actions, dependency management, troubleshooting via logs/metrics
  • Hadoop S3A connector: `fs.s3a.*` configuration, endpoint/path-style access, credential providers, S3 semantics
  • MinIO (S3-compatible) integration: bucket policies, TLS endpoints, signature/redirect troubleshooting
  • TLS/SSL & PKI with custom CA: certificate chains, SAN/hostname validation, diagnosing handshake/PKIX errors
  • Java truststores (JKS/PKCS12) & JVM SSL config: `keytool`, distributing truststores, setting driver/executor JVM options
  • PostgreSQL integration: Spark JDBC reads/writes at scale, indexing/performance basics, data type mapping
  • SAP HANA integration: JDBC/ODBC connectivity, driver management, calculation views vs tables, pushdown/performance tuning
  • ETL engineering: incremental loads/CDC concepts, idempotency, retries, backfills, data quality/reconciliation
  • Data Warehousing integration: strong SQL, staging-to-publish patterns, SCD concepts, bulk load strategies
  • Data modeling & governance basics: dimensional modeling, schema evolution, lineage/documentation practices



Requirements
What we Expect

  • Strong ownership of assigned tasks and responsibility for delivering reliable, production-ready solutions.
  • A hands-on engineering mindset with the ability to investigate complex issues across Spark jobs, Dataproc infrastructure, databases, networking, and security configurations.
  • Ability to work independently, troubleshoot problems systematically, and clearly communicate findings, risks, and proposed solutions.
  • Proactive identification of performance, reliability, data quality, and security issues before they affect production workflows.
  • Experience collaborating with technical leads, project managers, infrastructure teams, and client-side stakeholders.
  • Ability to understand existing architecture and business requirements before proposing technical changes.
  • Clear and structured technical communication, including maintaining documentation for pipelines, configurations, integrations, and operational procedures.
  • Commitment to clean, maintainable, testable, and reusable code.
  • Readiness to support existing data pipelines, investigate production incidents, and participate in root-cause analysis.
  • Comfortable working in an enterprise environment with security requirements, approval processes, legacy integrations, and multiple external dependencies.
  • Upper-Intermediate or higher level of English for regular communication with an international team.
  • Availability for full-time, long-term collaboration within a time zone compatible with the Middle East and European teams.


Professional Skill
Adaptability Communication Agile
Tools
Google Cloud Platform (GCP) Apache Spark Google BigQuery
Tech Stack
Python
Languages
English | Upper-Intermediate
Will be a plus

  • Linux + networking fundamentals: DNS, routing, firewall/LB/proxy basics; tools like `curl`/`openssl s_client` for validation
  • Secure secrets handling: GCP Secret Manager (or equivalent), least-privilege access, avoiding hardcoded credentials
  • KAFKA knowledge if we ever bring KAFKA into the architecture again


What we offer
What we Offer

  • A long-term international project
  • Opportunity to work on a national-scale digital platform used by thousands of users
  • Remote full-time collaboration
  • Professional and supportive team environment
  • Challenging technical tasks and a meaningful product with real-world impact

Hiring process

  • HR interview
  • PM/Technical interview with Implex
  • Dev Lead interview on the client side



About Company
Implex
https://implex.dev/

About Us

An A-people software development company that provides high-quality outsourcing services to the US and Europe. We work with the best people giving them the ability to work with interesting projects and modern technologies. We strive to build relationships with our clients that last for years. This makes trust, reliability, and respect the principle values of our culture. Our clients recognize us as a one-stop software development company that allows focusing on business development without having to worry about the tech side.

OUR SERVICES:

+ Product design: UX/UI design, functional & technical architecture, product roadmap

+ Product engineering: Full-cycle web & mobile app development

+ Product maintenance: Support of applications and cloud infrastructure