← Все вакансии/indrive
OfficeAlmaty

Data Engineer

I
indrive
Формат
Office
О роли

Описание вакансии

About the company

inDrive is a global ride-hailing company.

Responsibilities
  • Build and operate batch and streaming ingestion into a layered BigQuery DWH (raw → ODS → data marts) using Airflow, Debezium CDC over Kafka with protobuf, Pub/Sub, and Dataflow
  • Integrate external data sources end-to-end — marketing platforms (GA4, AppsFlyer, TikTok/Meta/Google Ads), payment providers, S3 buckets, and third-party APIs — including schema contracts, backfills, and reconciliation
  • Engineer the data platform itself in Python: custom Airflow operators and connectors in a shared ETL framework, Kafka Connect on Strimzi (K8s), Cloud Functions, and API integrations with external providers
  • Build CI/CD and change-management tooling for BigQuery: GitHub-based test-and-approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback
  • Own reliability and correctness of pipelines: idempotency, deduplication, late-data handling, backfill and replay, freshness monitoring and alerting; write integration and unit tests
  • Drive data governance and compliance: ITGC-compliant change management for BigQuery, IAM and least-privilege access, PII policy tags and DLP, Unity Catalog on Databricks, column-level lineage (OpenMetadata/Dataplex), and disaster-recovery planning
  • Build internal data tools and platform services for agentic workflows with data — Streamlit apps, Slack bots, LLM-based agents and MCP servers that help teams find and use data
  • Support analysts and business teams with data requests, fostering data-driven decision-making across the company
  • Contribute to system design and architecture with the development team
Стек и навыки

С чем работаем