Description
Overview
As the Senior Data Engineer, you will own the pipelines and platform that move Nerdy’s data from source systems into a governed lakehouse. You’ll build streaming and batch ingestion, run the Iceberg lakehouse on Starburst Galaxy, and ship reliable, well-tested dbt layers that analytics engineers and analysts build on—leveraging AI to accelerate development, QA, and operations.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, Math/Stats, or equivalent experience.
- 3+ years as a Data Engineer (or similar) in high-growth environments.
- Expert SQL and strong Python; experience building production pipelines that run unattended and recover on their own.
- Hands-on with open table formats and lakehouse engines: Iceberg on Trino/Starburst Galaxy, or similar (Delta Lake, Hudi, Spark, Databricks, Presto/Athena).
- Streaming experience with Kafka/Confluent or similar (Kinesis, Redpanda), including schema registry, CDC (e.g., Debezium), and idempotent loading into lakehouse tables.
- Proficiency with dbt and dbt Cloud (incremental models, macros, tests, docs) and orchestration tools such as Airflow or Dagster.
- Solid AWS experience (S3, IAM, Glue/REST catalogs) and infrastructure as code (Terraform or similar).
- Solid grasp of data modeling, partitioning, incremental processing, and backfill strategy on large datasets.
- Git-based development and CI/CD familiarity; disciplined approach to reviews, releases, and on-call.
- Experienced with AI-native tools that enhance productivity and speed (e.g., Cursor, Make, Supabase, Netlify, Claude Code, n8n, Firecrawl, ChatGPT, Grok, Bolt, Vercel, etc).
- Alignment with Nerdy’s apolitical, mission-focused culture.
Responsibilities
- Ingestion & streaming: Build and run streaming and batch ingestion from Kafka and third-party APIs into Iceberg tables; own CDC, schema evolution, and late-arriving data.
- Lakehouse platform: Operate the Iceberg lakehouse on Starburst Galaxy (Trino)—table layout, partitioning, compaction, retention, and cluster sizing.
- Layered modeling with dbt: Build and maintain the bronze, silver, and gold layers with incremental models, idempotent merges, and safe backfills; push shared logic upstream so downstream consumers stay simple.
- Orchestration & reliability: Design dbt Cloud job schedules, dependencies, retries, and alerting; keep production jobs within their freshness SLAs and drive incidents to root cause.
- Data quality & governance: Implement tests, contracts, freshness and volume monitoring, lineage, and access controls; handle PII within privacy and security guardrails.
- Performance & cost: Tune queries, file layout, and cluster usage to improve speed and manage compute and storage spend.
- Infrastructure as code: Manage clusters, storage, and access through Terraform and CI, with changes that are reviewed, reversible, and documented.
- Collaboration: Partner with Analytics Engineering, Product Engineering, and business teams to turn source changes and new requirements into dependable data products.
- Documentation & change control: Maintain versioned code (Git), review PRs, run CI/CD for pipelines and dbt, and document sources, models, and runbooks.
- AI-enabled workflows: Use AI for code generation/review, test scaffolding, pipeline debugging, data profiling, and anomaly detection—operating within privacy and security guardrails.