HangzhouFull TimeSeniorPosted Today
Data Engineer大数据工程师 HangzhouExperiencedFull-timeResponsibilitiesResponsibilities
1.Scale Lakehouse & Data Pipelines
Iterate on the Iceberg lakehouse architecture; build and maintain multi-source batch and streaming pipelines to guarantee stability, data consistency and timeliness.
2.Data Modeling & Metric System
Build layered data models and standardized unified metrics (user, transaction, traffic, commercialization) via dbt to support self-service analytics and cross-functional business collaboration.
3.Query & Cost Optimization
Optimize performance and cost on computing engines including Spark, Glue and ClickHouse to boost dashboard and ad-hoc analytics experience.
4.Data Engineering, Observability & AI-Powered Exploration
Improve CI workflows, documentation, data quality and alerting systems; iteratively build AI-native data engineering capabilities such as automated data quality diagnosis, anomaly root-cause analysis, and data asset search & Q&A.
你将负责
1. 湖仓与管道规模化:演进 Iceberg 湖仓;建设/维护多源批流链路,保障稳定性/一致性/时效性;
2. 建模与指标体系:用 dbt 沉淀分层模型与统一口径指标(用户/交易/流量/商业化),支撑自助分析与业务协作;
3. 查询与成本优化:在 Spark / Glue / ClickHouse等上做性能与成本优化,提升看板与分析体验;
4. 工程化与可观测 + 智能探索:完善 CI、文档、质量与告警体系;增量推进“AI 为内核”的数据工程能力(质量诊断、异常定位、资产检索/问答等)。QualificationsRequirements
1.Over 1 year of data engineering experience
Solid SQL proficiency; capable of abstracting business logic into data models; familiar with layered modeling, consistent metric standards, incremental loads, backfilling and idempotency.
2.Strong coding capability
Able to write production-grade code with Python/Java/Go or equivalent; strong engineering awareness covering logging, exception handling and basic unit testing.
3.Hands-on lakehouse & cloud experience
Familiar with mainstream databases and data lake architectures & features; practical experience with S3-compatible storage paired with distributed compute engines.
4.Strong ownership mindset
Able to break down and estimate tasks reasonably; quickly troubleshoot and resolve issues, then conduct post-mortem reviews to accumulate experience; align metric standards with cross-functional teams and drive deliverables to completion.
Nice to Have
1.Practical implementation and optimization experience with Iceberg, Spark, Athena, Trino, Glue, DuckDB or ClickHouse.
2.In-depth dbt expertise (tests, documentation, CI/CD, macros) or ClickHouse/OLAP performance tuning experience.
3.Experience building event tracking pipelines and growth analytics (e.g. PostHog, funnel analysis, retention, attribution, A/B experiment metrics).
4.Familiarity with third-party data schemas such as Stripe and advertising platforms; or experience prototyping data Copilot / intelligent data tools.
What You’ll Get
1.Hands-on experience with next-gen data architecture
Work with a modern data stack built on Iceberg + Glue + dbt + Dagster to build more stable, faster and cost-efficient data pipelines.
2.AI-first working methodology
Leverage state-of-the-art AI tools to streamline intelligent data quality checks, troubleshooting, metric standardization and data asset management with higher efficiency.
3.High impact & fast feedback for accelerated growth
High ownership scope with rapid outcome feedback and strong sense of achievement. Collaborative engineering culture that encourages standardized SOPs and shared best practices.
我们希望你具备
1. 1 年+ 数据工程经验:SQL 扎实,能把业务抽象为数据模型;理解分层建模、口径一致性、增量/回填/幂等等;
2. 代码能力扎实:生产级别的代码能力(Python/Java/Go等),工程思维强(日志、异常处理、基础测试意识);
3. 湖仓/云上实践:熟悉各种常见数据库与数据湖架构、特性;有 S3 类存储 + 计算引擎经验;
4. Ownership:对任务进行有效拆解与评估,遇到问题能快速定位修复,并复盘沉淀;跨团队对齐口径并推进交付。
加分项(Nice to have)
1. Iceberg / Spark / Athena / Trino / Glue / DuckDB / Clickhouse 的实际落地与优化经验;
2.深度 dbt(tests/docs/CI/CD/宏)或 ClickHouse/OLAP 优化经验;
3.埋点链路构建经验与增长分析经验(例如使用过PostHog;漏斗/留存/归因/实验口径);
4.熟悉 Stripe/广告渠道等第三方数据模型,或做过“数据 Copilot/智能工具”原型
你将获得(What you’ll get)
1.新一代数据架构实战:Iceberg + glue + dbt + Dagster 的现代数据栈,做出更稳、更快、更省的链路;
2.AI 驱动的方法论:用最先进的AI工具,把质量、排障、口径与资产沉淀做得更智能、更高效;
3.高影响力+即时反馈 = 高速成长:权限大、成果反馈快、成就感强;技术氛围好,鼓励沉淀 SOP 与最佳实践。Apply
