← ClaudeAtlas

building-feature-pipelineslisted

Build ML feature pipelines and feature stores — point-in-time-correct joins to avoid label leakage, offline/online parity, feature freshness and backfills, and materialization with tools like Feast. Use when engineering features for ML, preventing train/serve skew or data leakage, building a feature store, or backfilling historical features for training.
Unknown-333/awesome-data-engineering-skills · ★ 16 · Data & Documents · score 68
Install: claude install-skill Unknown-333/awesome-data-engineering-skills
# Building Feature Pipelines ## When to use - Engineering features for ML models from warehouse/stream data. - Preventing label leakage and train/serve skew. - Setting up a feature store, online serving, or historical backfills. - Do NOT use for general modeling/aggregation (use dbt/Spark skills) unless it feeds ML features. ## Workflow ``` - [ ] Define each feature with an entity key and an event timestamp - [ ] Build training sets with point-in-time-correct joins (as-of the label time) - [ ] Share ONE definition for offline (training) and online (serving) - [ ] Set freshness/materialization for online features - [ ] Backfill historical features idempotently for training ``` 1. **Point-in-time correctness.** Join features as of each label's timestamp — use only data that was known before the prediction time. This prevents **label leakage**, the most damaging feature bug. 2. **Offline/online parity.** Compute a feature the same way for training (offline, batch) and serving (online, low-latency). Divergent logic causes **train/serve skew** and silent production degradation. 3. **Freshness.** Online features must be materialized on a schedule that meets the model's staleness tolerance. 4. **Idempotent backfills.** Recomputing historical features must be repeatable (see `designing-backfills-and-replays`). ## Patterns **Point-in-time (as-of) join** — pick the latest feature value strictly before each label event: ```sql SELECT l.entity_id, l.label_ts,