building-feature-pipelineslisted
Install: claude install-skill Unknown-333/awesome-data-engineering-skills
# Building Feature Pipelines
## When to use
- Engineering features for ML models from warehouse/stream data.
- Preventing label leakage and train/serve skew.
- Setting up a feature store, online serving, or historical backfills.
- Do NOT use for general modeling/aggregation (use dbt/Spark skills) unless it
feeds ML features.
## Workflow
```
- [ ] Define each feature with an entity key and an event timestamp
- [ ] Build training sets with point-in-time-correct joins (as-of the label time)
- [ ] Share ONE definition for offline (training) and online (serving)
- [ ] Set freshness/materialization for online features
- [ ] Backfill historical features idempotently for training
```
1. **Point-in-time correctness.** Join features as of each label's timestamp — use
only data that was known before the prediction time. This prevents **label
leakage**, the most damaging feature bug.
2. **Offline/online parity.** Compute a feature the same way for training (offline,
batch) and serving (online, low-latency). Divergent logic causes **train/serve
skew** and silent production degradation.
3. **Freshness.** Online features must be materialized on a schedule that meets the
model's staleness tolerance.
4. **Idempotent backfills.** Recomputing historical features must be repeatable
(see `designing-backfills-and-replays`).
## Patterns
**Point-in-time (as-of) join** — pick the latest feature value strictly before
each label event:
```sql
SELECT l.entity_id, l.label_ts,