Self-Serve Feature Platform for 40 ML Teams
Cutting feature-to-production time from 6 weeks to 2 days with a unified offline/online feature store and contract-based ownership.
ML teams each maintained bespoke pipelines, duplicating features and silently drifting between training and serving. Production incidents tracked back to train/serve skew were the #1 root cause for the previous year.
Features are declared in a typed registry with owners, SLAs, and freshness contracts. Offline jobs compute historical values into a Parquet lakehouse; the same transformations run online against a low-latency KV store. A point-in-time correctness layer guarantees training data matches what serving would have returned.
┌────────────────────┐
│ Feature Registry │
│ (typed contracts) │
└─────────┬──────────┘
│
┌───────────┴───────────┐
▼ ▼
┌─────────────┐ ┌─────────────┐
│ Offline │ │ Online │
│ (Parquet, │ │ (KV store, │
│ Spark) │ │ <10ms) │
└──────┬──────┘ └──────┬──────┘
▼ ▼
┌─────────────┐ ┌─────────────┐
│ Training │ │ Serving │
└─────────────┘ └─────────────┘- 01
Single transformation definition compiled to both Spark (offline) and Flink (online).
- 02
Point-in-time joins enforced; train/serve skew alerts wired to PagerDuty.
- 03
Per-feature SLAs (freshness, availability) tracked like any other production service.
- 04
Onboarding reduced to a PR against the registry; no platform-team ticket required.
