Feature Store Architecture (D2)
Streaming and batch pipelines feeding an online and offline feature store, with a drift monitor closing the retraining loop.
Rendering…
Make it your own.
direction: right
sources: Source systems {
txn: Card transactions (Kafka)
crm: CRM (Postgres) {shape: cylinder}
bureau: Credit bureau file {shape: page}
}
pipelines: Feature pipelines {
stream: Streaming job (Flink)
batch: Batch job (Spark, nightly)
}
store: Feature store {
offline: Offline store (Parquet on object storage) {shape: cylinder}
online: Online store (Redis) {shape: cylinder}
registry: Feature registry and lineage
}
consumers: Consumers {
train: Training job
serve: Real-time scoring API
monitor: Drift monitor
}
sources.txn -> pipelines.stream
sources.crm -> pipelines.batch
sources.bureau -> pipelines.batch
pipelines.stream -> store.online: p99 write under 200 ms
pipelines.stream -> store.offline: log for training
pipelines.batch -> store.offline
store.offline -> consumers.train: point-in-time correct join
store.online -> consumers.serve
store.registry -> consumers.train
store.registry -> consumers.serve
consumers.serve -> consumers.monitor: logged features and scores
consumers.monitor -> pipelines.batch: retrain trigger on PSI breach