Summit 2026
Hopsworks · Feature Store Summit 2026
Beyond Feature Retrieval: Feature Views in Hopsworks
Transformations and composition in feature views

Feature stores · a brief history of shift-left transformations
Features were computed before anyone asked for them
SELECT cc_num, ts AS event_time, SUM(amount) OVER ( PARTITION BY cc_num ORDER BY ts RANGE INTERVAL 1 HOUR PRECEDING ) AS sum_trans_last_hour FROM credit_card_transactions
Feature stores began as a place to keep precomputed features. It starts with raw card transactions.
Shift left: precompute feature values in a batch or streaming job with the best engine for the workload.
Store the results, keyed by entity, with their event time. One row per transaction, history kept.
Training data is a batch of historical feature data, inference the latest. The same precomputed values, no skew.
This talk · shift-right transformations in SQL and Python
The feature view transforms on read
@hopsworks.udf(return_type=str, mode="pandas") def custom_transformation(longitude, latitude): return location(longitude, latitude) robust_scaler = fs.get_transformation_function(name="robust_scaler") cc_trans_fg = fs.get_feature_group("cc_trans_fg", version=1) query = (cc_trans_fg.select_features() .aggregate({"amount": ["sum"]}, window=timedelta(hours=1))) fv = fs.create_feature_view(name="fraud_fv", version=1, query=query, labels=["fraud"], transformation_functions=[custom_transformation("longitude", "latitude"), robust_scaler("amount")], ) # Training X_train, y_train, X_test, y_test = fv.train_test_split(test_ratio=0.2) # Online Inference fv.get_feature_vector(entry={"cc_num": "1111 2222 3333 4444"})
This talk is about the other side of the store. We compute features on-demand from the raw event data (just like batch and streaming jobs) but only when needed. The main challenges are performance and scale.
Shift right: the feature view retrieves data and transforms read data.
Feature views · SQL, then Python
Feature Views read with SQL, then Transform in Python
Feature groups hold the rows: users, their transactions, their churn labels.
A feature view query joins the result of CTEs whose operations are pushed down to the database / lakehouse (feature groups).
Feature views define the query for reading data, the schema for the model, and the transformations performed in Python/Pandas UDFs.
Feature transformations fit on training data statistics. Custom transformations can also be applied. Transformations are a dataflow computation in a DAG.
A model is trained on one version and served with the same transformations.
Transformation DAG · a dataflow graph, as in Apache Hamilton
The DAG is in the parameter names
@hopsworks.udf(return_type=[int]) def A() -> int: """Constant value 35""" return 35@hopsworks.udf(return_type=[float]) def B(A: int) -> float: """Divide A by 3""" return A / 3@hopsworks.udf(return_type=[int, float]) def C(A: int, B: float) -> tuple[int, float]: """Ceiling of B, and A squared times B""" return math.ceil(B), A**2 * Bfv = fs.create_feature_view(..., transformation_functions=[A(), B("A"), C("A", "B")], )
Dependencies are declared by matching parameter names to the names of other functions.
A takes no input. A source node.
B(A): the parameter named A makes B depend on A.
C(A, B) depends on both and returns two features. The DAG is the code. Nothing else declares it.
The feature view takes the DAG as a list: each function once, its inputs named. The order of evaluation follows from the names.
Feature views · precomputed to on-demand
Feature Views: a spectrum from precomputed to stateless on-demand features
A feature view can sit anywhere on the spectrum, depending on when its features are computed.
Precomputed: batch and streaming jobs compute every feature ahead of time. The feature view only retrieves.
Both: precomputed features from the data warehouse or event bus, joined with features computed on read.
Stateless, on-demand: every feature is computed from the raw data when it is read. No data warehouse or event bus is needed. Feature logging saves data for training or evals.
Data transformations
Useful Properties of Data Transformations for ML
Feature DSLs · why they failed
Why domain-specific languages failed for feature stores
- Operations
- Discovery
- Governance
- Analytics
Feature views · feature logging
Feature logging
fv = fs.get_feature_view(name="fraud_fv", version=1) fv.enable_logging()features = fv.get_feature_vector(entry={"cc_num": cc_num}, transform=False) encoded = fv.transform(features) pred = model.predict(encoded)fv.log(untransformed_features=features, transformed_features=encoded, predictions=pred, model=model)logs = fv.read_log(start_time=yesterday, transformed=False)
Log the feature values with every prediction. Without them, drift cannot be measured.
Transformations lose information: a scaled amount or a bucketed hour cannot be read back as what the customer did.
So log both: the untransformed values for drift, the transformed values for what the model saw.
Read the logs back and compare with the training data: the amount feature has drifted.
- Operations
- Discovery
- Governance
- Analytics
Feature views · feature monitoring
Feature monitoring
fv = fs.get_feature_view(name="fraud_fv", version=1) fv.enable_logging()config = ( fv.create_model_monitoring( name=CONFIG_NAME, model_name=model.name, model_version=model.version, cron_expression="0 0 * * * ? *", ) .with_detection_window(time_offset="1d", window_length="1d") .with_reference_training_dataset() .compare_on(feature_name=FEATURE, metric="PSI", threshold=0.2) .save() )
Monitoring compares what the model sees with what it learned from.
Log every read of the feature view.
A monitoring job tied to one model version, on a schedule.
Compare yesterday's logged features against the training data the model was fitted on.
Population stability index per feature. Alert above 0.2.
- Operations
- Discovery
- Governance
- Analytics
Feature views · helper columns
Training and inference helper columns in feature views
In training and/or inference pipelines, sometimes you need to read extra columns from tables - for example, for model evaluation reports or for enriching predictions output by inference. For this, you can add columns to a feature view that will only be returned when you get training data or inference data, respectively.
fv = fs.create_feature_view( name="fraud_fv", version=1, query=query, labels=["is_fraud"], training_helper_columns=["customer_segment"], inference_helper_columns=["merchant_name"])X_train, y_train, X_test, y_test = fv.train_test_split( test_ratio=0.2, training_helper_columns=True)df = fv.get_batch_data(start_time=t0, inference_helper_columns=True) helpers = fv.get_inference_helper({"cc_num": cc_num})
A feature view returns the features the model takes. Helper columns are extra columns, returned only to the pipeline that asked for them.
Training helper columns come back only with training data: customer_segment breaks the evaluation report down by segment, without being trained on.
Inference helper columns come back only with inference data: merchant_name enriches the prediction the caller gets back.
- Operations
- Discovery
- Governance
- Analytics
Feature views · vector index
A vector index in a feature view
A feature group can hold an embedding column with a vector index. The feature view that reads it serves similarity search next to its other features: the k nearest rows come back as feature vectors, from the same online store.
index = EmbeddingIndex() index.add_embedding( "merchant_vec", 384, SimilarityFunctionType.COSINE)merchants_fg = fs.create_feature_group(name="merchants_fg", version=1, primary_key=["merchant_id"], online_enabled=True, embedding_index=index)fv = fs.create_feature_view(name="merchant_fv", version=1, query=merchants_fg.select_all()) similar = fv.find_neighbors(embed(description), k=5)similar = fv.find_neighbors(embed(description), k=5, filter=merchants_fg.category == "travel")
Rows carry an embedding like any other feature. The vector index makes the feature view searchable by similarity, with no second system to keep in sync.
Declare the index on the embedding column: merchant_vec, 384 dimensions, cosine similarity. Rows are indexed as they are inserted.
find_neighbors on the feature view: the k nearest rows come back as feature vectors, with every other feature of the row.
A filter on the other features narrows the candidates first: the five nearest among travel merchants only.
- Operations
- Discovery
- Governance
- Analytics
Feature views · lineage
Lineage through a feature view
Every link is recorded when it is created: a feature group from its source, a feature view from its feature groups, training data from the view, a model from the training data. Nothing is annotated by hand and no code is scanned.
fv = fs.get_feature_view(name="fraud_fv", version=1)links = fv.get_parent_feature_groups() for fg in links.accessible: print(fg.name, fg.version) # cc_trans_fg 1, merchants_fg 1for model in fv.get_models(): print(model.name, model.version, model.training_dataset_version) # fraud 3 2model = mr.get_model("fraud", version=3) fv = model.get_feature_view() X, y = fv.get_training_data( training_dataset_version=model.training_dataset_version)
A feature view sits in the middle of the graph: sources and feature groups on one side, training data, models and deployments on the other.
Upstream: which feature groups a view reads, which versions, and the sources that fill them.
Downstream: every training dataset built from the view and every model trained on one of them.
From a model back: the view that fed it and the training data version it learned from, for an audit or a retrain.
- Operations
- Discovery
- Governance
- Analytics
Feature views · reproducible training data
Reproducible training data: time travel in feature groups
Every write to a feature group is a commit. When a feature view creates training data, it records the commit ID of each feature group it read, so the same training dataset version can be recreated later from exactly those commits.
cc_trans_fg.commit_details() # {c4: {"committedOn": "2026-09-24 23:10", "rowsInserted": 18420, ...}, ...}version, job = fv.create_training_data( start_time="2026-07-01", end_time="2026-09-30", description="Q3 fraud") # v1 records the window and each feature group's commit IDcc_trans_fg.insert(october_df) # c5, c6: new commits keep landingX, y = fv.get_training_data(training_dataset_version=1) # recreated from the recorded commits: cc_trans_fg @ c4, merchants_fg @ m2
Feature groups are versioned by commit. A feature view remembers which commit of each feature group it read.
Training data is cut from the view over a time window. The version records the commit ID of every feature group at that moment.
Data keeps arriving: new commits land on the feature groups after the training data was made.
get_training_data with the version recreates the materialized dataset from those commit IDs: the same rows, however many commits have landed since.
- Operations
- Discovery
- Governance
- Analytics
Feature views · schematized tags
Schematized tags on feature views
# tag schema "model_card", defined once, enforced on every write {"required": ["owner", "pii", "status"], "properties": {"owner": {"type": "string"}, "pii": {"type": "boolean"}, "status": {"enum": ["experimental", "production", "retired"]}}}fv.add_tag("model_card", {"owner": "fraud-team", "pii": True, "status": "production"}) hits = fs.search_feature_views( tag_filter={"model_card.status": "production"})fv.add_tag("model_card", {"owner": "fraud-team", "status": "prod"}) # rejected: "pii" is required, "prod" is not in the status enum# tag history, kept for every write: who, what, when # 06-02 experimental · 08-15 production · 09-01 pii false → true
A tag on a feature view carries a schema. A value that does not fit is rejected on write, so every tag means the same thing everywhere.
Discoverable: search feature views by tag value: every production view, every view that holds PII, every view a team owns.
Governable: required fields cannot be skipped and a status must be an allowed value. A pii flag can drive access policy.
Analytics: the tag history is stored. Who moved a view to production and when, and how many views hold PII this quarter against last.
Feature view online performance
A benchmark of Hopsworks vs AWS SageMaker.
Online inference · model deployment with a feature view
Online transformations with a feature view
@hopsworks.udf(float, drop=["amount"]) def amount_ratio(amount, sum_trans_last_hour): return amount / sum_trans_last_hour # an on-demand transformation is registered on the feature group aggs_fg = fs.get_or_create_feature_group( name="cc_trans_aggs_fg", version=1, primary_key=["cc_num"], event_time="event_time", online_enabled=True, transformation_functions=[amount_ratio], ) # history: computed on insert, amount dropped, amount_ratio stored aggs_fg.insert(aggs_df) fv = fs.create_feature_view(name="fraud_fv", query=aggs_fg.select_all()) # 10:03, card …4821 asks to spend 1,200.00: computed on read fv.get_feature_vector( entry={"cc_num": cc_num}, request_parameters={"amount": 1200.00}, )
The inference pipeline, opened up: a deployment wraps a model with its feature view.
10:03: a request for card …4821 brings the serving key and the amount only the caller knows, 1,200.00.
The serving key reads the precomputed 741.40 from the online store.
An on-demand transformation, amount_ratio, divides the request by the precomputed value: 1.62. On insert, the same function filled the training rows.
The model predicts on 1.62, and the response goes back.
Features and predictions are logged to the feature view's logging feature group: the input for monitoring.
Online inference · Hopsworks vs SageMaker, prediction logging on
Model serving benchmark: Hopsworks vs SageMaker
- Same churn model on both: a RandomForest of 50 trees over five features from two precomputed feature tables.
- Same request: ten serving keys. Features are looked up inside the deployment with the blocking
feature_view.get_feature_vector(). - 12,000 requests per pass at 40 rps for 300 s, zero errors. Client-observed latency, about 13 ms of network on both sides.
- SageMaker has no 1 vCPU instance, so it was pinned to one worker and one BLAS thread. Hopsworks: second of two passes; p99 varies up to 2× between passes.
With the blocking feature_view.get_feature_vector(), Hopsworks answers about 5 ms faster at the median on half the cores, and sustains about 15% more throughput.
SageMaker Data Capture logs only inputs and outputs: the features never leave the container. Hopsworks feature logging records the whole feature vector, with request id, training dataset and model version.
Online inference · async feature retrieval
Asynchronous feature retrieval and transformations
New support for async feature retrieval speeds up model deployments: feature_view.get_feature_vector_async(...)
async def predict(inputs): keys = [{"cc_num": i["cc_num"]} for i in inputs] vectors = await feature_view.get_feature_vector_async( entry=keys, request_parameters=inputs) return model.predict(vectors)
One worker, blocking, as in the benchmark: a request waits for the online store, and the next request waits for it.
Async: while one request awaits its features, the worker serves the next. Same core, more requests in flight, higher throughput.
Online inference · async model serving and async feature retrieval, Hopsworks vs SageMaker
Async model serving benchmark
- The same experiment as the blocking benchmark, run on one day: same churn model, same two online feature tables, ten serving keys per request, prediction logging on.
- Both predictors are async and await the feature lookup. SageMaker needs a one-process FastAPI container for it, since AWS's managed container has no async handler. Hopsworks:
async def predictawaitingfeature_view.get_feature_vectors_async(). - 12,000 requests at 40 rps for 300 s, zero errors. Client-observed latency, 10 to 11 ms of network on both sides.
- SageMaker still has twice the cores. One pass per cell, so the tail is noisy.
Async against async: Hopsworks answers about 2 ms faster at the median and 10 ms faster at p99, on half the cores. SageMaker's one-process container has the higher ceiling.
Awaiting the lookup lifted both ceilings: SageMaker's by 74%, where each lookup is a 15 ms call to a separate service, and Hopsworks' by 14%, where the RonDB lookup takes about a millisecond and the core is already busy predicting.
Online inference · get_feature_vector_async, measured
Async feature retrieval: p99 down 72%, throughput up 24%
- The same experiment as the SageMaker comparison benchmark, but changing the predictor function in FastAPI to be async, including all calls inside it, including
feature_view.get_feature_vector_async(...).
Async feature retrieval instead of blocking on it, on the same core: p99 89.6 → 25.1 ms, saturation 218.6 → 270.1 rps.
At p99, 71 of the 90 ms are the worker idle, waiting on the online store: 2.48 ms of wall time for 0.18 ms of CPU per request. Async overlaps that wait with the other requests in flight.
O'Reilly · Batch, Real-Time, and Agentic AI Systems
Building Machine Learning Systems with a Feature Store
A must-read for anyone serious about building efficient, real-world ML systems.
In this crazy industry of ours, Jim's the closest thing we have to a world-class expert. A detailed, practical, reusable manual on how to get a good-quality running AI system.
A lot of practical tips on how to navigate production ML deployments.
A great service to ML practitioners with best practices, a clear step-by-step guide.