Rc · Sales recommendation platform: all pages
What did the model actually know at the time?
Following an order from the source system into training and live recommendations, including the awkward case where the record arrives late.
Suppose a customer places an order on Monday. The recommendation platform does not receive the record until Wednesday, but a salesperson used its suggestions on Tuesday. Months later, we rebuild the training data and join that Monday order to the customer’s history, which can make Tuesday’s recommendation appear better informed than it actually was.
The dates all look reasonable until we ask a more precise question: are we reconstructing what happened in the business, or what the platform could have known when it made the decision? Both histories can be useful, but using one while claiming to evaluate the other gives us misleading evidence about model quality.
This is the problem behind the point-in-time feature pipelines I built for the recommendation platform. Features are the inputs a model uses, such as a customer’s recent order count, and the pipeline prepares them. For evaluation, those inputs need to reflect what was available when the prediction would have been made.
I used as-of joins to look up the relevant historical records and embargo gaps to leave space between training and evaluation periods where overlapping information could make a test misleading. Let’s follow that late order through the design to see why the dates need so much care.
Keeping two kinds of time
I would preserve the business event time and the time the record became available to the feature pipeline. An order’s transaction date explains when the business event happened; an availability timestamp explains when the platform could have used the particular record or correction. To replay what the live system could have seen, both dates matter. An order must have happened before the decision, and its record must also have reached the platform by then.
A point-in-time join retrieves a historical feature version for a prediction date. It still needs an explicit late-data policy: an event-time condition alone can admit a correction that only arrived later. Feast’s documentation describes this distinction and its support for constraining retrieval by a creation timestamp when that timestamp represents actual availability. Point-in-time joins are a useful reference for the underlying behaviour; the design is not dependent on using Feast.
For the Monday order, a replay of Tuesday’s live inputs would exclude the late record. A corrected business-history dataset could include it, but it would have a different purpose and a distinct version. I would record that distinction alongside the dataset, so someone returning to it months later can tell which history they are looking at.
Read the diagram as text
In this example, the order occurs Monday, a recommendation is made Tuesday, and the record reaches the feature pipeline Wednesday. An as-served Tuesday feature snapshot must exclude that order even though its business date is Monday. Later corrected analysis must use a separately identified dataset.
Defining the unit before calculating the features
The training row needs to represent a real decision. An account-health example might be one eligible account at a scoring time; a product-ranking example needs an account, a recommendation time and the candidates eligible to be shown together. Without that context, a collection of customer-product rows can look convenient to train on while no longer representing the list a salesperson actually received.
Next, I would write down what counts as a successful outcome and how long we need to wait to observe it. A purchase within a defined future window is different from an eventual order with no time limit, and a young open enquiry is different from a lead whose outcome window has closed. Records whose outcome period is incomplete should not silently become negative examples.
For ranking, an unpurchased product is also not automatically a strong negative. The account may never have seen it, it may have been out of stock, or it may have been placed where the salesperson rarely looked. That is why I would record which products could have been offered and which were actually shown. Before using an unpurchased item as a negative training example, we need to be clear about what the absence of a purchase tells us.
Giving each feature a contract
A feature such as “orders in the last 90 days” sounds simple until different jobs count cancelled orders differently or use different customer identities. Alongside the calculation, I would write down which customer identifier it uses, which orders count, where the time window starts and ends, and what happens when data is missing. Someone also needs to own that definition when questions arise. Changes to those meanings create a new feature version even when the column name would otherwise remain the same.
That contract also describes expected freshness and the response when it is not met. Missing recent-order data must not automatically become zero orders: zero is a business observation, while missing means we do not know. Depending on what was evaluated, the model might accept an explicit missing indicator, use a compatible reduced-input variant, or defer the suggestion.
The source and feature definitions can be shared across training and serving without making every input a live lookup. Slow-moving history can be prepared in batch, while only the small set of features whose freshness changes an interactive decision needs current serving storage. Stock and price remain operational facts that are checked again at commitment, even if recent snapshots were used to create the recommendation.
Read the diagram as text
Versioned raw records become cleaned facts retaining event time and availability time. Shared definitions generate historical feature snapshots for training and fresh features for serving. Mature outcomes train models, while current business truth is revalidated separately at commitment.
Rebuilding history without silently changing the model’s world
Raw supplier and business-system extracts should remain identifiable by source and version. Cleaning jobs produce validated facts, and feature jobs produce snapshots tied to their code, definition versions and source snapshots. The model dataset then records which feature and label snapshots it used, along with the eligible population and time windows.
If we find a bug in a calculation, I would recalculate the affected history into a new version and compare it with the old one. This is often called a backfill. Replacing historical feature values in place would make an old experiment difficult to reproduce and could change the apparent quality of a released model without changing the model itself. An approved backfill needs a deliberate path into a new training run or current feature publication.
Duplicate source events need similar care. A stable identifier for each event or transaction can prevent a retry from counting one order twice. If the source later corrects an order, I would preserve its identity and record the new version, so the correction is not counted as another purchase.
Checking that training and serving agree
Shared code reduces the chance of disagreement, but I would still test the two paths on matched examples. For the same account, decision time and known input snapshot, compare the offline feature result with the value the serving path would supply, including timezone boundaries, missing fields and late-arriving corrections. A test using today’s values on both sides would not establish historical correctness.
The model package needs to declare the feature versions and input schema it expects. If a feature definition changes from ordered quantity to delivered quantity, the service should not accept it merely because the type is still numeric. Schema checks catch the shape of the contract; version and semantic checks catch the meaning.
For the serving store, I would prepare a complete batch under a new snapshot identifier, validate its coverage and freshness, then publish that snapshot as a unit. Incrementally overwriting the current worklist while the job runs can expose a mixture of old and new scores. Readers should either use the completed new snapshot or the previous acceptable one, with its age visible.
Why this becomes an operating responsibility
A failed feature job needs an owner and a recovery decision before the morning worklist is due. It may be reasonable to keep yesterday’s recommendations within an agreed freshness window, while some inputs require the affected path to stop. I would agree those limits with the team before a failure happens, so whoever is on duty knows when yesterday’s list is still useful and when it needs to be withdrawn.
For the late Monday order, the value of this design is that we can explain Tuesday’s suggestion honestly, correct Wednesday’s view, and train the next model on a dataset whose meaning is clear. It gives model evaluation something dependable to stand on and makes a data incident much easier to distinguish from a modelling problem.
The monitoring article follows that distinction when a quality alert arrives.