A buyer can bid on several lots and leave an auction empty-handed. A sales-only recommender sees no new purchase. A bid-aware model sees intent, but also competition, abandoned plans, and outcomes that never became transactions. Making that signal useful requires deciding how much to trust it.
At Treverse, this became the central design problem for RecSys v3: how much should a recommendation depend on a buyer's history, and when should the system trust the inventory-level baseline instead?
V3 adds bid intent to retrieval, limits personalization when history is weak, and checks the recommendation list before publication. The production A/B test found 3.39% more buyers opening recommendations relative to a handpicked control, with a nominal 95% interval of +0.53% to +6.26%. Sales remained uncertain. A split between purchase-history cohorts made the next engineering priority more specific: investigate which lists improved, and for whom.
From completed sales to earlier intent
The first Treverse RecSys case study describes a temporal graph retriever trained on completed sales, with precomputed recommendations and a small online serving path. Bid history was not part of that tested model. Its experiment results belong to that sales-based version.
V3 adds a separate bid-intent channel and a candidate-aware ranker. Purchase history, bid history, and an auction prior now have distinct roles. The prior ranks currently eligible inventory using auction context without requiring personal history. The ranker controls how far personal evidence can move the decision away from that baseline.
The online service still reads precomputed, validated rankings. These are changes to the recommendation model and its release checks, with the lightweight serving boundary preserved. The v3 experiment used a contemporaneous handpicked control; it was not a head-to-head test against v1.
More signals need stronger time boundaries
A completed transaction records an outcome. A bid records intent under competition. Keeping separate purchase and bid representations lets retrieval use both without collapsing their meanings into one undifferentiated interaction count.
Purchase retrieval, bid retrieval, and the auction prior contribute to one candidate union. Every candidate receives scores from all three sources. An item discovered through bids can therefore rank highly because of purchase evidence or auction context. Its retrieval source does not determine which evidence the ranker is allowed to use.
Earlier signals also make leakage easier. Historical evaluation reconstructs what was eligible and what information was available at the recommendation time. A bid that occurred earlier but arrived in the data system later cannot become a feature retroactively. Information from the target auction is excluded from the historical graph path so the model cannot recover its own answer through neighboring interactions.
Whole auctions remain together in chronological data partitions. Model choices are made on tuning events, followed by an untouched promotion set. This keeps related outcomes from appearing on both sides of a split and aligns the offline comparison with the decision the model is meant to make.
Give personalization an evidence-dependent limit
A hard cold-start rule creates a discontinuity: one extra action can suddenly change a buyer from mostly generic recommendations to strong personalization. A fixed blend avoids that jump, but applies the same confidence to very different histories.
V3 separates two decisions. History strength and recency set a limit on personal influence. A learned router allocates influence within that limit. Any unused share stays with the auction prior. The source policy is shared across the candidates for a buyer query, so candidate features cannot silently change how much history is trusted from one item to the next.
With no usable history, the score is exactly the prior. A sparse history permits a limited personal contribution. Richer evidence allows more, while stale evidence can move the decision back toward the baseline. More history increases the permitted influence; it does not force the router to use all of it.
A subtle implementation requirement is that the rest of the ranker must respect the same boundary. Protecting only the source blend would be insufficient if a later neural correction could overwhelm it. V3 bounds that correction with the same personal influence, and training uses reliability-adjusted targets for the served intent score.
This is a structural cold-start guarantee. Whether buyers prefer the resulting list is a separate question for the online experiment.
Evaluate retrieval, ranking, and the final list separately
Retrieval and ranking solve different problems. The retriever needs to reach relevant inventory. The ranker needs to order the candidates it will actually receive, balancing intent, completed purchases, and sales value. It is possible to improve one stage while leaving the other as the bottleneck.
Training pairs observed outcomes with eligible, unobserved alternatives; lack of an interaction does not automatically imply dislike. Validation uses the naturally retrieved candidate union. A positive item the retriever missed stays missing. Inserting it for ranker evaluation would hide a retrieval failure and overstate the quality of the complete system.
Value needs a similar distinction. Predicting the value of a purchase, conditional on purchase, does not establish that the buyer will purchase. A ranker can otherwise learn to recommend expensive items that rarely convert. V3 learns value-aware ordering alongside behavioral objectives and checks the complete served score against its baseline. Multi-objective ranking research provides useful context for this tension; the practical question here is which ordering survives held-out and online evaluation.
Finally, individually attractive items can make a repetitive list. List construction enforces eligibility and uniqueness within known product families, manages category and event concentration, and accounts for expiring inventory. Supply matters: a narrow location-specific list and a broad discovery list need different handling when eligible inventory is limited.
Reject a weak model without freezing the catalog
Model freshness and inventory freshness are different requirements. If a new model fails validation, keeping the last accepted model should not leave the service recommending an old catalog.
V3 separates model admission from inventory refresh. A candidate must clear temporal-integrity, ranking-quality, cohort, value, and list-quality checks before publication. Inventory-dependent recommendations can then be refreshed using the previously accepted model. The online service receives a complete validated snapshot rather than an intermediate training output.
Production records show this boundary working: some retraining attempts were rejected by quality or data-contract checks, while subsequent inventory refreshes retained accepted model identities. A failed retrain therefore did not automatically become a served regression.
Those checks compare the candidate with a same-run auction baseline. They do not establish that every accepted model beats its predecessor online, or that an offline value gain will translate into the same sales lift. Their purpose is to exclude unacceptable releases before customer exposure. FORGE supplies the orchestration and deployment foundations; the RecSys workload defines what a valid recommendation release means.
What the production experiment showed
The production experiment ran from August 24 through September 22, 2026. Buyers were randomly assigned to v3 or the handpicked experience. The results below report relative changes within that experiment, rather than comparisons with an earlier version's test window.
The recommendation-open rate moved from 27.41% to 28.34%: about 0.93 percentage points, or 3.39% relative. The unit is a distinct buyer with at least one recorded recommendation open. Repeated opens by the same buyer do not turn this into a total-click metric.
Sales per exposed buyer had a +2.71% point estimate, with an approximate interval of −2.97% to +8.39%. More measurement is needed to distinguish a useful gain from a small loss. Bid and checkout estimates were slightly negative and also uncertain; returning-use estimates were near zero. These are separate buyer-level outcomes rather than successive steps in an enforced conversion funnel.
The experiment evaluates the combined treatment. It cannot assign the opening gain specifically to bid history, the ranker, or cold-start handling. The earlier sales-based experiment used different control lists and inventory, so its effect sizes cannot establish a v3-versus-v1 improvement.
The cohort result changed the next engineering question
An aggregate improvement can hide a weaker experience for a substantial group. In exploratory cuts defined by purchase history before measurement, recommendation opening improved most among buyers with richer histories. Buyers without a usable purchase-history profile showed a decline.
The last distinction matters. The reporting cohort describes the availability of a completed-purchase profile. The production model can also use bids. It would be incorrect to call every buyer in that group a zero-history user, or to attribute the entire decline to the fallback path. A tighter diagnosis needs the actual served route as well as the reporting cohort.
One concrete investigation concerns product-family identity. Different auction lots can represent essentially the same product. A family assignment based on normalized titles can miss equivalent products described differently. The known-family constraint can pass while a buyer still sees repetition.
That is a plausible mechanism supported by the implementation and the investigation recorded with the experiment, not an explanation established by the A/B test. The next useful comparison combines the buyer's actual serving route with the repetition in their exposed list. If repetition is concentrated in the weaker cohort, improve family resolution and isolate that change in a follow-up test. Structured product attributes provide a first option, with semantic matching for ambiguous cases. The family index can be computed per inventory refresh, outside the request path.
What carries into the next iteration
The durable design decisions are the boundaries. Purchases and bids remain distinct evidence. History controls the room for personalization. Retrieval misses remain visible in evaluation. A recommendation list must satisfy constraints beyond individual item scores. Inventory refresh can continue when a candidate model fails admission.
The experiment gives the next release a concrete agenda: preserve the measured opening gain, narrow the sales uncertainty, and investigate the weaker purchase-profile cohort using actual exposure and routing evidence. That makes the next model change testable, with a specific failure mode and an outcome that can confirm or reject the proposed fix.
For the complementary system-design details, the original RecSys article covers immutable serving state, capacity planning, and cost. The FORGE case study explains the shared delivery platform. Together, they describe the decision model, its runtime, and the infrastructure that lets the team iterate on both.