Skip to content

Goal-Conditioned Supervised Learning for Multi-Objective Recommendation

Conference: NeurIPS2026
arXiv: 2412.08911
Code: https://github.com/xiwenchao/MOGCSL
Area: Recommender Systems / Multi-Objective Learning
Keywords: goal-conditioned supervised learning, multi-objective recommendation, cumulative reward, conditional variational autoencoder, low-goal noise

TL;DR

MOGCSL preserves multiple future session rewards as a goal vector, trains a goal-conditioned next-item predictor with ordinary cross-entropy, and selects inference goals using statistics or two CVAEs, improving purchase prediction and training cost without dominating every click metric or guaranteeing that specified goals are attainable.

Background & Motivation

E-commerce recommendation must predict not only clicks but also behaviors with different values, such as purchases. Shared-Bottom and MMOE construct multi-task models using shared representations, task towers, or expert gates, while DWA, PE, Nash-MTL, and FAMO adjust objective weights or update directions. These approaches primarily address representation sharing and optimization, but generally do not condition on what happened after an interaction to distinguish training examples. A click may express genuine interest or merely an incidental choice from an unsuitable list; learning all next-item labels directly mixes these behaviors.

Goal-conditioned supervised learning offers another perspective: each offline trajectory demonstrates some outcome, allowing the model to learn which actions accompanied that outcome in a given historical state. However, combining cumulative click and purchase rewards into a weighted scalar loses the particular outcome combination. The same scalar may describe high clicks with low purchases or low clicks with high purchases; once weights are fixed, this mixing occurs before training. Preserving the full goal vector avoids that step but introduces an inference problem: future outcomes have not occurred and cannot be calculated as during training, nor is increasing every coordinate necessarily appropriate.

The paper therefore separates training from decision preferences: all observed outcomes serve as training conditions without additional multi-task loss weights, while inference selects goals aligned with application requirements. Its noise interpretation assumes that better long-term outcomes generally indicate more reliable preference signals, rather than establishing a universal claim about user behavior. Core Idea: condition offline behavior on multidimensional future outcomes so that one supervised predictor learns outcome-specific patterns, then select data-supported goals aligned with utility preferences at inference time.

Method

Overall Architecture

Inputs are the user's recent item-interaction history, the current session timestep, and a vector of objectives such as clicks and purchases; the output is a next-item distribution over the full candidate catalog. Training applies “Multidimensional Goal Relabeling” before fitting the “Goal-Conditioned Predictor”; inference can use timestep-based statistical goals or select personalized goals through “Dual-CVAE Goal Modeling” and “Utility-Based Goal Selection.”

The two CVAEs do not generate recommended items: one proposes candidate input goals, and the other estimates each candidate's achieved-outcome distribution. The latter's predicted mean is used only to evaluate candidates; the selected original input goal is what enters the trained predictor. Dashed arrows denote offline supervision or model conditioning, while solid arrows labeled as inference describe the goal-selection data flow.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Offline session logs"] --> B["Multidimensional<br/>Goal Relabeling"]
    B -->|Training: history, timestep, goal and next item| C["Goal-Conditioned<br/>Predictor"]
    B -.->|Training: logged state-goal pairs| D["Dual-CVAE<br/>Goal Modeling"]
    C -.->|Fixed policy as a condition| D
    S["New state"] -->|Inference: propose candidates and predict outcome means| D
    D -->|Inference: candidate goals and corresponding means| E["Utility-Based<br/>Goal Selection"]
    E -->|Original candidate goal enters the trained predictor| F["Next-item ranking"]

Key Designs

1. Multidimensional Goal Relabeling: preserve click and purchase outcomes before training

In a complete session, clicks and purchases form separate coordinates of the immediate reward vector. Once the session ends, rewards are accumulated backward from every position, replacing its immediate reward with a goal vector to construct state, next-item, and future-goal supervision. The goal includes the current position and subsequent outcomes, making it hindsight information; the history encoder still receives only preceding items, not future interactions that would be unavailable online.

\[ \boldsymbol{g}_{t}=\sum_{t^{\prime}=t}^{|\tau|}\boldsymbol{r}_{t^{\prime}}. \]

A high-click, low-purchase session and a lower-click, higher-purchase session therefore need not become the same condition simply because their weighted totals are similar. Compared with scalar-goal MOPRL, the model preserves relationships between each outcome dimension and the next behavior without first specifying reward weights. The goal vector is not a collection of task-specific output labels: the supervised label remains the next item, and the outcome vector conditions its prediction.

All examples still participate in ordinary supervised training; low-goal sessions are neither explicitly removed nor assigned smaller loss weights. Denoising is an indirect effect of conditioning: if noise is concentrated in low-goal sessions, the model can distinguish their behavioral patterns from high-goal patterns, and querying a higher goal at inference can favor the latter. “Automatically excluding noise” should thus mean reducing its influence on a specified conditional prediction, not identifying genuine instance-level noise labels.

2. Goal-Conditioned Predictor: learn outcome-conditioned behavior with one item-classification head

The model passes embedded historical items through a Transformer encoder to obtain a sequence representation, encodes the timestep through an embedding table, and maps the goal vector through a fully connected layer. These representations are concatenated and processed by self-attention to combine history, session position, and desired outcomes; an MLP and softmax then produce item probabilities. The appendix implementation retains the most recent 10 interactions and pads shorter histories; state and goal embedding dimensions are both 64.

The learning problem is to predict logged behavior associated with an outcome condition, rather than train an independent task tower for each objective. The training objective is simply next-item negative log-likelihood:

\[ \mathcal{L}(\theta)=\mathbb{E}_{(\boldsymbol{s}_{t},a_{t},\boldsymbol{g}_{t})\in\mathcal{D}_{tr}}[-\log\pi_{\theta}(a_{t}\mid\boldsymbol{s}_{t},\boldsymbol{g}_{t})]. \]

This explains the smaller model and simpler optimization, while clarifying the method's scope. Although the authors formulate interactions as a multi-objective MDP, the predictor has no Bellman target, TD error, or value-function optimization. Cumulative rewards provide additional supervision conditions, and the primary evaluation is next-item ranking rather than a proof of maximized long-term session returns.

3. Dual-CVAE Goal Modeling: separate proposed goals from their predicted outcomes

MOGCSL-S is the lower-cost default: it takes the training logs' mean remaining cumulative reward vector at the current timestep and multiplies it by a validation-selected scaling factor. It varies with the timestep but does not generate an individual distribution for every new state. MOGCSL-C instead uses CVAE2 to learn a state-conditioned goal prior and sample potential input goals, then uses CVAE1 to model achieved outcomes conditioned on the state, input goal, and fixed policy, estimating a mean through repeated sampling for each candidate. The experimental candidate count is 20; the cached text does not separately specify the inner sampling count for the achieved-outcome mean.

Both CVAEs are trained on the same offline state and hindsight-goal pairs. CVAE2 reconstructs logged goals while regularizing the latent variable; CVAE1 likewise uses the logged goal as the achieved-outcome target, additionally conditioning on that same goal and the policy. The paper motivates this by arguing that the policy imitates the demonstration, so its outcome can be treated as a policy-conditioned sample. This explanation must distinguish the outcome actually achieved by the logged trajectory from the outcome that a finite-data learned policy will achieve: the latter is not guaranteed, nor is the supervision obtained directly from fresh trajectories of the deployed policy.

Theorem 1 establishes only that achieved-outcome distributions depend on the initial state, input goal, and policy. Appendix A fixes the environment's transition and reward distributions, subtracts each observed reward from the remaining goal, factorizes trajectory probability into policy, reward, and transition terms, and treats cumulative reward as a deterministic trajectory function. Random session lengths are handled with an absorbing terminal state or integration over lengths. The theorem therefore motivates conditioning variables but does not prove correct CVAE estimation, goal fulfillment by offline conditional imitation, or membership in the true feasible region.

4. Utility-Based Goal Selection: evaluate predicted outcomes but execute the original input condition

After associating each candidate input goal with its predicted achieved-outcome mean, the algorithm selects the preferred outcome under a utility principle and retrieves its corresponding input goal. These vectors are not interchangeable: feeding the predicted outcome mean directly into the policy changes Algorithm 2. The following is a shorthand for candidate selection, with utility determining the selected index:

\[ b\in\arg\max_{k}U(\tilde{\boldsymbol{g}}^{a}_{k}),\qquad \text{output}=\pi(\cdot\mid\boldsymbol{s}^{\prime},\boldsymbol{g}^{\prime}_{b}). \]

The experiments use a nondominance condition among predicted candidate means: no other candidate is at least as high on every coordinate. A footnote in Appendix C.5 adds that when several candidates qualify, the one with the highest purchase goal is selected. This filters a finite sampled candidate set, rather than recovering the environment's complete Pareto frontier. The original Eq. (8) specifies coordinatewise nondecrease without an explicit strict improvement in at least one coordinate, leaving identical vectors and tie handling incompletely specified.

Training avoids preset multi-task loss weights, but preferences do not disappear: they move into the inference utility and goal selection. Applications can also use weighted utility at this stage; this differs from collapsing rewards into a training scalar, yet still prioritizes objectives. For continued session execution, Appendix C.1 says CVAE goal selection occurs only at the sequence start, after which obtained rewards are subtracted incrementally. MOGCSL-S's timestep-based statistical goals should be distinguished from that update procedure.

A Worked Example

This example illustrates Algorithm 2 and is not reported experimental data. Suppose the click and purchase coordinates of two candidate input goals are \((4,1)\) and \((2,2)\), while CVAE1 predicts achieved-outcome means of \((3,0.5)\) and \((2,1.2)\), respectively.

Neither predicted mean dominates the other: the former has higher clicks and the latter higher purchases. If utility prioritizes purchases, the second candidate is selected, and the next-item predictor receives \((2,2)\), not \((2,1.2)\); the latter vector is only an outcome estimate used for selection. The predictor then ranks items using the current history and selected goal, with the remaining goal updated only after actual feedback arrives.

The example also separates requesting more purchases from being able to achieve them. Candidate generation, outcome prediction, and policy generalization all involve error, so selecting an apparently better condition does not guarantee actual returns.

Loss & Training

After policy training, both CVAEs optimize reconstruction and latent-variable KL regularization; CVAE1's condition is denoted by \(\boldsymbol{c}=(\boldsymbol{s},\boldsymbol{g},\pi)\). The following retains the mechanisms in the original Eq. (6) and Eq. (7), with encoders \(Q_{1}\) and \(Q_{2}\) and decoders \(P_{1}\) and \(P_{2}\):

\[ \mathcal{L}_{CVAE1}=\mathbb{E}_{(\boldsymbol{s},\boldsymbol{g})\in\mathcal{D}_{tr},z\sim Q_{1}}[-\log P_{1}(\boldsymbol{g}\mid z,\boldsymbol{c})+D_{KL}(Q_{1}(z\mid\boldsymbol{g},\boldsymbol{c})\Vert P(z))]. \]
\[ \mathcal{L}_{CVAE2}=\mathbb{E}_{(\boldsymbol{s},\boldsymbol{g})\in\mathcal{D}_{tr},z\sim Q_{2}}[-\log P_{2}(\boldsymbol{g}\mid z,\boldsymbol{s})+D_{KL}(Q_{2}(z\mid\boldsymbol{g},\boldsymbol{s})\Vert P(z))]. \]

The paper assumes Gaussian encoders and decoders, with latent variables sampled from a standard normal distribution. The dual-CVAE variant consequently adds training and sampling costs; the support of its Gaussian samples should not be equated with the environment's feasible goal set.

Implementation uses Adam and a batch size of 256; the learning rate is selected from 0.0001, 0.0005, 0.001, and 0.005, while manually weighted baselines tune weights on a validation grid from 0.1 to 0.9. Models share a Transformer encoder and self-attention backbone to reduce representation differences. All results report means and standard deviations across 5 random seeds, not confidence intervals or formal significance tests.

Key Experimental Results

Main Results

Challenge15 and RetailRocket provide click and purchase labels, retaining sessions of length 3 to 50 after filtering and using an 8:1:1 training, validation, and test split. Challenge15 contains 200,000 sessions and 26,702 items; RetailRocket contains 195,523 sessions and 70,852 items. The cache does not specify whether splitting is chronological or how cross-split user overlap is handled, so strict temporal extrapolation cannot be claimed.

HR measures whether the true next item enters the top-ranked list, while NG is the original tables' abbreviation for NDCG and also accounts for its position. The following excerpts Table 1, with all metrics in percent and cells giving mean ± standard deviation. Its MOGCSL uses statistical goal selection, matching MOGCSL-S in Table 3 rather than the CVAE variant.

Dataset Config Purchase HR@10 Purchase NG@10 Click HR@10 Click NG@10
RetailRocket MMOE-FAMO 47.93 ± 0.32 46.42 ± 0.23 35.92 ± 0.21 26.14 ± 0.17
RetailRocket PMORS 63.14 ± 0.15 51.02 ± 0.17 34.16 ± 0.26 24.09 ± 0.23
RetailRocket MOGCSL-S 65.43 ± 0.15 52.92 ± 0.11 36.30 ± 0.25 25.24 ± 0.15
Challenge15 MMOE-PE 36.40 ± 0.36 24.66 ± 0.19 44.04 ± 0.09 27.44 ± 0.03
Challenge15 PMORS 54.98 ± 0.31 34.77 ± 0.28 42.36 ± 0.39 25.10 ± 0.42
Challenge15 MOGCSL-S 56.82 ± 0.25 35.93 ± 0.15 42.47 ± 0.15 25.64 ± 0.11

Purchase improvements coexist with click trade-offs. RetailRocket click NG@10 is 25.24, below MMOE-FAMO's 26.14; on Challenge15, MMOE-PE also exceeds MOGCSL-S on click HR@10 and NG@10. Stronger performance on the higher-value purchase objective does not mean dominance across every objective and ranking metric.

Ablation Study

The following excerpts the goal-selection comparison in Table 3 using the same metrics as the main table; Tenrec's value objective is likes rather than purchases. This compares selection strategies, not conventional removal of model components.

Dataset Config Value-objective HR@10 Value-objective NG@10 Click HR@10 Click NG@10
RetailRocket MOGCSL-S 65.43 ± 0.15 52.92 ± 0.11 36.30 ± 0.25 25.24 ± 0.15
RetailRocket MOGCSL-C 65.01 ± 0.07 52.89 ± 0.04 36.54 ± 0.02 25.41 ± 0.04
Challenge15 MOGCSL-S 56.82 ± 0.25 35.93 ± 0.15 42.47 ± 0.15 25.64 ± 0.11
Challenge15 MOGCSL-C 55.13 ± 0.07 35.04 ± 0.02 42.14 ± 0.04 25.37 ± 0.07
Tenrec MOGCSL-S 5.96 ± 0.17 2.15 ± 0.12 4.87 ± 0.13 1.52 ± 0.08
Tenrec MOGCSL-C 6.78 ± 0.07 2.99 ± 0.02 5.66 ± 0.05 2.14 ± 0.05

CVAE selection is not uniformly better on RetailRocket: purchase HR@10 is 65.01 versus the statistical variant's 65.43, and purchase NG@10 is also slightly lower, while click and some longer-list metrics improve. On Challenge15, the CVAE variant is lower on every original-table metric; Tenrec provides the clear pattern of improvement across all reported metrics.

Appendix B.2 injects random next-item labels into low-goal RetailRocket sessions while retaining the original test set. The text describes selecting 20% of the original training sequences from those with the lowest 20% average goal and replacing each selected sequence's final true item with a uniformly sampled catalog item. This deliberately correlates noise with low goals rather than corrupting arbitrary positions and goals. The following excerpts Table 4; drop ratios are relative percentages against uncorrupted training results, not percentage-point differences.

Config Corrupted purchase HR@20 Relative drop ratio Corrupted click HR@20 Relative drop ratio
Share-FAMO 46.32 ± 0.11 7.71% 36.97 ± 0.24 9.19%
PMORS 64.17 ± 0.18 4.86% 36.05 ± 0.30 9.56%
MOPRL 62.37 ± 0.11 3.69% 36.61 ± 0.21 6.08%
MOGCSL 67.92 ± 0.17 1.96% 39.91 ± 0.16 4.79%

Key Findings

  • Inference goals should not be increased indefinitely. Statistical goal scaling performs best in the range from 1 to 2, with larger conditions reducing performance; the paper attributes this to inadequate behavioral demonstrations in high-goal regions.
  • Dual-CVAE gains relate to coverage. Mean cumulative clicks and purchases are approximately 5.3 and 0.2 in the first two datasets, versus clicks and likes of approximately 28.3 and 1.2 in Tenrec. This supports a coverage explanation but is not a causal experiment varying coverage alone.
  • Costs must be distinguished by variant. For RetailRocket, Table 2 reports 9.1M parameters and 3.0Ks training time for the statistical variant, while Table 6 reports 10.7M and 4.8Ks for the CVAE variant; Share-FAMO uses 14.0M and 5.1Ks, MMOE-PE 14.1M and 60.5Ks, and RMTL 17.5M and 100.2Ks. Hardware is NVIDIA RTX 3090 and AMD 3960X. These are measured model sizes and times, not uniform asymptotic bounds, and online end-to-end latency is not directly reported.
  • Mean regression is also viable. Appendix C.6 reports RetailRocket purchase HR@20 of 69.37 ± 0.30 for MOGCSL-R and 69.34 ± 0.05 for the CVAE variant. The latter has lower cross-seed variation, but not every mean is higher, nor does sampling inherently establish a theoretical variance guarantee.

Highlights & Insights

  • Outcome conditioning preserves information while separating mixed supervision. Instead of adding a task tower, it allows different behavioral distributions at the same state under different long-term outcomes, particularly for complete session logs with multiple feedback types.
  • Preferences are applied at inference rather than fixed through reward scalarization. A change in application priorities can be expressed by another utility principle or goal without retraining for each reward weighting, although flexibility remains limited by observed goal coverage.
  • Separating candidate conditions from predicted outcomes is reusable. After proposing a desired outcome, the method evaluates the likely outcome before deciding which original condition to execute; state-conditioned proposals are more individualized than a single global high goal.

Limitations & Future Work

  • Achievability remains a model estimate, not a theorem-certified constraint. CVAE1 uses the same logged goal in both its condition and reconstruction target, without calibrating the learned policy's actual returns under new conditions; future work could add policy-execution trajectories, offline evaluation, and uncertainty constraints.
  • Denoising addresses only noise correlated with poor long-term outcomes. The synthetic experiment explicitly randomizes labels when any goal coordinate falls below a threshold, validating the constructed mechanism. High engagement need not indicate genuine satisfaction, and a short session can reflect successful task completion; engagement gains should not be equated with user welfare.
  • Main recommendation experiments remain dual-objective offline next-item evaluations. Although the synthetic noise dataset uses multidimensional goals, it cannot substitute for real high-dimensional recommendation, complete frontier recovery, or online long-term return validation; the authors explicitly report no live-user A/B testing.
  • Sparse high goals, negative feedback, and logging biases remain unresolved. The authors suggest additional high-goal trajectories, online iteration, and goal coordinates for quits, skips, and dislikes. Deployment also requires attention to exposure bias, privacy, fairness, and noncommercial dataset license restrictions.
  • The paper has wording and cross-reference inconsistencies. Section 4.2 describes Share-PE's NDCG advantage as purchase-related, although Table 1 places it on clicks; Section 4.4's statement that RetailRocket performs better does not cover short-list purchase metrics. Some relative gain percentages disagree with tables, and the appendix references nonexistent B.7 and B.8. This note follows the corresponding table values and actual sections rather than silently repairing the paper.
  • vs PRL / MOPRL: Both learn goal-conditioned next actions, but MOPRL accumulates weighted immediate rewards into a scalar. MOGCSL preserves a vector to avoid mixing outcome combinations before training.
  • vs MMOE / DWA / PE: These primarily address representation sharing or multi-task optimization, while MOGCSL supplies outcomes as input conditions. Lower costs and stronger purchase metrics are empirical findings in this setting, not evidence that gradient conflict is universally resolved in other tasks.
  • vs RMTL: RMTL adds a reinforcement learning head and TD optimization, whereas MOGCSL performs supervised item prediction. Using reward trajectories does not make their long-term return optimization objectives equivalent.
  • vs Decision Transformer / RVDT: The former offers a return-conditioned sequence-learning perspective, and the latter investigates return coverage in offline conditional sequence models. A transferable research direction is to select goals jointly using utility, data support, and predictive uncertainty rather than simply requesting larger conditions.

Rating

  • Novelty: 4/5 — Combining vector goals with inference-time selection is useful, but achievability modeling still relies on a strong offline-imitation interpretation.
  • Experimental Thoroughness: 3/5 — Multiple recommendation datasets, controlled noise, and cost analyses are included; real high-dimensional objectives and online long-term outcomes are absent.
  • Writing Quality: 3/5 — The main procedure and proof are understandable, but some percentages, metric descriptions, and appendix references require verification.
  • Value: 4/5 — A low-cost baseline for recommendation logs with multiple session feedback signals, with evidence clarifying when statistical and generative goals are useful.