Skip to content

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

Conference: NeurIPS2026
arXiv: 2609.35695
Code: https://github.com/Martin-qyma/Rethinking-Personalized-Generation
Area: Recommender Systems / Personalized Text Generation
Keywords: test-time alignment, candidate ranking, hidden-state recycling, personalized generation, ranking-guided decoding

TL;DR

The paper reframes personalized generation as selecting suitable candidates already produced by a generator, using an MLP with 2.8M parameters by default to rank frozen generator representations; it beats four generalist reward models across nine datasets but remains behind a dataset-finetuned 8B reward model, while single-trajectory decoding guidance recovers only part of the selection gain.

Background & Motivation

Generalist reward models typically judge helpfulness, correctness, and fluency, whereas personalized writing also requires a response to resemble what a particular user would write. Two candidates can both be linguistically acceptable yet differ in wording, emphasis, and consistency with the user's writing history. Common approaches fine-tune the generator for individual users or learn personalized reward models from their pairwise preference annotations. The former updates a large model; the latter adds another text encoder that scores each candidate. The paper first asks whether the generator lacks suitable responses or whether those responses already exist but are not selected.

The authors sample multiple candidates for the same prompt and let the evaluation metric select the highest-scoring response. As the pool grows from 1 to 64, the ROUGE-L oracle upper bound keeps increasing, whereas the best of four generalist reward models mostly stalls. This demonstrates selection headroom under reference-text similarity on these benchmarks, not an oracle measurement of real users' complete subjective preferences. Larger pools do not automatically improve output: a misdirected selector has more opportunities to pick a generally appealing response that is poorly matched to the target user.

The challenge is therefore to learn these fine distinctions without repeatedly paying for large-model text encoding. The generator already provides semantic representations of the query, user profile, and answer, so a separate 8B model need not understand the same text again. A small network can instead learn how closely the represented response matches the reference. Core idea: freeze language understanding and generation, train a shared lightweight ranker on generator hidden states using fine-grained supervision from the generator's own candidate pools, and reuse the ranking signal for test-time selection or local decoding guidance.

Method

Overall Architecture

The inputs are a task query, user history or summary, and sampled candidate responses; the output is either the candidate with the highest predicted utility or one response produced with ranking guidance. Frozen Qwen2.5-7B-Instruct supplies final-layer last-token representations for three texts, and a small MLP concatenates them to predict a scalar supervised by candidate ROUGE-L against the reference. Here, factorized primarily means decoupling language understanding from scoring, rather than explicitly learning low-rank parameters for each user.

Three connected designs define the pipeline: hidden-state recycling reduces redundant text understanding; candidate-pool supervision adapts scoring to the similar candidates encountered at deployment; test-time selection and guidance consume complete candidates or the current decoding state, respectively. Dashed edges indicate training supervision, and solid edges indicate deployment data flow. The output branches are alternative modes, not Best-of-N followed by guided decoding.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Query and user profile"] --> B["Frozen generator"]
    B --> C["Hidden-state recycling"]
    T["Training candidates and reference<br/>ROUGE-L labels"] -.-> D["Candidate-pool supervision"]
    C -.-> D
    D -.->|Train shared MLP| E["Test-time selection and guidance"]
    C -->|Inference representations| E
    E -->|Complete candidate pool| F["Select one with Best-of-N"]
    E -->|High-entropy single-trajectory steps| G["Adjust logits and continue"]

Key Designs

1. Hidden-state recycling: leave text understanding to the generator and matching to a small network

A conventional reward model rereads the prompt and answer for each candidate, duplicating expensive language understanding. The proposed ranker takes final-layer last-token hidden states of the task query, serialized user profile, and candidate answer. Each has dimension 3584; their concatenation enters an MLP with GELU and dropout. Query and profile representations are shared within a candidate pool and can be computed once per query, while each answer supplies its own representation. The MLP thus learns utility in a fixed feature space rather than learning linguistic and contextual semantics from scratch.

The same frozen model represents the user profile, with no separate user encoder or per-user trainable vector. LaMP uses three BM25-retrieved history items; LaMP-QA uses six historical questions; XRec uses an existing user summary and includes an item summary in the prompt. A shared ranker handles users through profile content rather than optimizing parameters for each new user. Requiring no user preference annotations does not mean requiring neither user history nor training reference texts.

The architectural proposal must be separated from the implementation. Appendix C.2 explicitly states that the current code obtains representations with vLLM last-token pooling in a pass separate from generation. A serving framework that exposes generation hidden states could eliminate that pass. The lowest scoring latency therefore describes an MLP whose representations are already available, not the current complete selection pipeline. The cached paper does not provide implementation detail sufficient to establish exact equivalence of the three representations across different text serialization schemes without an extra pass.

2. Candidate-pool supervision: learn distinctions among plausible answers from the same generation distribution

Randomly incorrect answers could teach the ranker fluency rather than which answer resembles this user's writing. The authors instead sample candidates from the frozen generator under real prompts. These candidates are often readable but differ in wording, content, and stylistic overlap with the reference. Such on-policy candidates provide hard negatives closer to the test-time selection problem. Low ROUGE-L does not establish low subjective preference, nor does the negative-construction process guarantee factual correctness.

Each candidate initially receives its ROUGE-L against the reference, then targets are standardized to zero mean and unit variance within the same prompt's pool. The ranker can focus on within-pool distinctions instead of predicting how easily a prompt attains a high score. Below, \(z_i\) is the standardized target, while \(\bar s\) and \(\sigma_s\) are the mean and standard deviation of raw pool scores. This notation combines the paper's standardization and MSE descriptions; treatment of zero-variance pools is unspecified.

\[ z_i=\frac{s_i-\bar s}{\sigma_s},\qquad \mathcal L_{\mathrm{MSE}}=\mathbb E_i\left[(f_\phi(h_x,h_u,h_{y_i})-z_i)^2\right]. \]

MSE preserves numerical target values and processes candidates individually, avoiding construction of all pairwise comparisons. It is not a new loss, however, and Bradley–Terry, RankNet, and listwise objectives perform similarly in the appendix. The gains cannot all be attributed to pointwise regression. Pool standardization also means that scores do not recover an absolute cross-prompt scale of raw ROUGE-L; the paper's language about comparable absolute utilities should be restricted to the training definition and actual within-pool selection use.

3. Test-time selection and guidance: complete-candidate ranking and local single-trajectory correction have different limits

Best-of-N fully generates \(N\) answers, predicts their utilities, and selects the maximum. Increasing \(N\) expands the available selection space but still incurs generation costs for all answers. The small MLP only makes additional scoring inexpensive. The paper summarizes selection complexity as decreasing from a reward model's \(\mathcal O(NL d_r^2)\) to a shallow network's \(\mathcal O(Nd^2)\), where \(L\) is sequence length, \(d_r\) is reward-model width, and \(d\) is generator width. This is a stage-level approximation, not an exact account of full Transformer computation, actual MLP layer widths, or total generation cost.

Ranking-guided generation constructs no candidate pool and continues along one decoding trajectory. It first measures next-token entropy, \(H_t=-\sum_v p_t(v)\log p_t(v)\). Steps at or below threshold \(\tau\) retain the original greedy choice. Only uncertain steps compute the gradient of the ranker with respect to the current answer hidden state and project it through the language-model output head into vocabulary space, adding it to the original logits:

\[ g_t=\nabla_{h_t}f_\phi(h_x,h_u,h_t),\qquad \ell'_t=\ell_t+\alpha_t Wg_t. \]

Here \(W\) is the language-model output head, with vocabulary-size by hidden-dimension shape. Neither MLP nor generator weights are updated. This is not back-propagation through the full generator for fine-tuning, nor a separate reward-model call for every possible continuation. Its key approximation is to predict a reward-increasing direction in the current representation and turn that direction into token biases through the output head. The paper does not guarantee that complete-answer supervision applies to every prefix.

Intervention strength increases with entropy beyond the threshold. Decoding selects the maximum of the adjusted logits and advances to the next step:

\[ \alpha_t=\alpha_0\,\operatorname{clip}\!\left(\frac{H_t-\tau}{H_{\max}-\tau},0,1\right). \]

Shared defaults are \(\tau=1\), \(H_{\max}=8\), and \(\alpha_0=5\), with no intervention during the first three tokens. \(H_{\max}=8\) is a scaling setting, not necessarily the vocabulary's theoretical maximum entropy. This single-trajectory method also differs from temperature-1.0 candidate sampling: the sensitivity analysis uses greedy decoding and at most 48 new tokens. Its results cannot be interpreted as obtaining Best-of-64 for free under an identical generation distribution.

A Worked Example

Consider one XRec user–item interaction. The prompt already contains user and item summaries, and the frozen generator can produce 64 natural-language explanations. Query and user-summary representations are shared, while each explanation supplies an answer representation. Only training uses the reference explanation to compute candidate ROUGE-L. At test time, the MLP predicts scores and chooses an answer without reading the reference. The user need not additionally annotate that explanation A is better than explanation B.

In ranking-guided generation, the system instead starts from an empty answer. After skipping the first three tokens, low-entropy steps proceed normally and high-entropy steps adjust vocabulary logits using the gradient above. It never first generates 64 explanations and cannot later recover missing content from other trajectories. This example illustrates the reported procedure; it is not an additional experiment or a fabricated numerical result.

Loss & Training

The default MLP has three layers, width 256, 2.8M parameters, and dropout 0.1. Training uses AdamW with learning rate \(5\times10^{-4}\), weight decay 0.01, batch size 64, 15 epochs, gradient clipping at 1.0, and one random seed. Only the ranker is trained; Qwen2.5-7B-Instruct remains frozen. Actual model sizes in the scaling experiment are 1.0M, 2.8M, 12.1M, and 30.4M, abbreviated as 1M, 3M, 10M, and 30M in the figures.

Candidate counts differ in the source descriptions. Section 4.1 says each training prompt contributes 64 candidates, whereas Appendix C.2 says 256 samples are stored per prompt, Best-of-N uses the first \(N\), and the appendix uses \(N=64\). These might refer to training usage versus generated storage, but the cache does not explicitly establish the complete correspondence. They should not be silently rewritten as one uniform protocol.

The finetuned 8B comparator uses the same pointwise objective and candidate source, but not an identical training budget. It takes eight candidates per prompt: the best, the worst, and six random samples. It uses 1,500 training prompts, or all 690–966 available prompts for LaMP-QA. Training uses learning rate \(2\times10^{-5}\), generally one epoch and two for LaMP-QA, 2,048-token truncation, one seed, and full fine-tuning on two B200 GPUs. The shorthand “same data and objective” does not replace these budget details.

Key Experimental Results

Main Results

Evaluation covers News, Scholarly, and Tweet from LaMP; Art, Lifestyle, and Society from LaMP-QA; and Amazon, Yelp, and Google from XRec. LaMP has 1,889, 2,498, and 1,495 evaluation prompts, respectively; LaMP-QA has 77, 99, and 108; each XRec domain has 3,000. LaMP-QA originally provides rubrics rather than answers. This work uses matching reference answers from Personalized RewardBench and splits questions 90/10 for ranker training and evaluation. These answers should not all be described as texts personally authored by the asking user.

The following held-out BLEU scores come from Appendix Table 4. All selectors remain fixed; the proposed ranker is trained only on ROUGE-L, and the Best-of-64 selected answer is evaluated with BLEU. The table supports cross-lexical-metric advantages among these selectors, not a claim that the ranker always finds the pool's true BLEU-maximizing answer.

Dataset Ours BLEU Skywork InternLM2 URM ArmoRM
News 0.0387 0.0330 0.0297 0.0326 0.0316
Scholarly 0.0671 0.0528 0.0503 0.0593 0.0544
Tweet 0.1150 0.1050 0.0990 0.1105 0.1083
Art 0.0259 0.0202 0.0205 0.0214 0.0225
Lifestyle 0.0217 0.0166 0.0148 0.0168 0.0170
Society 0.0277 0.0248 0.0235 0.0251 0.0247
Amazon 0.0822 0.0498 0.0446 0.0500 0.0370
Yelp 0.0417 0.0347 0.0332 0.0347 0.0347
Google 0.0337 0.0274 0.0243 0.0272 0.0256

In the main ROUGE-L comparison, the default ranker beats all four generalist models at Best-of-64 on all nine datasets, by an absolute 0.010–0.048 over the strongest generalist model. However, finetuned Skywork 8B beats the 30M ranker in every domain: gaps are 0.012 / 0.019 / 0.026 for News / Scholarly / Tweet, 0.006 / 0.004 / 0.011 for Art / Lifestyle / Society, and 0.004 / 0.006 / 0.008 for Amazon / Yelp / Google. The main text summarizes long-form QA gaps as below 0.01, which conflicts with Society's 0.011. The appendix's explicit value is retained here.

Ablation Study

Appendix Table 3 fixes the 2.8M architecture, data, and training budget while changing the loss. The table below retains MSE, one pairwise objective, and one listwise objective. Scores are Best-of-64 ROUGE-L, not the BLEU scores above.

Dataset MSE (default) Bradley–Terry ListMLE
News 0.1591 0.1545 0.1551
Scholarly 0.3962 0.3980 0.3967
Tweet 0.4120 0.4122 0.4135
Art 0.1547 0.1558 0.1557
Lifestyle 0.1488 0.1492 0.1492
Society 0.1495 0.1512 0.1527
Amazon 0.2650 0.2664 0.2670
Yelp 0.1920 0.1928 0.1930
Google 0.1686 0.1694 0.1691

Across all six objectives, the largest spread is at most 0.012 and below 0.003 in each XRec domain; MSE is not universally best. Increasing capacity from 1M to 30M changes scores by at most 0.006 per domain without consistent improvements. These observations support the importance of features and supervision over scoring-head capacity, but the paper does not separately provide enough training-candidate-count ablations to quantify the contribution of data scaling.

Appendix Table 8 measures costs for 64 candidates on one RTX A6000. These stages must not be treated as interchangeable end-to-end timings.

Stage Per candidate Per query (64 candidates) Peak GPU memory
Short-form candidate generation 10.2 ms 0.65 s Approximately 40 GiB
Separate embedding pass 2.08 ms 0.13 s 25–39 GiB
MLP scoring with available embeddings 0.0015 ms 0.098 ms 0.93 GiB
Average scoring by four large reward models 37.1 ms 2.38 s 20.6 GiB

Only MLP scoring excluding embedding extraction achieves the four-orders-of-magnitude latency advantage. Including the current separate pass, selection costs about 2.1 ms per candidate, with an approximately 18-fold advantage reported by the authors. Long-form QA pool generation costs about 11 s per query and remains substantial. Generation and embedding extraction use vLLM, whereas reward models use HF transformers at batch size 32; timing differences therefore also include serving implementation differences.

Key Findings

  • ROUGE-L and BLEU select the same pool maximum in only 12%–37% of pools, making cross-metric improvement informative. Both remain lexical metrics and cannot substitute for user evaluation.
  • Single-trajectory guidance does not consistently beat sampled selection: Tweet reaches 0.426 versus Best-of-64's 0.412, while News reaches 0.136, below Best-of-1's 0.144. Most datasets fall within the Best-of-1 to Best-of-4 range.
  • The guidance hyperparameter grid changes ROUGE-L by at most 0.0052 per domain. The default gate fires on 4%–39% of decoding steps; this is not a random-seed confidence interval.
  • XRec's seen-user comparison retrains the ranker: Amazon / Yelp / Google scores are 0.2641 / 0.1906 / 0.1685, versus GPO's 0.2690 / 0.1957 / 0.1743. The ranker is close to VPL but does not beat every personalized method, and this protocol cannot be mixed directly with Table 3.

Highlights & Insights

  • An oracle first tests whether the candidate pool has selection headroom before effort goes into a lightweight selector. This diagnostic can transfer to other generation tasks with user-specific reference texts rather than immediately attributing every limitation to generator capacity.
  • Representation recycling turns existing language understanding into shared infrastructure. Actual engineering gains depend on whether the serving stack exposes suitable states, not only on MLP parameter count.
  • In the appendix, swapping user conditions for another user's conditions changes ROUGE-L by at most 0.0008 for trained methods. Candidates may already carry substantial personalization from generation; this does not establish a strong independent preference mechanism in the ranker's explicit user branch.

Limitations & Future Work

  • Evaluation relies on lexical reference similarity, without human experiments on subjective preferences or style consistency. LaMP-QA reference provenance is also more complex than the main-text summary. Real pairwise choices, semantic correctness, and style assessment would strengthen validation.
  • Both rankers and finetuned reward models use one seed. Figure ranges span the minimum and maximum over three datasets, not variability across repeated training runs. Small decimal differences should not be interpreted as statistically significant.
  • Main experiments use one generator backbone. Frozen-representation quality, sparse history, transfer across generators, and shifts to new user populations require further study. History buckets are observational comparisons, not controlled causal ablations of history length.
  • Complete-answer training and prefix guidance have a distribution mismatch, and the greedy, 48-token setting limits guidance evidence. Longer outputs, additional decoding settings, and prefix-calibration experiments are needed.
  • Two source discrepancies are retained: 64 training candidates in the main text versus 256 stored samples in the appendix; a below-0.01 long-form QA fine-tuning gap in the main text versus 0.011 for Society in the appendix. Current separate embedding extraction and ideal direct recycling should also be reported separately.
  • vs generalist reward models: Skywork, URM, ArmoRM, and InternLM2 see the same profile-bearing prompts, but generic training is not calibrated to reference-text matching. The work changes supervision and scoring architecture; it does not show that large reward models cannot personalize.
  • vs PAL / VPL / PReF / LoRe: These methods fit or infer user parameters from target-user pairwise labels, while the proposed ranker receives a content profile only. The appendix instantiates some comparators on frozen representations at a matched budget of about 2.8M; conclusions are limited to that setting.
  • vs GPO / SynthesizeMe: GPO uses labeled user context, while SynthesizeMe runs a persona-prompted judge to eliminate candidates. GPO's higher raw scores show that additional supervision may still help; lightweight scoring is not a lossless replacement.
  • vs FUDGE and related decoding guidance: The proposed differentiable score reads generator states instead of using another network to re-encode each prefix. Better prefix-specific training objectives are a research direction, not a validated contribution of this paper.

Rating

  • Novelty: 4/5. The personalized-matching diagnosis and representation-recycling combination are clear; MSE itself is not novel.
  • Experimental Thoroughness: 3/5. Nine datasets and diverse comparators provide coverage, but one seed, lexical evaluation, and implementation scope limit conclusions.
  • Writing Quality: 3/5. The core question is clear, but candidate counts, reference provenance, and efficiency descriptions require tighter consistency.
  • Value: 4/5. Relevant to generation services with user history and accessible hidden states; deployment gains must include generation and embedding-extraction costs.