Skip to content

Personalizing MLLMs via Reinforced Multimodal Reference Game

Conference: ECCV 2026
Paper: ECCV Official Page
Area: Multimodal VLM
Keywords: Multimodal Large Language Model / Personalization / Reinforcement Learning / Reference Game / GRPO

TL;DR

Reframing MLLM personalization into a collaborative speaker-listener multimodal reference game, this paper optimizes the speaker using GRPO with verifiable probability-normalized contrastive rewards derived from hard positives and hard negatives, producing state-invariant and discriminative concept descriptions that substantially advance personalization benchmarks.

Background & Motivation

Multimodal Large Language Models (MLLMs) demonstrate impressive generic vision-language reasoning, yet they lack prior knowledge regarding user-defined personal concepts such as a specific shirt, pet, or custom souvenir. To enable models to recognize private visual concepts and respond to user queries, early personalization paradigms borrowed inversion techniques from generative modeling (e.g., MyVLM and Yo'LLaVA) by optimizing custom tokens or concept vectors at test time. However, such test-time fine-tuning approaches are computationally prohibitive and require retraining every time a new concept is introduced, making them unsuitable for dynamic deployment. Recent retrieval-based methods (e.g., R2P) circumvent retraining by using the model's own generated textual fingerprints for cross-modal matching, but standard MLLM captioning is heavily biased by generic pretraining distributions.

Generic MLLMs suffer from three core limitations when describing specific personal concepts. First, descriptions frequently emphasize mutable physical states (such as "hanging on a wardrobe" or "folded on a table"), which mislead identification when the item's state changes or when distractors share that state. Second, they capture extraneous background and spatial context that vary across camera viewpoints. Most critically, descriptions generated independently for a single item lack comparative focus, failing to emphasize the distinctive features that separate the target from visually similar alternatives within the same semantic class (such as subtle floral patterns on a shirt). Meanwhile, naive reinforcement learning formulations (like RePIC) optimize the model directly to output the concept identifier, incentivizing short-cuts rather than faithful and robust conceptual representations.

The core insight of this paper is that if a textual description allows a listener to reliably identify a target object among subtle distractors and across extreme viewpoint variations, that description necessarily captures the target's invariant and discriminative essence. Core Idea: Reframe MLLM personalization as a multimodal reference game where the model alternates between speaker and listener roles, leveraging hard positives and hard negatives to compute a verifiable contrastive reward that trains the speaker via GRPO to produce fine-grained, state-invariant, and highly discriminative concept descriptions.

Method

Overall Architecture

Reinforced Reference Game (RRG) decouples personalized concept learning into an offline reinforced communication game and an online retrieval-augmented reasoning stage. During training, without requiring ground-truth human descriptions, the same backbone MLLM is instantiated as both a speaker \(\Phi_\text{speak}\) (equipped with trainable LoRA parameters \(\theta\)) and a listener \(\Phi_\text{list}\). Given an image of a target concept, the speaker generates a descriptive text; the listener, given only this description, must guess the target image from a candidate pool comprising another view of the target concept (a hard positive) and visually similar items from the same semantic category (hard negatives). The listener's guess outcome and confidence provide a verifiable contrastive reward to train the speaker via online GRPO. At inference time, the frozen speaker generates clean referential descriptions for user concepts to build a multimodal database, enabling reliable retrieval and personalized question answering.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Target Image & Candidate Set Input"] --> B["Dual-Role Multimodal Reference Game<br/>Speaker generates compact discriminative description"]
    C["Hard Positives & Negatives Construction<br/>Cross-view/state positive & same-class distractors"] --> B
    B --> D["Listener Binary Probing Match<br/>Evaluate candidate matching probabilities"]
    D --> E["Probability-Normalized Verifiable Reward<br/>Online GRPO updates speaker LoRA"]
    E --> F["Fingerprint Retrieval & Closed-Set Reasoning<br/>Multimodal DB search & downstream task answering"]

Key Designs

1. Dual-Role Multimodal Reference Game: Personalization as Collaborative Communication Prior approaches either rely on heavy test-time parameter adaptation or directly prompt frozen MLLMs to generate generic dense captions that lack discriminative power. RRG reformulates concept personalization as a collaborative reference game inspired by the classic game "Guess Who?". The speaker \(\Phi_\text{speak}\) observes an image \(v_i^\text{tgt}\) of target concept \(c_\text{tgt}\) and generates description \(d\); the listener \(\Phi_\text{list}\) cannot access \(v_i^\text{tgt}\) and must identify the target within a candidate set \(V\) using solely description \(d\). Because the speaker generates each description independently without needing access to all user concepts simultaneously, concept representations remain strictly modular and do not require re-computation when new items arrive. Crucially, the training game concepts \(C\) and the user personalization concepts \(P\) are completely disjoint (\(C \cap P = \emptyset\)), driving the speaker to learn a generic, transferable mechanism for articulating discriminative attributes.

2. Hard Positives and Hard Negatives Construction: Filtering State Fluff and Isolating Subtle Differences The quality of the learned description directly depends on the hardness of the reference game. If candidates consist of coarse, dissimilar categories, the speaker can succeed by outputting trivial category labels (e.g., "a shirt"), failing to learn fine-grained attributes; conversely, if candidates share the identical background or viewpoint, the speaker easily overfits to background noise. RRG constructs candidate set \(V = \{v_j^\text{tgt}\} \cup \bar{V}_\text{tgt}\) using a dual-hardness scheme. Hard positives \(v_j^\text{tgt}\) (\(i \neq j\)) capture the same concept under variations in viewpoint, lighting, or physical state (e.g., a garment hung vs. folded), forcing the model to discard transient environmental cues. Hard negatives \(\bar{V}_\text{tgt}\) comprise \(U\) distinct concepts from the exact same semantic category, compelling the speaker to bypass broad class names and pinpoint instance-specific details such as unique textures, cuts, or accents.

3. Probability-Normalized Verifiable Contrastive Reward: Driving Policy Optimization via GRPO Because current MLLMs exhibit instability and attention dispersion when consuming multi-image inputs simultaneously, RRG decomposes the listener's matching into independent binary probing evaluations. For each candidate \(v_u \in V\), the listener assesses the matching probability of the "Yes" token: $\(\rho_u = \text{Prob}_{\Phi_\text{list}}(\text{Yes} \mid v_u, d, \text{prompt}_m)\)$ The listener selects the candidate with maximum probability, \(\text{guess} = \arg\max_u \rho_u\). Rather than adopting a crude binary reward (1 for correct, 0 for incorrect) that discards confidence gradations, RRG designs a normalized posterior verifiable reward: $\(r_\text{match} = \begin{cases} \frac{\rho_\text{guess}}{\sum_u \rho_u}, & \text{if } \text{guess} = \text{tgt} \\ 0, & \text{otherwise} \end{cases}\)$ This formulation penalizes ambiguous descriptions where the listener's probability distribution is flat across distractors. Optimizing with Group Relative Policy Optimization (GRPO), the model samples \(S\) rollouts \(\{d_s\}_{s=1}^S\) per sample, computes relative advantages within the group, and regularizes divergence from the reference policy \(\pi_\text{ref}\) using a KL divergence penalty, enabling the speaker to discover optimal descriptive policies without manual text supervision.

4. Fingerprint Retrieval and Closed-Set Reasoning: Seamless Downstream Personalization At test time, the speaker pre-computes descriptions \(d_p = \Phi_\text{speak}(v_p)\) for each personalized concept \(p \in P\), populating a multimodal database \(\mathcal{D} = \{t_p, v_p, d_p, f_p^V, f_p^T\}_{p \in P}\), where \(f_p^V\) and \(f_p^T\) are CLIP visual and textual embeddings respectively. Given a query image \(Q\), the system retrieves the Top-\(K\) relevant candidates via joint image-to-image and image-to-text cosine similarities, converting open-ended personalization into a tractable closed-set problem. The candidate options \(O = \{(t_p, d_p)\}\) are formatted into the listener's prompt alongside task-specific instructions \(q\) to produce personalized captions or answers. If the user explicitly mentions the concept name \(t_p\) in the prompt, the system directly fetches \(d_p\) and bypasses retrieval entirely.

Loss & Training

The speaker LoRA parameters \(\theta\) are trained using the online GRPO objective over game concept set \(C\). For each prompt, \(S\) rollout descriptions are sampled from \(\Phi_\text{speak}\), generating contrastive rewards \(r_\text{match}\) that yield normalized advantages \(\hat{A}_t^s\). The objective optimizes clipped policy ratios alongside a KL divergence constraint: $\(\mathcal{R}(\theta) = \mathbb{E}\left[ \frac{1}{S}\sum_{s=1}^S \frac{1}{|d_s|}\sum_{t=1}^{|d_s|} \min\left(r_t^s(\theta)\hat{A}_t^s, \, \text{clip}(r_t^s(\theta), 1-\epsilon, 1+\epsilon)\hat{A}_t^s\right) - \beta D_\text{KL}(\pi_\theta \parallel \pi_\text{ref}) \right]\)$ where \(r_t^s(\theta)\) is the importance sampling probability ratio and \(\beta\) controls the regularization weight. The model trains on 30 concepts from PerVA and directly evaluates zero-shot generalization on the remaining 269 PerVA concepts as well as the external MyVLM and Yo'LLaVA benchmarks.

Key Experimental Results

Main Results

On MyVLM, Yo'LLaVA, and PerVA benchmarks, RRG is compared against test-time optimization methods (MyVLM, Yo'LLaVA), retrieval-augmented baselines (RAP-LLaVA, RAP-Qwen), reinforcement learning on concept identification (RePIC), and training-free fingerprinting (R2P-Qwen).

Dataset Method Recog. Pos. (↑) Recog. Neg. (↑) Recog. Wtd (↑) Caption Prec. (↑) Caption Rec. (↑) Caption F1 (↑)
MyVLM MyVLM 96.6 90.9 93.8 - - -
Yo'LLaVA 97.0 95.7 96.4 - - -
R2P-Qwen 86.9 95.4 91.5 94.4 94.1 93.4
RAP-Qwen 97.2 94.8 95.5 88.6 85.2 84.9
RePIC 98.1 95.6 96.8 84.2 73.2 76.7
RRG (Ours) 91.5 92.0 91.8 95.7 95.4 95.0
Yo'LLaVA Yo'LLaVA 94.9 89.8 92.4 - - -
R2P-Qwen 86.2 95.3 90.8 89.8 87.5 86.3
RAP-Qwen 97.2 87.5 92.3 76.3 70.7 69.4
RePIC 91.9 91.1 91.5 79.5 57.7 62.7
RRG (Ours) 97.0 92.3 94.7 91.9 89.3 88.6
PerVA MyVLM 66.0 58.5 62.2 - - -
Yo'LLaVA 75.1 69.0 72.0 - - -
R2P-Qwen 88.2 91.8 90.0 73.0 67.7 67.3
RAP-Qwen 93.2 92.9 93.1 66.0 50.4 52.8
RePIC 88.8 96.5 92.6 62.4 44.3 48.9
RRG (Ours) 94.0 88.3 91.1 89.5 87.9 86.9

In personalized VQA (Yo'LLaVA benchmark), RRG achieves an answering accuracy of 95.3%, surpassing R2P-Qwen (94.1%), large-scale pretrained RAP-LLaVA (93.2%), Yo'LLaVA (92.9%), and GPT-4V+Vprompt (86.6%).

Ablation Study

The authors evaluate alternative training paradigms and reward configurations on MyVLM and Yo'LLaVA datasets.

Table 1: Effect of description optimization and training targets (MyVLM & Yo'LLaVA) | Configuration | MyVLM Recog. Wtd | MyVLM Caption F1 | Yo'LLaVA Recog. Wtd | Yo'LLaVA Caption F1 | Note | |---|---|---|---|---|---| | zs-speaker | 89.3 | 89.3 | 89.3 | 84.6 | Descriptions from untuned zero-shot speaker | | listener-training | 89.7 | 92.0 | 90.4 | 86.3 | Freeze speaker, train listener via RL | | speaker-training | 92.2 | 92.3 | 93.9 | 87.0 | Train speaker with standalone MLLM verifier | | RRG (Full Model) | 91.8 | 95.0 | 94.7 | 88.6 | Multi-candidate reference game with normalized reward |

Table 2: Ablation on reward formulations | Reward Strategy | MyVLM Recog. Wtd | MyVLM Caption F1 | Yo'LLaVA Recog. Wtd | Yo'LLaVA Caption F1 | Note | |---|---|---|---|---|---| | binary reward | 92.3 | 93.4 | 93.9 | 85.5 | Discrete reward: 1 if correct guess, 0 otherwise | | \(r_\text{match}\) (Normalized Probability) | 91.8 | 95.0 | 94.7 | 88.6 | Soft confidence-weighted probability reward |

Key Findings

  • Speaker training outshines listener training: Enhancing description quality (RRG) yields much stronger personalization gains than directly training the listener on noisy zero-shot captions (listener-training), boosting Yo'LLaVA captioning F1 by +2.3%. Low-quality concept descriptions create an insurmountable bottleneck for listener reasoning.
  • Reference game outperforms isolated verification: While speaker-training improves over zero-shot, RRG's multi-candidate competitive setting provides superior contrastive pressure, forcing the speaker to discard extraneous context and articulate subtle discriminating details.
  • Continuous normalized reward prevents gradient stagnation: Transitioning from discrete binary rewards to posterior probability normalization (\(r_\text{match}\)) provides rich confidence feedback, delivering substantial boosts in captioning F1 (from 85.5 to 88.6 on Yo'LLaVA).
  • Circumventing dual-image hallucination: In MyVLM recognition, presenting both query and reference images simultaneously can trigger hallucination biases where models assume presence. In captioning, where only textual descriptions serve as context, RRG shines brightest, outperforming the prior state-of-the-art on PerVA by nearly 20 F1 points (86.9 vs. 67.3).

Highlights & Insights

  • Formulating concept identification as a cognitive communication game: By converting the complex goal of learning invariant discriminative attributes into an automated guessing game, the authors eliminate the need for expensive ground-truth text supervision.
  • Probability-normalized contrastive reward design: Leveraging the listener's binary token probability distribution across candidates neatly bypasses MLLM multi-image input limitations while preserving fine-grained confidence gradients.
  • Zero-shot generalization across visual domains: Trained on only 30 daily objects from PerVA, the speaker policy generalizes seamlessly to diverse categories such as landmarks, buildings, and human characters on Yo'LLaVA.

Limitations & Future Work

  • Author-acknowledged limitations: Communication is currently unidirectional within each round; speaker and listener do not engage in multi-turn adaptive dialogue. In addition, inference still relies on retrieval over the concept database, which may scale poorly if millions of concepts are registered.
  • Unstated potential limitations: The listener model remains frozen during training; its intrinsic multimodal perceptual blind spots could inadvertently shape the speaker's generation style. Furthermore, mining hard negatives requires semantic class labels, which may be non-trivial in completely uncurated real-world streams.
  • Future directions: Exploring bidirectional co-evolution where both speaker and listener co-adapt; incorporating active query mechanisms where the listener can ask clarifying questions about ambiguous attributes.
  • vs R2P (ICCV 2025): R2P uses frozen MLLMs to generate template descriptions and relies on costly pairwise image matching. Its descriptions retain transient states and background clutter. RRG uses RL in a reference game to generate state-invariant, discriminative descriptions, eliminating expensive pairwise inference and achieving superior accuracy.
  • vs RePIC (NeurIPS 2025): RePIC applies RL directly to train the model to output personalized concept names, which easily leads to shortcut learning and compromises general reasoning. RRG restricts RL to descriptive language generation, preserving and enhancing broader vision-language reasoning.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Elegant combination of cognitive reference games and GRPO for training-free personal concept description]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across three benchmarks, multiple tasks, ablations, retrieval-oracle settings, and qualitative analysis]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, rigorous mathematical formulation, and well-structured exposition]
  • Value: ⭐⭐⭐⭐⭐ [Offers a practical and performant blueprint for lifelong MLLM personalization without parameter retraining]