Re:Cognize: Open-Set Comic Character Re-Identification¶
Conference: NeurIPS 2026 โ ED Track (not the main track; venue attribution supplied in the task package)
arXiv: 2609.34032
Paper: https://re-cognize.vercel.app
Code: https://github.com/eternal-f1ame/Re-Cognize
Area: Human Understanding (comic character re-identification)
Keywords: open-set re-identification, streaming evaluation, identity maintenance, gallery updates, commit condition
TL;DR¶
Re:Cognize separates identity emergence from identity maintenance through four gallery protocols over the same reading-order stream, finds that growth under predicted labels harms recognition with random seeds, and uses Re:Cast's character averages, page evidence, and cross-page binding to recover part of the benefit of correctly labelled updates under explicit conditions.
Background & Motivation¶
Comic character re-identification is not merely nearest-neighbour search over photographs of a known cast. A reader meets a character a few times and later recognises that character despite changes in pose, occlusion, and drawing style; a complete cast list is usually unavailable in advance. Traditional closed-set Re-ID instead prepares a fixed gallery for every identity before evaluating whether a query retrieves the correct reference. That setting measures visual similarity but does not establish how a system builds and maintains character records while reading a story for the first time.
Existing comic-native encoders provide usable appearance representations, making it possible to move beyond whether there are enough reference crops to whether new crops should be accepted. The main difficulty is that a gallery consumes its own predictions: an incorrectly filed crop can become a reference for later queries, not merely an error on the current one. Comparing a static gallery, growth under predicted labels, and growth under true labels separates representation limitations from acceptance failures. First appearances are diagnosed through a separate unlabelled protocol rather than presented as a problem solved by maintenance results.
Core idea: a gallery update is valuable not when its added crops are accurately labelled in isolation, but when it is more accurate than the original gallery on the queries whose answers it actually changes; that condition should be measured before deciding which streaming evidence to commit.
Method¶
Overall Architecture¶
The input is a stream of character crops ordered by each dataset's page numbering. Outputs are query identities or, in the unlabelled protocol, assignments to existing or newly created clusters. Re:Cognize is the evaluation framework; Re:Cast is the gallery mechanism for maintaining known identities. They should not be conflated into a single new network. Evaluation uses ground-truth boxes and does not cover a complete system from whole-page detection to final character naming.
The framework fixes query order while changing the initial gallery and its update permissions. P1, P2, and P4 measure known-identity maintenance; P3 starts with an empty gallery and only diagnoses identity emergence. Random seeds, Seq-R, and first-appearance seeds, Seq-T, have equal label budgets but different temporal placement. In particular, Seq-R can initialise the gallery with labelled crops from anywhere in a volume, so it is not strictly first-reading supervision.
With random seeds, Re:Cast uses character averages, commitment under page evidence, and optional seed expansion. First-appearance seeds lack temporally distributed label anchors, so cross-page binding additionally propagates a known identity. The paper primarily contributes protocols and gallery-decision analysis: the four protocols are not a serial network, and oracle updates are not supervision available during deployed inference.
Key Designs¶
1. Four-protocol comparison: separate few-shot recognition, gallery growth, and identity emergence
P1 randomly places approximately one fifth of each character's crops in the gallery and uses the rest as queries, ranking by cosine similarity and reporting mAP and Rank-k. Single-crop characters are excluded and counted. P2 keeps 1 to 5 labelled seeds per character in an unchanged gallery and queries the remaining crops; characters with no more crops than seeds appear only in the gallery and produce no queries.
P3 supplies no identity labels and processes crops sequentially from an empty gallery. Each query is compared with the normalised running centroid of every cluster. It joins the most similar cluster above a fixed novelty threshold or starts a new cluster otherwise. The reference threshold is 0.55 for all representations. This is a measurement instrument, not a proposed optimal clustering algorithm. Cluster count, Purity, NMI, ARI, and Hungarian-matched accuracy are reported together to prevent excessive fragmentation from masquerading as high-quality clustering.
P4 starts from P2's seeded gallery, predicts and scores a query against the current gallery, and only then considers adding it. Predicted growth files the crop under its top-1 predicted identity; oracle growth uses its true identity solely to measure headroom; the static policy never updates. Original seeds are protected, and each character's added buffer is capped at 50, evicting the oldest added entry when full.
P4 reports identity Rank-1, the fraction of queries whose top-ranked identity is correct. Exemplar ranks and average precision change with gallery size, so changes in exemplar mAP cannot directly establish better character naming. Metrics are computed within each series and then macro-averaged across series; frequent characters still dominate queries within a series.
2. Commit condition: compare captured queries rather than added-crop label precision
A gallery operation affects only some queries: an added or changed entry becomes their nearest match, taking over answers previously supplied by the original gallery. Let \(c\) denote the fraction of all queries captured this way, \(p_{\mathrm{eff}}\) the updated gallery's identity accuracy on those queries, and \(a^{+}\) the original gallery's accuracy on the same queries. The overall accuracy change obeys the paper's Equation (1):
For a change that captures queries, an update therefore pays only when \(p_{\mathrm{eff}}>a^{+}\). Even a correctly labelled crop can be closer to other characters, so added-crop label precision is not its accuracy on captured queries. Similarly, the original gallery's overall accuracy is not its accuracy on this subset. A strong gallery often already answers queries near new crops correctly, making it harder for an update to outperform it.
Under the page constraint on POPCharacters, added-crop labels are 86.9% correct, but effective accuracy is only approximately 31% to 48%. For MagiV2 on Manga109, effective accuracy is 54.418%, the original gallery's accuracy on captured queries is 67.799%, and the capture rate is 46.53%. The resulting change is -6.23 percentage points: accurate added labels alone do not justify growth.
The equation is an exact decomposition of an accuracy difference, not an online confidence score computable without labels. The authors estimate its terms on half of the target corpus's labelled series or volumes and decide whether to grow on the other half. Estimates cannot simply be reused across corpora. Without additional evidence, it would be incorrect to claim that an inference system already knows whether each update satisfies the condition.
3. Character averages and commitment under page evidence: refine records instead of multiplying competing entries
The cast sheet represents each character by an L2-normalised running average of its filed features rather than retaining every crop as an independent nearest neighbour. A new crop refines a character prototype instead of creating another retrieval entry that can take over other characters' queries. Averaging is a standard prototype representation; the contribution is its connection to streaming acceptance analysis, not a new averaging operation.
Commitment under page evidence permits a crop to join a character only when another crop in its same-page group is already filed under that character. Groups come from MagiV2's character-to-character affinity head, run once per page without retraining. Among 3,977 within-page pairs, 89.9% depict the same character. The commitment rule consults only crops already read and abstains when a page supplies no identity evidence, rather than treating every query prediction as a label.
This constraint raises effective accuracy at the cost of coverage and does not guarantee a benefit for strong galleries. With five random seeds, the cast sheet plus commitment gains 3.34 to 6.51 percentage points over the static gallery. However, MagiV2's cast sheet alone is slightly better at +6.81. On Manga109, its +2.78 from the cast sheet must likewise be distinguished from the cumulative +1.33 after commitment; adding a mechanism is not necessarily a monotonic improvement.
Before the query stream begins, seed expansion uses the same page groups to add crops grouped with a seed, averaging 1.43 added crops per seed with 90.6% correct labels. It changes the query set because expanded crops must be removed from the queries. The paper therefore evaluates it against its own corresponding static reference, not a directly additive comparison with columns that retain those queries.
The source has a text-table inconsistency for single-seed results. Section 6 states that the cast sheet and commitment are inert and attributes benefits to seed expansion. Table 5 nevertheless places k=1 gains such as +4.71 to +6.84 in the "++ commitment" column, while its caption says the first two changes coincide. This note preserves that uncertainty rather than relabelling those results or inferring an implementation for commitment.
4. Cross-page binding: bring identity references closer to queries with first-appearance seeds
Seq-T concentrates labels near the beginning, leaving same-page commitment with few usable anchors. With five seeds, the labelled share of the last twenty crops is 20.5% in the first quarter and below 5% afterwards. Commitment coverage drops from 9.5% under Seq-R to 2.3%. Changing recent candidates or gallery capacity alone cannot move a correct identity reference closer to later queries.
Binding first groups a page's crops by similarity in a frozen representation. It then merges page groups, in reading order, into their best earlier match or opens a new group. A group receives an identity only when it encounters its first labelled seed, after which later members can join the gallery under that identity. The shared binder uses fine-tuned MagiV2 features, with within-page and cross-page thresholds of 0.7 and 0.5, for all five retrieval backbones. It needs no detector or panel geometry.
This does not constitute a proposed new-character discovery algorithm: identity names still come from supplied seeds, and cross-page groups propagate labels already obtained. On POPCharacters with Seq-T, the shared binder adds 12.51 to 16.94 percentage points. Four backbones benefit on Manga109, but MagiV2 loses 7.44 because its stronger static gallery exceeds the quality of the binder's evidence.
Each backbone can alternatively bind using its own representation. Its threshold is selected using the change predicted by the commit condition, not the final measured gain. Gains on POPCharacters are 1.73 to 8.46, below the shared binder. The latter's advantage includes the external capability of a comic-native representation and does not establish that every backbone can independently produce equally reliable cross-page groups.
Loss & Training¶
Re:Cast's gallery rules fit no new parameters to target data, but their representations and affinity head are not untrained. The paper separately evaluates an encoder-side maintenance baseline: a frozen backbone with a trained BNNeck, optionally augmented with the memory block and LoRA. Here, "Finetuned" normally means training BNNeck rather than fully fine-tuning the backbone; the memory block is not LoRA.
Working Memory holds a FIFO of 8 recent features per identity, while Episodic Memory stores 5 non-parametric prototypes per identity. Cross-attention from both branches produces residuals, combined through a gate and a small-initialisation projection into the backbone descriptor. Inference first searches all prototypes for a predicted identity, then routes Working Memory by that prediction while Episodic Memory still searches all prototypes in the second pass. True query labels must not route deployed inference. P3 bypasses the memory block because there is no initial gallery and evaluates only the BNNeck trained alongside it.
Training combines prototype classification, batch-hard triplet, and InfoNCE memory-consistency losses with weights 1.0, 1.0, and 0.1. Auxiliary CE has weight 0.3 without memory and 0 with it. Training uses 200 epochs, AdamW, a learning rate of \(10^{-4}\), 5 epochs of warmup, and cosine annealing. PK sampling normally uses 8 identities with 4 crops each, reduced to 4 identities for MagiV3. Optional LoRA has rank 8, adapts only the last 4 attention layers, and uses a learning rate of \(10^{-5}\).
Encoder memory brings small gains, and not every component contributes. Removing Working Memory reduces MagiV2/MagiV3 P1 mAP by 0.59/0.98 relative to the full block; removing Episodic Memory instead improves it by 0.26/0.06. Appendix A.6 says prototypes stay in clean BN-normalised feature space, whereas Algorithm 1 updates EM with memory-enhanced features. These specifications conflict and should be checked against the implementation before reproduction, rather than resolved by reconstructing an equation.
Key Experimental Results¶
Main Results¶
In-domain training uses POPCharacters: 13 training, 2 development, and 8 test series. The test split has 70 characters and 4,058 crops. Manga109's 27 held-out volumes contain 784 characters and 29,315 crops and are evaluated only through zero-shot transfer. Re:Verse contains one series, 12 characters, and 1,825 crops and is actually evaluated only under P1.
The following selects the Finetuned rows without memory or LoRA from Table 2: POPCharacters, k=1, Seq-R, an added-buffer cap of 50, and means over three training runs. P1 is closed-set mAP; the last three columns are P4 identity Rank-1. All use a 100-point scale, but differences across metrics are not meaningful.
| Backbone | P1 mAP | P4 static | P4 predicted growth | P4 oracle growth |
|---|---|---|---|---|
| TransReID | 37.4 | 17.4 | 15.0 | 41.7 |
| MagiV2 | 51.2 | 36.8 | 36.3 | 59.1 |
| MagiV3 | 41.3 | 22.7 | 19.9 | 49.9 |
| InstructReID | 37.8 | 17.5 | 14.5 | 42.6 |
| ReID5o | 38.7 | 20.8 | 17.9 | 45.6 |
Predicted growth is below the static gallery in all five rows, while oracle growth reveals substantial headroom. This conclusion is specific to random seeds. Some predicted-growth rows with first-appearance seeds improve slightly, so degradation should not be asserted for every streaming setting.
Ablation Study¶
The sequential analysis in Table 5 uses POPCharacters, Seq-R, k=5, one training run, and three seed draws. Change columns are cumulative percentage-point gains over the static gallery; "cast sheet + commitment" is not an additional increment to the preceding column.
| Backbone | Static identity Rank-1 | Cast sheet only gain | Cast sheet + commitment gain | Oracle gain |
|---|---|---|---|---|
| TransReID | 23.82 | +2.86 | +4.27 | +10.74 |
| MagiV2 | 44.35 | +6.81 | +6.51 | +15.97 |
| MagiV3 | 32.88 | +1.47 | +3.34 | +11.83 |
| InstructReID | 25.38 | +1.49 | +3.74 | +9.66 |
| ReID5o | 28.68 | +1.71 | +3.74 | +10.47 |
The following selects the Seq-T, k=5 binding analysis from Table 6. POPCharacters static scores retain the source's one-decimal precision. Manga109 changes are relative to its own static galleries, not to the POPCharacters static column.
| Backbone | POPCharacters static | Shared binding: POPCharacters | Self binding: POPCharacters | Shared binding: Manga109 |
|---|---|---|---|---|
| TransReID | 20.7 | +14.32 | +4.96 | +10.32 |
| InstructReID | 20.4 | +15.45 | +4.60 | +10.34 |
| ReID5o | 20.9 | +16.94 | +2.52 | +10.35 |
| MagiV3 | 26.5 | +16.45 | +1.73 | +8.00 |
| MagiV2 | 37.4 | +12.51 | +8.46 | -7.44 |
Key Findings¶
- One seed recovering approximately P1 mAP does not mean that character naming is nearly solved. With memory configurations, P2 mAP reaches 102% to 107% of P1, but Rank-1 reaches only 43% to 66%. Gallery size and identity-frequency distributions change metric baselines.
- At P3's 0.55 threshold, TransReID/MagiV2 produce 95.8/38.9 clusters per series against 8.8 true identities on average, with ARI on a 100-point scale of 2.0/24.0. High purity cannot conceal excessive identity fragmentation.
- Across 200 random splits that label half of the target corpus's series or volumes, the commit condition correctly predicts the other half's gain direction for six of Table 4's seven settings. The remaining MagiV2/POPCharacters effect is only -0.04. Reusing POPCharacters terms on Manga109 instead predicts +0.33 against an actual -6.23.
- Table 21 retains paired differences rather than simply subtracting displayed means. For example, MagiV2's full block is 51.75 and removing Working Memory gives 51.15, but the reported difference is -0.59. The displayed precision is inconsistent; the reported value should be preserved rather than changed to -0.60.
Highlights & Insights¶
- Updates should compete with the existing system, not merely pass their own confidence test: the same commitment evidence can help a weak gallery and harm a strong one. Focusing on captured queries explains why 86.9% correct added labels still do not ensure a positive gain.
- Label placement is part of the few-shot problem: Seq-R and Seq-T supply the same number of labels at different distances from queries. Cross-page binding propagates a known identity's reference to later pages rather than merely increasing gallery capacity.
- Negative results are not hidden in an aggregate score: memory affects ranking and top-1 naming differently, Episodic Memory does not contribute, and updates can degrade strong backbones. Reporting these distinctions is more useful for system design than selecting only the highest mAP.
Limitations & Future Work¶
- Methodological claims concern identity maintenance only. P3's fixed-threshold diagnostic does not solve new-character emergence, cluster merging, or identity naming; "open-set" in the title should not be interpreted as solving all of these problems.
- Evidence is restricted to Japanese comics and ground-truth boxes, without early-versus-late-arc appearance-evolution splits or a complete detection pipeline. Synthetic displacement and blur do not replace real detection errors, missed detections, or incorrect page ordering.
- The commit condition requires estimates from a labelled slice of the target corpus. Reliably estimating effective accuracy and the original gallery's counterfactual accuracy during unlabelled deployment remains unresolved.
- The shared binder relies on a stronger MagiV2 representation. Negative gains on strong galleries require joint validation of binding quality, original-gallery quality, and corpus distribution rather than unconditional activation.
- At 768 dimensions, the memory block adds 10.1M trainable parameters. Its implemented two-pass path repeats backbone inference and costs 2.5 to 3.1 times the backbone alone. This overhead belongs to the encoder-memory baseline and must not be attributed to Re:Cast's character averages.
- The source contains inconsistencies in single-seed commitment/expansion naming and EM update-feature specifications. Reproduction should consult the public implementation; this note does not erase uncertainty through reconstructed formulas or relabelled tables.
- Comic pages and crops remain subject to their respective research licences; the paper's licence does not authorise redistribution of underlying comic images. This note summarises mechanisms and results without reproducing panels or dialogue, and does not generalise fictional-character recognition into real-person identification capability.
Related Work & Insights¶
- vs TransReID/InstructReID/ReID5o: these backbones supply representations from different pretraining regimes. This paper evaluates each representation under four gallery conditions; small encoder-side gains cannot replace an acceptance rule.
- vs MagiV2/MagiV3: comic-native representations exhibit stronger identity structure under the same clustering rule and supply important resources for page grouping and shared binding. The paper reuses these capabilities rather than training a new comic-vision system from scratch.
- vs prototypical networks: character averages remain standard class prototypes, but their support set changes as a story is read, and acceptance errors affect later queries. The transferable lesson is to measure the net benefit of support-set updates, not merely reuse mean pooling.
- vs pseudo-label self-training/template updates: conventional rules often accept examples by individual confidence. Here, the analysis compares an update's conditional accuracy with the original gallery's on captured queries. It suggests a useful analysis for tracking or growing reference sets, but cross-task effectiveness is not established experimentally.
Rating¶
- Novelty: 4/5 โ The unified four-protocol evaluation and conditional acceptance analysis contribute more than mean prototypes themselves.
- Experimental Thoroughness: 4/5 โ Five backbones, cross-corpus evaluation, and extensive negative results, but real detection and long-term appearance evolution are absent.
- Writing Quality: 3/5 โ A clear main argument, with inconsistencies in single-seed attribution and memory-update specifications.
- Value: 4/5 โ A useful measurement framework for comic retrieval and identity-record maintenance, not yet a complete open-set reading system.