MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models¶
Conference: NeurIPS2026 โ Evaluations & Datasets Track
arXiv: 2609.38543v1
Area: Medical Natural Language Processing (medical_nlp)
Keywords: medical knowledge updates, knowledge integration, generalization evaluation, temporal updates, knowledge retention
TL;DR¶
MedKIT converts 6,196 timestamped oncology evidence updates into fact-centered, multi-task probes and compares 12 knowledge integration methods across 5 models, showing that recalling an updated fact usually does not imply using it in open-ended reasoning or generation.
Background & Motivation¶
Medical knowledge is not a static question-answer collection: new clinical studies can change treatment comparisons for a particular condition, clinical context, and endpoint. Knowledge editing, continual post-training, and retrieval-augmented generation (RAG) can all expose deployed models to new evidence, but accessing evidence, remembering a label, and using evidence correctly in a different task are distinct abilities. Evaluating only an update question and its paraphrases can mistake query matching for usable knowledge.
Existing benchmarks cover paraphrases, multi-hop questions, document updates, and sequential editing, but often evaluate these forms of transfer separately or use synthetic facts without a realistic temporal structure. MedKIT does not introduce another editor. Instead, it fixes an underlying clinical comparison and changes the wording, comparison direction, and response format, while monitoring past updates, neighboring knowledge, and general capabilities. Curated HemOnc comparisons and associated publication dates enable this granular, temporal evaluation.
Core Idea: build a progressively transformed probe chain around each sourced knowledge update, measuring factual recall, relational transfer, open-ended use, and knowledge preservation separately rather than treating high update-task performance as successful knowledge integration.
Method¶
Overall Architecture¶
MedKIT is a benchmark construction and evaluation protocol, not a new model architecture. Its input is a set of HemOnc.org treatment comparisons linked to clinical studies; its outputs are canonical update facts, associated PubMed titles and abstracts, and evaluation tasks generated from manually designed templates. Evaluation first measures the original model, then accumulates method state in publication-date order and compares task performance before and after updates.
The process consists of evidence canonicalization, fact-centered probes, stateful temporal evaluation, and tier-specific scoring validation. Canonicalization establishes what the evidence actually compares; probes change only how that knowledge is used; temporal evaluation separates immediate learning from later retention; scoring distinguishes explicit labels from semantic consistency. This note does not depict benchmark sections as a neural architecture or assume that all methods share a training objective.
Key Designs¶
1. Evidence canonicalization: define each atomic update as a conditional clinical comparison
The data come from a fixed HemOnc.org snapshot dated 2026-03-12. Each fact contains treatment A, relation, treatment B, condition, clinical context, and endpoint, along with the study date and PubMed title and abstract. Relations are mapped to superior, inferior, or no difference. Condition and context are essential qualifiers: the same treatment pair can have different conclusions under different settings or endpoints, so dropping them would turn conditional evidence into an unwarranted global preference.
Preprocessing is entirely rule-based and uses no generative model. It removes incomplete entries, self-comparisons, unresolved supporting evidence, and labels that cannot be mapped reliably. For the same study-level comparison, it retains one canonical endpoint, prioritizing Primary, Co-primary, Secondary, and Undesignated endpoint types, followed by clinically relevant endpoints such as OS and PFS. Rare symmetric comparisons use seed 42 to select the canonical direction; reverse comparisons swap treatments and flip the relation. Exact duplicates and inconsistent tuples within the same evidence setting are removed.
The released benchmark contains 6,196 updates from 4,098 studies, covering 329 conditions, 2,135 treatment regimens, and 649 condition-context pairs across 1960โ2026 and 12 oncology groups. The appendix gives exact label proportions of 43.4% no difference, 28.7% inferior, and 27.9% superior; the main text uses rounded integers. These are all eligible comparisons after rule-based filtering, not a randomly sampled subset of available records.
Different outcomes across studies or dates are not automatically treated as errors. The release retains 154 updates marked conflicting_edit, but excludes them from the main experiments. Headline results therefore test unambiguous factual integration rather than adjudication of contested evidence. The benchmark preserves genuine conflicts, but its main results do not establish whether models can handle them.
2. Fact-centered probes: hold the fact fixed while changing how knowledge must be used
All tasks use fixed templates manually designed with an MD clinician, rather than LLM-generated questions. Each fact has seven probes: one Update, two Lexical, one Relational, one Compositional, one Operational, and one Locality. Thus, four generalization categories do not mean four probes, and seven probes do not mean seven independent facts. This is controlled variation around an individual update, not standard generalization across disjoint training and test samples.
Update asks for the original comparison using one of three relation labels, checking whether integration succeeded. The two semantically equivalent Lexical paraphrases test robustness to surface wording. Relational reverses the treatment order: superior and inferior are exchanged, while no difference remains unchanged. Reverse questions expose models that associate a question with a label without learning the underlying relative relationship.
Compositional provides the condition, context, endpoint, and two treatments and requests an open-ended comparison and explanation. It tests whether the updated fact participates in generated reasoning, but is not a comprehensive test of arbitrary multi-hop or multi-fact composition. Operational provides only condition and context, without explicitly requesting a comparison of the target treatments, and tests whether free-text generation implicitly respects the updated relationship. This is closer to using a fact than finding it, but the score does not establish complete clinical correctness.
Locality uses an independent comparison involving a different condition, preferably within the same oncology group and typically with a different context and label. If suitable candidates are unavailable, selection constraints are relaxed while maintaining factual independence. It measures spillover into neighboring knowledge and is not a generalization tier. Although the Operational template requests treatment descriptions, this note does not reproduce recommendations, dosing, or administration instructions; all conclusions concern research evaluation only.
3. Stateful temporal evaluation: separate immediate integration, past-update retention, and capability drift
The main experiments use 283 non-conflicting updates published from 2025 onward, ordered chronologically into 48 weekly batches. Construction statistics additionally report 98 daily and 14 monthly batches. Date filtering reduces the likelihood that evidence was encountered during pretraining, but does not prove that every model's training data excluded it. Parameters, auxiliary memories, or retrieval indices persist across batches rather than being reset for independent batch evaluations.
The six knowledge editing methods are AlphaEdit, MEMIT, WISE, GRACE, MEMOIR, and IKE; the three continual post-training methods are LoRA-Merge, O-LoRA, and SEEKR; the three retrieval methods are BM25-RAG, Dense-RAG, and Agentic-RAG. Parameter editing changes weights directly, augmented editing uses overrides, routing, or side memories, continual post-training accumulates LoRA updates, and retrieval methods freeze the base model while extending external evidence stores. Additional DPO and GRPO experiments are diagnostic baselines, not part of the 12 main methods.
RAG starts with pre-2025 evidence and adds newly observed evidence after each batch. BM25 uses lexical retrieval, Dense-RAG uses a biomedical encoder and FAISS HNSW, and all settings retrieve a top-k of 3. Agentic-RAG generates a search query, optionally retrieves again after inspecting snippets, and then produces an answer. Frozen base weights do not imply invariant answers: context and indices still change, so parameter-level capability preservation and retrieval-conditioned output quality must be considered separately.
Past-update retention is estimated with a small sentinel pool: two earlier cases are added after each batch, the pool is capped at 100 and maintained with replacement, and evaluation occurs every two batches. The appendix explicitly limits retention evaluation to closed-form tasks to control judge costs. It therefore does not establish long-term retention of open-ended knowledge use. Locality separately measures neighboring medical knowledge, while CapTrack measures out-of-domain capabilities; neither is synonymous with sentinel retention.
The five models are Gemma-3-4B-IT, Qwen-3-4B-Instruct, MedGemma-4B-IT, Llama-3.1-8B-Instruct, and Bio-Medical-Llama-3-8B. Medical-adapted variants help assess whether domain adaptation changes the observed pattern, but the study remains limited to 4B and 8B scales and cannot rule out different outcomes for larger models or other integration mechanisms.
4. Tier-specific scoring validation: measure update-induced changes without conflating score types
Closed-form tasks first normalize responses and match labels exactly. Longer responses are accepted when their sole extracted label matches the ground truth; a binary judge decision is used only when extraction fails. For open-form Compositional, Operational, and Locality tasks, gpt-4o evaluates semantic consistency with the target comparison. Intermediate points on the five-point scale allow neutral or partially aligned answers, so higher initial open-form scores do not imply higher factual accuracy.
The source defines the principal reporting quantity and normalization as follows:
In the first equation, the score is closed-form accuracy or a normalized open-form score; in the second, it is the raw 1โ5 judge score. Changes are multiplied by 100 when reported in percentage points (pp). The protocol aggregates within batches and then averages across batches, which is not necessarily equivalent to an overall instance-weighted mean when batch sizes differ. Normalization aligns numerical ranges; it does not turn semantic consistency into clinical accuracy.
Judge validation uses seven models and evaluates performance tiers against a six-judge leave-one-judge-out consensus, rather than forcing fine-grained rankings among methods near baseline. All seven judges recover the three Compositional tiers, and five recover the two Operational tiers; the remaining two each show one boundary-level disagreement. Because tiers are derived from the judge ensemble itself, this validates comparative stability rather than an external ground-truth ranking.
Two independent clinicians score the same 100 held-out responses. Their Spearman correlations are 0.76 and 0.62 for Compositional and Operational, respectively; gpt-4o's correlations with their consensus are 0.83 and 0.64. Each judge is queried five times on a deterministic 10-item subset, with median per-item variance equal to zero in nearly all settings. This limited validation supports broad comparisons but cannot guarantee precise, reliable assessment of every open-ended response.
Loss & Training¶
There is no single new loss function. Editing methods are adapted to chat templates and multi-token labels, continual post-training updates LoRA sequentially, and SEEKR uses a capped replay buffer. Additional DPO experiments build preference pairs from correct and alternative labels, while GRPO uses correctness and a smaller valid-label reward. These are evaluated integration strategies, not learning components of MedKIT itself.
Hyperparameters are selected on historical held-out updates to improve closed-form Update accuracy while checking Locality. Appendix D.3 specifies tuning three representative architectures and transferring configurations to their medical counterparts, which is more specific than the main text's per-method ร model summary. All runs use seed 42, greedy decoding for closed-form tasks, and default sampling for open-form tasks. Cross-model error bars are not multi-seed confidence intervals.
The study comprises 225 runs requiring approximately 7,200 A100-hours, plus approximately 1,000 hours for tuning. Standard weight artifacts generally support vLLM, whereas custom routing or side-network forward passes require HuggingFace. Deployment cost therefore depends not only on editing speed but also on compatibility with efficient inference backends.
Key Experimental Results¶
Main Results¶
The following selected results come from appendix Table 8 for Llama-3.1-8B. All values are postโpre changes in pp, not absolute scores or cross-model means. Oracle (Abs) directly provides the correct abstract and is a retrieval diagnostic, not a deployable method or a proposed model.
| Method | Update | Lexical | Relational | Compositional | Operational |
|---|---|---|---|---|---|
| MEMIT | -8.4 | -12.6 | -18.2 | -66.3 | -29.9 |
| GRACE | +59.4 | +1.8 | +0.3 | -0.7 | +0.6 |
| MEMOIR | +57.7 | +51.5 | +13.7 | -1.6 | -1.1 |
| LoRA-Merge | +65.0 | +54.8 | +31.0 | +2.3 | +1.1 |
| SEEKR | +66.1 | +52.5 | +31.7 | +2.6 | +1.9 |
| BM25 RAG | +10.8 | -0.5 | -0.7 | +1.2 | +1.0 |
| Oracle (Abs) | +29.5 | +20.1 | +26.7 | +19.8 | +47.3 |
The main text summarizes MEMOIR across models with rounded values of Update +60, Lexical +51, Relational +14, Compositional -1, and Operational -2 pp. That aggregation differs from the single-model row above. The central observation is that preserving recall across wording changes does not guarantee gains in more demanding knowledge use. Some combinations show small positive gains; โno meaningful overall improvementโ should not be rewritten as โevery individual cell is non-positive.โ
Ablation Study¶
The retrieval analysis below comes from appendix Table 7 and measures recall@3 under the realistic full-corpus setting for Llama-3.1-8B. It measures evidence retrieval, not generation accuracy.
| Tier | BM25 | Dense |
|---|---|---|
| Anchor | 0.40 | 0.27 |
| Lexical | 0.49 | 0.33 |
| Relational | 0.40 | 0.29 |
| Compositional | 0.32 | 0.35 |
| Operational | 0.00 | 0.03 |
Operational queries contain only condition and context, without treatment names, so topical relevance cannot substitute for exact evidence alignment. Starting with an empty corpus improves retrieval, and directly supplying the correct abstract improves open-form results. These diagnose corpus competition and evidence use separately; approximately 40% recall at one tier does not establish that RAG is universally ineffective.
The following results come from main-text Table 2. WikiBigEdit evaluates Llama-3.1-8B and Qwen-3-4B and reports their mean exact-match containment changes in pp. Compositional here corresponds to multi-hop questions, not MedKIT's judge scale, so magnitudes should not be compared directly across benchmarks.
| Method | Update | Lexical | Compos. | Locality |
|---|---|---|---|---|
| GRACE | +75.5 | +0.0 | +0.0 | +0.0 |
| MEMIT | +38.1 | +22.8 | 0.0 | -1.9 |
| MEMOIR | +52.2 | +38.1 | -0.8 | -2.2 |
| SEEKR | +41.9 | +20.2 | -1.9 | -0.9 |
Key Findings¶
- Gains do not automatically propagate along the probe chain. GRACE improves the original question substantially but stays near baseline after paraphrasing; MEMOIR transfers across paraphrases but shows much weaker relational and open-ended gains.
- Continual post-training generally yields smoother transfer but still faces forgetting. The main retention analysis reports approximately +48 pp immediate gains and a subsequent drop of approximately 32 pp for MEMOIR; this is not the same metric as its Update entry in the main table.
- Daily and weekly batching show similar qualitative trends, and medical-adapted models do not change the overall degradation pattern. These observations support robust trends at the evaluated scales, not a universal law for all models.
- Parameter editing can damage both neighboring facts and general capabilities. Zero CapTrack values for frozen-weight methods partly follow from evaluation construction and do not establish absence of output drift in every retrieval-conditioned application.
Highlights & Insights¶
- Fact-centered evaluation localizes failure. Success on the original question but failure on its mirror more directly exposes form matching rather than relational learning than a single overall accuracy score.
- Time, generalization, and retention are independent dimensions. Separating them prevents successful editing from concealing later forgetting or spillover.
- Oracle evidence paired with realistic retrieval provides a mechanism diagnostic. If correct evidence supplied directly is substantially better, evidence matching should be examined before attributing all failures to the generator's reasoning ability.
Limitations & Future Work¶
- Three qualitative labels and one canonical endpoint omit effect sizes, uncertainty, multi-arm relationships, and subgroup differences. No difference should not be interpreted as equivalence on every clinical attribute. Extensions could incorporate quantitative evidence and contested settings rather than turning results into treatment advice.
- The benchmark inherits HemOnc and literature coverage biases, and the main experiments use only 283 clean updates. Date filtering reduces exposure risk without providing a per-model training-data audit. Runs use one seed, and broad error bars primarily reflect model differences.
- Operational scoring checks consistency with a comparison, not dosage, contraindications, or complete safety. This is a research benchmark, not a clinical guideline, decision-support system, or medication recommendation. Clinician validation includes only two annotators and 100 responses.
- The source contains a reporting tension: the main text states that correct context substantially improves all task types, and Table 8 reports clearly positive oracle open-form gains, whereas appendix E.5 says oracle evidence does not produce strong open-form gains. This note retains verifiable table values rather than reconciling the authors' differing strength assessments.
- Cost section E.9 separately describes eight probes and five portability variants, inconsistent with the seven-probe composition in A.4. The main evaluation is described according to A.4; the cache does not support reconstructing an unambiguous cost-probe inventory.
- Refusal handling also differs across descriptions: B.2 scores closed-form refusals as zero, F.2 describes closed-form refusals as incorrect and open-form refusals as neutral, while D.2 summarizes refusals as incorrect. The implementation for all open-form refusals should not be inferred. CapTrack omits long-context tasks and subsamples some augmented editors, requiring caution around zero values and small differences.
Related Work & Insights¶
- vs CounterFact / ZSRE: These emphasize edited-fact recall, paraphrases, and locality. MedKIT connects relation reversal and open-ended use within fact-centered evaluation and adds real publication-date sequences.
- vs MQuAKE / WikiBigEdit: Multi-hop and lifelong editing already investigate update propagation. MedKIT adds a granular hierarchy for clinical comparisons. WikiBigEdit supports a similar recallโuse gap outside medicine, but does not use identical tasks or scoring.
- vs RAG and continual post-training: The paper does not establish a universally superior family. It exposes different failure mechanisms: related but imprecise retrieval, memory activation tied to the original wording, and interference from sequential weight updates. Future work could target comparison-direction consistency and evidence qualifiers in training and diagnostics.
- Resource boundary: The supplied cache does not preserve verifiable code or dataset URLs, so none are invented. The appendix declares CC BY-NC-SA 4.0 for the dataset and Apache 2.0 for evaluation code; these are distinct assets from the paper text's CC BY 4.0 license.
Rating¶
- Novelty: 4/5 โ Combining fact-centered, multi-level transfer with real temporal sequences has clear diagnostic value, without proposing a new editor.
- Experimental Thoroughness: 4/5 โ Broad methods, models, and judge validation, but model scale, a single seed, and limited clinician annotation constrain extrapolation.
- Writing Quality: 3/5 โ The central argument is clear, while oracle strength claims, cost-probe counts, and refusal handling require clarification.
- Value: 4/5 โ Provides a more granular acceptance standard for knowledge maintenance than recall alone, without establishing clinical readiness.