Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State¶
Conference: NeurIPS2026 (Accepted, according to the supplied metadata)
arXiv: 2609.31354
Code: https://github.com/QxLabIreland/ReChat
Area: LLM (Other)
Keywords: mutable transcripts, context pollution, conversational state, natural language editing, humanโAI interaction
TL;DR¶
The paper lets users rewrite the state of an entire conversation through natural language rather than append more corrective messages; a controlled study with 17 participants supports usability, and three representative cases show shorter retained transcripts, without establishing improvements in long-term task performance or total API cost.
Background & Motivation¶
Conventional chat systems accumulate user and assistant messages chronologically and use the retained history as context for subsequent generation. This records the interaction, but does not necessarily represent current intent: a user may initially name the wrong country and later correct the year, or introduce a budget constraint after discussing a plan, while the original questions and answers remain in the transcript. The model must infer which conditions have expired from conflicting history, and users rereading the conversation face the same problem. The paper calls outdated, contradictory, or irrelevant information that can still affect subsequent responses context pollution, rather than attributing the problem simply to an insufficient context window.
Truncation, summarization, and retrieval-based memory can change the context presented to a model, but usually leave retention decisions to the system. Users may have no direct way to state that historical assumptions should now be replaced. Editing a message and creating a branch preserves alternative exploration paths, but does not jointly update related user messages and assistant responses within one continuing working transcript. Document editors allow revision of an artifact, but do not fully preserve the chat interaction model. The paper therefore seeks an interface that retains chat while giving users control over current conversational state, not a stronger memory model.
The tension is that original history supports traceability, whereas current state supports subsequent collaboration; they are not the same object. Users explicitly initiate a revision, and the system propagates it through the transcript instead of requiring manual message-by-message editing. Core idea: turn the transcript from an append-only interaction log into a user-revisable state representation, update affected historical turns through natural language editing, and continue chatting from the revised transcript.
Method¶
Overall Architecture¶
The ReChat prototype maintains a JSON transcript of user and assistant messages and switches between normal chat and history editing. In normal chat, the system appends the user input and generates an assistant response using the full history. In history editing, text entered in the same input box becomes an edit instruction, and the model outputs an entirely revised transcript. The interface replaces the displayed conversation, and subsequent chat uses the revised transcript rather than retaining the old one as active context.
The process consists of mode routing, full-transcript rewriting, and state replacement. Both ordinary responses and edits use Gemini 2.5 in the prototype, whose interface was built with Google AI Studio. The architecture depends on structured instruction following, not a newly trained model. The diagram shows inference-time data flow, with no training-supervision branch.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Current transcript + user input"] --> A["Mode routing"]
A -->|normal chat| N["Append input<br/>generate and retain response"]
A -->|edit history| B["Full-transcript rewriting"]
B --> C["State replacement"]
N --> O["Current conversational state"]
C --> O
O -->|next input| I
Three objects must be distinguished: the interaction that actually occurred, the transcript currently displayed, and the context supplied to the model on the next turn. The prototype primarily changes the latter two together. A rewritten assistant response is not necessarily what the assistant originally said, and a rewritten user message is not the original input. This distinction enables removal of obsolete conditions, but also creates the traceability and trust issues acknowledged by the authors.
Key Designs¶
1. Mode routing: distinguish continuing a conversation from revising its history
Append-only chat does not naturally distinguish a new question from an instruction to change the premises of an entire earlier discussion. ReChat provides an โEdit Historyโ mode selector near the input box and visual cues indicating that the next input will be interpreted as a revision request. Users still communicate in natural language and need not learn separate editing controls for individual messages.
In normal chat, the input becomes a new user message, and the model generates a response from the accumulating transcript. In history editing, the input specifies the desired historical change, and the system selects the editing prompt template. The mode determines the scope of the same text: either a current-turn question or a transformation over the entire history.
This is not automatic classification of correction intent, nor autonomous model selection of what to forget. Users initiate editing and specify its scope; the system executes the revision. Retaining the familiar input mechanism lowers interaction overhead, but users must understand the active mode to avoid confusing ordinary questions with global changes.
2. Full-transcript rewriting: propagate changes through user messages and assistant responses
Changing an early question while leaving its answers untouched would preserve contradictory context. Editing therefore supplies the complete JSON transcript and the natural language edit instruction to the model, which generates a new transcript satisfying the instruction. The paper expresses this mechanism as the following state transformation:
Here, \(T\) is the original transcript, \(e\) is the natural language edit instruction, \(T^{\prime}\) is the revised transcript, and \(f\) is the model-executed transformation. The equation describes a change in textual state, not a learning objective or a proof that every edit preserves factual correctness.
The editing template in Appendix A assigns the model a conversation-editor role, represents historical messages by role and text, and requests a JSON array as output. The model can modify, delete, or add messages according to the instruction, while preserving Markdown lists, emphasis, headings, and code blocks. These formatting requirements aim to avoid disrupting document structure during revision; the paper does not report formatting preservation rates or JSON parsing success rates.
The transformation supports three primary operations. Retroactive correction changes earlier errors and related answers; constraint injection applies new conditions to the existing discussion; context pruning removes irrelevant detours. Their shared mechanism is not placing a correction at the end, but ensuring the retained transcript no longer simultaneously presents obsolete and current conditions.
For example, changing a GDP discussion to a different country and year should update both the initial question and later category analysis. Removing an unrelated discussion should instead preserve unaffected content. Both operations require reasoning about relationships between messages rather than simple keyword replacement. Full regeneration gives the model access to these dependencies together.
However, the paper demonstrates feasibility rather than providing a formal consistency checker, an edit-correctness guarantee, or external factual verification. The revised transcript may be cleaner while losing useful information. Extensions such as style or language changes in Appendix E are additional possible uses of transcript transformation, not separately completed controlled experiments.
3. State replacement: make the edit result the active context for subsequent interaction
If a revision were merely appended as a new assistant message, obsolete information would still be sent to the model, reverting to ordinary correction. ReChat instead replaces the original transcript with the generated one and rerenders the interface. Subsequent normal chat continues from the revised transcript. This connects a visibly updated conversation with the state actually used on the next turn.
The interface indicates that previous turns are being updated and then displays the complete revised discussion. Users can read the current state directly instead of mentally combining scattered corrections with original answers. For tasks with a defined goal and repeatedly changing constraints, this presentation reduces the burden of identifying currently active conditions.
Replacing current state also means it is no longer an immutable audit record. The authors do not claim the prototype already implements comprehensive versioning, undo, visual diffs, or provenance tracking; these appear as proposed improvements in the discussion and limitations. Workflows requiring preserved exploration branches or evidence of the original interaction should not rely solely on rewritten transcripts.
A Worked Example¶
The task in Appendix B first asks users to request Irelandโs GDP in 2021 and then its largest categories. The transcript now contains two user questions and two assistant answers, all based on the initial country and year.
Under normal chat, Condition A, the user appends a country correction and then a year correction, each producing another assistant answer. The final transcript contains original questions, obsolete answers, and later revisions. The corresponding representative analysis retains 8 messages and 422 tokens.
Under history editing, Condition B, the user activates editing to change the country and activates it again to change the year. Each operation rewrites the retained exchange so the current transcript concerns Iceland in 2022, rather than preserving a sequence of apologies and corrections. The representative example ultimately retains 4 messages and 169 tokens.
Those 4 messages do not mean only 4 interactions occurred or that the system generated only those texts. Both edits still require processing and generating revised transcripts. The table counts the context that remains afterward. The GDP example illustrates propagation of revisions, not verification that the regenerated economic figures are accurate.
Loss & Training¶
The paper introduces no loss function, weight fine-tuning, or reinforcement learning training. What changes is the external JSON transcript, not model parameters; updating conversational state does not permanently insert new constraints into model knowledge.
The implementation regenerates the complete transcript, so output length grows with transcript size. Structured diffs, local updates, and retrieval-based recomposition are proposed as future optimizations, without ablations comparing them. Conceptual model independence should not be described as experimentally established cross-model performance.
Key Experimental Results¶
Main Results¶
The study uses a within-subjects design: 17 participants complete retroactive correction, constraint injection, and context pruning under both conditions. Ten follow AโB and seven BโA; A is normal chat and B uses mutable transcripts. Of the participants, 15/17 work in research, data science, or software development. Participation is voluntary and uncompensated.
After each task, participants assess both conditions using 1โ7 scales. The table records directions reported in the paper. The local text contains neither Figure 3โs numerical means and error bars nor the per-participant CSV, so precise scores are not inferred from its caption.
| Evaluation dimension | What it measures | B relative to A | Reported statistical result |
|---|---|---|---|
| Transcript cleanliness | Absence of obsolete information that could mislead later responses | Higher | \(p<0.001\) |
| State clarity | Understanding the currently active assumptions and constraints | Higher | \(p<0.001\) |
| User confidence | Subjective confidence in subsequent adherence to updated constraints | Higher | \(p<0.001\) |
| Restart intent | Inclination to start another conversation; lower is better | Lower | \(p<0.001\) |
| Ease of use | Ease of applying corrections | Higher | \(p<0.001\) |
The authors use paired t-tests for each question, pairing at the participantโtask level, and compute confidence intervals over those responses. Each person performs three tasks, so these pairs cannot be treated as independent participant samples. The results provide early evidence about subjective experience, but participant-level aggregation or a repeated-measures model should further assess the robustness of statistical significance. Consistent directions in both order groups do not completely rule out learning or carryover effects.
Ablation Study¶
The paper reports no component-removal ablation. The following structural analysis preserves its three representative transcript pairs and should not be interpreted as aggregate performance across every task from all 17 participants. โObsolete tokensโ counts tokens in entire user or assistant turns superseded by later corrected turns, rather than judging semantic obsolescence word by word. Counts use tiktoken.
| Scenario | A messages | B messages | A retained tokens | B retained tokens | A obsolete tokens | A obsolete share |
|---|---|---|---|---|---|---|
| Retroactive correction | 8 | 4 | 422 | 169 | 236 | 56% |
| Constraint injection | 6 | 4 | 608 | 479 | 353 | 58% |
| Context pruning | 10 | 4 | 712 | 66 | 655 | 92% |
| Mean as reported in the table | 8 | 4 | 581 | 238 | 415 | 71% |
The original table marks Bโs obsolete counts and shares with dashes, rather than listing 0 for each case; the accompanying text states that revision eliminates obsolete retained context. The table reports a 50% reduction in mean message count and a 59% reduction in retained tokens. The prose separately gives Aโs mean obsolete share as 71.4%, whereas the table gives 71%. Both representations are retained here rather than silently normalized.
These counts describe final retained transcripts, not complete request and generation bills. Editing still reads and regenerates the full transcript, and output cost grows with conversation length. The 59% reduction must therefore not be called a reduction in total API tokens or total fees. Counts from tiktoken also do not equal Geminiโs actual billing measurements.
For interaction-demand analysis, Appendix C reports correction habits and feature-adoption intent. The first five rows below are multi-select responses with 17 participants as the denominator, so their percentages should not sum to 100%.
| Survey item | Count | Share | Interpretation |
|---|---|---|---|
| Send a follow-up correction | 16 | 94.1% | Most common current correction strategy |
| Modify an old prompt and resend it | 12 | 70.6% | Revision still occurs through a new message |
| Directly edit an earlier message | 6 | 35.3% | Depends on existing interface support |
| Start a new conversation | 4 | 23.5% | Rebuild context to avoid obsolete information |
| Usually do not correct inputs | 1 | 5.9% | A minority usage pattern |
| Would use Edit History | 15 | 88.2% | 9 Likely and 6 Very Likely responses in a single-select adoption question |
Key Findings¶
- Reported experience improves on all five subjective dimensions, but increased confidence does not replace measurement of long-term constraint adherence or factual accuracy.
- Structural benefits vary substantially across the three examples: constraint injection retains considerable useful content, whereas pruning removes a large detour. This does not establish that pruning is generally the best operation.
- Appendix feedback describes both cleaner, easier-to-review transcripts and loss of original information. Requests for diffs, version history, and clearer mode cues show that recoverability is a practical cost of state cleanup.
Highlights & Insights¶
- Separate records from state: a complete interaction log need not be the ideal working context. The contribution makes currently valid information a user-manageable object, rather than only a background compression decision.
- Revise assistant responses as well: editing only user questions can preserve obsolete deductions. Joint historical revision makes constraint propagation an interface capability, but these texts must be identified as revised state rather than original interaction evidence.
- Retain the chat entry point: familiar natural language specifies revisions, while a mode switch changes their scope. This idea could support current-assumption management in writing and planning, although transfer benefits require separate evaluation.
Limitations & Future Work¶
- The sample is small and technically concentrated, with prescribed tasks; findings do not directly generalize to ordinary users, long-term open-ended exploration, or multi-user collaboration.
- There are no direct comparisons with branching, summarization, or retrieval-based memory, and no objective task-success, edit-accuracy, or long-term constraint-adherence results.
- Pairing at the participantโtask level creates correlated responses within participants. The local cache lacks the individual raw data needed for reanalysis.
- Full regeneration adds editing latency and output cost. Three transcript pairs do not establish lower end-to-end costs; evaluation should include both editing and subsequent interactions.
- Natural language edits can overmodify content, lose useful information, or alter a speakerโs intended meaning. Previews, undo, version history, and visual diffs require evaluation.
- Mutable transcripts introduce provenance, accountability, and security concerns that the prototype does not resolve. Audit-sensitive workflows should not retain revised state alone.
Related Work & Insights¶
- vs Lost in the Middle and LLMs Get Lost in Multi-Turn Conversation: these studies characterize long-context or multi-turn difficulties. This paper changes how users revise retained state, without establishing a solution to all model limitations in using long contexts.
- vs recursive summarization: summarization primarily reduces length and extracts memory; mutable transcripts change the semantic premises of active history. The approaches could be combined, but preserving edit provenance would require design work.
- vs retrieval-based memory: retrieval selects what to read, whereas mutable transcripts determine what remains currently valid. Retrieving a superseded version can still create conflicts.
- vs branching and Canvas-style document editing: branching preserves exploration paths, document editing changes an artifact, and this paper globally revises the chat transcript itself. A promising direction is distinguishing working state from original history, rather than treating revision as universally preferable to preserving branches.
Rating¶
- Novelty: 4/5. Clearly frames user-directed full-transcript revision as a chat interaction paradigm.
- Experimental Thoroughness: 4/5. The prototype, controlled user study, and case analysis complement each other, but support feasibility-level conclusions only.
- Writing Quality: 4/5. Operations and implementation boundaries are understandable; statistical independence and measurement scope need greater clarity.
- Value: 4/5. Offers a reusable approach to current-state management in ongoing collaboration, with deployment dependent on recoverability and provenance mechanisms.