Social Choice Foundations for Simulation-Augmented Generation¶
Conference: NeurIPS2026
arXiv: 2609.38287
Area: AI Safety / Pluralistic Alignment / Social Choice Theory
Keywords: proportional representation, simulation routing, viewpoint embeddings, expanding approvals rule, pluralistic alignment
TL;DR¶
This paper formalizes representative routing in simulation-augmented generation as proportional clustering: predict simulation-response viewpoint embeddings, then select a small set of simulators with SEAR; it establishes approximate representation guarantees under explicit reward-factorization and error assumptions and outperforms clustering and random-routing baselines on two simulation pools, without studying final-answer synthesis.
Background & Motivation¶
Simulation-augmented generation (SAGE) aims to collect different viewpoints from a target population before answering, rather than compressing all preferences into an average reward. Querying every simulator produces a complete candidate-response collection but increases generation cost, latency, and context demands. A practical system can query only a few simulators from its existing pool, and the selection must depend on the current prompt: a fixed set of representatives may not cover the preference structures of different questions.
The challenge is not generic diversity. Distance-minimizing clustering can assign separate centers to sparse outliers while leaving only one position for a large, cohesive group. Random sampling also lacks deterministic guarantees for every sufficiently large, cohesive group. The paper therefore draws on proportional representation axioms from social choice theory: if a group accounts for a sufficient fraction of the population and collectively values an unselected response, it should receive a corresponding number of positions in the selected response slate.
The paper distinguishes three scales: the human population, the maintained simulation pool, and the simulators actually queried per prompt. Reducing the pool is a sampling problem; reducing current generation calls is a routing problem. Their assumptions and errors must remain separate. Core Idea: translate proportional representation claims into budget constraints in viewpoint space, replace exhaustive generation with inexpensive embedding prediction, and apply axiomatically justified SEAR routing to the predicted geometry.
Method¶
Overall Architecture¶
The inputs are a prompt, an existing simulation pool, and its contexts. The outputs are selected simulator indices and the responses generated by those simulators. The theory first considers ideal access to individual rewards; the practical Predictive Route–then–Simulate (PRS) procedure uses only black-box simulators and a viewpoint embedding model, without requiring online access to actual individual rewards.
The four key components are, in order, “Reward Geometry,” “Viewpoint Proxy,” “Budget Routing,” and “Selective Simulation.” Offline, simulation responses to training prompts are embedded to provide proxy supervision. At inference time, the proxy predicts vectors for the entire pool, budget routing selects indices, and only those simulators generate responses. Dashed edges below denote offline supervision, not exhaustive simulation for the current prompt.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Prompt + simulation pool"] --> B["Reward Geometry"]
B --> C["Viewpoint Proxy"]
T["Offline simulation responses<br/>and viewpoint embeddings"] -.->|Training supervision| C
C -->|Inference: predict for all| D["Budget Routing"]
D -->|Selected indices| E["Selective Simulation"]
E --> F["Representative response slate<br/>No final synthesis"]
Key Designs¶
1. Reward Geometry: replace average quality with proportional group claims
Let the simulation pool correspond to \(n\) individuals, with \(k\) response positions to select. For every \(\ell\in[k]\), a group of at least \(\ell n/k\) individuals has a claim to \(\ell\) positions. This claim also depends on preference cohesion: group members must collectively value an unselected candidate that serves as a comparison. Candidates remain indexed by their sources; two identical texts or vectors are still two candidates rather than being deduplicated before computing proportionality.
Reward proportionally justified representation (rPJR+) requires that, for every such group \(S\) and unselected candidate \(c\), the selected slate \(Y\) contains at least \(\ell\) responses satisfying:
The right-hand side measures the group's lowest reward for its common candidate; the left-hand side requires only some group member to value a selected response. This is a coalition-level representation condition, not individual satisfaction. Nor must the same person value all \(\ell\) responses. It does not guarantee a position to every small group: groups below the population threshold corresponding to one position may receive no protection from the axiom.
To route before generating responses, the paper assumes that rewards factor through a shared response embedding and a prompt-conditioned individual preference embedding on the unit sphere. It additionally assumes that each individual has a response whose embedding exactly matches their preference vector:
Inner products of unit vectors are monotonically related to Euclidean distance. Thus the reward axiom is equivalent to metric proportional justified representation plus (mPJR+): at least \(\ell\) selected points must have a distance to their nearest group member no greater than the unselected candidate's distance to its farthest group member. This nearest–farthest asymmetry is the geometric counterpart of the group maximum and minimum rewards above. Ordinary semantic similarity does not automatically justify this reward interpretation; the shared basis, unit normalization, and realizability of optimal responses are theoretical assumptions.
Population sampling is a separate layer. For a fixed prompt, optimal-response candidates, and a uniform sample without replacement from a population of size \(n_H\), Theorem 3 transfers exact sampled rPJR+ to an approximate population guarantee with probability at least \(1-\delta\). The sample size and population-level group eligibility become:
If the group's common comparison candidate receives reward at least \(\cos\beta\), the population guarantee gives at least \(\ell\) selected responses reward of at least \(\cos(\min\{\pi,4\beta\})\) from some group member. This is neither lossless transfer nor a claim that one sample simultaneously covers all future prompts. The sample must also be feasible and support the selected slate size. The paper gives \(n_H=8\) billion, \(k=10\), \(\delta=0.05\), and \(\varepsilon=0.03\) as an example requiring roughly \(9\times10^3\) samples, but that approximation is inconsistent with the displayed sample-size formula. The two statements should remain distinct rather than being treated as a verified deployment budget.
Appendix B replaces a union bound over population candidates with VC uniform convergence to obtain a dimension-dependent result. Its sample requirement can be independent of population and slate sizes, but Theorem 8 restricts comparison candidates to unselected sampled points and protects arbitrary cohesive groups with a factor-3 distance approximation. It is not simply a stronger version of Theorem 3 with identical scope. The stratified-sampling result likewise requires a specified sampling mechanism and group-size slack; a sample composition that resembles the population is insufficient by itself to establish fairness.
2. Viewpoint Proxy: predict simulator-expression vectors, not actual human preferences
Without reward access, the direct approach generates every simulation response and embeds it with a viewpoint embedding (VPE) model. This geometry is observable offline but saves no generation calls. PRS instead trains a proxy conditioned jointly on the prompt and simulator context to predict the embedding of the response the simulator would produce. The proxy neither generates that response nor claims to recover the unobservable human preference vector.
The paper fine-tunes a copy of an existing preference embedding model with LoRA. Supervision consists of offline simulation responses embedded by the original model. At inference time, the entire pool requires only proxy forward passes. The online workload therefore changes from exhaustive autoregressive generation to pool-wide encoding, geometric selection, and a few autoregressive generations. Predictions remain prompt-dependent rather than permanently representing a population through fixed clusters.
3. Budget Routing: approval budgets prevent one group from repeatedly funding all positions
SimulationRouter treats predicted pool vectors as both voters and candidates and applies the Spatial Expanding Approval Rule (SEAR). Every voter starts with budget 1. Approval balls expand around all candidates at a common radius. When an unselected candidate's ball first contains remaining budget of at least \(n/k\), that candidate is selected, and its approvers pay a total of \(n/k\). Selection continues until \(k\) positions are filled. The corresponding reward-space rule, REAR, progressively lowers the approval reward threshold.
Budgets gradually remove the funding power of members who have already obtained representation, preserving resources for large groups that remain underrepresented. The axiom's proof relies on the following observation: if a sufficiently large group's common candidate remains unselected, that group must already have spent enough budget on a corresponding number of selected responses. The guarantee therefore follows from group size, cohesion, and budget conservation, not average squared-distance minimization.
The accelerated implementation first sorts voters by distance for each candidate and stores inverse ranks, the current approval prefix, and its remaining-budget sum. Payments are collected in sorted order. Except for the last contributor in each round, every positive-budget contributor exhausts their budget, yielding at most \(n+k\) positive payment events overall. Prefix pointers only advance, so maintenance after sorting has amortized cost \(O(mn)\), and total selection time is \(O(mn\log n)\). Simulation routing uses \(m=n\), giving \(O(n^2\log n)\).
This is neither the entire system cost nor a linear-time algorithm. It requires pool-wide prediction and candidate–voter distances; explicit \(d\)-dimensional distance computation adds its own cost. The cited previous \(O(k^2n^4)\) bound refers to an earlier version of prior work. A footnote explains that the prior paper was updated in July 2026 to \(O(n^2(\log n+k))\), so the old bound should not be presented as the only current baseline.
4. Selective Simulation: select indices before generating and account for two distinct errors
If every simulator returns an optimal response for its corresponding individual, simulation-response embeddings equal human preference embeddings. With shared tie-breaking, simulating before routing and routing before simulating then select identical indices. Online generation calls fall from \(n\) to \(k\) while preserving exact rPJR+. Outside this ideal setting, commuting the operations preserves only approximate representation, not identical outputs.
There are two error sources. The simulation fidelity gap \(\alpha\) bounds the maximum angle between human preference vectors and simulation-response embeddings. The VPE prediction error \(\eta\) bounds the maximum angle between proxy vectors and simulation-response embeddings. Both are maxima over the entire pool for a fixed prompt, not training losses or average test errors. The PRS theorem relaxes a group's comparison-candidate angle threshold \(\beta\) to \(\widetilde\gamma_\beta=\beta+2\alpha+4\eta\).
This result also depends on the reward-geometry assumptions. Proxy accuracy alone cannot establish fairness for a real human population. Theorem 5 states an uncapped angular bound in the main text, whereas its appendix proof explicitly caps it at \(\pi\). Theorem 6 and its proof both state the uncapped error sum. Beyond \([0,\pi]\), cosine cannot continue to serve as a monotonic reward lower bound: the angular guarantee is then uninformative, rather than implying higher reward through cosine periodicity. This note records the source's differing formulations without inventing a replacement theorem.
The observable corollary is narrower: a sufficiently large group collectively close to an unselected candidate in predicted geometry receives a corresponding number of selected responses that lie within \(\beta+2\eta\) of some group member in exhaustive simulation-response geometry. This does not establish exact mPJR+ for every possible group in the exhaustive geometry. Empirical failures therefore do not contradict the approximate theory.
A Worked Example¶
Appendix A.2 gives a geometric counterexample without a substantive prompt: among \(n=6\) indexed points, 4 coincide at the origin, and the other 2 lie at \((1,0)\) and \((0,1)\), with \(k=3\). Optimal clustering places one center at each occupied location with squared-distance cost 0, yet gives the origin group only one selected candidate.
The origin group has 4 members, exactly the proportional threshold for \(\ell=2\). Its distance to another unselected origin candidate is 0, so it requires two selected representatives at distance 0. Clustering provides only one and violates mPJR+. SEAR permits two distinct, co-located indexed candidates to receive positions: representative counts reflect population shares, not the number of visually distinct locations.
Loss & Training¶
Proxy training fits individual vectors while preserving inter-individual geometry. A batch samples \(s\) simulation responses to the same prompt. The loss combines elementwise mean-squared error with a Pearson correlation loss between predicted and target pairwise Euclidean distances:
The paper uses \(\lambda=1\) and batches of 16 user responses to one prompt. LoRA uses rank 8, scaling parameter 16, and dropout 0.05 on query/value attention modules; that scaling parameter is not the fidelity gap. AdamW uses learning rate \(10^{-5}\) and weight decay 0.01, with 1 epoch and a 2,048-token sequence limit. The appendix reports 8 H100 GPUs. Reducing online generation calls does not remove this offline training and data-collection cost.
Key Experimental Results¶
Main Results¶
Remesh uses simulators corresponding to 310 participants and expands its original 4 questions into 1,000 training/evaluation prompts. Its simulators use Llama-3.1-8B-Instruct without additional fine-tuning. Reddit AITA contains 992 posts and 4,567 users and uses HumanLM simulators; because axiom verification is expensive, evaluation subsamples 300 users. Both domains are evaluated through simulation-response embeddings, not actual human rewards as ground truth.
The central metric is the proportion of test prompts satisfying exact mPJR+, with voters and candidates both taken from exhaustive simulation-response embeddings. The cached text does not contain pointwise values from Figure 3's curves. The table therefore records only comparisons explicitly stated in the text, without estimating plotted values.
| Method | Online generation calls | Results in both domains for \(k\in\{3,\dots,10\}\) | Evaluation boundary |
|---|---|---|---|
| Simulate–then–Route, oracle SEAR | \(n\) | 100% mPJR+ satisfaction | Guarantee in exhaustive simulation-response geometry |
| PRS, predicted vectors + SEAR | \(k\) | Better than both routing baselines at every tested \(k\) | Nonzero prediction error does not guarantee the exact axiom |
| PRS, predicted vectors + \(k\)-means | \(k\) | Lower satisfaction than PRS-SEAR | Selects nearest actual candidates, not arbitrary centers |
| Uniform random routing | \(k\) | Lower satisfaction than PRS-SEAR | Sampling provides no per-prompt proportional guarantee |
Ablation Study¶
The paper does not report systematic ablations removing loss terms or changing LoRA settings. The table below summarizes axiom-verification labels for the six illustrated appendix cases at \(k=4\). Case numbers replace substantive prompts and advice content. These are selected illustrations, not a random test set, and cannot establish overall win rates.
| Illustrated case | PRS-SEAR | PRS-\(k\)-means | Random routing | Oracle SEAR |
|---|---|---|---|---|
| Remesh, Table 1 | Satisfied | Violated | Violated | Satisfied |
| Remesh, Table 2 | Satisfied | Violated | Satisfied | Satisfied |
| Remesh, Table 3 | Violated | Violated | Satisfied | Satisfied |
| Reddit, Table 4 | Satisfied | Violated | Satisfied | Satisfied |
| Reddit, Table 5 | Satisfied | Satisfied | Violated | Satisfied |
| Reddit, Table 6 | Satisfied | Satisfied | Violated | Satisfied |
Key Findings¶
- Increasing \(k\) does not make proportional representation easier to satisfy. The text reports declining satisfaction for non-oracle methods as \(k\) increases. The authors hypothesize that smaller eligibility thresholds make smaller groups more sensitive to prediction error; this is not an established causal ablation result.
- Better aggregate PRS performance does not imply per-prompt dominance. Appendix Table 3 explicitly shows PRS-SEAR failing while random routing satisfies the axiom.
- Exact mPJR+ verification costs \(O(mn\log n\cdot2^k)\), unlike polynomial-time routing selection. This cost motivates the smaller Reddit evaluation pool; the full 4,567-user pool was not evaluated under the same verification protocol.
- Reddit data and simulation judgments are class-skewed. Appendix Figure 5 reports that HumanLM outputs NTA for 87% of classifiable response pairs whose human judgment is YTA. Only pairs where both outputs contain NTA or YTA are included; this is not an error rate over all simulation responses.
Highlights & Insights¶
- The axiom distinguishes how many viewpoints exist from how many individuals those viewpoints represent. Indexed multisets permit co-located candidates to occupy multiple positions, exposing how geometric deduplication can destroy proportional information.
- Proxy supervision targets the geometry needed for routing rather than requiring an inexpensive model to reproduce long responses. Pairwise-distance correlation also aligns with approval-ball structure, although no ablation independently establishes this loss term's benefit.
- Separating simulation fidelity from prediction error is essential. The former determines whether representation can be interpreted in terms of human preferences; the latter determines whether inexpensive routing preserves the simulation pool's structure. Observing only the latter cannot audit the former.
Limitations & Future Work¶
- Shared, normalized, realizable reward factorization is a strong structural assumption. Experiments do not establish that practical viewpoint embeddings satisfy it; mathematical guarantees and actual model assumptions should be reported separately.
- Experiments cover only two domains and target simulation-pool representation, not human preference ground truth or final-answer quality. Response synthesis remains open, so representative candidate slates cannot be equated with fair final answers.
- Maximum angular errors can be dominated by individual simulators. The paper does not measure fidelity gaps for the real population or report complete deployment latency, cost, or fairness confidence intervals.
- Groups below the size threshold may have no guarantee. Selecting \(k\) changes the protected group scale, not merely system throughput. The authors also identify heterogeneous fidelity across groups and domains as future work.
- The sample-size example conflicts with the formula, and some angular error bounds omit a capping qualification. Applications should check their valid ranges rather than interpreting loose bounds as nontrivial reward guarantees.
Related Work & Insights¶
- vs the overall SAGE architecture: this paper provides axioms and algorithms for routing, not population identification, simulator validity, or final synthesis. Its arXiv ID is 2609.38287 and should not be conflated with the separate SAGE paper.
- vs \(k\)-means / balanced \(k\)-means: clustering optimizes distances or cluster sizes without automatically ensuring proportional representation. Separate appendix counterexamples show that balanced cluster sizes cannot replace common-candidate group comparisons.
- vs EAR / SEAR / REAR: the contribution is not inventing expanding approvals from scratch. It connects rewards and geometry, accelerates an implementation with specific payment rules, and proves approximate commutation between routing and simulation.
- vs fixed community models and pre-deployment preference aggregation: this method maintains a larger simulation pool but dynamically selects a few indices for each prompt. The transferable insight is to audit whether resource allocation protects cohesive groups rather than reuse one representative list for all inputs.
Rating¶
- Novelty: 4/5. Connects proportional representation, sampling transfer, and reduced generation calls in a routing theory with explicit conditions.
- Experimental Thoroughness: 3/5. Two simulation domains support comparative results, but human reward evaluation, loss ablations, and complete system-cost measurements are absent.
- Writing Quality: 4/5. Reward-access and black-box settings are clearly separated; the sample-size example and angular-bound formulations still require checking.
- Value: 4/5. Provides a debatable and verifiable routing criterion for pluralistic alignment, with practical value dependent on simulator and viewpoint-geometry reliability.