Skip to content

PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations

Conference: NeurIPS2026
arXiv: 2609.34571
Code: https://github.com/airaer1998/PersonaManifold
Area: Interpretability
Keywords: persona representations, manifold geometry, anisotropy, geodesic steering, behavioral similarity

TL;DR

PersonaManifold models LLM persona activations as a curved, anisotropic low-dimensional manifold, using graph shortest paths weighted by local metrics to measure similarity and interpolate personas; combined BST triplet consistency improves over Euclidean distance by 5.1–6.1 percentage points across three 7–8B models, primarily benefiting high-deviation persona pairs.

Background & Motivation

Inference-time activation steering lets a model switch personas without fine-tuning: extract an internal vector associated with a personality trait or role, then add it to the residual stream. Methods such as PERSONA and Persona Vectors can therefore adjust trait strength or combine traits, but usually assume a sufficiently linear persona space: averaging two role vectors should yield an intermediate role, and equally long displacements in different directions should have comparable semantic costs. Reported trait coupling, asymmetric steering effects, and composition deviations suggest that this default geometry may be inappropriate.

The problem is not simply high vector dimensionality: a coordinate midpoint need not correspond to a behaviorally plausible intermediate persona. If two roles lie on a curved data-supported surface, their connecting straight line may traverse regions with almost no observed persona support. If local variation is directional, moving along common professional differences and moving along rare value combinations should not have identical costs. Another personality questionnaire alone cannot establish whether this geometry matters, because a model's statements about its traits can diverge from its situational behavior.

The paper addresses this through two complementary evidence streams: characterize intrinsic dimensionality, directional costs, and discrete curvature, then test whether the resulting distances and paths improve behavioral similarity prediction and persona interpolation. Core idea: replace straight-line operations in ambient activation space with movement along a manifold supported by observed personas, and assess their utility through situational behavior rather than self-report questionnaires alone.

Method

Overall Architecture

The inputs are approximately 2,000 persona descriptions and a frozen LLM. Offline, “Persona Activation Extraction” produces vectors, and “Manifold Geometry Estimation” constructs a neighborhood graph with local metrics. At inference time, “Geodesic Steering” queries the path between two personas and produces intermediate directions for residual-stream injection. “BST Behavioral Evaluation” independently constructs behavioral similarity labels and evaluates distances; it does not train LLM weights.

This is neither a newly trained generative network nor a system that reconstructs the entire geometry every dialogue turn. The persona pool and graph are prepared for each model; path querying, intermediate sampling, and dialogue generation follow endpoint selection. Curvature describes and analyzes the space. Local metrics, graph shortest paths, and spline smoothing directly implement steering; curvature values themselves are not injected into the residual stream.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Persona pool and frozen LLM"] --> B["Persona Activation Extraction"]
    B --> C["Manifold Geometry Estimation"]
    C -->|Inference: endpoint path| D["Geodesic Steering"]
    D --> E["Intermediate-persona dialogues"]
    C -.->|Distance evaluation| F["BST Behavioral Evaluation"]
    G["Independent situational questions<br/>and behavioral responses"] --> F

Key Designs

1. Persona Activation Extraction: remove generic response components using shared questions

The pool combines approximately 1,500 stratified PersonaHub personas with 500 literary characters from CoSER. The former cover professions, demographics, and value descriptions; the latter supplement trait combinations that are uncommon in the synthetic pool. Each persona conditions the system prompt while answering the same 40 probes. These are open-ended situations adapted from BFI-44, HEXACO-60, and TIPI items and manually reviewed. The representation comes from a single middle-layer residual-stream activation at the last token of the generated response, not a direct encoding of the persona description.

For each question, subtract the activation obtained without an additional persona prompt from the persona-conditioned activation, then average over the 40 questions. This difference is intended to remove shared question and generic-response components, yielding the persona's deviation from a neutral state; averaging reduces individual-question effects. However, it also compresses variation across contexts within a persona, so the vector should not be interpreted as a complete psychological model.

The authors select the layer using a validation set of 200 personas. The main text describes scanning candidate layers 12–20; Llama and Qwen ultimately use layer 16, and Mistral uses layer 14. PCA then retains 99% cumulative variance, reducing approximately 4,096 dimensions to 200–500. This retained dimensionality differs from the subsequently estimated intrinsic dimensionality: PCA preserves the representation space, while local geometry estimates the number of effective variation directions within it.

2. Manifold Geometry Estimation: model both directional costs and supported paths

The authors cross-check intrinsic dimensionality with a Marchenko–Pastur spectral test on local covariance matrices and Levina–Bickel maximum likelihood estimation. Across reported layers and models, estimates fall between 15 and 23 dimensions. This is far below activation dimensionality but clearly exceeds the five Big Five dimensions: low dimensionality does not imply that a five-dimensional personality inventory describes all persona variation.

Each local metric is derived from neighborhood covariance. Retain the leading intrinsic directions and weight each by the inverse of its variance plus a regularizer. Common directions have larger variance and lower movement costs; rare combinations have smaller variance and make equally long movements more expensive. Unlike a single global Mahalanobis matrix, the method uses different directional weights at different locations. The default neighborhood size is five times the intrinsic dimensionality, with regularization \(10^{-4}\).

After connecting points into a neighborhood graph, measure each edge under both endpoint metrics and average the two lengths. The following mechanism is Equation (4) from the paper: \(g(x_i)\) denotes the local metric above, and the norm measures displacement under that metric.

\[ w_{ij}=\tfrac{1}{2}\big(\|x_j-x_i\|_{g(x_i)}+\|x_i-x_j\|_{g(x_j)}\big). \]

The accumulated edge length along a Dijkstra shortest path defines the approximate geodesic distance \(d_g\). Local metrics determine which directions are cheaper, while graph connectivity prevents shortcuts across gaps without neighborhood support. These effects jointly shape distance rather than merely rescaling Euclidean distance. Finite sampling and neighborhood size still affect the result: it is a graph estimate, not an exact geodesic on a known continuous manifold.

Ollivier–Ricci curvature provides another analysis quantity. For adjacent points, take their uniform neighborhood measures, divide their Wasserstein-1 distance by the endpoint geodesic distance, and subtract the ratio from 1. Node curvature averages incident-edge curvature. Positive values indicate neighborhood convergence and negative values indicate divergence. The authors additionally divide points into five density bins and compare persona types to test whether curvature differences between professional and literary personas merely reflect sampling density.

3. Geodesic Steering: form intermediate roles along data-supported paths

Directly averaging two persona vectors may traverse positions unsupported by persona samples. The method first finds a graph shortest path, smooths it with cubic splines in local tangent spaces, and samples 10 intermediate points at equal arc-length intervals. Smoothing reduces behavioral jumps caused by moving between discrete neighbors. It does not mathematically guarantee that every spline point lies on the true manifold, so adherence still requires evaluation.

Inverse PCA maps intermediate points back to the original activation space for injection at the selected layer. Equation (6) first normalizes the mapped direction, then multiplies it by the interpolated endpoint norm \(s(t_k)\) and steering coefficient \(\alpha\):

\[ h_{\ell}\leftarrow h_{\ell}+\alpha\cdot\frac{P^{\dagger}\gamma(t_k)}{\|P^{\dagger}\gamma(t_k)\|}\cdot s(t_k),\qquad s(t_k)=(1-t_k)\|x_1\|+t_k\|x_2\|. \]

Here \(\gamma\) is the smoothed path and \(P^{\dagger}\) is the inverse PCA mapping. Path shape determines the intermediate persona direction, while intensity varies smoothly with endpoint norms, avoiding confounding direction effects with changing intermediate norms. The implementation table gives an \(\alpha\) range of 0.5–2.0; no LLM parameter updates are required.

The authors stratify persona pairs by \(d_g/d_E\) into high- and low-deviation groups and suggest that ratios above 1.2 make geodesic steering more worthwhile. This ratio is influenced by both the anisotropic metric and path structure, so it should not be equated with pure curvature. The 1.2 threshold and approximately 60% coverage are empirical findings for this pool, not universal geometric theorems.

4. BST Behavioral Evaluation: ground similarity in situational choices and actions

BST covers six constructs: Moral Foundations, DOSPERT risk attitudes, the Interpersonal Circumplex, Decision-Making Style, Schwartz Values, and Communicator Style. GPT-5.2 generates 10 forced-choice and 20 open-ended questions per construct, totaling 180. Pilot testing on 500 personas removes questions with choice entropy below 0.8 nats or open-response embedding variance below 0.05, retaining 48 forced-choice and 102 open-ended questions. These differ from the 40 probes used for activation extraction.

Behavioral distance combines forced-choice signals and open-response embedding signals, with a \(1/3\) weight on the former. For each anchor, positives come from the nearest 5% and negatives from the farthest 20%, requiring anchor–negative behavioral distance to be at least three times anchor–positive distance. This yields 10,000 triplets. A sample of 200 triplets is assessed by three annotators each, with 86.1% human–system agreement; this sampled validation does not mean that all 10,000 triplets were manually reviewed.

Triplet Consistency Rate (TCR) evaluates whether a representation distance follows these behavioral labels. For an anchor \(a\), behavioral positive \(p\), behavioral negative \(n\), and candidate representation distance \(d\):

\[ \mathrm{TCR}=\Pr[d(a,p)<d(a,n)]. \]

The paper separately reports TCR using labels derived from choices, open-response embeddings, and their combination. TCR-choice does not depend on open-response embeddings and provides an additional check against validating one embedding distance with another.

Section 3.4 describes the behavioral distance components as “option agreement” and “embedding cosine” but does not explicitly state the sign convention converting similarities into distances. This note preserves the weights and triplet rules without supplying an exact distance formula that the authors did not clearly specify; reproduction requires checking the code.

A Worked Example

Appendix N uses a conflict-averse pediatrician and an adversarial trial lawyer as endpoints, with a deviation ratio of 1.73. When a colleague publicly criticizes their work, the linear midpoint mixes appreciation, strong disagreement, and hesitation. The geodesic midpoint instead calmly acknowledges valid concerns and arranges an evidence-based response. The transition follows plausible intermediate roles rather than combining two styles of phrasing.

For this pair, geodesic versus linear IC is 0.71 versus 0.54, TS is 0.74 versus 0.51, and MA is 0.65 versus 0.40. The example illustrates the mechanism but remains an individual case, not a substitute for the model- and deviation-stratified results below.

Key Experimental Results

Main Results

The models are Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Mistral-7B-Instruct-v0.3. The following selection from Table 2 reports percentages. “Choice / Combined” uses forced-choice and combined behavioral labels, respectively.

Distance method Llama Choice / Combined Qwen Choice / Combined Mistral Choice / Combined
Euclidean 60.2 / 62.1 59.5 / 61.5 58.8 / 60.8
Mahalanobis 62.5 / 64.3 61.8 / 63.9 61.0 / 63.0
Isomap-Euclidean 63.8 / 65.8 63.0 / 65.1 62.5 / 64.3
Geodesic 65.2 / 68.2 64.3 / 67.0 63.9 / 65.9

Absolute combined-TCR gains are 6.1, 5.5, and 5.1 percentage points; forced-choice gains are 5.0, 4.8, and 5.1 points. The paper reports bootstrap significance against Euclidean distance at \(p<0.01\).

Steering evaluation uses 200 persona pairs: 50 high-deviation, 50 low-deviation, and 100 random pairs, with 10 intermediate points per path and at least 20 dialogue turns per point. The following Table 3 results concern only high-deviation pairs, not averages over all pairs.

Model Method IC TS MA
Llama Euclidean-Linear 0.58 0.52 0.41
Llama Graph-NN Chain 0.67 0.58 0.57
Llama Geodesic 0.72 0.69 0.62
Qwen Euclidean-Linear 0.55 0.49 0.38
Qwen Geodesic 0.69 0.65 0.58
Mistral Euclidean-Linear 0.53 0.47 0.36
Mistral Graph-NN Chain 0.66 0.53 0.56
Mistral Geodesic 0.66 0.63 0.55

IC averages pairwise cosine similarities among response embeddings across 20 turns at each intermediate point, then averages across intermediates. TS is the fraction of adjacent-step changes across five Big Five traits whose direction matches the overall endpoint change. MA divides each intermediate's geodesic distance to its nearest observed persona by the 95th percentile of inter-persona distances, averages these values, and subtracts from 1. It measures proximity to observed support, not independent behavioral correctness.

In external Table L.1, Llama's PersonaGym PersonaScore increases from 3.12 with linear steering to 3.71, BFI-44 Monotonicity from 58.4% to 72.8%, and Trait Expression Score from 61.3 to 76.5. A separate naturalness evaluation covers 100 dialogues and three annotators: geodesic steering receives 3.82/5 versus 3.41/5 for linear, with Krippendorff agreement coefficient 0.73.

Ablation Study

The following selection comes from Table 4. TCR is Llama combined TCR; IC/TS concern Llama high-deviation pairs. Steering variants without a reported distance result are marked “not applicable.”

Config TCR (%) IC TS
Default: local PCA metric, cubic spline 68.2 0.72 0.69
Isotropic Euclidean edge weights 63.8 0.65 0.61
Neighborhood size: 3 times intrinsic dimensionality 66.9 0.70 0.67
Neighborhood size: 10 times intrinsic dimensionality 67.4 0.71 0.69
Shallow layers (4–8) 63.1 0.63 0.58
PCA retaining 95% variance 67.1 0.71 0.68
Fixed steering-intensity variant Not applicable 0.69 0.65
Without spline smoothing Not applicable 0.67 0.58

Removing local directional weighting lowers TCR by 4.4 percentage points. Removing splines lowers TS from 0.69 to 0.58, indicating that graph paths and path smoothing address different problems.

Source values and reporting boundaries must be preserved. In Table 3, Mistral Geodesic and Graph-NN Chain tie on IC at 0.66, while MA is 0.55 versus 0.56; the method does not strictly lead every metric. Table 4's isotropic variant gives Qwen/Mistral TCR of 63.1/62.0, unlike the 65.1/64.3 Isomap-Euclidean values in Table 2. The paper does not sufficiently explain the setting difference, so these rows should not be merged as identical configurations.

Key Findings

  • Gains depend on deviation: Table K.1 reports top-quartile combined-TCR gains of 11.7, 11.3, and 11.1 percentage points, compared with only 1.4, 1.2, and 1.1 points in the bottom quartile.
  • Low deviation does not guarantee a win: in Table K.2, Mistral SLERP achieves TS 0.64 versus Geodesic's 0.63. Geometry estimation is most useful where linear assumptions clearly fail.
  • Persona-type curvature differences remain significant after density control, \(p<0.001\). This supports the geometric interpretation under the current sampling setup, but does not exclude all graph-construction or representation-extraction effects.

Highlights & Insights

  • Similarity prediction and actionable intervention support each other. Useful geometry should not only improve distance rankings but also generate interpretable behavioral transitions when traversed.
  • Anisotropy and path support are integrated in one framework. The former concerns local directional costs, the latter whether a direct cross-persona shortcut is supported, avoiding attribution of every gain to generic “nonlinearity.”
  • Deviation can guide whether additional geometric computation is worthwhile. Transfer to other attribute-steering tasks should validate distance scaling and behavioral gains rather than directly reuse the 1.2 threshold.

Limitations & Future Work

  • The authors test only 7–8B models. Mean persona vectors omit context-dependent variation, and BST's six constructs do not provide complete behavioral coverage. Larger models and context-conditioned representations remain to be examined.
  • Offline computation and sampling coverage constrain scalability. The appendix reports approximately 45 minutes for extraction and 15 minutes for geometry per model, but shortest paths and tangent-space estimates can be unstable in sparse regions. Synthetic validation also explicitly shows degradation under higher noise and ambient dimensionality.
  • Factor separation is not a strict two-factor controlled experiment. Comparing Mahalanobis with Isomap changes both the metric and graph path, so this staircase alone cannot establish a fully independent causal contribution from curvature. Isotropic and anisotropic comparisons on the same graph would help.
  • IC may mistake repeated wording for consistency, while MA relies on the geometry constructed by the method itself. External evaluation and human ratings mitigate these issues, but broader behavioral tasks and cross-context annotation are still needed.
  • Persona control should remain limited to authorized, transparent role applications with content-safety constraints. Better behavioral naturalness does not establish that deployment risks have been adequately evaluated.
  • vs PERSONA / Persona Vectors: these methods use trait-vector algebra for flexible activation control; this paper replaces straight endpoint interpolation with a data-supported path. The trade-off is needing a persona pool and graph rather than composing a few trait directions alone.
  • vs Isomap / Diffusion Maps: the paper borrows neighborhood structure from classical manifold learning and adds location-dependent directional weights, curvature analysis, and residual-stream intervention. Its main value is connecting geometry with persona behavior, not inventing a new shortest-path algorithm.
  • vs self-report personality questionnaires: BST focuses on situational behavioral choices, reducing dependence on model self-description. BFI-44 still participates in transition evaluation, so the method does not abandon questionnaire information entirely.

Rating

  • Novelty: 4/5 — Connects local geometric analysis with inference-time persona interpolation through a clear problem formulation.
  • Experimental Thoroughness: 4/5 — Three models, distance and steering baselines, and external evaluations provide broad evidence; factor controls and reporting differences need clarification.
  • Writing Quality: 4/5 — The narrative is clear, but behavioral-distance conventions and some cross-table results are insufficiently specified.
  • Value: 4/5 — Offers a testable geometric account of when activation-vector algebra fails, with scope still requiring expansion.