AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning¶
Conference: NeurIPS2026
arXiv: 2609.34188
Code: https://github.com/BiQiBaoWinner/AlphaPareto
Area: Reinforcement Learning
Keywords: formulaic alpha discovery, multi-objective reinforcement learning, alpha-pool semantic encoding, Pareto optimization, quantitative investing
TL;DR¶
AlphaPareto encodes the current alpha pool with a frozen LLM, conditions token-by-token MaskPPO formula search on existing signals, and uses adaptive multi-objective rewards for predictive power, ranking stability, perturbation robustness, and diversity, achieving out-of-sample ICs of 3.92%, 5.70%, and 10.10% on CSI300, CSI800, and the full market.
Background & Motivation¶
A formulaic alpha is a mathematical expression that transforms historical prices, trading volume, and other market features into a stock-selection signal. Such expressions are easier to interpret and audit than black-box predictors, but an individual factor rarely remains effective across market conditions. Practical prediction therefore often uses an alpha pool: each factor is standardized cross-sectionally and combined with fitted linear weights into a mega-alpha. Discovery should optimize not just the strength of an isolated expression, but whether adding it covers information missing from the existing pool. Genetic programming typically searches for high-IC expressions independently; AlphaGen advances synergistic discovery through pool-level rewards, but retains a mismatch between state and reward.
Specifically, the conventional RL state contains only the formula prefix under construction, whereas its reward depends on the current pool. The same completed formula can receive different evaluations when the pool is empty, already contains similar momentum factors, or lacks price-volume signals. The pool changes across episodes, but the policy cannot observe this context and must infer a search direction indirectly from changing rewards. Meanwhile, a single IC measures only average predictive correlation: a high-IC pool may still change stock rankings sharply between adjacent days, be sensitive to perturbations, or accumulate redundant factors. Higher average predictive power does not automatically resolve these structural problems.
The paper changes both what the policy observes and what it optimizes. It includes the current pool in the state and uses frozen LLM hidden representations to summarize expression structures and combination weights. It then expands the reward to four dimensions, expresses preferences as minimum target levels, and derives scalarization weights from recent reward distributions. The LLM is not a direct formula generator here; the RL policy still constructs the formulas. Core idea: make formula search explicitly pool-aware and turn multidimensional pool-quality feedback into a trainable reward through online Pareto regularization.
Method¶
Overall Architecture¶
The inputs are historical market data, future-return labels, and the current alpha pool; the outputs are a capacity-limited formula collection and its linear combination weights. Each episode first performs alpha-pool semantic encoding, then generates a candidate through pool-conditioned formula search. After a valid candidate enters the pool and the combiner is refitted, four-dimensional pool rewards are computed and adaptive Pareto scalarization drives PPO updates.
During training, future returns fit combination weights and provide the IC reward; no gradients propagate into the LLM. Prediction after discovery executes only the selected formulas, standardization, and linear combination. It requires neither future returns nor new LLM-generated factors. Dashed edges in the diagram denote training feedback rather than deployment data flow.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Current alpha pool<br/>expressions and weights"] --> A["Alpha-pool<br/>semantic encoding"]
A --> B["Pool-conditioned<br/>formula search"]
H["Historical market data"] --> B
B -->|Valid formula enters pool; refit| C["Four-dimensional<br/>pool rewards"]
Y["Training future returns"] -.->|Combiner fitting and IC| C
C --> D["Adaptive Pareto<br/>scalarization"]
D -.->|Training only: PPO update| B
C -.->|Next episode: updated pool| A
C -->|Discovery ends: retain pool and weights| O["Prediction: execute formulas<br/>standardize and combine linearly"]
Key Designs¶
1. Alpha-pool semantic encoding: expose the hidden context behind changing rewards
Conceptually, the state is augmented to include both the current formula prefix and the current alpha pool. Reward drift caused by pool evolution then becomes a state change rather than a changing reward function under the same observed state. This addresses the state–reward mismatch in discovery; it does not establish that actual financial markets become stationary.
In practice, a pool consists of complex expressions and fitted weights, and concatenating numerical inputs does not readily communicate overlapping signals, time scales, or missing information. A structured prompt supplies each expression and weight and directs the frozen LLM toward factor styles, data-source coverage, weight concentration, and potential blind spots. The appendix prompt also requests construction-style recommendations for the next factor, but the policy interface uses the last-layer hidden representation rather than directly treating these textual recommendations as new formulas.
The default encoder is Qwen-Embedding-4B, with a 2560-dimensional pool representation. It remains fully frozen during RL. The representation is recomputed from the current pool at the beginning of each episode and reused during that formula's token search. This supplies pool context without expensive joint LLM training. However, sufficient information retention in the compressed representation remains an empirical assumption, not a lossless encoding of the complete state.
2. Pool-conditioned formula search: make the next token depend on existing signals
Formulas use reverse Polish notation (RPN), arranging features and operators into token sequences in stack-evaluation order. Starting with BEG, the policy selects prices, volume, constants, time windows, or operators, while invalid action masking excludes syntactically inadmissible next steps. Choosing SEP or reaching the maximum length ends the episode, after which the full expression is parsed. Interpretability therefore comes from a restricted formula grammar and length constraint, not from post-hoc explanations of a black-box predictor.
A two-layer LSTM encodes the prefix under construction. The LSTM representation and frozen LLM pool representation are separately projected into a shared 256-dimensional space, multiplied element-wise, and mapped by another FFN to a 128-dimensional representation for the MaskPPO policy and value heads. Multiplicative fusion lets pool context modulate prefix features: the same construction progress can produce different action distributions in different pools.
Candidates are not simply ranked by standalone IC and appended to a list. A valid formula is temporarily added, combination weights are refitted for all members, and at most the specified number of factors are retained according to fitted contributions. A candidate can replace an existing member, and its reward evaluates the entire updated pool. This connects formula generation to complementarity while preventing unbounded pool growth.
Available inputs include Open, High, Low, Close, Vwap, and Volume. Operators include cross-sectional Rank, arithmetic operations, and time-series means, standard deviations, and correlations. Individual expressions can be nonlinear, but the final mega-alpha is still a linear combination; these two levels of nonlinearity should not be conflated.
3. Four-dimensional pool rewards: evaluate the updated combination beyond average predictive power
IC is the training-period average of daily cross-sectional Pearson correlations between the mega-alpha and future returns. It is neither the candidate's standalone correlation nor the IC improvement over the previous pool. The other three components assess temporal stability of stock rankings, fidelity of combined outputs under perturbations, and dispersion of pool signals across different directions.
Below, \(S\) denotes the number of evaluation days, \(M\) the updated pool size, \(\widehat{\boldsymbol\alpha}_s\) the daily mega-alpha, \(\boldsymbol p_s\) the probability vector formed by dividing stock ranks by their sum, and \(\widetilde p_i\) the share of the factor-signal covariance matrix's total eigenvalue mass assigned to its \(i\)-th eigenvalue. The three definitions can be summarized as:
RRE (Relative Rank Entropy) computes KL divergence between ranking distributions on adjacent dates, transforms it through a reciprocal, and averages the results. More consistent rankings yield higher scores. It does not measure the variance of daily IC or directly measure return stability. The initial 1 in the formula is the authors' treatment of the first evaluation day.
In PFS (perturbation fidelity score), \(\rho\) denotes Spearman rank correlation. Perturbations multiply the mega-alpha output directly; raw prices are not perturbed and factors are not recomputed. Noise is drawn with equal probability from Gaussian or Student-\(t_3\) distributions and rescaled to match the empirical cross-sectional variance of stock returns. The former approximates routine fluctuations and the latter heavy-tailed shocks. This is a proxy for output-ranking robustness, not a guarantee of reliable returns during real policy shocks.
DH (diversity entropy) normalizes the eigenvalues of the covariance matrix of individual factor signals, computes their entropy, and divides by the logarithm of pool size. More evenly distributed variance across orthogonal directions yields higher DH; redundant factors concentrate variance in fewer directions. This differs from the prompt's weight-concentration diagnostic and from simply counting expressions in a pool.
The four components form a vector reward evaluated only after a valid formula terminates and updates the pool. Incomplete formulas receive a zero vector; terminated invalid expressions receive a vector whose components are all negative one. The original DH normalization has a boundary issue for a pool of size 1, and the cache does not describe its numerical treatment during initialization; reproduction should check the implementation.
4. Adaptive Pareto scalarization: replace fixed manual weights with minimum target levels
Predictive power, stability, robustness, and diversity do not always improve together. MAP (Multi-Human-Value Alignment Palette) uses a preference vector \(\boldsymbol c\) to specify minimum desired levels for each objective, then derives nonnegative multipliers from sampled four-dimensional rewards rather than requiring users to assign fixed weights directly.
With \(\boldsymbol r_t\) denoting the four-dimensional reward and \(T\) the recent sample count, the empirical dual optimization and resulting PPO reward are:
Unlike assigning a large IC weight at the outset, this design interprets target thresholds against the current reward distribution and produces an online-adaptable scalar reward. It does not eliminate subjective preferences; it expresses them as target levels in the original units of each metric. Infeasible thresholds and multiplier stability remain relevant application concerns.
The authors adapt MAP from an offline setting with fixed transitions to an online procedure. A random policy first collects warm-up rewards for initial multiplier estimation. Recent reward vectors then periodically update the empirical dual solution alongside standard PPO policy and value updates. The paper draws on MAP's Pareto argument, but this does not imply that finite-sample online PPO necessarily recovers the complete efficient frontier, nor does it demonstrate a policy collection spanning all preferences.
A Worked Example¶
An episode can be understood as finding the next missing piece for an existing pool. Suppose the pool mainly contains price transformations. Semantic encoding compresses those expressions and weights into context, and the policy searches token by token rather than asking the LLM to supply a price-volume formula. This illustrates the mechanism, not a specific discovery trajectory reported in the paper.
When the policy completes a valid expression, the system temporarily adds it, refits linear weights, and may remove an older factor with a smaller contribution. It then evaluates correlation between the updated mega-alpha and future returns, changes in stock rankings between adjacent dates, ranking fidelity under mixed noise, and spectral entropy of the members' signal covariance matrix. These jointly assess the quality of the pool update.
The current multipliers turn this vector into a PPO reward. In the next episode, pool encoding changes with the members and weights, so the policy receives both training feedback and direct context for its next search. Deployment retains formulas and combination weights, not training labels as prediction inputs.
Loss & Training¶
Training uses MaskPPO: the policy head selects tokens under action masking, the value head estimates scalarized returns, and the LLM stays frozen. The finite-episode discount factor is 1, maximum expression length is 15, and pool capacity is selected on validation data from 10 and 20.
The appendix specifies LSTM hidden dimension 128, dropout 0.1, batch size 128, Adam learning rate \(5\times10^{-4}\), and combination-model \(\ell_1\) regularization strength \(5\times10^{-3}\). Total steps are 250,000 and 300,000 for capacities 10 and 20, respectively. The separately listed optimization steps/tolerance of 10,000/500 should not be mistaken for the overall search budget.
IC preference thresholds are 0.090, 0.085, and 0.250 for CSI300, CSI800, and Market; RRE/PFS/DH thresholds are 0.95/0.93/0.80. These are training targets, not out-of-sample results, and do not establish test ICs of 9%, 8.5%, or 25%.
Key Experimental Results¶
Main Results¶
The target is the future 20-day return. Training, validation, and testing cover 2013/07/01–2023/06/30, 2023/07/01–2024/06/30, and 2024/07/01–2025/06/30, respectively. The table selects representative baselines from original Table 1; parentheses contain standard deviations across random seeds, all expressed as percentages. The cache's “55” is duplicated mathematical text; Appendices F and K explicitly state 5 seeds.
| Method | CSI300 IC | CSI800 IC | Market IC |
|---|---|---|---|
| Alpha158 | 2.99% (—) | 4.77% (—) | 4.04% (—) |
| AlphaAgent | 0.51% (0.34%) | 0.48% (0.46%) | 1.10% (0.59%) |
| R&D-Agent-Quant | 2.39% (0.96%) | 2.45% (0.73%) | 1.48% (0.91%) |
| AlphaGen | 3.69% (0.66%) | 5.07% (0.63%) | 8.44% (1.09%) |
| AlphaQCM | 0.50% (0.95%) | 4.79% (0.34%) | 9.16% (4.22%) |
| AlphaPareto | 3.92% (0.46%) | 5.70% (0.87%) | 10.10% (1.01%) |
Mean IC improves by 0.23, 0.63, and 0.94 percentage points over the strongest baseline in each market. However, the paired comparison against AlphaQCM on Market is not significant: the paper reports \(t=0.65\) and \(p=27.60\%\). On CSI300, \(p=8.34\%\) against AlphaGen meets only the authors' 10% significance level; the comparisons should not all be described as significant at 5%.
Ablation Study¶
Table 4 compares the two components on the same prediction task; parentheses again denote standard deviations across seeds. LLM means pool semantic encoding, while MORL combines four-dimensional rewards with MAP optimization rather than changing an isolated loss term.
| Config | CSI300 IC | CSI800 IC | Market IC |
|---|---|---|---|
| No LLM, no MORL (AlphaGen) | 3.69% (0.66%) | 5.07% (0.63%) | 8.44% (1.09%) |
| LLM only | 2.37% (0.68%) | 5.04% (0.42%) | 9.68% (0.83%) |
| MORL only | 3.99% (1.27%) | 5.54% (1.03%) | 9.63% (0.80%) |
| Full AlphaPareto | 3.92% (0.46%) | 5.70% (0.87%) | 10.10% (1.01%) |
Key Findings¶
- Pool encoding is not universally beneficial: LLM-only underperforms AlphaGen on CSI300 but increases Market IC from 8.44% to 9.68%. The full model does not have the highest mean in every ablation either: on CSI300 it trails MORL-only by 0.07 percentage points while reducing standard deviation by 0.81 percentage points.
- Appendix H reports 4.92% (0.52%) for equal-weight four-objective rewards on CSI800 versus 5.70% (0.87%) for MAP. Adaptive weights improve the mean, not variance in this comparison. Appendix I's matched random-noise representation reaches only 4.22% (1.62%), arguing against attributing encoding benefits simply to noise regularization.
- In the full-market monthly long-only backtest, AlphaPareto achieves annualized return 58.92%, IR 2.62, Sharpe 2.59, and maximum drawdown −18.12%. AlphaGen's drawdown is −15.91%, so AlphaPareto is not best on every risk metric. The appendix specifies transaction costs of 15 basis points per side and a one-year test period.
- On S&P 500, AlphaPareto reaches 7.65% (0.13%) versus AlphaQCM's 7.18% (0.84%). This is a reevaluation in another market, not zero-shot transfer of an A-share-trained policy to U.S. equities.
Highlights & Insights¶
- Reward non-stationarity can arise from omitted state, not just uncontrollable environmental drift. Exposing the pool context behind the reward is more direct than only adding an uncertainty-based exploration bonus.
- An LLM need not generate the final answer to be useful. Here it supplies frozen context encoding, while an RL searcher with executable syntax and numerical feedback constructs expressions.
- Multi-objective preferences can be specified as targets in the metrics' original units. MAP still requires subjective thresholds, but these are easier to check against application requirements than fixed weights without a clear scale interpretation.
Limitations & Future Work¶
- The authors do not systematically study prompt sensitivity, and the linear combiner may miss richer factor interactions. Stronger encoders are not necessarily better: Table 3 shows a non-monotonic relationship with model scale.
- Frozen hidden representations do not guarantee complete pool information; PFS perturbs only combined outputs, and RRE constrains only ranking changes. These are context and robustness proxies, not substitutes for market-state, raw-input perturbation, and live-trading evaluation.
- MAP receives an additional comparison only against simple equal-weight scalarization, insufficient to establish superiority over mature MORL optimizers. Online solving, threshold feasibility, and DH boundary handling during initialization warrant further inspection.
- Five seeds and a one-year test period limit statistical and economic extrapolation. LLM-generation baselines also undergo candidate-count and expression-format adjustments, so results rank methods under this experimental protocol rather than under all possible configurations.
Related Work & Insights¶
- vs AlphaGen: Retains the RPN and MaskPPO search backbone, but replaces prefix-only state and IC-only reward with pool-conditioned state and four-objective feedback.
- vs AlphaQCM: Uses distributional value estimation and exploration bonuses to mitigate changing rewards; AlphaPareto directly supplies the pool context behind the changes. Its Market mean advantage does not pass the authors' paired significance test.
- vs AlphaAgent / R&D-Agent-Quant: These generate candidate factors with LLMs; AlphaPareto freezes the LLM for pool encoding and lets RL select formula tokens. LLM parameter count alone does not fairly measure search capability or computational budget across these roles.
- vs MAP / AlphaEval: MAP provides preference thresholds and dual scalarization, while AlphaEval emphasizes multidimensional factor evaluation. AlphaPareto connects these ideas to online pool-level search, but modifies PFS's perturbation location and aggregation.
Rating¶
- Novelty: 4/5 — Clearly combines pool-context completion with multi-objective feedback; the optimizer primarily builds on existing MAP.
- Experimental Thoroughness: 4/5 — Includes component ablations, noise controls, backtesting, and U.S. evaluation, but the test window and MORL comparisons remain limited.
- Writing Quality: 4/5 — Clearly motivates state mismatch; compressed-state sufficiency, initialization boundaries, and online theoretical guarantees need fuller explanation.
- Value: 4/5 — Offers reusable context encoding and preference optimization for interpretable formula search, without guaranteeing investment returns.