Skip to content

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

Conference: NeurIPS2026
arXiv: 2609.38487
Area: Image Restoration
Keywords: image quality assessment, multidimensional observer model, perceptual space, behavioral modeling, ventral visual stream

TL;DR

The paper represents images as Gaussian distributions in a low-dimensional perceptual space and fits human choices through noisy candidate–reference distance comparisons; the CORnet-S Multi observer reaches 95% of its own above-chance performance at 3, 9, and 96 dimensions on BAPPS, PieAPP, and NIGHTS, respectively, suggesting task-dependent perceptual structures.

Background & Motivation

Image quality assessment (IQA) involves more than measuring pixel errors: it helps denoising, super-resolution, generation, and rendering systems select outputs. PSNR and SSIM measure distortions through hand-designed relationships, while LPIPS, DISTS, and DreamSim use learned visual features to improve agreement with human judgments. Predicting which image is better, however, does not establish what humans compare. Whether color, texture, and object structure occupy distinct dimensions, how many dimensions support a judgment, and how these dimensions change with the task remain questions requiring a testable model.

Assigning a single quality score to each image collapses these questions in advance. Moreover, the same observer may change their answer across repetitions, and different observers may disagree. Treating the average preference as an unequivocally correct label discards this reproducible probabilistic structure. The paper therefore starts from triplet two-alternative forced choice (2AFC): a reference and two candidates appear together, and an observer selects the candidate perceptually closer to the reference. The target is the distribution of repeated choices, not an individual response.

To make the latent space interpretable, the authors do not train an arbitrary deep embedding from pixels. They freeze feature encoders aligned with the ventral visual stream, learn only an affine projection and image-dependent noise, and systematically vary the latent dimensionality. Core idea: separate multidimensional noisy image representations from reference-relative distance decisions, constrain the perceptual space through behavioral likelihood, and diagnose task-specific representational complexity through predictive saturation rather than a prescribed dimensionality.

Method

Overall Architecture

The model takes a reference and two candidate images and outputs the probability of choosing the first candidate, together with means, covariances, and distance parameters available for analysis. Each image passes through a frozen visual encoder and an affine projection into a shared low-dimensional space. A covariance predictor generates its noise distribution from the projected mean. Each simulated trial independently samples the three image representations, compares candidate distances to the same reference sample, and repeated simulations yield a choice probability.

This is a statistical observer model, not a new image restoration network or a direct measurement of neural activity. Human choice frequencies supervise the model head during training; inference requires only the image triplet and random sampling, without human labels. The research procedure additionally trains models across dimensionalities, encoders, and tasks, identifies where performance saturates, and examines the visual content associated with learned dimensions.

The central mechanism is distributional geometry and stochastic decision making rather than coordination among multiple neural modules, so the parameter relationships are not forced into a network diagram. The following four designs explain the representation source, noise structure, choice mechanism, and dimensionality diagnosis.

Key Designs

1. Visually constrained affine projection: keep latent dimensions traceable to visual features

A low-dimensional projection learned directly from pixels would be simple but difficult to interpret as a human visual quality representation. The primary encoder is CORnet-S, whose four computational blocks have hierarchical correspondence with neural responses in V1, V2, V4, and IT. Spatial average pooling produces vectors of 64, 128, 256, and 512 dimensions; the Multi variant concatenates them into 960 dimensions. The encoder remains frozen, and either individual layers or their concatenation can supply the observer model.

An affine projection, rather than another arbitrary nonlinear encoder, directly controls latent dimensionality while preserving a traceable relationship between each latent dimension and the original channels. With frozen features denoted by \(B(x)\), the central relationship is:

\[ \boldsymbol{\mu}(x)=\mathbf{W}B(x)+\mathbf{b},\qquad \mathbf{W}\in\mathbb{R}^{N\times D}. \]

Varying \(N\) compares representational capacity under the same visual source. Each projection row specifies how a perceptual dimension weights visual channels; BRODEN channel concepts can subsequently indicate whether it emphasizes color, texture, or objects. A contrast pyramid approximating retinal/LGN spatial processing and VGG features test whether the findings depend on CORnet-S. Here, biological grounding constrains feature sources and projection form; it does not establish that the learned coordinates are actual neural coordinates.

2. Image-dependent covariance field: represent perceptual ambiguity with distributions rather than uniform noise

Images with nearby average perceptual positions may still differ in judgment stability. The model therefore learns a Gaussian distribution for each image: its mean specifies location, and its covariance describes directional uncertainty and noise correlations. Covariance is not a constant shared by all images but a function of the projected mean. Consequently, noise prediction depends on the compressed representation and cannot independently recover features discarded by the projection.

To ensure valid covariance matrices, an MLP outputs the lower-triangular entries of a Cholesky factor. Its two hidden layers each have 64 units and use ReLU; diagonal entries undergo softplus followed by addition of \(10^{-4}\). This avoids negative eigenvalues from direct matrix prediction and makes the representation vary continuously with the input.

\[ Z(x)\sim\mathcal{N}\!\left(\boldsymbol{\mu}(x),\boldsymbol{\Sigma}(x)\right),\qquad \boldsymbol{\Sigma}(x)=\mathbf{L}(x)\mathbf{L}(x)^{\top}. \]

This noise jointly models within-observer variability and between-observer differences without separately identifying their contributions. Low-dimensional visualizations display covariance ellipses at different mean-space positions to examine content-dependent ambiguity. These changes are not direct measurements of human neural noise.

3. Noisy distance comparison: turn multidimensional representations into human choice probabilities

Each simulated trial draws one sample from the reference distribution and each candidate distribution. Both candidate distances must use the same reference sample: reference noise jointly affects the comparisons, so the two distances cannot be treated as fully independent random variables. The model uses a learned weighted Minkowski distance, with weights compensating for dimensional scales and an exponent controlling how dimensional contributions combine.

\[ d_i=\left(\sum_{m=1}^{N}\alpha_m\left|z(x_i)_m-z(r)_m\right|^p\right)^{1/p},\qquad \alpha_m\geq0,\quad p\geq1. \]

At \(p=1\), a dimension's local contribution does not depend on the magnitudes of other dimensions; \(p=2\) gives a more integrated Euclidean combination. The exponent is learned for each task rather than fixed. Candidate 0 is selected when its sampled representation is closer to the reference, not necessarily when its mean is closer. A probability estimate follows from \(K\) joint simulations:

\[ P=\Pr(D_0<D_1)\approx\frac{1}{K}\sum_{k=1}^{K}\mathbf{1}\!\left\{d_0^{(k)}<d_1^{(k)}\right\}. \]

Nearby candidates can therefore elicit unstable choices, whereas widely separated candidates produce more consistent decisions. Crucially, probability emerges from repeated noisy comparisons rather than an additional sigmoid applied to a deterministic distance difference. The appendix reports similar prediction with a sigmoid alternative, but that smooth comparator can account for probability itself, encouraging narrow, nearly uniform covariance fields and weakening the interpretability of the noise structure.

4. Above-chance saturation diagnosis: operationalize how many dimensions are needed

Dimensionality cannot be established from visualization or a final prediction score alone. For each encoder–dataset pair, the authors sweep \(N\) and use the model's own best performance as the reference. They identify the smallest dimensionality explaining more than 95% of above-chance performance. With a 2AFC chance baseline of 0.5, the definition is:

\[ n^{*}=\min\left\{N:\frac{m(N)-0.5}{m^{*}-0.5}>0.95\right\}. \]

Here, \(m(N)\) is performance at the corresponding dimensionality, and \(m^{*}\) is the best swept performance without restricting dimensionality. This criterion measures the dimensions required for behavior explainable by this model. It is not 95% of the human theoretical ceiling or a direct measurement of the brain's absolute intrinsic dimensionality. A weaker encoder may saturate early because it lacks information, so its peak prediction score must also be considered.

To test whether low dimensionality merely reflects natural-image compressibility, the authors introduce a frequency-space control. After conversion to YUV, they retain the largest \(K\) DCT coefficients per channel, reconstruct the images, and train an observer head. Three channels retain \(3K\) coefficients in total, so matching the budget of an \(N\)-dimensional perceptual representation requires \(K=N/3\). The control uses a square projection without further dimensionality reduction. It tests whether sparse frequency reconstruction supports comparable behavioral prediction, not whether every possible image compression scheme fails.

Loss & Training

For triplet \(i\), \(y^{(i)}\) is the human fraction choosing candidate 0 and \(J^{(i)}\) is the observation count. The model maximizes a count-weighted binomial log likelihood, equivalently minimizing:

\[ \mathcal{L}=-\sum_i J^{(i)}\left[y^{(i)}\log P^{(i)}+(1-y^{(i)})\log\left(1-P^{(i)}\right)\right]. \]

Training updates the projection, covariance MLP, distance weights, and exponent without fine-tuning the visual encoder. Sampling uses reparameterization: standard Gaussian noise is transformed by the Cholesky factor and added to the mean. A straight-through estimator (STE) propagates gradients through hard distance comparisons. Defaults are Adam with learning rate \(10^{-3}\), cosine annealing, batch size 256, 50 epochs, and \(K=64\) Monte Carlo samples per comparison.

The appendix also extends the observer to MOS/DMOS ratings. Expected candidate–reference distribution distance supplies the quality signal, and a monotonic score head maps it to a badness score. Training uses the full KADID-10k dataset with an L1 plus correlation loss and 512 samples, followed by cross-dataset evaluation. This differs from triplet likelihood training: rating results are not direct outputs of the same model without retraining.

Key Experimental Results

Main Results

BAPPS and PieAPP concern low-level distortion judgments. NIGHTS uses generated images with identical prompts and different seeds, introducing structural and semantic differences. Agreement on the first two datasets averages consistency between human choice fractions and the model's binary choices; it is not ordinary accuracy. For fair comparison with deterministic metrics, model probabilities are binarized by majority vote. KL evaluates unbinarized choice distributions and is lower-is-better; NIGHTS has binary labels and uses accuracy.

The following results come from main-text Table 1 with CORnet-S Multi as the observer encoder. Representative baselines are retained without implying universal superiority.

Dataset and metric Ours DreamSim DISTS PieAPP metric CVVDP
BAPPS agreement ↑ 0.688 0.683 0.678 0.629 0.645
PieAPP agreement ↑ 0.710 0.719 0.690 0.717 0.591
NIGHTS accuracy ↑ 0.859 0.957 0.860 0.629 0.703
BAPPS KL ↓ 0.170 Not applicable Not applicable 0.236 0.265
PieAPP KL ↓ 0.088 Not applicable Not applicable 0.083 0.985

The BAPPS agreement ceiling is 0.797 and independent human inter-observer agreement is 0.731; the model's 0.688 reaches neither. On PieAPP, inter-observer agreement is 0.687 versus model agreement of 0.710. This is not contradictory: binary model choices can agree more often than two independent noisy observers. On NIGHTS, the model still trails DreamSim by 0.098, so competitive performance with several metrics should not be described as the best semantic judgment performance.

Ablation Study

Main-text Table 2 reports peak performance and saturation dimensionality for different visual sources. This is an encoder and dimensionality analysis, not a module-removal ablation. Each cell is “peak performance / \(n^{*}\)”; the first two dataset columns use agreement and the final column uses accuracy.

Feature encoder BAPPS PieAPP NIGHTS
Contrast Pyramid 0.645 / 3 0.647 / 2 0.703 / 7
CORnet-S V1 0.685 / 4 0.701 / 6 0.747 / 8
CORnet-S V2 0.689 / 3 0.712 / 5 0.793 / 14
CORnet-S V4 0.688 / 4 0.714 / 8 0.855 / 48
CORnet-S IT 0.677 / 6 0.702 / 14 0.853 / 96
CORnet-S Multi 0.688 / 3 0.710 / 9 0.859 / 96

The following additional CORnet-S Multi controls are read under their respective protocols rather than mixing metrics across settings.

Analysis config Dataset and metric Result Control and interpretation
Perceptual space \(N=6\) BAPPS agreement 0.686 Main-text Table 3
DCT \(K=2\) per channel BAPPS agreement 0.566 Matched coefficient budget to \(N=6\); lower by 0.120
DCT \(K=200\) per channel BAPPS agreement 0.684 600 coefficients in total; close to the 6-dimensional perceptual model
PCA projection BAPPS agreement / \(n^{*}\) 0.664 / 12 Learned projection: 0.688 / 3; appendix Table 5
BAPPS training → PieAPP PieAPP agreement 0.707 In-domain training: 0.710; appendix Table 6
BAPPS training → NIGHTS NIGHTS accuracy 0.825 In-domain training: 0.859; appendix Table 6

Key Findings

  • Low-level tasks primarily need early-to-intermediate visual features; higher-level IT is not always better. BAPPS IT agreement is 0.677 versus V2's 0.689, whereas NIGHTS V1 accuracy of 0.747 is substantially below V4's 0.855.
  • Low dimensionality does not mean that retaining the same number of frequency coefficients works equally well. The matched-budget Multi agreement gap is 0.120, although this finding only rules out the paper's DCT control.
  • The two-dimensional BAPPS visualization mainly reveals texture/spatial-frequency and color directions. Higher-dimensional BRODEN channel analysis shows more IT part and object information for NIGHTS, while BAPPS and PieAPP remain oriented toward color and texture.
  • The appendix reports average off-diagonal dimensional correlation coefficients below 0.25. This supports low redundancy, but low Pearson correlation is not statistical independence and does not establish strict orthogonality of the projection matrix.
  • The source contains numerical and descriptive inconsistencies: the PieAPP ceiling is 0.766 in Table 1 but 0.765 in the metric paragraph; Multi BAPPS KL is 0.170 in Table 1 but 0.167 in appendix Table 5; the main-text “3–8 dimensions” low-level summary does not fully match Table 2 entries of 2, 9, and 14. This note retains the respective table values rather than reconciling them without evidence.

Highlights & Insights

  • Make disagreement a research target: the covariance field is not merely an error term added to improve scores; it describes content-dependent judgment ambiguity. Stochastic comparison lets behavioral data jointly constrain mean locations and uncertainty.
  • Use frozen hierarchical features for comparable explanatory questions: connecting the same observer head to V1, V2, V4, and IT distinguishes the information needed for low-level distortions and high-level similarity. Quality metric designers should not assume object-recognition features are optimal for every restoration task.
  • Read saturation together with peak performance: Contrast Pyramid saturates at only 7 dimensions on NIGHTS but peaks at 0.703. Low-dimensional saturation alone is not evidence of greater human fidelity or an intrinsically simple task.

Limitations & Future Work

  • The authors acknowledge scale non-identifiability from jointly learning means and covariances: proportional changes in location and noise scale can preserve choice probabilities, making absolute distribution scales difficult to compare across triplets. Neural recordings or cross-content judgments could add constraints.
  • Gaussian distributions, smooth covariances, Minkowski distance, and feature pooling are modeling assumptions. Competitive behavioral fitting supports the model as an analytical tool, not proof that human vision executes these computations.
  • Saturation dimensionality depends on encoder, dataset, sweep range, and the 95% criterion; it is not a universal neural constant. NIGHTS semantic judgments also concern a particular generated-image distribution, not every natural-scene task.
  • The model currently combines between-observer differences with repeated-trial noise, and BAPPS validation triplets have only 5 observations. Denser repeated measurements and individual-level models could test covariance stability.
  • The appendix Bradley–Terry control retains a multidimensional embedding and changes both noise and distance, so it is not a pure covariance-removal ablation. Its average BAPPS → NIGHTS cross-domain advantage is 0.022, showing that the observer mechanism does not guarantee stronger generalization.
  • vs LPIPS / DISTS / DreamSim: these metrics primarily optimize perceptual prediction, whereas this paper emphasizes the supporting space's dimensionality and noise structure. It leads the BAPPS main table but substantially trails DreamSim on NIGHTS; interpretability does not remove performance boundaries.
  • vs Bradley–Terry / Thurstone: classical psychometric models place randomness in a one-dimensional quality score; this observer generates score distributions from multidimensional noisy representations and distance decisions. The appendix explicitly notes that an embedding-augmented BT model can also study dimensionality, so this analytical capability is not exclusive to the proposed observer.
  • vs frequency-sparse representations: DCT controls preserve information for image reconstruction, whereas behavioral projections preserve information needed for judgments. Stronger compression baselines and task-conditioned projections could test whether low-dimensional findings persist in restoration evaluation; this is a proposed extension, not an established result of the paper.

Rating

  • Novelty: 4/5 — Combines multidimensional distributions, visual hierarchy constraints, and behavioral dimensionality diagnosis rather than only improving a quality metric.
  • Experimental Thoroughness: 4/5 — Includes three triplet settings, frequency controls, projection comparisons, cross-domain evaluation, and rating extensions, but lacks direct neural validation.
  • Writing Quality: 4/5 — The mechanism is clear, with a few numerical and summary-boundary inconsistencies between the main text and appendix.
  • Value: 4/5 — Offers an interpretable tool for perceptual quality modeling and informs restoration and generation evaluation, but should not be treated as a universal brain model.