Skip to content

Towards Unified Dynamic Face Landmark Detection

Conference: NeurIPS2026
arXiv: 2608.10346
Area: Human Understanding
Keywords: face landmark detection, unified annotation templates, part-anchored positions, dynamic queries, cross-template generalization

TL;DR

The paper describes landmarks using a face part and a normalized sequence position, then uses image-conditioned queries and iterative decoding to train one model across annotation formats and predict landmarks on demand; the unified ViT-B achieves full-set NME of 4.05, 2.80, and 1.02 on WFLW, 300W, and AFLW-19, respectively, offering unified training and interfaces rather than superiority on every metric.

Background & Motivation

Face landmark detection usually hard-codes its output format: AFLW-19 requires 19 points, 300W requires 68, and WFLW requires 98. Methods such as SLPT and DTLD improve coordinate accuracy but typically still require separate models or dataset-specific regression heads. When a downstream application needs only the eyes and mouth corners, a fixed model still produces its predefined outputs; when it needs denser contours, the existing head cannot simply add positions. Here, dynamic means changing the number and layout of queries at runtime, not modeling video sequences.

This separation of formats does not fully reflect annotation semantics. Most landmarks across datasets lie on the same structures, including eyes, eyebrows, lips, and the face contour, but differ in sampling density, endpoints, and local definitions. LAB already shares visual representations through part boundaries, while CLD queries 2D landmarks using positions on a canonical 3D face. This paper aims to establish a queryable interface directly from existing sparse 2D annotations, without canonical 3D mappings. The challenge is not merely variable output length: points from different templates must acquire shared meanings that can be supervised.

Core Idea: register heterogeneous landmarks as semantic positions on face-part curves, encode each position as a query, and localize the requested points directly from the image rather than maintaining multiple fixed output heads or relying only on geometric interpolation of existing predictions.

Method

Overall Architecture

The inputs are a cropped face image and a set of requested face-part-name/FPALP-position pairs; the outputs are their 2D coordinates. A unified template and mappings from dataset landmarks are established before training. At runtime, the model applies part-position encoding, image-conditioned initialization, and cross-modal iterative refinement. The image encoder provides shared features, so changing the query layout does not require retraining an output head.

The mechanism separates two questions: the unified template specifies how supervision from different datasets is represented, while dynamic queries specify which points to predict in a particular inference call. A training image supervises only the points annotated by its source dataset; combining datasets does not provide ground-truth coordinates for every template on every image. Solid arrows below represent inference data flow, and dotted arrows represent training supervision. FPALP registration is a predefined semantic mapping, not a measurement of curve length in each image.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    T["Dataset templates or target layout"] --> A["Part-Anchored Positions"]
    A --> B["Part-Position Encoding"]
    I["Face image"] --> F["Image encoder"]
    B --> C["Image-Conditioned Initialization"]
    F --> C
    C --> D["Cross-Modal Iterative Refinement"]
    F --> D
    B --> D
    D --> O["2D coordinates of queried points"]
    G["Training ground-truth coordinates"] -.->|Gaussian heatmaps and PossLoss| C
    G -.->|Wing Loss at each stage| D

Key Designs

1. Part-Anchored Positions: replace fixed landmark IDs with shared semantic addresses on curves

A dataset's landmark 17 generally has no direct meaning in another dataset, whereas the middle of the face contour can be interpreted across templates. The authors align and cluster dataset face templates into a unified template, then assign points to practitioner-defined part curves. Appendix A.15 explains that clustering is initialized from a medium-density template, with restrictions on large center movements. Newly introduced points that do not fit existing clusters are flagged as new clusters or parts rather than forcibly merged. The reported mean intra-cluster distance across parts is 2.22 pixels, indicating reliance on approximate alignment rather than inherently identical semantic coordinates across all datasets.

A Face Part-Anchored Landmark Position (FPALP) normalizes a landmark's position in an ordered part-template sequence:

\[ \mathit{FPALP}_{l,p}=\frac{\mathit{pos}_{l,p}}{N_p-1}. \]

Here, \(N_p\) is the part-sequence length, with zero-based ordering. Appendix A.9 additionally discusses the offset between a dataset's starting point and the unified template's starting point, so cross-template registration should not be reduced to independently numbering and normalizing each dataset's points. This quantity describes semantic progression in a template sequence, not the cumulative arc-length fraction of the actual face curve in the current image. Perspective, expressions, and nonuniform sampling can make those quantities differ.

For closed curves such as eyes and lips, the starting point is duplicated at the sequence end, making 0 and 1 correspond to the same physical endpoint and explicitly representing closure. The practitioner jointly defines the part name and its endpoints. One physical point may belong to overlapping parts and receive different FPALPs under those definitions. Non-interpolatable points such as pupils are also listed as separate parts in the mapping tables, so the interface should not be interpreted as placing every point on a single continuous face contour.

The representation decouples training labels from output-head size, but only when target points can be reliably mapped to parts and their ordering. A new FPALP can be encoded without guaranteeing accurate localization at every new position; registration error and training-template coverage still matter.

2. Part-Position Encoding: specify both the part and the position within it

A scalar between 0 and 1 cannot distinguish the middle of the left eye from the middle of a lip contour. The model encodes the FPALP with a ReLU MLP, encodes the part name with pretrained SentenceBERT, and adds the equally sized vectors to obtain an image-agnostic landmark encoding. The position provides a local address and the part text provides its identity; together they distinguish queries on different structures.

The authors choose text encodings over freely learned part embeddings and hypothesize that pretrained language representations may capture part layouts and expression-related relationships. The ablation does show better results with SentenceBERT, but it establishes the effectiveness of this representation choice in the evaluated setting, not that the encoder demonstrably contains those specific geometric relationships. The current framework is not an open-vocabulary part-discovery system: Appendix A.10 explicitly assumes that queried parts are defined in training data, leaving automatic discovery of new parts to future work.

The interface can load different numbers of part-position pairs or query only a local structure. Nevertheless, text descriptions must correspond to established part definitions and registration conventions. An unaligned natural-language phrase cannot be treated as an annotation protocol with already guaranteed accuracy.

3. Image-Conditioned Initialization: use one spatial attention map for query features and coarse coordinates

The image-agnostic encoding specifies the target but not its location on the current face. The authors take its dot product with image features and apply softmax over spatial positions, producing a separate attention map for each query. Attention-weighted averaging of image features gives the initial landmark query; the same weighting of image-space feature-grid coordinates gives the initial 2D coordinate. The decoder therefore starts with image-conditioned features and coarse positions rather than fixed coordinates or random vectors.

This initialization links semantic and visual localization: the same left-eye progression query can acquire different coarse coordinates under different face poses. During training, ground-truth points are converted into 2D Gaussian heatmaps, and PossLoss supervises the attention maps, rather than learning spatial weights solely through final coordinate regression. Heatmap supervision is training-only; inference does not require ground-truth coordinates.

Weighted-average coordinates can still be displaced by occlusion or multiple high-response regions. Initialization is therefore not the final detection result, but an adjustable starting point for structural reasoning and local image sampling.

4. Cross-Modal Iterative Refinement: alternate landmark relationships, local visual evidence, and query semantics

Each decoder layer first applies self-attention among landmark queries to exploit relationships between eyes, nose, lips, and other structures. Deformable attention then samples local evidence from image features around the preceding coordinate predictions. Queries subsequently cross-attend to the image-agnostic part-position encodings to realign with their target definitions. A feed-forward network produces updated queries, and an MLP regresses offsets from the previous coordinates. The default model repeats this process for three layers, carrying queries and coordinates from one layer to the next.

This order explains why the method is more than a learned interpolation curve: a new query can read current image evidence and draw on other landmarks' structural information, while renewed access to its part-position encoding helps limit semantic drift. Each deformable-attention head samples 4 features per image-feature level for each query, avoiding repeated dense retrieval over every image position.

Dynamic querying still has computational limits. Landmark self-attention becomes more expensive as query count grows, with quadratic pairwise-relation costs in a standard dense implementation; image sampling and coordinate regression also require computation for added queries. The paper's arbitrary-number claim means that a fixed output-head size does not constrain the interface, not that infinitely many points, constant computation, or unlimited accuracy are possible. Adding or removing other queries also changes the self-attention context, so a point's prediction should not be assumed invariant to the query set without evaluation.

A Worked Example

In the appendix mappings, 300W's face contour uses \(0/32,2/32,\ldots,32/32\), whereas WFLW uses \(0/32,1/32,\ldots,32/32\). Thus, the ninth contour point in 300W and the seventeenth in WFLW both map to \(16/32\), sharing the meaning of a middle-contour query. This demonstrates alignment of template addresses, not identical annotation geometry for every face in both datasets.

If inference requires only five contour points, the interface can load five face-contour queries at 0, 0.25, 0.5, 0.75, and 1. The model encodes their addresses, produces individual attention maps and coarse coordinates, lets the five queries interact and refine their locations, and returns only five coordinate pairs. This illustrates interface use; it is not a separately measured five-point accuracy result from the paper.

For a closed eye contour, endpoint duplication requires care: 0 and 1 represent the same endpoint and should not be counted as different physical landmarks. Intermediate positions without direct training supervision can be queried, but their errors require held-out-point experiments or actual new annotations; interface operability does not establish accuracy.

Loss & Training

Coordinate supervision applies to initialization and every decoder output. The paper gives the following coordinate loss:

\[ \mathcal{L}=\sum_{\mathit{\mathrm{dec}_{i}}=0}^{n_{\mathrm{dec}}}\text{WingLoss}(C_{\mathrm{dec}_{i}},C_{\mathrm{GT}}). \]

Attention initialization additionally uses PossLoss, with weighting and temperature parameters of 2 and 0.1. The displayed equation is the paper's multistage coordinate objective. It is not expanded into a complete joint objective with unknown coefficients, nor are exact Wing Loss or PossLoss expressions reconstructed from damaged cache formatting.

Unified training applies dataset-level oversampling to give each template approximately equal sample exposure per epoch. Each batch comes from one dataset, ensuring consistent query counts and tensor shapes. The default feature dimension is 256, with 3 decoder layers and 8 attention heads per layer. ViT-B uses FaRL pretraining and 224ร—224 inputs; ResNet uses 256ร—256 inputs. Face boxes are enlarged by 10%, with rotation, scaling, flipping, and translation augmentation.

Training runs for 32 epochs on an A100 40GB GPU with batch size 16, Adam learning rate \(10^{-4}\), and weight decay \(10^{-5}\). The learning rate becomes \(10^{-5}\) from epoch 25, while image and text encoders use one-tenth of the current learning rate.

Dataset Adapters are an additional specialization comparison, not a default component of every unified-model experiment. The unified network is trained and frozen first; rank-4 LoRA modules are then trained separately for each dataset, only in the final decoder layer's cross-attention and FFN, for 5 epochs at learning rate \(10^{-5}\). These adapters are used only for the main comparisons in ยง4.1, not for the remaining experiments.

Key Experimental Results

Main Results

NME is the average Euclidean landmark error divided by a normalization scale and multiplied by 100. WFLW and 300W use inter-ocular distance; AFLW-19 uses the face-box diagonal. FR10 is the proportion of WFLW samples whose NME exceeds 10%. Lower is better for every value below, but values with different normalizations should not be directly compared across datasets.

Method WFLW NMEio WFLW FR10 300W Common 300W Challenge 300W Full AFLW-19 NMEdiag
STAR Loss 4.02 2.32 2.52 4.32 2.87 Not reported
PossLoss 4.07 2.12 2.51 4.21 2.84 Not reported
MCUDN 4.15 Not reported 2.42 4.33 2.79 1.47
Unified ViT-B 4.05 2.38 2.47 4.25 2.80 1.02
Ours + Dataset Adapters 4.02 2.19 2.43 4.19 2.76 1.01

Source: Table 2. Prior-method results are taken from their authors, with different backbones, pretraining, and training conditions, so this is not a fully controlled architectural comparison. Faceptor additionally uses data from other tasks and reports AFLW-19 NME of 0.95; the paper includes such generalist methods for reference only.

The unified model does not win on every metric: WFLW NME of 4.05 is higher than STAR Loss's 4.02, FR10 of 2.38 is higher than PossLoss's 2.12, and 300W Full NME of 2.80 is slightly higher than MCUDN's 2.79. Even with adapters, WFLW FR10 of 2.19 does not beat 2.12. The paper's claim of consistently surpassing prior methods is therefore stronger than the table supports. The defensible conclusion is competitive accuracy alongside unified training and dynamic outputs.

Ablation Study

Training data AFLW-19 NMEdiag 300W NMEio WFLW NMEio COFW NMEio WFLW68 NMEio COFW68 NMEio
300W 2.18* 3.01 6.32* 3.81* 6.08 4.40
WFLW 2.21* 4.03 4.09 3.71 3.89 4.61
300W + WFLW 2.20* 2.89 4.11 3.64 4.41 4.36
300W + WFLW + AFLW-19 1.02 2.80 4.05 3.52 4.38 4.27

Source: the numerical table in Figure 6; no Dataset Adapters are used. An asterisk indicates exclusion of landmarks undefined in the template. Starred and unstarred entries therefore cover different evaluation-point sets, and their differences should not be interpreted directly as improvements on an identical full task.

Combining all datasets helps most columns, but WFLW68 changes from 3.89 with WFLW-only training to 4.38, an increase in error rather than a decrease. The authors suggest that other datasets dilute challenging samples and emphasize distribution and annotation-quality alignment. This is an interpretation, not a separately established causal mechanism. Appendix A.13 describes this deterioration as a decrease in NME, inconsistent with the table's error direction; the numerical results and their actual direction are retained here.

Appendix A.6 withholds native annotated points during training and computes NME exclusively on those positions. Cubic splines are fitted to retained-point predictions from the same reduced-supervision model, using periodic cubic splines for closed contours.

Dataset and held-out split Full supervision Direct queries with reduced supervision Cubic-spline interpolation
300W: 37 retained, 31 held out 3.56 4.08 5.04
WFLW: 73 retained, 23 held out 4.51 4.92 5.88
WFLW: 50 retained, 46 held out 5.32 5.76 6.83

Source: Table 4. WFLW splits use 96 contour points after excluding the two eye-center points. The full-supervision model has seen the evaluated points during training and is evaluated on the same designated subset. Direct querying under reduced supervision lowers error relative to cubic splines by 19.0%, 16.3%, and 15.7%, respectively, but these are neither full-template test NME values nor accuracy guarantees for arbitrarily denser new points.

Key Findings

  • With 300W-only training, ViT-B achieves 6.08 on WFLW68, better than DTLD's 7.23. Its COFW68 result is 4.40, still worse than DAG's 4.22, so cross-dataset generalization is not uniformly best.
  • Table 3(b) isolates 28 WFLW points absent from 300W. ViT-B achieves WFLW_E NME of 6.52, versus 7.63 for ResNet18 and 7.39 for ResNet101. This supports unseen-position generalization more directly than dense-output visualizations alone.
  • Replacing learned embeddings with SentenceBERT changes 300W NME from 2.99 to 2.80 and WFLW from 4.19 to 4.05. The representation choice has empirical benefits, while the source of its specific semantic knowledge remains a hypothesis.
  • Increasing decoder depth from 1 to 3 changes WFLW NME from 4.41 to 4.05 and 300W from 3.12 to 2.80. At 5 layers, the respective values are 4.06 and 2.82, so additional refinement does not automatically improve results.

Highlights & Insights

  • The output interface is central to the contribution. Shared query addresses connect joint training and variable prediction layouts through one representation instead of repeatedly adding specialized heads to a common backbone.
  • Dynamic localization is distinct from geometric densification. Queries access image evidence directly, and held-out-point comparisons outperform cubic splines under identical reduced supervision, indicating more than an increase in output count.
  • Template heterogeneity and image diversity should not be conflated. Added templates can expand position coverage, but added images also change difficulty distributions; the WFLW68 deterioration shows why these effects should be controlled separately.

Limitations & Future Work

  • Unified templates depend on approximate semantic alignment. Endpoints, visible contours under pose changes, and annotation noise can differ across datasets; the mean intra-cluster distance of 2.22 pixels describes the evaluated templates, not a universal error bound.
  • Arbitrary dense new positions lack ground truth. The authors do not use interpolation-derived Enriched 300W to validate high-density accuracy. Held-out native points offer more reliable evidence but do not cover all continuous query positions.
  • Parts with lower training FPALP diversity, such as the nose boundary, exhibit weaker densification quality than contours or eyes. Distribution constraints or more real annotations may help, but these changes are not validated in this paper.
  • Queried parts must be explicitly defined during training, and text-encoder experiments are limited to English. New-part discovery, multilingual consistency, and 2D part-surface parameterization remain future directions.
  • Comprehensive latency, memory, and mobile-deployment measurements across query counts are absent. Reducing storage for multiple templates does not imply fixed inference costs for large query sets.
  • The discussion refers to 14-, 68-, and 98-point templates, whereas dataset definitions, main experiments, and the AFLW mapping use AFLW-19. This point-count conflict is not used to redefine the experimental dataset. COFW is excluded from training because of annotation-quality concerns, further illustrating sensitivity to label quality.
  • vs LAB / LDDMM-Face: These methods already use part boundaries or semantic flows to represent cross-annotation relationships. This paper turns within-part positions into dynamic query addresses for direct localization, while still requiring template registration.
  • vs CLD: CLD queries 2D landmarks using positions on a canonical 3D face. This method uses part text and FPALPs derived from 2D annotations, reducing reliance on canonical 3D mappings. Its current curve-based representation is not equivalent to querying an entire 3D facial surface.
  • vs FreeEnricher: FreeEnricher refines offsets from base detections and interpolated positions, whereas each target query here directly reads the image. Controlled held-out experiments support direct querying over cubic splines, but do not provide a fully matched comparison against FreeEnricher.
  • Research direction: Model template-registration noise separately from image-localization uncertainty, and hold image distributions fixed while varying template coverage to identify whether generalization comes from new-position supervision or additional samples.

Rating

  • Novelty: 4/5. Combining part-semantic positions with dynamic queries contributes a unified detection interface, rather than introducing boundary representations or continuous landmarks for the first time.
  • Experimental Thoroughness: 4/5. Cross-template, fused-data, representation, and controlled held-out experiments are included, but dense continuous-position ground truth and query-scale efficiency remain insufficient.
  • Writing Quality: 3/5. The main mechanism is clear, but some point counts, metric directions, and comprehensive-superiority claims need tighter distinctions.
  • Value: 4/5. Useful for systems requiring multiple landmark formats, with practical benefits depending on registration quality and deployment query budgets.