PCD

class torchjd.aggregation.PCD(tau=0.02, beta=0.999, eps=1e-08)[source]

Stateful GramianWeightedAggregator implementing Priority-Constrained Descent (PCD), from Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization (TMLR 2026, arXiv:2606.29521).

The first row \(g_1\) of the input matrix is the gradient of the primary objective, and the other rows \(g_2, \dots, g_m\) are the gradients of the secondary objectives. PCD follows the primary gradient as closely as possible, subject to each secondary objective receiving at least a fraction \(\tau\) of the first-order progress that a step along its own normalized gradient would give:

\[\tilde d = \mathop{\mathrm{arg\,min}}_{d} \ \frac{1}{2} \left\| d - \tilde g_1 \right\|^2 \quad \text{subject to} \quad \tilde g_j^\top d \geq \tau \left\| \tilde g_j \right\|^2 \quad \text{for } j = 2, \dots, m,\]

where:

  • \(\tilde g_i = s_i g_i\) is the normalized gradient of objective \(i\),

  • \(s_i = 1 / \sqrt{\hat v_i + \epsilon}\) is its scale,

  • \(\hat v_i\) is the bias-corrected exponential moving average of \(\|g_i\|^2\) over the successive calls, with decay \(\beta\).

The output is \(\tilde d\) rescaled to the norm of \(g_1\). Because of the normalization, \(\tau\) is a scale-free fraction: multiplying a secondary row of the input by the same positive constant at every call leaves the output unchanged, up to the effect of \(\epsilon\).

Special cases:

  • If \(g_1 = 0\), the output is zero.

  • If the constraints cannot be satisfied simultaneously (which requires \(m \geq 3\), e.g. two anti-parallel secondary gradients), they are dropped and the output is \(g_1\).

  • If \(\tilde d = 0\), up to the rounding errors of the Gramian, the output is zero.

  • If the input contains nan or inf, the output is nan and the moving average is left unchanged.

Parameters:
  • tau (float | Tensor) – The fraction \(\tau \in [0, 1]\) of normalized first-order progress guaranteed to each secondary objective. Either a float, shared by all secondary objectives, or a vector with one value per secondary objective (i.e. of length \(m - 1\)).

  • beta (float) – The decay \(\beta \in [0, 1)\) of the exponential moving average of the squared gradient norms.

  • eps (float) – The non-negative constant \(\epsilon\) added to the moving average before taking its inverse square root.

Note

PCD is not symmetric in the objectives: the first row of the input matrix (e.g. the first loss given to backward()) is the primary objective.

Note

This aggregator is stateful: it keeps the moving average of the squared gradient norms across calls. Use reset() to clear it. It is also cleared automatically when the number of rows changes.

Note

The reference implementation is available at github.com/DaraVaram/priority-constrained-descent.

__call__(matrix, /)[source]

Computes the aggregation from the input matrix and applies all registered hooks.

Parameters:

matrix (Tensor) – The Jacobian to aggregate.

Return type:

Tensor

reset()[source]

Clears the moving average of the squared gradient norms.

Return type:

None

class torchjd.aggregation.PCDWeighting(tau=0.02, beta=0.999, eps=1e-08)[source]

Stateful Weighting [PSDMatrix] giving the weights of PCD.

The first row and column of the Gramian correspond to the primary objective.

Parameters:
  • tau (float | Tensor) – The fraction \(\tau \in [0, 1]\) of normalized first-order progress guaranteed to each secondary objective. Either a float, shared by all secondary objectives, or a vector with one value per secondary objective (i.e. of length \(m - 1\)).

  • beta (float) – The decay \(\beta \in [0, 1)\) of the exponential moving average of the squared gradient norms.

  • eps (float) – The non-negative constant \(\epsilon\) added to the moving average before taking its inverse square root.

Note

The quadratic program is solved exactly with the dual active-set method of Goldfarb and Idnani (1983), expressed in terms of the Gramian only. The reference implementation instead enumerates the working sets of constraints (Appendix B.2 of the paper). Both methods find the same minimizer, but the enumeration is exponential in \(m\).

reset()[source]

Clears the moving average of the squared gradient norms.

Return type:

None

__call__(gramian, /)[source]

Computes the vector of weights from the input Gramian and applies all registered hooks.

Parameters:

gramian (Tensor) – The Gramian from which the weights must be extracted.

Return type:

Tensor