PCD¶
- class torchjd.aggregation.PCD(tau=0.02, beta=0.999, eps=1e-08)[source]¶
StatefulGramianWeightedAggregatorimplementing Priority-Constrained Descent (PCD), from Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization (TMLR 2026, arXiv:2606.29521).The first row \(g_1\) of the input matrix is the gradient of the primary objective, and the other rows \(g_2, \dots, g_m\) are the gradients of the secondary objectives. PCD follows the primary gradient as closely as possible, subject to each secondary objective receiving at least a fraction \(\tau\) of the first-order progress that a step along its own normalized gradient would give:
\[\tilde d = \mathop{\mathrm{arg\,min}}_{d} \ \frac{1}{2} \left\| d - \tilde g_1 \right\|^2 \quad \text{subject to} \quad \tilde g_j^\top d \geq \tau \left\| \tilde g_j \right\|^2 \quad \text{for } j = 2, \dots, m,\]where:
\(\tilde g_i = s_i g_i\) is the normalized gradient of objective \(i\),
\(s_i = 1 / \sqrt{\hat v_i + \epsilon}\) is its scale,
\(\hat v_i\) is the bias-corrected exponential moving average of \(\|g_i\|^2\) over the successive calls, with decay \(\beta\).
The output is \(\tilde d\) rescaled to the norm of \(g_1\). Because of the normalization, \(\tau\) is a scale-free fraction: multiplying a secondary row of the input by the same positive constant at every call leaves the output unchanged, up to the effect of \(\epsilon\).
Special cases:
If \(g_1 = 0\), the output is zero.
If the constraints cannot be satisfied simultaneously (which requires \(m \geq 3\), e.g. two anti-parallel secondary gradients), they are dropped and the output is \(g_1\).
If \(\tilde d = 0\), up to the rounding errors of the Gramian, the output is zero.
If the input contains
nanorinf, the output isnanand the moving average is left unchanged.
- Parameters:
tau (
float|Tensor) – The fraction \(\tau \in [0, 1]\) of normalized first-order progress guaranteed to each secondary objective. Either a float, shared by all secondary objectives, or a vector with one value per secondary objective (i.e. of length \(m - 1\)).beta (
float) – The decay \(\beta \in [0, 1)\) of the exponential moving average of the squared gradient norms.eps (
float) – The non-negative constant \(\epsilon\) added to the moving average before taking its inverse square root.
Note
PCD is not symmetric in the objectives: the first row of the input matrix (e.g. the first loss given to
backward()) is the primary objective.Note
This aggregator is stateful: it keeps the moving average of the squared gradient norms across calls. Use
reset()to clear it. It is also cleared automatically when the number of rows changes.Note
The reference implementation is available at github.com/DaraVaram/priority-constrained-descent.
- class torchjd.aggregation.PCDWeighting(tau=0.02, beta=0.999, eps=1e-08)[source]¶
StatefulWeighting[PSDMatrix] giving the weights ofPCD.The first row and column of the Gramian correspond to the primary objective.
- Parameters:
tau (
float|Tensor) – The fraction \(\tau \in [0, 1]\) of normalized first-order progress guaranteed to each secondary objective. Either a float, shared by all secondary objectives, or a vector with one value per secondary objective (i.e. of length \(m - 1\)).beta (
float) – The decay \(\beta \in [0, 1)\) of the exponential moving average of the squared gradient norms.eps (
float) – The non-negative constant \(\epsilon\) added to the moving average before taking its inverse square root.
Note
The quadratic program is solved exactly with the dual active-set method of Goldfarb and Idnani (1983), expressed in terms of the Gramian only. The reference implementation instead enumerates the working sets of constraints (Appendix B.2 of the paper). Both methods find the same minimizer, but the enumeration is exponential in \(m\).