From Labels to Computation: HRP, ERC and Inverse-Volatility Weights in Practice
Several indices in the OLTA catalogue carry weighting labels that name an algorithm — Hierarchical Risk Parity, Equal Risk Contribution, Inverse Volatility. A label of that kind is a methodological claim, and the desk's standard is simple: a weighting label must describe code that runs. This paper documents the computation stack now running behind those labels — a returns pipeline on a 180-day lookback, Ledoit-Wolf shrunk covariance, HRP by recursive bisection, ERC by fixed-point iteration, inverse-volatility weighting, and a constraint pass for cash sleeves, single-name caps and floors. It publishes the first five candidate target vectors next to the committee-set weights currently in production, reads the gaps between them, and sets out the pre-registered walk-forward discipline that will govern adoption. No adoption is announced here. The evaluation gate has not yet been executed, every production weight remains committee-set as of this date, and each candidacy is pending index-committee review. The outcome set is fixed in advance and contains exactly two exits: adopt the computed weights, or relabel the index. Keeping a computed-sounding label on weights no algorithm produced is not among them.
- Published
- Jul 24, 2026
- Read
- 20 min read
- Sections
- 06
- Words
- 4,056
- Author
- OLTA Research Desk
Contents12
1. The label-honesty problem
An index's weighting label is not decoration. When a methodology page says an index is weighted by Hierarchical Risk Parity, it asserts that a specific, citable procedure was applied to data and produced the weights the customer holds. That assertion is checkable, and sophisticated allocators check it: the first diligence question against any systematic claim is to ask for the run — the covariance estimator, the lookback, the constraint set, the date the computation was last refreshed.
An internal review of the catalogue in June 2026 found a gap between claim and practice. Fourteen indices carried weighting labels that name a computation, and every production weight vector in the catalogue was a hand-authored literal, set by the committee at design time. The algorithms themselves were not missing — HRP and Ledoit-Wolf shrinkage were implemented, tested against synthetic fixtures, and documented in the methodology paper — but they were wired to nothing an index actually shipped. The labels described intent, not process.
This is a common failure mode in index products, and it usually arises innocently: a design team prototypes with an optimizer, adjusts the output by judgment, ships the adjusted vector, and the label keeps the optimizer's name. The result is still a misstatement. Discretionary weights can be excellent — most of the catalogue is deliberately and openly committee-constructed, and says so — but a label that names an algorithm the pipeline never ran is a claim the desk cannot stand behind in diligence.
There are only two honest exits. Make the computation real, so the label describes the process; or make the label modest, so it describes the judgment. The desk committed to resolving every computed-sounding label through one of those two doors, and to closing the third door — the silent status quo — permanently. Labels whose claimed inputs do not exist in the data the desk actually holds (fee shares, staking yields, fundamental quality scores) take the second exit by construction: they are being reworded as disclosed committee scoring rather than left implying a live feed. Labels whose claimed computation is a pure function of price history can take the first, and five indices now do. This paper documents that first path: the stack, the outputs, and the discipline that separates a candidate vector from an adopted one.
2. The computation stack
The pipeline is deliberately small: one pure library that turns aligned price history into constrained target weights, one operator-run script that applies it per index, and one versioned JSON artifact that records every output with its diagnostics. It composes the estimators that already existed (lib/hrp.js, lib/ledoit-wolf.js) and adds only what did not (lib/computed-weights.js, scripts/compute-weights.mjs). Nothing in the published-statistics pipeline was modified; the computation writes to its own artifact, data/analysis/computed-weights.json, and adoption into production is a separate, deliberate act (section 4).
2.1 Returns
Each run builds a matrix of daily log returns over a trailing 180-day lookback ending at the as-of date, using the same alignment rules as the backtest engine: UTC-day bucketing, forward-fill of missing bars, and an intersection clamp so the window covers only days on which every constituent has history. An adaptive guard caps the effective lookback at 80 percent of available history with a hard floor of 90 observations; any shrink is printed in the run output, never applied silently. The 180-day choice balances responsiveness against estimator noise for daily crypto data and was fixed once, globally — it is a hyper-parameter, and section 5 returns to what that means.
Mixed baskets that hold both crypto and tokenised-equity legs raise a calendar problem: crypto trades seven days a week, equities five. For those baskets the pipeline builds the matrix on the weekday trading-day intersection, folding crypto weekend returns into the Monday observation so that no leg contributes phantom zero-return days to the covariance. The residual bias — weekend crypto variance compressed into Monday rows rather than dropped — is disclosed in the artifact's notes.
2.2 Covariance and shrinkage
A sample covariance estimated from 180 daily observations across six to twelve assets is noisy, and portfolio constructions that consume it inherit the noise. The pipeline therefore estimates covariance with Ledoit-Wolf single-target shrinkage toward a constant-correlation target: the estimator is a convex blend of the sample matrix and a structured target whose off-diagonals carry the average pairwise correlation, with the blend intensity computed in closed form to minimise expected distance to the unknown population matrix, then clamped to the unit interval. Intensity near zero means the data was trusted largely as observed; intensity near one means the sample's correlation structure was largely replaced by the structured target.
The intensity is not an internal detail. It is published per run as a diagnostic, because it measures how much of the resulting allocation rests on measured structure versus imposed structure. On the July 23 run the two mixed-asset baskets produced interior intensities of 0.22 and 0.24 — ordinary readings for windows of this length. The two all-crypto baskets produced an intensity of exactly 1.0, the clamp boundary. That reading is anomalous, it is disclosed on every surface that consumes these vectors, and it is an open investigation item discussed in section 5.
2.3 Hierarchical Risk Parity
HRP, per López de Prado (2016), allocates in three steps. First it clusters the assets by similarity, using single-linkage hierarchical clustering on a distance metric derived from correlations, so that assets that move together sit in the same branch of a tree. Second it reorders the covariance matrix along the tree's leaf order, placing similar assets adjacent. Third it allocates by recursive bisection: starting from the full list, it splits the ordered assets in half, assigns each half a share of the budget inversely proportional to that half's inverse-variance-weighted cluster variance, and recurses into each half until every asset has a weight.
The property that matters is what HRP does not do: it never inverts the covariance matrix. Classical mean-variance optimization requires that inversion, and on noisy estimates the inversion acts as an error amplifier — small estimation errors in near-collinear assets become large, unstable weight swings, which is the documented source of most out-of-sample underperformance in optimized portfolios. HRP substitutes the clustering tree for the inversion, which is why it degrades gracefully on exactly the short, fat-tailed windows crypto provides. The pipeline runs HRP on the shrunk covariance and publishes the cluster order alongside the weights, so a reader can see which assets the tree paired before seeing where the budget went.
2.4 Equal Risk Contribution
ERC weights a basket so that each constituent contributes equally to total portfolio risk, rather than holding equal capital. The pipeline solves for that point with a multiplicative fixed-point iteration on the shrunk covariance: compute each asset's risk contribution at the current weights, scale each weight by the square root of the ratio of target contribution to actual contribution, renormalise, and repeat to tolerance, starting from inverse-volatility weights. Iteration count and convergence status are recorded in the artifact. The scheme follows the risk-budgeting literature (Roncalli 2013) and differs from HRP in using the full covariance directly rather than a cluster tree — appropriate for the small, single-asset-class basket that carries the label.
2.5 Inverse volatility
The simplest member of the stack: each asset is weighted proportionally to the inverse of its annualised realised volatility over the lookback. No correlation input, no optimizer — just the statement that quieter assets carry more capital. Its transparency is the point: for the basket labelled this way, any reader with the price series can reproduce the vector to the decimal.
2.6 Sleeves, caps, floors, and rounding
Constraints are applied after the scheme, in a fixed order, and every binding constraint is recorded rather than hidden. Stablecoin sleeves are structural, not optimized: the sleeve is stripped before the computation, the scheme runs on the risk assets alone, and the result is rescaled to the non-sleeve budget so the sleeve weight ships unchanged. A conservative default band — a 40 percent single-name cap and a 2 percent floor — is enforced by water-filling: breaching names are clipped to the bound and the excess is redistributed proportionally across in-bounds names, iterating until the vector is feasible. Final targets are published on a one-decimal grid with largest-remainder repair so each vector sums to exactly 100.0. The cap and floor are script-level defaults pending a schema decision that will make them versioned, per-index data; when they bind, the artifact says so by name.
3. Hand weights versus computed targets
The July 23, 2026 run produced headline target vectors for the five candidate indices, each computed as of that date on the trailing 180-day window, together with a walk-forward history of what the same scheme would have targeted at each declared rebalance date. The table reports the single largest per-name gap between the committee-set production weights and the computed targets; the full vectors, both arms, are published in the artifact.
| Index | Scheme (per label) | Cadence | Largest single-name gap (hand → computed) | Max gap (pp) | Constraints binding |
|---|---|---|---|---|---|
| OSHARP6 | HRP | quarterly | GLDon 20.0 → 31.4 | 11.4 | none (headline) |
| OBAL6 | HRP | drift-25, walk-forward proxied quarterly | SPYon 15.0 → 40.0 | 25.0 | cap SPYon · floor ETH |
| OHRP12 | HRP | monthly | PAXG 8.0 → 23.0 | 15.0 | none · USDC sleeve fixed at 8.0 |
| ORP6 | ERC | monthly | BNB 17.0 → 20.8 | 3.8 | none |
| OLV8 | Inverse volatility | quarterly | TRX 11.0 → 26.6 | 15.6 | none |
Three of the five gaps are large, and each is legible once the scheme is taken seriously.
The optimizer reallocates toward the diversifying legs. On OSHARP6, HRP moves roughly eleven points into the tokenised gold leg (GLDon 20.0 to 31.4) and eight into the pharmaceutical leg (LLYon 15.0 to 23.0), funded by trims to the semiconductor pair (NVDAon 25.0 to 19.2, AVGOon 15.0 to 8.2) and the crypto pair (BTC 13.0 to 9.7, ETH 12.0 to 8.5). The cluster order explains it: the tree pairs the two semiconductor names and the two crypto legs into tight clusters that share one budget each, while gold and pharma sit apart and collect the diversification premium. OHRP12 tells the same story inside crypto: PAXG rises from 8.0 to 23.0 because gold clusters away from everything else in the basket, while the DeFi names — highly correlated with one another — split a single cluster's allocation rather than each drawing a full one.
A vector pinned to its constraints is a design signal. OBAL6 is the sharpest reading in the set. The computed target puts the broad-equity leg at the 40 percent cap and the ETH leg at the 2 percent floor — two of six names pinned to the constraint set — and the walk-forward history shows the binding name rotating rather than relaxing: the gold leg held the cap through late 2025, the broad-equity leg has held it since January 2026. On a risk-parity criterion this universe wants to be mostly broad equity plus gold with a crypto sliver; the hand vector encodes a different intent, a roughly balanced six-way split holding 35 percent crypto. When an optimizer spends the entire window pressed against its bounds, the message is not that the caps are wrong; it is that the algorithm and the basket disagree about what the index is. Either the mandate is what the hand weights express — in which case the honest label is a balanced, committee-constructed one — or the mandate is risk parity, in which case the composition deserves review. Both are legitimate. The current label-weight pair is the one combination that is not.
Inverse volatility concentrates in whatever is quiet. OLV8's gap is pure mechanics. TRX realised 19.1 percent annualised volatility over the window — roughly half of BTC's 41.0 percent and a quarter of the basket's most volatile member at 72.9 percent — so the scheme's largest weight lands on TRX at 26.6 against a hand weight of 11.0, while the hand vector's ordering tracks market prominence (BTC 22.0 at the top). The walk-forward makes the mechanism visible: the TRX target climbed from 8.2 in April 2025 through each quarterly recompute to about 25 as its realised volatility compressed. An inverse-volatility label commits the index to following the estimator wherever it leads, including when the quietest asset is not the most prominent one. That is what the label means; the computation makes the meaning concrete.
And one near-agreement. ORP6 is the counterexample that calibrates the exercise. The ERC targets land within 3.8 points of the hand weights everywhere, and the walk-forward is stable — the BTC target moved inside a 20.6-to-30.5 band across eighteen monthly recomputes, drifting rather than jumping. The committee's discretionary vector was, in substance, already risk-balanced; the computation largely confirms it. The point of the exercise is not that hand weights are wrong — sometimes they are close to the algorithm's answer. The point is that until the computation runs, nobody can say which case an index is in. A 3.8-point confirmation and a 25-point disagreement carry the same kind of information, and the desk wanted both on the record.
A caution on reading the gaps: a large delta is not evidence that the computed vector would have performed better, and this paper makes no such claim. The deltas measure one thing only — the distance between what the label promises and what the promised procedure produces. Performance enters through the discipline of section 4, not through this table.
4. The adoption discipline
A pipeline like this fails in a predictable way: compute many alternatives, adopt whichever backtests best, and publish the winner's number. The literature on backtest overfitting — surveyed in the desk's own literature review, following Bailey, Borwein, López de Prado and Zhu — is unambiguous that a best-of-N selection inflates the reported statistic by an amount that grows with N, and that the inflation is silent unless the number of trials is disclosed. The adoption discipline here was designed against that failure from the start, and its defining property is that almost every degree of freedom was removed before any result existed.
One scheme per index, pre-registered by the label. The HRP-labelled indices get HRP, the ERC-labelled index gets ERC, the inverse-volatility index gets inverse volatility. There is no scheme sweep and no per-index selection among alternatives: the number of trials per index is one, fixed by a label that predates the pipeline. The lookback is likewise single-valued and global — 180 days, chosen once with its rationale documented, never swept per index in search of a winner.
Walk-forward, against an incumbent that saw the future. The comparison runs both arms at the index's declared cadence over the full available window: the hand arm holds the static production weights; the computed arm follows the walk-forward targets, where the vector applied at each rebalance date was computed from data available up to that date and no further. The asymmetry deserves stating plainly, because it sets the bar. The hand weights are in-sample-tainted — they were authored with full knowledge of the history they are now tested on — while the computed arm is out-of-sample by construction. A computed arm that merely matches the incumbent is therefore a win: it matched, with no hindsight, an incumbent that had all of it.
Pre-registered gates. The adoption rule was fixed in the design record of July 21, 2026, before the first comparison output existed, and it is a guardrail, not a horse race. Because the label already claims the algorithm, the default is adoption; the gates exist to block adoption only where computing materially hurts. The computed arm must satisfy all three: its Sharpe may trail the hand arm by no more than 0.10, and any deficit must sit inside the hand arm's own ±15-day endpoint-robustness spread, so that it cannot be distinguished from start-date noise; its maximum drawdown may exceed the hand arm's by no more than five points; and its additional turnover, priced at the desk's on-record cost anchors, must not erase the margin — the engine is gross of costs, so the cost-adjusted delta is recorded explicitly. Both arms are published whatever the verdict, with endpoint-dispersion bands rather than point estimates alone.
Two exits, no third. An index that passes the gates adopts the computed weights, and its label becomes a description of running code. An index that fails them is relabelled to what its weights honestly are, with the comparison published as the recorded justification. Where a label is so central to an index's identity that no modest relabel exists, the committee's choice is adoption with full disclosure, or retirement — a truthful label on modest numbers is defensible research, while the reverse is not defensible at all. What is excluded, permanently, is the silent third option of keeping a computed-sounding label on hand-set weights.
Status at publication. The target-generation stage has run: the five candidate vectors and their walk-forward histories were computed on July 23, 2026 and are published in the artifact. The comparison stage has not: the paired simulations and the gate evaluation are scheduled ahead of the next quarterly review and have not yet been executed. No production weight has changed, no label has changed, and nothing in this paper is an adoption announcement. All five candidacies are pending index-committee review, and the committee will decide on the gate output — under the rule fixed before the output existed.
5. Limits
Window length. A 180-day estimation window is short by asset-allocation standards, and the walk-forward histories built on it are shorter still: the long-window candidates carry five to eighteen recompute points, spanning roughly twelve to eighteen months of out-of-sample path. That is enough for a guardrail decision — is the computed arm materially worse than the incumbent — and not enough for fine performance ranking, which is precisely why the gates are calibrated as tolerances rather than as a contest. The shortest-history candidate, OHRP12, has only four monthly recompute points; its walk-forward was pre-registered as context only, incapable of passing or failing the gates on its own, with scheme-level evidence borne by the longer-window HRP candidates instead.
Regime dependence. Every number here is conditioned on a single macro regime — the BTC-sideways, equity-bull, gold-bull window the whole backtest programme shares. The volatility ranking that hands TRX the largest inverse-volatility weight, and the correlation structure that sends HRP's budget toward gold, are regime facts, not constants. A different tape would produce different vectors, and an adopted scheme would follow it there at the declared cadence. That is the design, not a defect, but an allocator should read the current vectors as the scheme's answer for this regime, not its answer in general.
The shrinkage clamp at 1.0. On the two all-crypto candidates, the closed-form shrinkage intensity clamped at exactly 1.0 — full replacement of the sample correlation structure by the constant-correlation target — while the two mixed-asset baskets produced interior readings of 0.22 and 0.24 on the same run. The clamp reading is under open investigation. Candidate explanations include heavy-tailed daily crypto returns inflating the estimator's dispersion term relative to its structure term over 180 all-days observations, and the implementation's documented omission of a small cross-term that biases intensity upward. The downstream consequences are disclosed rather than waved away: at intensity 1.0, HRP's clustering operates on constant-correlation structure, so the tree ordering is driven by variances rather than by measured correlation pairing; and the ERC solve's single-iteration convergence on that basket is a symptom of the same flattened matrix, not an independent validation of the solver. Resolving the anomaly — or documenting it as a genuine property of short crypto windows with the estimator behaving as specified — is a logged precondition of the committee packet.
A continued series inside the OLV8 basket. One OLV8 constituent, TON, stopped trading under that ticker on June 30, 2026 and continues as GRAM following a verified one-for-one rebrand. The series the July 23 run consumed is the spliced TON–GRAM history documented in Index Continuity Through Constituent Migrations: Delistings, Rebrands, and the Stale-Tail Problem, with the two-day listing gap at the boundary left unfilled. The composition rename itself is pending an index-committee decision, and the candidacy carries an open concentration question on the computed vector's largest position; the proposed sequencing for both sits in The OLTA Rebalancing Calendar: Effective Dates, Announcements, and Drift Monitoring.
Proxied cadence for the drift-declared candidate. OBAL6 declares a drift-triggered cadence, but drift triggers are path-dependent between recompute points, so its walk-forward was built on a quarterly grid as a cadence-isolation device. The proxy is disclosed in the artifact and the comparison will run both arms on the same grid, so the weights effect is isolated; the drift interaction remains untested in this phase.
No transaction costs in the simulator. The engine is gross of costs. The discipline compensates at the gate — computed-arm turnover is priced against on-record cost anchors before a verdict — and per-event turnover is recorded in the artifact so costs can be layered onto any result later without regeneration. The residual limitation is that cost anchors are estimates, not measured execution.
Calendar folding for mixed baskets. Folding crypto weekend returns into Monday rows compresses weekend variance into one observation. The bias is shared by both arms of the comparison and disclosed in the artifact notes; it slightly understates all-in crypto volatility inside mixed-basket covariances relative to an all-days treatment.
Constraint defaults. The 40 percent cap and 2 percent floor are conservative script-level defaults, pending the schema decision that will turn caps into versioned, per-index data reviewed by the committee. Where they bind — and on one candidate they bind decisively — the binding is itself part of the finding, and section 3 treats it as such.
These limits are the price of the property the pipeline was built for: every number above is the output of code that ran, on data the desk holds, under a rule fixed before the results were seen. The catalogue's weighting labels are converging on the same property — by computation where the data honestly supports it, by plainer language where it does not.
References
- López de Prado, M. (2016). "Building Diversified Portfolios that Outperform Out of Sample." The Journal of Portfolio Management 42 (4): 59–69.
- Ledoit, O., and M. Wolf (2003). "Honey, I Shrunk the Sample Covariance Matrix." The Journal of Portfolio Management 30 (4): 110–119.
- Roncalli, T. (2013). Introduction to Risk Parity and Budgeting. Chapman and Hall/CRC.
- Bailey, D. H., J. M. Borwein, M. López de Prado, and Q. J. Zhu (2014). "Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance." Notices of the American Mathematical Society 61 (5): 458–471.
- Bailey, D. H., and M. López de Prado (2014). "The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality." The Journal of Portfolio Management 40 (5): 94–107.
- OLTA Research, "OLTA Backtest Methodology" (2025-09-11).
research/backtest/01-methodology.md. - OLTA Research, "Backtest Overfitting and the Pseudo-Mathematics Critique" (2026-01-08).
research/literature/05-backtest-overfitting.md. - Implementation:
lib/hrp.js,lib/ledoit-wolf.js,lib/computed-weights.js,scripts/compute-weights.mjs. Published artifact:data/analysis/computed-weights.json(run of 2026-07-23).
Baskets referenced in this paper
05- OSHARP6Sharpe MaxLive1.58Sharpe
- OBAL6BalancedLive1.29Sharpe
- OHRP12HRP 12Under study0.27Sharpe
- ORP6Risk Parity 6Conviction1.00Sharpe
- OLV8Low Vol 8Under study1.03Sharpe
Sharpe figures are backtested on simulated funds and shown only where the measured window clears 90 days. Baskets with a shorter track record read "Not graded".
Simulated funds, backtested results. Past performance is not a guarantee and nothing here is an offer or a recommendation. OLTA is in public preview: mainnet is planned for H1 2027.