Correlation and Volatility Aren’t Enough: The New Science of Portfolio Diversification

Written by: Ben Walsh

I made a claim a few weeks ago that looking at asset classes is looking at labels, and that looking at correlation and volatility is looking at the genome underneath them. It was a good line. It was also, I've since decided, only half the argument. Two statistics don't make a genome. They make a blood pressure reading. Useful, cheap to take, and dangerously easy to mistake for a full health check.

The chart that prompted this piece is simple enough to describe in a sentence. Australian shares and bonds, monthly total returns, rebalanced, measured across two eras. From 1994 to 2020, the correlation between the two sat at essentially zero, plus 0.01. A fifty-fifty mix bought an investor 2.58 percentage points less volatility than a straight line between the two asset classes would have predicted. That gap is the diversification benefit, drawn as the space between a curve and a straight line on a risk-return chart, and for nearly three decades it was wide and dependable. From 2021 to now, the correlation has jumped to plus 0.45, and the same fifty-fifty mix only buys 1.11 points of that benefit. Less than half of what it used to.

Every commentator who has picked up this chart, and plenty have, reaches for the same conclusion: the 60/40 portfolio is broken, correlations have risen, diversification doesn't work like it used to. All true. All also the least interesting thing the chart tells you, because it's a two-gene test, and portfolios have a great deal more DNA than that.

Correlation Is Not a Trait, It Is an Expression

Start with the number itself. A correlation of plus 0.01 averaged over 324 months is not a description of how shares and bonds behaved in any particular month. It is a description of how they behaved on average, which is a completely different thing. Average enough calm months together with enough stressed ones and you can produce almost any number you like, including one that looks reassuringly close to zero while hiding long stretches where the relationship spiked violently in one direction or the other.

This matters because diversification is not something you need on average. You need it specifically in the months when growth assets are falling, which is precisely when correlations have a documented habit of moving against you. The academic literature on dependence structures describes this properly: markets shift into an asymmetric regime under stress, where a downside move in one asset raises the probability of a similar move in the other, even while the long-run average correlation looks perfectly tame. A portfolio genome test that reports a single correlation figure for a twenty-seven-year period is like taking an average heart rate once a year. It tells you almost nothing about what happens during the sprint.

The honest way to think about correlation is as an epigenetic trait rather than a genetic one. The underlying relationship between two assets isn't fixed in the way a gene sequence is fixed. It is switched on and off by the surrounding environment, the regime, in ways that a static number obscures entirely. Two assets can carry a correlation gene that expresses as near zero for years and then switches on hard the moment a common shock, say an inflation surprise, hits both legs of the portfolio at once. That is exactly what appears to have happened since 2021. Inflation stopped being a tailwind that pushed bond yields down while growth assets did their own thing, and became a shared threat that hurts both simultaneously. The gene didn't change. The environment did, and the gene expressed differently as a result.

The Label Hides Different Genomes Entirely

The second gap in a two-statistic test is that the label sitting on top of the numbers can conceal genuinely different underlying portfolios. Two funds can both be marketed as 60/40 and carry almost nothing in common genetically. One holds long-duration government bonds with no currency hedge; the other holds short-duration investment-grade credit hedged back to the base currency. One holds large-cap growth-tilted equities, the other holds a value and quality blend. Both wear the same label. Neither has the same genome.

The chart above shows nine different funds, all carrying variations of the "Australian bond" label, all trading on the same exchange, all ostensibly serving the same defensive purpose in a portfolio. Yet look at the spread of outcomes. At one extreme, 5GOV — the five-to-ten-year government bond ETF — sits up 1.97% over the period shown. At the other, LEND — the listed private credit fund — has fallen 23.70%. Better than twenty-five percentage points between two investments that many investors would casually describe as "the bond portion" of their portfolio.

The divergence isn't random, and it isn't a credit-quality story of the kind the labels imply. The 2022 inflation shock hit every line on this chart through the same duration channel; what happened afterwards is where the genomes separated. The shorter-duration government fund recovered its losses and finished positive. The long-duration government funds are still down double digits, because long duration is a one-way bet on the rate path. The composite funds sit in between, carrying duration and credit spread together. And the private credit fund did what none of the others did: it bled slowly, then broke, as illiquid marks met a genuine liquidity bid. Nothing on any of these labels predicts that ordering. The factor exposures do.

Notice also the Ulcer Index in the bottom panel — a measure of the depth and persistence of drawdowns, which captures what standard volatility cannot. It does not record the 2022 shock as a single spike and forget it. It records it as a broad hill that takes years to decay, because the metric charges you for every month spent below the previous peak. That hill is the investor's lived experience of the regime shift: time underwater, not dispersion around a mean. And note the turn upward at the right-hand edge, accumulating again just as the private credit line breaks.

This is the genome test that matters. Not "what happened to bonds in 2022" but "which bonds, with which exposures, under which conditions?" The investor who held 5GOV through that period experienced a painful but ultimately temporary drawdown. The investor who held LEND experienced something closer to a structural re-pricing of credit risk and liquidity premium, with no guarantee of mean reversion. Same asset class label. Genetically unrelated outcomes.

The chart also reveals something subtler: the convergence and divergence patterns over time. From 2013 to 2020, these lines moved in a relatively tight band, reinforcing the illusion that "bonds are bonds." It was only when the regime shifted—when inflation became the dominant shock, when real yields repriced, when liquidity evaporated—that the genetic differences expressed themselves fully. This is precisely the epigenetic behaviour the essay described: the same underlying exposures that looked similar in one environment revealed their true differences in another.

What should an investor have done with this information? Not abandoned bonds. Not concluded that diversification is dead. But asked harder questions: Which bonds? What duration? What credit exposure? What liquidity profile? What happens to this position if real yields rise 200 basis points in six months? The chart doesn't answer those questions, but it makes them impossible to avoid.

This is where factor-based thinking earns its keep over asset-class thinking. The actual drivers of return and risk in a portfolio are exposures like equity beta, duration, credit spread sensitivity, inflation sensitivity, currency exposure and liquidity, not the accounting classification of "shares" or "bonds" sitting on top of them. Two portfolios with identical asset-class labels can have almost opposite factor DNA, and it's the factor DNA, not the label, that determines how a portfolio actually behaves when the regime shifts. A genome test that only checks the label is testing the wrapper, not the contents.

Article content

Nine ASX-listed fixed income funds, one mental bucket. Cumulative total return, with the Ulcer Index below.

Skew Is the Trait Volatility Cannot See

The third gap is arguably the most dangerous one in practice, because it's invisible to every mean-variance chart, including the one that started this whole conversation. Volatility measures the spread of outcomes around an average. It says nothing about whether that spread is symmetric. Two portfolios can carry identical volatility and identical correlation to every other asset in the room and still have entirely different shapes to their return distribution. One might have positive skew, a return profile with a capped, well-understood downside and open-ended upside, behaving something like a long option position. The other might have negative skew, the profile that looks perfectly fine for long stretches and then produces the single catastrophic month that erases years of steady gains.

Negative skew is the genetic trait responsible for most of the portfolio failures that actually end careers and retirement plans, and it is completely absent from a correlation-and-volatility genome test. You can build two 50/50 portfolios that sit at the exact same point on that risk-return chart, and one of them is a far more dangerous animal than the other. Mean-variance analysis, the entire framework underneath the chart that started this piece, is structurally blind to that distinction. It was built at a time when returns were assumed to be roughly symmetric and well-behaved, an assumption that has never been particularly true and is becoming less true as more of the return stream in modern portfolios comes from strategies and instruments with genuinely asymmetric payoffs.

What Compounds Is Not What Is Quoted

There is a further wrinkle worth adding here, because it changes how you should read any claim that returns held steady while volatility rose. What compounds into an investor's actual wealth is the geometric return, not the arithmetic one, and the geometric return is approximately the arithmetic mean minus half the variance. If volatility genuinely rises while the arithmetic average holds constant, the realised compound growth rate falls, mechanically, with no correlation subtlety required at all.

This is really an argument about ergodicity, a concept borrowed from physics that has quietly become one of the more useful lenses in finance. The average return everyone quotes is an ensemble average, the return across many hypothetical parallel paths. No investor lives in the ensemble. Each one lives a single path through time, and for a compounding, non-ergodic process, the time average that single path actually delivers is systematically worse than the ensemble average, the higher the variance runs. A genome test that reports an unchanged average return alongside a higher volatility figure isn't describing an unchanged outcome with more noise attached. It's describing an outcome that has already gotten worse for the one investor who has to live through it, whether or not the quoted average admits it.

The Genome That Writes Itself

There is a further layer, and it's the one I find most interesting because it turns the whole exercise reflexive. Correlation is usually described as something markets have, a fact about the world that a portfolio manager discovers through measurement. But a meaningful share of the correlation shift since 2021 looks less like a fact being discovered and more like a fact being manufactured by the very act of large numbers of investors holding similar portfolios and reacting to the same signals at the same time.

When enough capital sits in structurally similar exposures, whether that's the classic 60/40 template, risk parity, or the same crowded systematic factor trades, a shock that causes one cohort to de-risk causes all of them to de-risk simultaneously, because they're built the same way and respond to the same triggers. That correlated selling is not a fundamental relationship between the underlying assets. It's a flow artefact created by the population of investors holding the same genome and behaving the same way when frightened. The correlation shows up in the data identically either way, but the causal story, and therefore the right response, is completely different. A macro-driven correlation shift calls for a genuinely different asset mix. A crowding-driven correlation shift calls for genuine differentiation from the crowd, which is a much harder thing to sell and a much harder thing to buy.

Inflation Is the Narrative, Not the Noise

It's worth naming plainly what has actually rewritten the genome, because it isn't a mystery. Stock-bond correlation is the net result of two competing forces: growth shocks, which push shares and bonds in opposite directions, and inflation shocks, which push them in the same direction. Through the 2010s, growth shocks dominated, and inflation was largely dormant, which is why bonds hedged so reliably and why a near-zero correlation held for so long. Since 2021, inflation volatility has dominated growth volatility, and that alone is enough to flip the sign of the relationship, without any change in how the two asset classes are fundamentally constructed.

Two further characteristics compound that shift. Real yields repriced sharply higher from 2022, which hits long-duration equities and long-duration bonds through the identical discounting channel, tying the two together mechanically rather than emotionally. And the Reserve Bank moved from a largely passive stance, where an inflation surprise got absorbed rather than fought, to an active one, where it gets met with tighter policy that damages both bond prices and equity valuations at the same time. The term premium itself, the extra yield long bonds pay over the expected path of short rates, spent most of the 2010s suppressed near zero by quantitative easing and dormant inflation. It turned positive again once inflation became the dominant shared risk, and research on the relationship shows the causality runs both ways: once shares and bonds start moving together, investors demand a higher term premium precisely because bonds have stopped diversifying equity risk, which then pushes yields higher still. None of that is noise sitting beside an unchanged portfolio. It is the portfolio's genome being rewritten by a change in which macro variable is doing the driving.

Passive Investing Is a Bet on a Regime, Not an Escape From One

A market-capitalisation-weighted index is presented and sold as the neutral option. Own the market, avoid the risk of picking winners, keep costs low, let the index do the diversifying. But a cap-weighted index is not neutral. Its weights are the direct product of which assets performed best under the conditions of the recent past. The largest holdings earned their size by winning under whatever regime, whatever combination of rates, growth, inflation and risk appetite, happened to prevail while the index was forming its current shape. Buy the index today, and you are not buying a diversified slice of the future. You are buying yesterday's regime, sized by trailing price performance, on the implicit assumption that the conditions which produced those weights are still the conditions in force.

That assumption holds, often for long stretches, which is exactly why passive investing has performed so well for so long. It is precisely wrong at the moment the regime turns, and a cap-weighted index has no internal mechanism that notices the turn has happened. The weights don't update on a change in correlation, a change in inflation regime, or a change in the relationship between growth and value. They update on price, and price is the trailing record of the last regime, not a forecast of the next one. By the time the weights have fully adjusted to a new environment, through years of relative price movement reallocating capital toward the new winners, the bulk of the transition has already happened. The index is a lagging description of the regime, not a hedge against its change.

Basket trading intensifies this rather than diluting it. When a stock is added to a major index, or when passive flows into an index rise generally, every fund tracking that benchmark has to buy the same names in the same trade, on the same day, for reasons that have nothing whatsoever to do with the underlying businesses. The research on this is now reasonably deep: stocks with higher levels of passive and index ownership show measurably higher correlation with the index and with the broader market, without a corresponding rise in their own company-specific volatility. In plain language, becoming more heavily owned by passive vehicles doesn't make a company intrinsically riskier on its own merits. It makes the company move more like everything else around it, because the flow mechanism, rather than the fundamentals, increasingly sets the price. That is a genome mutation happening for entirely mechanical reasons, unrelated to anything happening in the real economy.

So the honest description of a passive allocation isn't "diversified exposure to the market." It's closer to "a bet, sized by trailing price data, that the regime responsible for today's weights is still the regime we're operating in, held alongside enough other capital doing the same thing that the flow itself now shapes the correlation everyone is measuring." That's a very different product to the one being sold on cost and simplicity, and it's worth advisers being honest with themselves about which one they're actually recommending, because the client experience of the two is not remotely the same when the regime turns.

From Diagnosis to Construction

Everything so far is diagnosis. It's fair to ask, having spent this much time arguing that the standard genome test is too crude, what an actual construction method that respects the genome looks like, rather than one that measures it and then falls back on the same flat mean-variance machinery regardless.

There is a useful answer emerging from the hierarchical portfolio construction literature that Marcos López de Prado began with Hierarchical Risk Parity and later extended with Nested Clustered Optimisation. The core insight of that lineage is that a raw correlation matrix, the input to almost every mean-variance model, is not a flat, structureless object. It has a nested, tree-like architecture, clusters of assets that behave alike sitting inside larger clusters, and traditional optimisers, which treat every pairwise correlation as equally informative, are effectively ignoring the very structure that a genome metaphor is trying to describe. Hierarchical clustering methods build that tree directly from the data instead of assuming a flat structure or relying on a pre-existing asset-class label. That is, quite literally, sequencing the genome rather than reading the label glued on top of it.

A recent extension of that tradition, Hierarchical Core-Orbital Allocation, applied to global fixed income, sharpens this further in a way I think is genuinely useful for the argument this essay has been building. It splits the construction problem into two deliberately decoupled levels. At the strategic level, a small number of representative "core" assets are identified within each cluster the hierarchy reveals, and those cores are optimised using whatever objective function suits the mandate- minimum variance, maximum Sharpe, risk parity; the choice is left open. At the tactical level, capital is then redistributed among the remaining "orbital" assets inside each cluster, using an allocation rule that can differ entirely from the one used at the core level, while still exactly reproducing the upper-level decision in aggregate.

Read against everything above, that architecture maps almost exactly onto the layered genome this essay has argued for. The clusters themselves are the real family tree, the structure that a flat correlation-and-volatility snapshot averaged over decades completely fails to reveal. The core assets, in a fixed income context, are capturing the systematic factors, duration and credit risk, the very same term premium and growth-inflation dynamics responsible for the correlation regime shift discussed earlier. They are the macro-level genome, the part that moves when the Reserve Bank's reaction function changes or when inflation replaces growth as the dominant shock. The orbital assets are capturing something closer to the instrument-level idiosyncrasies, the skew, the liquidity quirks, the issuer-specific traits that a portfolio genome test focused purely on asset-class correlation would never surface. Applied to a universe of sovereign, supranational and corporate credit, where duration and credit genuinely dominate the factor structure, the method reproduces the hierarchy that fixed income markets already have, rather than imposing a flat one that was never really there.

The wider point isn't that every adviser needs to run a clustering algorithm before the next asset allocation review. It's that a construction method exists which takes the genome metaphor seriously as an engineering problem rather than leaving it as a diagnostic complaint. It separates what the data's structure genuinely tells you- the clusters, the hierarchy- from what a human still has to decide: the objective function applied within and across them. That separation of statistical structure from economic judgement is exactly the discipline missing from a correlation-and-volatility genome test that quietly assumes the label was the right level of resolution all along.

What a Genuine Genome Test Would Ask

None of this is an argument for abandoning diversification, or shares, or bonds, or even passive vehicles, which remain a perfectly reasonable component of a well-constructed portfolio. It's an argument for testing the genome properly rather than settling for the two easiest statistics to compute. A genuine test would ask how a correlation behaves conditionally, in the specific months of stress that matter, rather than as a single number averaged across decades of calm and chaos alike. It would look through the asset-class label to the actual factor exposures sitting underneath it, because two portfolios wearing the same label can be genetically unrelated. It would check the shape of the return distribution, not just its spread, because two portfolios with identical volatility can have opposite risk profiles once the tails are examined. It would distinguish the ensemble average being quoted from the time average an actual investor will live through. And it would ask honestly how much of any measured correlation is a fact about the economy and how much is an artefact of enough capital being built the same way and reacting to the same triggers at the same time.

Most portfolio reviews still run the two-gene test, and most conversations about the death or survival of the 60/40 portfolio are conducted entirely within its limits. The chart that started this piece is a genuinely useful data point. It's just not the diagnosis. It's the symptom that tells you it's time to run the full sequence, and increasingly, the tools to actually do that are no longer theoretical.