ETF & Index Data Insights, News & Analysis | Ultumus

ETF Flows Have Memory But Leveraged Products Remember Differently

Written by Bernie Thurston | Jul 23, 2026 12:33:45 PM
Testing flow persistence across ~1,000 ETFs with two independent estimators, and what their disagreement taught us.

We tested whether daily ETF flows show statistical “memory” — a tendency for inflows to beget more inflows, or outflows to reverse into inflows — and whether that memory differs between ~205 mainstream, megacap ETFs (“Pareto”) and ~860 leveraged/inverse ETFs, using two independent estimators of the Hurst exponent: rescaled‑range (R/S) analysis and Detrended Fluctuation Analysis (DFA). Flows trend more than they revert in both groups. R/S found mainstream ETFs significantly more persistent than leveraged/inverse ones (p < 0.001, robust to same‑issuer product clustering); DFA found no significant difference between the two groups. The disagreement between the two methods turned out to be the more interesting result: leveraged/inverse flows are measurably spikier than mainstream flows, and R/S — unlike DFA — is known to mistake that kind of spikiness for genuine trend.

Why we looked at this

There’s an established line of research on why fund flows would have “memory” in the first place. In market microstructure, work going back to Lillo and Farmer (2004) and Bouchaud et al. (2004) found that the sign of market orders (buy vs. sell) shows long‑range autocorrelation, generally attributed to institutional investors splitting large trades into smaller pieces executed over days or weeks to reduce market impact. Separately, mutual fund research (Cashman, Nardari, Deli, and Villupuram, 2014) finds that fund‑level gross flows are themselves highly persistent and partly predictable from their own past, consistent with investors turning over positions gradually rather than all at once.

Leveraged and inverse ETFs are a natural place to ask whether that “memory” looks any different. They’re built and marketed for short‑term, tactical use rather than buy‑and‑hold positions, a different usage pattern from the gradual institutional trade‑splitting that motivates the market‑microstructure literature. So we wondered: does that different way of using the product show up in the shape of its flows? Do leveraged/inverse ETFs’ daily inflows and outflows look more random and less trending than a conventional ETF’s, or do they behave the same way?

To find out, we measured flow persistence across two universes: our regular “Pareto” universe of ~205 mainstream megacap ETFs (iShares, Vanguard, SPDR, and similar), and a separate universe of ~860 leveraged and inverse products (2x, 3x, and “short” funds) with enough trading history for a reliable read. For each product, we took a trailing year of daily flow as a percentage of the prior day’s AUM, split‑adjusted, and estimated its Hurst exponent two ways.

What we found

Both groups trend more often than they mean‑revert. 84% of Pareto products showed persistent (trending/momentum) flow behaviour; so did 77% of leveraged/inverse products. Very few products in either group showed mean‑reverting (“bounces back”) behaviour — 2.4% and 1.3% respectively.

  Pareto (n=205) Leveraged/Inverse (n=860)
Mean Hurst exponent (R/S) 0.651 0.611
Mean DFA α 0.730 0.747
% persistent (Hurst > 0.55) 84% 77%
% anti‑persistent (Hurst < 0.45) 2.4% 1.3%

The two estimators disagreed on whether the gap between groups is real. R/S says mainstream ETFs are significantly more persistent than leveraged/inverse ones. DFA says there’s no significant difference at all — if anything, leveraged/inverse products score higher on DFA, just not significantly so.

Method Mann‑Whitney p Cluster‑robust bootstrap (95% CI of the gap)
Hurst (R/S) 3.3 × 10⁻⁷ −0.040 [−0.057, −0.024] — gap holds
DFA α 0.15 +0.017 [−0.018, 0.051] — straddles zero

The cluster‑robust check matters here: many leveraged/inverse products are effectively siblings — the same issuer offering 1x/2x/3x, long/short variants on the same underlying asset — so treating each as a fully independent data point overstates how much evidence there really is. Even after collapsing ~860 leveraged/inverse products down to roughly 630 independent product families, the R/S gap survives; the DFA non‑result doesn’t change either.

Why they disagreed: leveraged/inverse flows are spikier

The surprise

R/S is known to be sensitive to short, sharp bursts in a time series — a handful of huge single‑day moves can push its persistence estimate up even when the series isn’t really trending in a sustained way.

DFA explicitly removes local trend from each window before measuring fluctuation, which makes it considerably more robust to exactly that kind of burstiness.

So we checked: are leveraged/inverse flows in fact spikier?

  Pareto Leveraged/Inverse Mann‑Whitney p
Excess kurtosis (fat‑tailedness) 38.4 mean / 19.7 median 66.1 mean / 41.3 median 1.1 × 10⁻¹¹
Largest single‑day move (own‑series z‑score) 7.8 mean / 7.2 median 9.2 mean / 8.9 median     1.8 × 10⁻⁹

Yes — decisively. Leveraged/inverse flows have roughly 70% higher excess kurtosis and noticeably larger single‑day outliers than mainstream flows. That’s consistent with more tactical, short‑burst trading activity, and it’s a plausible mechanical explanation for why R/S sees a persistence gap that DFA doesn’t: R/S is likely picking up spikiness, not a cleaner difference in trend.

Takeaways

  • ETF flows trend more than they mean‑revert, across the board. That’s the most robust finding here and it holds for both mainstream and leveraged/inverse products.

  • The “mainstream trends more than leveraged/inverse” claim is real but estimator‑dependent, and modest in size. Treat any single‑number Hurst comparison with some caution — which estimator you use can change the conclusion.

  • Leveraged/inverse flows are measurably burstier, independent of the persistence question. That’s arguably the more solid, actionable finding, and it’s consistent with these products’ reputation for short‑term, tactical use — though it’s a statement about flow shape, not trading volume itself.

Want more? Get in touch with our team for additional charts and data from our full ETF universe.

 

Next step: bring in real trading‑volume data to test directly whether leveraged/inverse products see disproportionate trading volume relative to AUM, rather than inferring it indirectly through flow persistence.


Methods (for the curious)

The Hurst Exponent and DFA, in plain terms. Both take a year of daily flow data for one ETF and return a single number summarising whether its ups and downs have “memory”:
  • Around 0.5 — flows are close to a coin flip; today tells you nothing about tomorrow.
  • Above 0.55 — “persistent”: a run of inflows tends to be followed by more inflows (and the same for outflows). This is the trending/momentum case.
  • Below 0.45 — “anti‑persistent”: inflows tend to be followed by outflows and vice versa — a mean‑reverting, “bounces back” pattern.
R/S (rescaled range) and DFA (Detrended Fluctuation Analysis) are two different recipes for computing that same underlying idea. They usually agree closely, which is why a disagreement between them is itself worth investigating rather than dismissing.
Why two significance tests, not one. A standard significance test (we used the Mann‑Whitney U test, which doesn’t require flows to be normally distributed) treats every product as an independent data point. That’s a shaky assumption for the leveraged/inverse group in particular: a lot of those ~860 products are the same issuer’s 1x/2x/3x and long/short variants on the same underlying asset, so their flow behaviour isn’t fully independent of each other. We addressed this with a cluster bootstrap: we grouped products into “families” (same issuer, same underlying, different leverage dial) and resampled whole families rather than individual products, which gives an honest answer to “is this difference bigger than we’d expect from chance, once we stop double‑counting near‑duplicate products.”
Kurtosis and z‑scores, in plain terms. “Excess kurtosis” measures how much of a series’ variation comes from occasional big jumps rather than a steady simmer of small moves — zero means it looks like a normal bell curve, higher means fatter tails and more outlier days. “Z‑score” measures how many standard deviations a single day’s flow was from that fund’s own typical day — a bigger number means a bigger, more unusual spike, standardised so it’s comparable across products of very different sizes.
Data and caveats. Daily flow, as a percentage of the prior day’s AUM, split‑adjusted, over a trailing year; ~205 Pareto and ~860 leveraged/inverse products with enough valid history for a reliable estimate. The “family” clustering used for the bootstrap only merges products from the same issuer (e.g., WisdomTree’s own Copper 1x/2x/3x lineup) — it doesn’t merge the same underlying asset across different issuers (e.g., a ProShares Bitcoin fund and a WisdomTree Bitcoin fund would still count as two independent families), so if anything it understates the true clustering, making our significance result conservative rather than inflated. This is a first look at flow behaviour, not trading volume, which would need a separate, direct test.
 
References
Bouchaud, J.‑P., Gefen, Y., Potters, M., and Wyart, M., 2004, Fluctuations and response in financial markets: the subtle nature of ‘random’ price changes, Quantitative Finance 4(2), 176–190.

Cashman, G., Nardari, F., Deli, D., and Villupuram, S., 2014, Investor behavior in the mutual fund industry: evidence from gross flows, Journal of Economics and Finance 38(4), 541–567.

Lillo, F., and Farmer, J. D., 2004, The long memory of the efficient market, Studies in Nonlinear Dynamics & Econometrics 8(3), article 1.