Strategy

How the capital is managed, what a proposal has to pass before it gets money, and what we have measured that does not work.

Evidence, selection, measurement and risk

What this is

This page is not a presentation of a strategy. It is a description of what has to be true before something becomes a strategy here.

The order is deliberate. The arithmetic comes before selection, because it decides what is worth looking for at all. Measurement comes before risk, because a risk number produced by an apparatus whose weaknesses you do not know is not a risk number.

The figures below come from our own work and can be checked against the report linked at the bottom. Where they point to something not working, they are here for the same reason they are there.

01

The premise

Passive exposure is the default. Everything else has to earn its place.

Hulc manages its own capital. That gives one advantage money cannot buy, namely that nobody can ask for the money back at an inconvenient moment. We have no redemptions to cover, no quarterly measurement against a benchmark, and no reason to trade because a period has passed without anything happening. It also gives one constraint that cannot be negotiated away, namely that fixed costs per order make up a larger share the smaller the order is. Both shape the strategy more than any market view does.

The assumption we work under is that the market is largely efficient over time, but that deviations arise and can in principle be exploited. The second half of that sentence is the interesting one, and it is also the one that has to be tested rather than believed. The literature is unambiguous on one point. Published anomalies shrink after publication. McLean and Pontiff (2016) measure the returns of published strategies as 26 per cent lower out of sample and 58 per cent lower after publication. A strategy that is known is therefore not automatically a strategy that works.

The consequence is a burden of proof that runs one way. The capital sits passively in EEA-domiciled UCITS funds under the Norwegian participation exemption until something else has passed a set of criteria written down before the measurement was run. Active exposure is the exception that must be argued for, not the default that must be defended.

That is not a theoretical position. We have tested five strategy families, namely intraday mean reversion on RSI, filtered variants of it, factor ETF rotation in European markets, absolute momentum as trend following, and two single-stock strategies on the US market over 27 years with point-in-time data. None of them delivered risk-adjusted returns above a passive reference after costs and tax. No configuration has been promoted to production.

Systematic trading.
Rule-based strategies that can be fully specified in advance and rerun by someone else. This is the area where the burden of proof can be settled in numbers, which is why it was tested first.
Fundamental analysis.
Assessment of individual companies on accounts, return on capital and pricing. Here the conclusion cannot be measured the same way, so the requirement moves onto the assumptions, which are written out and tested against alternatives.
Risk management and portfolio construction.
How positions are combined, how concentrated the portfolio is allowed to be, and how much friction the setup can carry. This area inherits the assumptions from the other two, and it is where most actual outcomes are decided.

The measurements on this page were run on our own backtesting framework, with point-in-time fundamental data, historical index membership, an explicit cost model and 22 per cent tax on realised gains where it applies. Every measurement was run against a full test suite before the number was written down.

02

The arithmetic

What actually determines terminal value, and why it is rarely the return.

The most common mistake in asset management is to treat return as the thing you optimise and risk as the thing you tolerate. It is the wrong way round. The terminal value of a portfolio follows the geometric return, not the arithmetic one, and the difference between the two is a function of variance. Approximately, the geometric return equals the arithmetic return minus half the variance. Two portfolios with identical average returns therefore end up in different places solely because one of them moves around more than the other.

That is not a textbook claim here. It is the main finding of our own comparison of two single-stock strategies over 27 years. One ranks on momentum alone. The other ranks on a combined score of quality, value and momentum. At 100 000 kroner of starting capital the arithmetic means are 8.06 and 8.04 per cent, that is, identical. The entire difference in terminal value comes from the second moment.

Annualised meanAnnualised volatilityCAGRVolatility drag
Pure momentum7.68 %26.13 %4.32 %3.36 pp
Multifactor7.89 %18.21 %6.40 %1.49 pp
S&P 5009.28 %15.11 %8.46 %0.82 pp
Monthly returns, 30 June 1999 to 30 June 2026, 50 000 kroner of starting capital. Multifactor does not earn more than pure momentum. It loses less on the way.

The mechanism behind the drag is the asymmetry of losses. A 50 per cent fall requires 100 per cent to recover. An 86 per cent fall requires 620. The momentum strategy above had three falls of that kind, and its old high from March 2000 was not regained until June 2025, more than 25 years later. The multifactor strategy fell 58 per cent from July 2007 to March 2009 and spent six and a half years under water. In both cases the annual return is a summary of something that in practice was unbearable to sit through.

The practical consequence is that we look harder for lower variance than for higher returns. That is not caution. It is that variance reduction is the one of the two that can actually be achieved with reasonable confidence, while excess return in our own material disappeared every time it was measured properly.

Friction is a function of size, not of strategy

The literature that motivates systematic strategies is written mainly on institutional assumptions, where the order is large enough that fixed costs per order do not matter. That assumption does not hold for a small company. Our cost model follows actual terms at Interactive Brokers Pro, that is, 0.005 dollars per share with a floor of one dollar per order, three basis points of spread per side and five basis points of slippage per side. That comes to roughly 16 basis points per round trip plus a floor that binds on small orders.

Cost componentPure momentum, 50kPure momentum, 100kMultifactor, 50kMultifactor, 100k
Commission0.840 pp0.404 pp0.272 pp0.137 pp
Spread0.057 pp0.057 pp0.050 pp0.050 pp
Slippage0.636 pp0.635 pp0.423 pp0.424 pp
Total1.533 pp1.095 pp0.745 pp0.611 pp
Annual cost drag in percentage points. Spread and slippage are proportional and do not move with capital. Commission halves, because the one dollar floor is twice as large a share of a 2 500 kroner order as of a 5 000 kroner one.

Add tax on top and friction costs pure momentum 2.5 percentage points a year and multifactor 2.1. Both therefore spend roughly a quarter of the gross return on existing. A strategy can be profitable gross and unprofitable net purely because of portfolio size, without anything about the strategy having changed. That is the single most important difference between us and the literature we read.

Tax is a structural parameter, not a closing entry

A Norwegian limited company pays 22 per cent of realised gains on shares outside the EEA. Inside the EEA, gains are largely exempt for corporate investors under the participation exemption in section 2-38 of the Norwegian Tax Act. In isolation that means 12 per cent gross becomes 9.36 per cent net if the whole gain is realised continuously outside the exemption. That makes tax a parameter in the choice of strategy rather than a line entered afterwards.

The picture is more layered than the simple split suggests, and we write out why rather than round it off. The United Kingdom is no longer an EEA member and Switzerland never was. The two make up a substantial share of a broad European equity universe, which therefore carries an effective rate in the order of 7 to 8 per cent rather than zero. At the same time the minimum commission on European exchanges is several times higher than on American ones, and UK shares carry 0.5 per cent stamp duty on purchase. The ranking between markets therefore depends on order size, and it is not settled here. The European commission disadvantage has not been measured, and the choice of market stands as an open item.

Low tax paid is not in itself a good sign. The momentum strategy in our material paid 3 075 dollars in tax where the multifactor strategy paid 6 934. The reason is that the first carried losses forward in 23 of 28 years. The tax was low because it lost money.

03

Selection

What has to be true before something gets capital, and why the best result in a grid is the least credible one.

Selection is two different jobs here, and they carry different standards of evidence. A rule-based strategy can be fully specified and rerun, so the criteria can be set in numbers before the measurement. An assessment of a single company cannot, so the requirement moves onto the assumptions.

The gate for a systematic strategy

The criteria are written into the code before the first run and are not changed afterwards. That is not formalism. Without predefined criteria every measurement becomes a story, because the threshold can always be moved to wherever the result landed. Four requirements apply:

A parameter sweep, not a point.
A strategy is run across the whole space of reasonable parameter choices. A single result says nothing about whether it is a property of the construction or of the number that happened to be picked.
A split into first and second half.
The relationship between what worked in the first part of the period and what worked in the second is measured explicitly. If it is weak or negative, the grid is measuring regime rather than quality.
Costs, tax and settlement inside the simulation.
Not as a haircut at the end, but in each individual trade, with T+1 settlement and tax computed per year.
An explicit judgement on whether the finding exceeds noise.
With a monthly standard deviation of 3 percentage points in the difference, far more than 326 observations are needed to establish a couple of percentage points. 27 years is a lot in practice and little statistically.

The first strategy of ours that passed a formal robustness test also shows how little that is worth. The factor rotation gave a positive Sharpe in 12 of 12 parameter combinations, with a standard deviation of 0.148 against a mean of 0.411, and the worst drawdown inside the threshold. Robust by the definition. But only 4 of 12 beat buy-and-hold on Sharpe, average return was lower than the benchmark, and the correlation between Sharpe in the first and second half was 0.25. Our default choice had a Sharpe of 0.82 in the first half and minus 0.09 in the second. Robust here meant that the strategy does not collapse across the parameter space, not that it has a demonstrated edge.

The multifactor strategy was run through the same apparatus, with 30 combinations of five weight sets and three portfolio sizes, with and without sector neutralisation. Three of four criteria were met. The fourth was breached by half the grid. And the decisive measurement is this. The correlation between Sharpe in the first and second half is minus 0.72.

Minus 0.72 is not an absence of relationship. It is an inverse one. The combination that did best in the first half ranked 29 of 30 in the second. A choice made halfway through the period would have pointed systematically the wrong way, and that is the strongest single argument we have against selecting a configuration on historical data.

The mechanism is not noise, and it is visible. Value and quality led from 1999 to 2012. Momentum led from 2013 to 2026. The grid largely measures which factor regime each half belonged to. One combination in the entire material beats the index, namely quality plus value with 25 names, at 9.27 against 8.46 per cent a year and a shallower drawdown. It has not been promoted, and the reason is that it was chosen after all thirty had been run. Picking it would have been doing exactly what the material documents that you should not do.

This is the problem the literature calls multiple testing. Harvey, Liu and Zhu (2016) review hundreds of published factors and conclude that a t-statistic of 2.0 is no longer a meaningful threshold when the number of trials is large, proposing roughly 3.0. Bailey and co-authors (2014) formalise the same thing as a deflated Sharpe ratio, subtracting the expected maximum Sharpe from N random trials. The point is simple once stated. The best cell in a grid is not an estimate. It is an order statistic.

How the score itself is built

The multifactor score uses eight measures across three blocks, namely quality measured by return on equity, gross profitability over assets following Novy-Marx (2013) and leverage, value measured by four different yields, and momentum measured as return over twelve months excluding the most recent one. A company must have current values in twelve fundamental fields to be ranked at all, and qualification happens in the data layer before the strategy sees the universe. A company that drops out lands in an exclusion list with a reason.

Value is written as a yield, not as a multiple.
A multiple with a negative denominator flips sign. A company losing money gets a negative price to earnings, and in a sort on lowest first it ends up on top, as the cheapest company in the universe. The error is silent, and it hits precisely the companies a value filter is supposed to remove. The same applies to leverage on negative equity and to operating profit over enterprise value when a company holds more cash than its market value. Both are set to missing rather than being read as a sign of quality.
The measures become rank percentiles, not z-scores.
Fundamental ratios have heavy tails. A single return on equity of 400 per cent, from a company with almost no book equity, moves both the mean and the standard deviation enough to change the scores of the whole block. A ranking cares about order and not about distance.
Ranking happens within industry group.
Without it the value block becomes a sector bet in practice. Banks are structurally cheap on price to book and software structurally expensive, and a plain cross-sectional ranking would have filled the portfolio with financials in every period.
All measures are oriented to higher is better in one place in the code.
Scattered sort direction is the most likely source of a sign error that no test catches, because a flipped sign produces a perfectly valid result.

The choice of industry classification is grounded in a measurement rather than a preference. None of the available classifications carry history, so stability over time cannot be measured directly. What can be measured is how much independent judgement each of them contains, measured against the industry code the company itself reports to the authorities. Two of the schemes are in practice deterministic functions of that code. The third is not, and 58.4 per cent of its codes map to more than one sector, affecting 86 per cent of companies. It therefore carries editorial judgement that can be revised without a trace, and a revision would change a historical backtest. The choice fell on Fama-French twelve industries, which inherits the stability of the code and can be cited.

One finding from the construction is worth carrying, because it is the kind of error that does not announce itself. The block weights are equal in the code. They are not equal in the portfolio. Momentum is a single measure while quality and value are averages of three and four measures each, and an average of several rankings pulls toward the middle. Measured across all ranked names every individual measure is equally dispersed, with a standard deviation of 0.289, while the blocks are not, roughly 0.19 for quality, 0.20 for value and 0.289 for momentum. On dispersion alone that corresponds to an effective weight of about 42 per cent to momentum against 29 per cent to each of the other two. Equal weight in the code was not equal weight in the portfolio, and nobody had decided that it should not be.

Individual companies

For an individual company the question is not whether a statistical edge exists, but whether the assumptions in the valuation are the ones we actually believe. The starting point is return on capital against the cost of capital. Growth only creates value when the return on invested capital exceeds the cost of capital. If it is lower, growth destroys value, and the company is better off paying the money out. Growth itself is not an assumption you set freely, it is a consequence of how much is reinvested and how well it is reinvested.

The practical implication is that a discounted cash flow valuation is mainly a claim about how long an excess return lasts, not a claim about next year revenue. The terminal value typically makes up most of the value in such a model, and the terminal value is entirely an assumption about how fast competition erodes the advantage. We therefore write out the fade profile as a separate choice and show what the value becomes under alternative profiles, rather than delivering one number.

The balance sheet is read before the income statement. Net interest-bearing debt, maturity profile, covenants and working capital tie-up decide whether the company gets to choose the timing of its own decisions at all. And earnings quality is read against cash flow. Sloan (1996) documented that the part of earnings made up of accruals rather than cash is systematically less persistent than the cash part. Earnings not followed by cash flow are earnings with a shorter life, however good they look.

04

Measurement

Why most backtests are wrong, and what the error looks like when it does not announce itself.

A backtest result is largely a property of the measuring apparatus. That is the most transferable lesson from our work, and it is also why this section is longer than the one about results. An apparatus whose weaknesses you know is worth more than a result you cannot check.

Look-ahead, that is, the strategy seeing information that did not exist at the time of the trade, has three sources for a strategy operating on individual stocks. Price data is the obvious one. The other two are restatements, where accounting figures are corrected later and the corrected version is read back in time, and publication lag, where a figure is made available from period end rather than from the actual publication date. A fourth and related source is universe construction, where companies later delisted are left out of the periods in which they actually were index members.

Our framework enforces this through a contract. Every request must state a decision date, only observations available on or before that date are returned, only originally reported values are visible, and the universe is built on historical index membership. Missing data raises an explicit error instead of returning an empty result, so that absence cannot be mistaken for a valid observation.

What the contract does, on one company

Lehman Brothers was a member of the S&P 500 through the 30 June 2008 snapshot and filed for bankruptcy on 15 September the same year. On the decision date of 30 June the contract answered with a return on equity of 15.4 per cent, earnings of 3 461 million dollars and equity of 24 832 million dollars. The two preceding years gave 20.5 and 21.2 per cent. A quality ranking that day should have placed Lehman among the better names in the universe.

That is the correct answer, because it was the information available. The example also shows the publication lag concretely. The figures available on 30 June 2008 covered the fiscal period ending 29 February and became available on 9 April. They were 82 days old on the decision date. Filtering on period end instead of availability date would have given the strategy numbers it could not have seen.

The price of free data, measured

The claim that free data sources are not good enough for this has been tested directly against a commercial point-in-time dataset. The sample is 120 randomly drawn index members per membership date, 311 unique symbols in total. The share of the members of the day that the free source today has nothing at all on falls monotonically with the age of the membership.

Membership dateSampleMissingShare
30.06.20051205445.0 %
30.06.20151202420.0 %
30.06.2022120108.3 %
Between 92 and 96 per cent of the missing companies are delisted. This is survivorship in its purest form.

A backtest from 2005 built on that source would be missing almost half the universe, and the missing ones are systematically those that did worst. On top of that the source shows the wrong version of the accounting figures. Among the 42 observations where the commercial dataset has itself recorded a correction, the free source shows the corrected value in 27 cases, that is 64.3 per cent, and the originally reported one in 11. The matches are exact. The source delivers precisely the number the company reported later.

Two errors at once, and both pull the same way. The universe is missing the losers, and where a correction exists, the revised version is shown. The first produces returns that are too high. The second produces factor values that are too good. Neither announces itself in the result. Both look like a better strategy.

How much restatements matter has been measured on the dataset own as-reported dimension. Of 91 740 fiscal periods, 1 321 have more than one filing, that is 1.44 per cent. Measured per company, 724 of 1 159 are affected at least once, that is 62.5 per cent. The contrast between those two shares is the point. Corrections are rare per period and nearly universal over a company lifetime. A backtest reading corrected figures therefore does not get a small and evenly distributed error. It gets correct numbers for most periods, and systematically favourable numbers for precisely those periods where something went wrong.

Tyco International reported earnings for the quarter ending 31 March 2002 as plus 1 400 million dollars on 15 May, corrected on 12 June to minus 3 113 million. A quality factor computed on corrected figures would have correctly ranked the company among the weakest in the universe in the spring of 2002. That ranking did not exist when the decision had to be made.

Green tests that test nothing

A test suite can be green on a rule it never exercises. That is not a theoretical worry, and the only way to detect it is to introduce known faults into the code deliberately, one at a time, and verify that the right test actually turns red. We have run the exercise in three rounds, with 6, 13 and 11 injected faults. In the last round, 4 of 11 passed with no red test at all on the first attempt. None of the eleven raise an error. All produce a valid but wrong result.

Wrong date.
The restatement test passed even when the version filter was removed entirely, because it queried a date on which the correction had not yet been published. The rule was untested, with a green suite saying the opposite.
Wrong resolution.
The mutation computing momentum on unadjusted prices produced zero red tests. The test existed, but the tolerance was 0.03 while the actual difference for the test company was 0.026. The right quantity measured at a resolution that could not tell right from wrong.
Wrong layer.
The sector neutralisation test showed that the function neutralises correctly when given sector groups. It said nothing about whether the strategy actually passes groups. A test on a pure function tests the function, not that it is used.
Documentation against code.
The year-end tax routine took the realised gain out of the bucket and then sold enough to cover the liability. The gain from those sales was written back to a bucket that had already been emptied. More than 200 000 dollars went untaxed. The module own documentation described the correct behaviour precisely. No final number looked wrong.

Two lessons from the exercise are worth more than the faults it found. The number of red tests is a poor measure of how serious a fault is. Removing the version filter produced one red test, while filtering on period end instead of availability date produced nine, and the two faults are equally destructive. Prioritising by coverage counts means prioritising wrongly.

And an aggregate check cannot replace a concrete case chosen because it is extreme. The mutation computing market capitalisation on unadjusted prices passed the broad test across 400 companies with an 85 per cent match requirement, because large companies rarely reverse split. Only the narrow test turned red, the one checking a single company chosen precisely because it reverse split by a factor of 240. The tail is not random. It is exactly the small and distressed companies that reverse split, and the error would consistently have made them look cheap.

The most instructive fault was found by no test at all. The portfolio engine matched the rebalancing date by exact equality against the trading calendar, and if the date fell on a weekend or a holiday the rebalancing was skipped without a log line. That applied to 16 of 55 dates, that is 29 per cent. The fault was found because the number of rows in a result file did not match the number of dates stated in the text. Two things about it are worth more than the fault itself. It had direction, and made the momentum strategy 1.16 percentage points better than it is. And it flipped a conclusion, because the effect of the ranking band was measured at minus 0.12 percentage points with 39 rebalancings and plus 0.45 with all 55. A finding about a mechanism is also a finding about how often the mechanism gets to act.

The limit of what the controls can say is written down as well. The contract guarantees that a number was known on the decision date, not that it is correct. A wrong sign, a wrong currency, a wrong scaling or a ratio computed on a different definition than assumed passes unhindered. Quality control of the values themselves is a separate task and a standalone risk.

05

Risk

What actually produces the large falls, and how little three stress episodes can carry.

Risk is not the same as volatility, but volatility has a measurable price, and it is calculated in section 02. What ends a strategy in practice is nevertheless not the swings. It is the fall, and the length of the time spent under water.

In our material the source of the large falls has been identified, and it is not the one you would expect. We ran the same strategy as a ladder, adding one mechanism at a time, so that the contributions can be separated.

StepVolatilityMax drawdownSharpeCAGR
Pure momentum, no industry spread31.63 %−86.23 %0.2934.32 %
With ranking inside industry24.00 %−58.23 %0.3295.11 %
With quality and value added21.91 %−58.00 %0.3946.40 %
The sector band alone takes volatility down 7.63 percentage points and the drawdown up 27.99. The fundamental blocks add 2.09 and 0.24. Industry spread accounts for 99 per cent of the improvement in drawdown and 79 per cent in volatility.

That is a correction to the picture you would have drawn without the ladder. The risk reduction separating multifactor from pure momentum comes mainly from spreading the portfolio across industries, not from selecting companies on accounting figures. A concentrated momentum portfolio of fifteen names ends up in one industry at a time, and that is what produces falls of 86 per cent.

We also tested whether it is the actual industry classification that works, or merely dividing the portfolio across twelve groups whatever they are, by shuffling the group labels at random with five seeds and otherwise running identically. Random groups reproduce 19 per cent of the effect on drawdown and 30 per cent of the effect on volatility. Four fifths of the drawdown reduction therefore requires that the groups really are industries.

The practical consequence is the most usable single piece of information in the whole exercise. Sector neutralisation requires no fundamental data. It requires only an industry classification derived from a publicly reported code, and it delivers almost the entire drawdown reduction. The commercial dataset, by far the largest cost in the setup, pays for the remaining few per cent.

Every strategy has its own way of breaking

The holdings in the three stress episodes show that the strategies did exactly what they are built to do. In June 2000 the momentum strategy owned AMD, Sun, Micron, Oracle and Siebel, the dot-com names right before the fall. In June 2008 it owned energy and coal right before the commodity collapse. In December 2020 it owned growth names right before the 2021 correction. Momentum buys what has risen, and the concentration turns each episode into a full breakdown rather than a setback. The pattern is known in the literature as momentum crashes (Daniel and Moskowitz, 2016).

Value and quality fail in the opposite direction, which is equally instructive. In June 2000 the multifactor portfolio owned railways, chemicals, insurance, industrials and a newspaper, and fell 5.5 per cent through dot-com against the index 41.6. That is the one episode it handled well, and it explains almost the entire excess return over 27 years. In June 2008 it owned energy and materials alongside four insurers, that is, companies that were cheap on accounting figures right before both the commodity collapse and the financial crisis hit precisely those sectors. The value block pointed into the crisis rather than away from it.

All the drawdown measures in our material rest in reality on three observations, namely 2000, 2008 and 2020. The multifactor advantage was won in one of them. An argument resting on a single event cannot be strengthened by testing more parameters on that same event. That is why our next step is an independent sample rather than more variants.

Concentration and position size

The parameter sweep produced two patterns that recurred across weight sets. More names is better. 25 names gave the highest average Sharpe in all five weight sets and the lowest drawdown in four of five, and the ordering 25 ahead of 15 ahead of 10 was strict in four of five. Concentration into fifteen names costs, and it costs most in volatility.

That points to a general stance on position size. The Kelly criterion (Kelly, 1956) gives the stake fraction that maximises long-run growth when the distribution is known. The distribution is never known. And the error is asymmetric, since overbetting destroys the geometric return faster than underbetting costs in forgone upside. We therefore sit well below any computed optimum and treat the distance down as payment for estimation uncertainty.

The distinction we care most about is between permanent loss of capital and a temporary fall in price. The first comes from debt, forced selling, fraud or a business model that disappeared. The second passes if you can sit still. We do not use debt to increase exposure, and there is nobody who can force us to sell. That is the whole reason a fall is something we can absorb rather than something we must avoid.

06

Monitoring

What triggers a change, and why it is not the price.

Trading is a cost. That follows from section 02, and it makes monitoring an exercise in not acting. Under proportional costs the optimal rebalancing rule is not a calendar but a band. You trade when a position falls outside an interval, and not otherwise (Constantinides, 1986).

We use that concretely. A held position is kept as long as it stays inside the top 23, even though the portfolio holds only 15 names. The effect has been measured. The band gave 0.45 percentage points of higher annual return in the multifactor strategy and 0.15 in the momentum strategy, and turnover fell from 313 to 284 per cent. 72 positions were kept solely because they sat inside the buffer.

One detail stands here rather than being smoothed away. In the momentum strategy the band increased tax, from 2 625 to 3 075 dollars, while also increasing return. That is the opposite of the mechanism you would expect. A plausible explanation is that the variant without a band realises more losses, which do not trigger tax, but the mechanism has not been measured, and should not be presented as though it had.

The timing of a trade is a rule as well, not a habit. If a rebalancing date falls on a weekend or a holiday, it rolls forward to the next trading day. Forward and not backward, because the ranking is computed per snapshot date, and executing it on the last trading day before would be trading on information that did not yet exist. The consequence is that year-end rebalancings execute in January, so the gain is realised in the following tax year. That is correct, since the trade actually happens then.

When something is retired

The criterion for exiting is written down at the same time as the position is taken, and it is phrased as something that can occur rather than as a price level. Without that every decline becomes a story about why you should sit still, and every rise a confirmation. A fall in price is not in itself new information about the company. What is new information is that an assumption that was written down no longer holds.

The same applies to how we report our own work. A conclusion that changes gets a note about what changed and why. When the rebalancing fault was fixed, the conclusion about the ranking band flipped sign, and both the old and the new are written down. A number without its history is harder to trust than a number that has been wrong once and corrected.

07

Open items

What we have not measured yet, and what comes next.

A method is only credible if it also states what it has not done. Four items stand open, and they are written here in the same form as in the report.

The factor regression has not been run.
It would show whether what remains of the multifactor return is exposure to known factors, which can be bought more cheaply through index funds. It cannot change the main result, since the strategy already loses to the benchmark, but it is missing.
The European commission disadvantage has not been measured.
The minimum charges and the stamp duty are known, but no simulation has been run on a European universe with European rates. The choice of market therefore stands as an open methodological item, and the threshold for revisiting it should be computed rather than assumed.
The values in the dataset have not been quality checked.
The contract guarantees the timing, not the content. A wrong sign, currency or scaling, or a ratio computed on a different definition than assumed, passes unhindered.
The industry classification has no history.
The table is a present-day snapshot, so a company that changed industry in 2007 carries today value in a rebalancing from 1999. The direction of the bias is known, and it is documented rather than corrected.

The next step is not more variants on the same 27 years. It is an independent sample, that is, a different universe, and for us that is also the relevant universe, namely European listings under the participation exemption. The most informative test there is not a new multifactor variant, but running the same ladder again, namely pure momentum, momentum with industry spread, and multifactor. If the split between the steps holds, industry spread is a general finding and the fundamental data an expensive addition. If it reverses, this sample is explained.

What comes along regardless is the framework. The point-in-time contract, the portfolio engine with cost and tax modelling, and the validation methodology are built to carry strategies other than the ones tested. Any strategy that ranks companies on accounting figures, which is the core of fundamental analysis, can be run through the same layer and thereby inherits no-look-ahead, historical index membership and explicit friction. That is the lasting deliverable from this work.

The basis

Every measurement on this page is taken from our own report on systematic trading strategies, which sets out the methodology in full, states every parameter choice and lists the sources for what is borrowed from others.

Read the report →