- Europe’s loan-level securitisation disclosures cover over €2.5 trillion of credit, reported monthly — but remain substantially uncomputed by the market.
- The obstacle isn’t access, it’s history: inconsistent templates and evolving standards mean the data must be harmonised before it’s usable — roughly nine months of engineering in our experience.
- Once cleaned, the tape reveals borrower, collateral, and legal-process behaviour at the individual loan level — observed joint behaviour across more than 100 million loans, not a forecast plucked from the data.
- The barrier is fixed and paid upfront, but the advantage compounds: every reporting cycle extends the validated history the models are trained on.
European securitisation regulation obliges issuers to report standardised loan-level data to repositories such as the European DataWarehouse—over €2.5 trillion of securitised credit, reported monthly. It is easily read as a compliance artefact: filed, sampled for diligence, rarely computed upon. We read it differently. It is Europe’s equivalent of the US mortgage tape datasets that generated significant alpha for systematic credit investors after the global financial crisis.
Extracting signals from it is still harder than the equivalent task in the US market: Europe’s reporting regime is the younger of the two, and its templates and quality standards have changed repeatedly since inception. Turning the raw disclosure into modelling-ready data was a substantial engineering undertaking—and that difficulty is precisely the point: the preparation layer is where the barrier to entry sits. Disclosure was designed for transparency; systematically processed, it prices risk.
A dataset hiding in plain sight
After the crisis, European regulators concluded that securitisation had failed partly for informational reasons: investors could not see through pools to the loans beneath them. The response was structural. The ECB’s loan-level initiative, and later the EU Securitisation Regulation, obliged issuers to report standardised, loan-by-loan data to designated repositories—monthly, in prescribed templates, covering residential mortgages, SME loans, consumer credit, auto finance, and commercial real estate. The European DataWarehouse now holds that record.
The precedent for what such a record is worth comes from the United States. In the years that followed, loan-level mortgage tapes became the raw material for a generation of systematic credit investors: granular borrower histories, processed at scale, repriced a market that had been trading on pool-level assumptions. Europe has since assembled a dataset that is broader in asset-class coverage and more standardised in format—and it remains, in our assessment, substantially uncomputed.
Why the tape resists use
The obstacle is not access; the repositories are open to qualifying market participants. The obstacle is history. Loan-level reporting in Europe only started becoming the norm after the crisis, and quality standards and expectations have had to evolve since—across template generations, jurisdictions, and reporting practices. A dataset assembled under an evolving standard arrives as regulation shaped it, not as analysis needs it: histories must be consolidated, conventions harmonised across issuers and vintages, and every field validated before it can carry a model.
That work—schema harmonisation, entity resolution, plausibility validation, and the construction of a dedicated feature store—consumed roughly nine months of engineering before a single predictive model could be estimated. That figure is worth stating plainly, because it is the price of admission, and it is paid in a currency most investment organisations do not budget: patient, unglamorous data work with no deployment to show for it.
From disclosure to signal
Once the tape is clean, it stops being a compliance record and becomes a behavioural one. Monthly loan histories reveal how borrowers in specific segments respond to rate shocks; how arrears roll toward default or cure by jurisdiction and servicer; how prepayment behaviour varies with equity and seasoning; how recoveries actually time out through different legal systems. Enriched with variables the templates do not carry—court backlogs, collateral liquidity, probabilistic valuations—these histories support probability distributions at the level of the individual loan.
This is the substance of the signal: not a forecast plucked from the data, but the empirical joint behaviour of borrowers, collateral, and legal process, observed across more than 100 million loans and multiple cycles. Pool-level pricing cannot see it, because pool-level pricing consumes the summary and discards the sample.
The economics of the barrier
An informational advantage built on public data may appear fragile. We would argue the opposite. The cost of entry is fixed, upfront, and certain, while its payoff is deferred and uncertain—an unattractive trade for any organisation that measures progress in capital deployed. And the advantage compounds: every additional reporting cycle extends the cleaned history, refines the calibrations, and enlarges the set of realised outcomes against which models are validated. Replicating a modelling technique takes months. Replicating years of accumulated, validated loan-level history is a different order of problem.
Implications for institutional investors
For allocators, the attraction of this approach is not only the return stream but its auditability. The inputs are regulatory disclosures—standardised, timestamped, and independently held. Every underwriting assumption can be traced to observable loan histories rather than to management estimates. Mandated transparency, in other words, does more than level the informational field; properly computed, it makes the informational edge itself verifiable.
The advantage belongs to whoever computes it.