A block explorer will tell an institution, with total confidence, which addresses moved how much value at what block height. That certainty is real, and it is also narrower than most valuation memos, compliance files and diligence decks assume. The gap between what a public ledger records and what an analyst reports as fact is where an institution's actual risk sits.
Bitcoin, Ethereum and most of the assets built on them publish every transaction to anyone who wants to look. But a raw ledger is a list of addresses, not a list of names. Turning one into the other is the job of the blockchain analytics industry, and that job is done by inference, heuristics and clustering, not by reading a name off the chain. An institution that treats a vendor's attribution as equivalent to the underlying transaction data is treating a conclusion as if it were a fact.
What a block explorer actually proves
Directly, a public ledger proves that a quantity of an asset moved from one address to another at a given time, under a valid cryptographic signature. That is strong evidence: it cannot be altered after the fact, and it does not depend on any single party's records. On its own, it is silent about who controls either address, why the transfer happened, whether the sender had authority to make it, or whether the assets are encumbered by a loan, a lien or a side letter no blockchain will ever record. The same limit bounds what a proof of reserves can establish. Analytics platforms can be very good at describing flows between addresses; they cannot, from chain data alone, tell an institution who is legally entitled to what.
The attribution problem: from address to entity
Turning addresses into entities is called clustering. It rests on a small number of public heuristics, which commercial vendors extend with proprietary methods and off-chain intelligence. Two foundational heuristics do much of the work on Bitcoin and other UTXO-based chains.
Two foundational heuristics
The first, the common-input-ownership or "co-spend" heuristic, assumes that when a transaction spends from several addresses at once, one entity controls the keys for all of them, because each input needs its own signature. Bitcoin's white paper itself acknowledged that multi-input transactions leave some linking unavoidable. The assumption has a known exception: CoinJoin and other collaborative transactions are built to combine inputs from different owners, and they break it. The second, change-address detection, tries to identify which output in a transaction is "change" returning to the sender rather than a payment, so that the sender's cluster can be extended forward. It is far weaker, because it relies on guessing from patterns in how wallets build transactions, and naive versions of it cause "cluster collapse", where unrelated users are merged into a single, wrong, oversized entity. Malte Möser and Arvind Narayanan's paper "Resurrecting Address Clustering in Bitcoin", presented at Financial Cryptography and Data Security 2022, built a ground-truth set of transactions with known change and developed ways to detect and prevent that collapse.
How the more defensible co-spend heuristic performs against verified data was measured in a preprint first posted to arXiv in July 2026 and revised on 3 September 2026, by researchers including Bernhard Haslhofer and Christian Rückert. Their ground truth was address-to-entity mappings that European crypto-asset service providers are legally required to report to financial authorities. Restricted to those reported addresses, the heuristic looked strong: it merged no two reported services, and it recovered 71% of the true same-service address pairs. The authors note that this result is driven by one large service. Measured across the full clusters, precision fell to 0.36 and recall to 0.44, with an average error rate of 0.51. By entity, the spread was stark: on the largest service the heuristic scored an F1 of 0.825, and on the second-largest it scored 0.000, with an error rate of 0.999. The authors did not filter out CoinJoin or mixing transactions, and say the picture may change if that is done. Their conclusion is that multi-input clustering "should be treated as an investigative lead rather than an absolutely reliable attribution mechanism", and they warn that "dataset-level metrics can appear reassuring even when MIH fails almost completely for individual services".
What a Daubert ruling actually decided
The clearest test of how far a US court will trust this kind of evidence came in the Bitcoin Fog prosecution. On 29 February 2024, Judge Randolph Moss of the US District Court for the District of Columbia denied a defence motion to exclude expert testimony built on Chainalysis's Reactor software, finding it reliable under Federal Rule of Evidence 702. It is worth being precise about what the ruling did and did not establish. The court rejected the "black box" argument because the defence had received extensive material on how the clustering was done, including a confidential supplemental production. It accepted the co-spend heuristic in part because the weakness it exploits was recognised in Bitcoin's own white paper. It also recorded testimony that Chainalysis had not compiled false positives and false negatives in a central place, because the tool is designed to be conservative, and held that the lack of a compiled error rate did not alter its finding of reliability. What carried that finding was corroboration outside the software: FBI testimony that agents check attributions against exchange subpoena returns routinely; undercover transactions whose five addresses were traced by hand, four of which Reactor had attributed correctly while leaving out the fifth; and agreement from other tools, including TRM Labs and CipherTrace. A federal court found the evidence admissible. It did so on the strength of disclosure and corroboration, not a quantified error rate.
Confirmed labels are scarce
Part of the reason error rates are hard to compile is that confirmed labels are scarce even for researchers. The Elliptic Data Set, released in 2019 and described by its authors as, to their knowledge, the largest labelled transaction data set publicly available in any cryptocurrency, covers 203,769 Bitcoin transactions. Of those, 4,545 (about 2%) are labelled illicit and 42,019 (about 21%) licit; the remaining 157,205, about 77%, carry no label, according to the breakdown in a 2025 Scientific Reports paper built on it. That data set was built to classify transactions rather than to test attribution, but it shows how thin the confirmed ground truth is. The 2026 preprint makes the same point about clustering directly: despite its broad adoption, the multi-input heuristic "has rarely been evaluated against reliable ground truth data."
Where regulators say the limit sits
Standard-setters and supervisors that expect firms to use this technology have been candid about where it stops. The Financial Action Task Force's October 2021 Updated Guidance for a Risk-based Approach to Virtual Assets and Virtual Asset Service Providers, which is non-binding guidance for national authorities, states that "to date, the FATF is not aware of any technically proven means of identifying the VASP that manages the beneficiary wallet exhaustively, precisely, and accurately in all circumstances and from the VA address alone", as Notabene's reading of paragraph 197 sets out. That sentence sits inside guidance that still expects firms to identify counterparty providers for Travel Rule purposes, as Skadden's summary of the update describes. New York's Department of Financial Services, in guidance dated 28 April 2022 to its licensed virtual currency businesses, set out the use of blockchain analytics for know-your-customer controls, transaction monitoring and sanctions screening, and added its own caveat: such tools "may not be able to identify underlying owners, including ultimate beneficial owners," "may have limited attribution capability, absent further 'off-chain' verification methods," and their effectiveness "can vary depending on the particular virtual currency in question." Both describe the same tool the same way: useful enough to expect, limited enough to need corroboration. Pseudonymity is already thin once a regulated intermediary has linked a wallet to a verified customer, a point we take up in privacy in regulated finance; a third-party analytics firm without that link works with less information than the exchange the funds passed through.
That tension gave a Proof of Talk Paris 2026 panel its title. "Privacy and Compliance: Two Sides of the Same Coin", held on the Taostats Stage at the Louvre Palace on 2 June 2026, brought together James Smith, billed as Co-Founder and CSO of Elliptic, with speakers from BitGo, Deutsche Börse Group and Midnight Foundation, moderated by Nicola Massella of Storm. The full programme is on the agenda.
What the ledger cannot show at all
Some of the most consequential facts for an institution are not obscured by clustering error; they are never written to a public chain at all. Loan agreements, rehypothecation, side letters and other off-chain liabilities leave no onchain trace, which is why solvency cannot be inferred from address balances. Exchange internal ledgers are a related blind spot: venues commonly hold customer assets in pooled, omnibus wallets, so the chain shows the venue's aggregate position, not any customer's entitlement within it. New York's Department of Financial Services addressed this in guidance dated 23 January 2023 on custodial structures, which expects custodians holding customer assets in omnibus onchain wallets to keep records so that "each individual customer's beneficial interest is always evident and up-to-date." That division lives in the custodian's books, not on the chain. The same applies to institutional custody arrangements generally: a chain explorer can confirm that an address a custodian has identified as its own holds a given balance, and it cannot confirm which client that balance belongs to, in what proportion, or whether it has been pledged elsewhere. The same gap shapes how digital asset treasury companies are scrutinised: analysts can watch disclosed wallets, but cannot tell from the chain whether undisclosed wallets, credit lines or derivatives exposure sit alongside them.
Using the data without overtrusting it
None of this makes blockchain analytics a poor tool. It makes it a tool with a documented failure mode that varies by chain, by heuristic and by the specific entity being examined. For compliance functions, that argues for treating a clustering-based alert as the start of an investigation rather than its conclusion, which is how the Bitcoin Fog court's corroboration-heavy reasoning reads. For valuation, diligence and audit functions, it argues for asking a vendor for the confidence behind an attribution, not just the attribution, and for checking anything material against an independent, off-chain source before it enters a balance sheet or an investment memo. This is not legal, accounting or investment advice; it is a description, drawn from what courts, regulators and researchers have said, of where a useful technology stops being able to answer the question an institution is asking.