A blockchain indexer can answer every request and still produce an incomplete position report. Availability measures whether the service responds. Data quality concerns what the response includes, how it was derived and whether another process can reproduce it. Institutional users need both, with separate evidence for each.

The distinction becomes material when indexed data feeds holdings reports, monitoring or transaction reconciliation. An empty result can mean no activity occurred. It can also mean the wrong contract was indexed, the relevant historical blocks were excluded or a mapping failed. The response code alone cannot tell the consuming business which explanation applies.

The indexer interprets the chain

An indexer usually transforms selected chain data into a queryable model. Selection and transformation introduce assumptions. The Graph's subgraph manifest documentation shows how a deployment specifies data sources, contract interfaces, handlers and a starting block. That is useful primary evidence of why an indexed dataset has a defined scope rather than automatic coverage of everything onchain.

The institutional record should state the chain, included addresses, event or call coverage and historical range. It should also explain the business interpretation applied by the mapping. A transfer event might be relevant to ownership movement, but it does not establish every economic right or offchain adjustment associated with the asset. Data lineage needs to preserve those boundaries instead of letting a convenient field name imply a complete position.

A provider can make this scope reviewable by supplying a versioned dataset definition. The buyer can then compare that definition with the reporting obligation. If the business needs all wallet movements but the indexer covers only selected token contracts, the gap is a coverage decision. Adding another availability endpoint will not repair it.

Freshness needs a block reference

A query result needs a reference point that identifies the chain state used. Wall-clock response time does not show whether the indexer has caught up. The Graph's GraphQL API documentation describes metadata including deployment, block information and an indexing-error flag. These fields help consumers distinguish the request's success from the state of the dataset behind it.

For a business report, block number and hash are useful because they give reconciliation a concrete target. The report can also record the indexer's deployment identity and retrieval time. Those references do not prove that every mapping is correct. They make disagreement diagnosable: two results may have different cutoffs, different mappings or different interpretations of the same data.

Freshness thresholds should follow the task. A monitoring view and a period-end statement have different needs, and the institution may choose different rules for provisional and settled information. The report should describe those choices. Without them, an analyst may treat a recent-looking response as final simply because the interface does not expose the chain reference.

Compare at the same chain state

Ethereum's JSON-RPC documentation sets out methods for retrieving blocks, receipts and logs, along with block references used by relevant methods. These provide a route to selected underlying chain evidence. Comparing an indexer's total with an independently retrieved sample is useful only when chain, address, range and reference state agree.

A common analytical mistake is to compare a current wallet balance with an indexed transfer total measured at an earlier cutoff. A mismatch may be a timing difference rather than a mapping defect. The reverse mistake is to dismiss every mismatch as lag. A quality process first aligns the references, then examines any residual difference against the dataset's documented scope.

The relation between chain state and settlement confidence is covered in Blockchain finality and provisional settlement. An indexer adds another layer: even correctly chosen finality rules do not establish that all relevant records were ingested and interpreted correctly. Chain confidence and dataset completeness need separate checks.

Mapping changes can alter history

A mapping converts technical events into entities or calculated fields. A revision may fix an error, add coverage or change a definition. Each can alter reported history when the dataset is rebuilt. Institutional users need to know whether a new report differs because the chain changed, a previous interpretation was corrected or the business adopted a new definition.

A change record can state the affected fields, historical interval, reason and comparison results. It should preserve the old report as an artefact rather than silently replacing it. That makes corrections explainable to downstream users who have already relied on the earlier figures. A new deployment identifier is helpful, but it does not by itself describe the business effect of the change.

Tests should use examples that challenge the interpretation: repeated events in one transaction, contract upgrades, unusual decimals or an event that reverses an earlier state. The exact examples depend on the indexed application. The meaningful requirement is that expected results are established independently of the transformation being tested, rather than simply copied from its current output.

Historical queries require retention

The Graph's manifest documentation also describes historical pruning controls. The institutional implication is a question for the service agreement: can the required historical states still be queried, and for how long? A service able to answer today's holdings query may be unable to reproduce last quarter's mutable entities. Those capabilities should not be bundled under one generic data access promise.

The buyer can specify the artefacts it needs retained. They may include raw extracts, mapping versions, report outputs and the block references used for a close. Retaining a result allows review of what was delivered; retaining the ingredients supports reconstruction. The two serve different purposes, and the organisation should know which one its provider actually supports.

A reconstruction exercise can ask for a past report using only approved retained material. Missing history, changed definitions and unavailable dependencies become visible immediately. The exercise is more informative than an assurance that the indexer has operated for several years. Service longevity says little about the ability to reproduce one defined historical view.

Quality checks need meaningful samples

A sample should test the areas where the business could be misled. For transaction reconciliation, that might mean tracing a recorded transfer back to a receipt and then following its transformation into the report. For an exposure view, it might mean choosing a known address and checking whether every relevant event in a fixed interval is included. The method follows the decision being supported.

Negative controls also matter. A dataset can pass checks on familiar positive examples while failing to distinguish absent activity from excluded activity. The reviewer can test an address outside scope, a block before the configured start and an intentionally invalid query. Expected responses should make those conditions legible to the consumer, rather than giving the same empty result for each.

The broader interpretation problem is discussed in What onchain analytics can tell you. Indexer quality is the operational part of that problem: the dataset needs evidence for its coverage, transformation and omissions before its interpretation can be trusted for a particular institutional task.

Procurement should separate service and data measures

An availability agreement can describe successful requests and response times. A data agreement can describe maximum lag for defined workloads, coverage, correction procedures and evidence retention. These measures interact, but a buyer needs to see them separately. A fast response with an indexing error is not interchangeable with a complete response after an acknowledged delay.

Keep economic labels separate from raw events

A dataset may attach labels such as exchange, treasury or related counterparty to an address. Those labels are interpretations maintained outside the raw transaction record. The quality process should identify their source and effective date, because a corrected label can change an exposure report without changing any chain data. Readers need to know which layer produced the difference.

A sampled report can preserve the raw address alongside its interpreted label and the label version used. That allows a reviewer to reproduce the classification while testing the underlying movements independently. If a provider cannot disclose its attribution method, the buyer should state that limitation in the use of the dataset. An otherwise complete event history does not remove uncertainty about who controls the addresses appearing in it.

Responsibility for downstream correction also needs an owner. If an indexer discovers a historical omission, who identifies affected reports and informs the teams using them? A silent data rebuild may repair the current endpoint while leaving old decisions and exports unchanged. The correction process needs to follow the data into its institutional uses.

The result is a more precise procurement conversation. The question is not whether an indexer is reliable in the abstract. It is whether a defined dataset, at a defined chain state, supports a defined business decision and can explain exceptions. Availability keeps the pipe open. Coverage and reproducibility establish what passed through it.