
Probabilistic Lineage: How Much Uncertainty Was Already Baked Into the Input?
Upstream models classify, repair, and infer before your main model ever sees the data. Nobody tracks how much probability shaped the input. The research landscape, the traps, and a workable design.
Here’s a question I keep coming back to, and I don’t think the industry has an answer yet.
We’re embedding probabilistic components throughout the stack — increasingly upstream of the main inference. Data gets classified by a model, repaired by a model, enriched by a model, summarised by a model, and then handed to the agent that actually decides something. By the time the consequential model sees its input, that input may have been probabilistically transformed four or five times.
How much probability has already shaped the input before the next model sees it?
And the follow-ons that matter operationally: should the downstream agent know? Should it know how reliable those earlier decisions were? At what point does accumulated probabilistic depth warrant different treatment from deterministically produced data?
This is usually filed as an ML reliability question. It’s a governance question. My last post argued that the harness is where intent becomes an execution contract, and the third property of that contract is evidence a third party can verify. Conventional lineage tells a third party what transformed what. It says nothing about what the result is worth — and a derived artifact that gets treated as a source is exactly where confidence is manufactured out of nothing.
What already exists
More than I expected when I went looking, spread across four communities that mostly don’t cite each other.
Systems-level uncertainty propagation. Xia, Zhu, Gao, Lu, Xue and Sejdinovic’s Uncertainty Propagation in LLM-Based Systems (arXiv 2604.23505, April 2026) is the closest direct treatment. Their framing: uncertainty is usually studied at the level of a single output, but deployed applications are compound systems where uncertainty is transformed and reused across model internals, workflow stages, component boundaries, persistent state, and human processes. Their open-problems section names the hard part precisely — composition requires two distinct properties, calibration preservation (does the composed signal stay calibrated against outcomes) and scope preservation (does it retain a valid relationship to the thing it was originally about). Current handoff mechanisms assume both and provide formal conditions for neither.
The snowball finding, which is the most operationally useful result I’ve seen here (arXiv 2608.14588). Hallucinations injected at stage one of a multi-agent pipeline don’t merely persist — they transform. A raw numerical fact becomes a derived computation, becomes narrative prose, becomes an editorially approved conclusion. At each transformation, detectability degrades near-irreversibly. Their prescription is a placement rule: put your verification at the first handoff, where roughly three-quarters of hallucinations are still catchable, rather than at the end.
That reframes the whole question. Depth isn’t a counter, it’s a change of form. A number you can check against a source becomes a sentence you can only check against a judgement.
Step-level propagation methods exist too — SAUP (ACL 2025) propagates uncertainty through each step of an agent’s reasoning; UProp (Duan et al., 2026) addresses multi-step decision-making; there’s a 2026 taxonomy survey mapping estimation methods to pipeline stages at arXiv 2609.07395.
Older prior art the LLM literature mostly ignores, and which I’d argue is more relevant than any of the above:
- Uncertainty-lineage databases. Trio, out of Stanford in the mid-2000s, made uncertainty and lineage jointly first-class in the data model. It is the closest thing to a solved version of this problem — twenty years early, in a deterministic-query setting.
- Sculley et al., “Hidden Technical Debt in Machine Learning Systems” (2015). Names correction cascades and undeclared consumers. What we’re discussing is that paper’s failure mode with a generative model in the loop.
- Taint tracking and information-flow control. I think this is the right mental model — better than error bars. Probabilistic origin behaves like a taint label: it propagates through derivation, it’s difficult to remove, and the interesting design question is where you place a declassifier.
- Conformal prediction, because it’s distribution-free and composes under explicit risk budgets, unlike verbalised confidence.
And the transport layer, which omits all of this by design. A2A deliberately leaves out model-level metadata, quality signals, provenance and trust-domain enforcement; MCP operates at the tool level and doesn’t address delegation semantics. On the memory side the gap is explicit — a wire schema for a shared memory entry would need authorship, time, confidence and provenance, and neither A2A’s Parts and Artifacts nor MCP’s Resources type that. Meanwhile the identity side has moved faster: invocation-bound capability tokens chain verifiable delegation with provenance binding across MCP and A2A. We can prove who transformed something before we can express how confidently.
Three traps to avoid before you build anything
Naive composition is worse than nothing. Multiplying stage confidences assumes independence, and pipeline errors are strongly correlated — shared training data, shared retrieval corpus, shared prompt conventions, shared model family. You get a number that looks rigorous and is systematically overconfident. False precision laundering is a worse failure than an absent signal, because it licenses downstream reliance rather than merely failing to restrict it.
Verbalised model confidence is poorly calibrated. Propagating it faithfully propagates a miscalibration with a decimal point attached.
Scope drift. “85% confident this field is a date” and “85% confident this conclusion is sound” are not the same object. Composing them produces a number about nothing. This is the scope-preservation problem in practice, and it’s why a single scalar is the wrong data structure.
The implication: don’t build a confidence number. Build a provenance class. Categorical, honest, and hard to misuse.
A workable design
Six parts. None require research breakthroughs; all of them require someone deciding to do it.
1. Track derivation as a label, not an error bar
Attach to every material input a small, categorical record: what kind of transformation produced it, by what component, with what verifiability. Borrowed straight from taint tracking — the label propagates through derivation and can only be cleared deliberately.
Minimum viable classes, ordered by how far they sit from checkable ground truth:
| Class | Meaning | Checkable against |
|---|---|---|
| Observed | Recorded directly from a system of record | The source system |
| Deterministic-derived | Computed by rule from observed data | Recomputation |
| Model-classified | A label or extraction produced by a model | The underlying artifact |
| Model-repaired | Missing or malformed values inferred by a model | Nothing, usually |
| Model-synthesised | Prose, summary, or narrative generated from the above | Judgement only |
| Model-asserted | A conclusion presented as fact, origin no longer visible | Nothing |
Note the bottom two rows are where the snowball lands. Once something becomes model-asserted, it will be consumed as observed unless the label travels with it.
2. Record the form transition, not just the hop count
Depth measured as “three models touched this” is nearly useless. Depth measured as “this crossed from classification into synthesis” is actionable, because that’s the transition after which detectability collapses. Log the transition, not the tally.
3. Verify at the first handoff
The placement economics are counterintuitive and well-evidenced: verification is cheapest and most effective early, because the artifact is still in a form you can check against a source. A validation gate between extraction and computation earns more than a review step before publication. Most pipelines do the opposite — they put the human at the end, where the only thing left to check is plausible prose.
4. Carry the label across handoffs yourself
The protocols won’t do it for you, so put it in your own envelope. A minimal handoff record per material input:
| Field | Purpose |
|---|---|
derivation_class | The categories above |
producing_component + version | Attribution, and change detection |
form_transitions | Ordered list of class changes, with the component that caused each |
ground_truth_anchor | Identifier of the last checkable source, if any |
verification_state | Unverified / machine-verified / human-verified, with timestamp |
scope | What the confidence, if any, is about |
policy_version | The rules in force when this was produced |
That’s six fields and a version string. It is not a research programme; it’s a JSON object that most teams could add to their orchestration layer this quarter. And it plugs directly into the execution contract from my previous post — provenance class becomes a declared property of every input, the way tool worst-case is a declared property of every action.
5. Escalate treatment by depth — as a tier escalator, not a threshold
Don’t look for the magic number of transformations at which data becomes untrustworthy; there isn’t one. Instead raise the required treatment when any of these hold:
- The input crossed a form transition (classification → inference → synthesis → assertion)
- Part of the chain runs through a component you don’t control — a vendor’s model, an external agent, an MCP server
- The artifact will be relied on as fact by a party who can’t see its lineage
- The consequential action is irreversible, regulated, or rights-affecting
That last combination — a model-asserted input feeding an irreversible action — is the one I’d treat as a hard gate requiring human authorisation, regardless of how confident anything claims to be.
6. Place declassifiers deliberately
This is the part that makes the whole thing tractable, and it comes straight from information-flow control. A taint label that can never be cleared eventually paints everything, and a system where all data is tainted has the same information content as one where none is. So you need explicit declassification points: places where a derived artifact is re-anchored to checkable ground truth and its label is legitimately reset.
Three that work in practice: re-derivation from the system of record; human verification with attribution and a timestamp; and a conformal or statistical check with a stated coverage guarantee. What doesn’t work: a downstream model asserting that the upstream output looked fine. That’s not declassification, it’s laundering — and it’s the default behaviour of most multi-agent pipelines today.
Where this meets the frameworks
Worth being precise, because the gap is real and worth naming to a client or an auditor.
- ISO/IEC 42001 — A.7.3 to A.7.6 cover data acquisition, quality, and provenance, and A.6.2.8 covers event logs sufficient for investigation and accountability. A.10.3 covers allocating responsibilities across the AI value chain. All of these support what I’ve described. None of them require uncertainty to be propagated across component boundaries.
- NIST AI RMF — MAP covers documenting provenance and intended use; MEASURE covers uncertainty quantification. Again, no composition requirement.
- EU AI Act — Article 10 imposes data governance and quality obligations on high-risk systems; Article 13 requires instructions for use that let deployers interpret output. Arguably an input whose provenance class is unknowable makes Article 13 hard to satisfy honestly, but nothing states that.
So: every major framework requires you to document provenance, and none requires you to carry uncertainty across a handoff. That’s a genuine gap between what the standards ask and what compound systems need, and it’s the kind of gap that gets closed by practice first and by revision later. Organisations that start now will be ahead of the requirement rather than retrofitting under it — which is the same position ISO 42001 early adopters found themselves in.
Five questions I’d ask in an assessment
- For your most consequential model, can you state the derivation class of each material input?
- Where in your pipeline does something stop being checkable against a source and start being checkable only against judgement?
- Where is your verification placed — at the first handoff, or before publication?
- When an artifact crosses into a component you don’t control, what travels with it?
- Where are your declassification points, and is any of them just a model saying it looks fine?
If question five turns up a model verifying a model, you’ve found the highest-leverage fix in the system.
From the practitioner’s chair
Two observations from the audit side.
The recurring lesson from leading ShareVault through ISO 42001 Stage 2 certification and then serving as internal auditor: controls are rarely the failure point, evidence is. Probabilistic lineage is the same lesson arriving one layer deeper. It isn’t enough to show that a decision was made and logged — at some point a counterparty will ask what the decision rested on, and “a model said so, and before that another model said so” is not an answer that survives the question.
And from auditing a client’s MCP Governance Standard, where nearly all 27 changes in my v1.1 redline reduced to one idea — authority must be bound to a specific action rather than held ambiently by a component — there’s a direct analogue here. Confidence must be bound to a specific claim rather than held ambiently by an artifact. An unaudienced token is authority without a destination; a floating confidence score is reliability without a referent. Same failure, different currency.
What to do in the next 90 days
- Map one consequential pipeline end to end and mark every point where a model touched the data before the final decision. Most teams find more than they expect, and the surprises are usually upstream.
- Classify the inputs using the six derivation classes. No numbers yet — categories only.
- Find the form transitions and move one verification step earlier.
- Add the envelope to your orchestration layer. Seven fields.
- Name your declassification points, and remove any that are a model checking a model.
- Add a hard gate where model-asserted input meets an irreversible action.
The honest summary of the field: there are four partial answers, none of them compose, and the protocols deliberately don’t carry the field. That’s not a reason to wait. It’s an argument for building the categorical version now — because the teams that can say what their inputs are made of will be the ones able to answer the counterparty who eventually asks.
#ProbabilisticLineage #UncertaintyPropagationLLM, #Agenthandofftrustsignals, #MCPA2Aprovenance, #ISO42001 #Dataprovenance, #Hallucinationsnowball, #Conformal prediction
Work with DISC InfoSec
DISC InfoSec helps B2B SaaS and financial services organisations govern compound AI systems — AI and agent inventories, data provenance and lineage design, risk tiering, harness and MCP permission review, human oversight placement, and the evidence architecture that maps to ISO/IEC 42001, NIST AI RMF, and EU AI Act Articles 10, 13, 14 and 26.
I led ShareVault through ISO 42001 Stage 2 certification on the first audit attempt as the internal practitioner, served as internal auditor, and authored their MCP Governance Standard.
Readiness path: free 15–20 minute readiness call → ISO 42001 gap assessment or ISO 27001 gap assessment → 7–10 day AI Governance Quick-Start → full AIMS/ISMS implementation → vCISO / vCAIO retainer.
DiscInfoSec — Principal Consultant, DISC InfoSec (Deura Information Security Consulting LLC), Petaluma, CA CISSP, CISM | ISO/IEC 42001 & ISO/IEC 27001 Lead Implementer | PECB Authorized Training Partner
calendly.com/hd-deurainfosec info@deurainfosec.com (707) 998-5164 deurainfosec.com
References
- Xia, B.; Zhu, L.; Gao, E.; Lu, Q.; Xue, M.; Sejdinovic, D. — Uncertainty Propagation in LLM-Based Systems, arXiv:2604.23505, April 2026
- The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines, arXiv:2608.14588, 2026
- Uncertainty Propagation on LLM Agent (SAUP), ACL 2025; Duan et al., UProp, 2026; Uncertainty Quantification for LLM Agents: A Taxonomy, an Evaluation Protocol, and an Empirical Study, arXiv:2609.07395
- Widom, J. et al. — Trio / uncertainty-lineage databases, Stanford, mid-2000s
- Sculley, D. et al. — Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015
- Vovk et al., conformal prediction; conformal risk control literature
- Model Context Protocol; Agent2Agent (A2A) specification and Agent Cards; agent identity and delegation-provenance work (invocation-bound capability tokens)
- ISO/IEC 42001:2023 — A.6.2.8, A.7.3–A.7.6, A.10.3; NIST AI RMF 1.0 — MAP, MEASURE; Regulation (EU) 2024/1689, Arts. 10, 13, 14, 26
























