Contexxt Perspective · PSP-001

The Rehearsal Layer

Imagine coming to a fork in the road. In life you choose one path and go down it. But what if you could go down every single pathway, then come back to the beginning with the knowledge of what lies along each of them, and choose the best one? What if we also asked why you chose that pathway, and why the others weren't the ones you chose? Now you have knowledge of the path you've taken, plus all the paths you didn't. This is the rehearsal layer we've built in: any decision can have a multitude of pathways, and learnings from all of them.

One person at the head of six lit pathways through a dark landscape — each path a simulated route through the same decision
EXECUTIVE SUMMARY
  • The Problem: Organisations make decisions, commit to them, then wait months to evaluate outcomes. This slow feedback loop creates bias, fear of risk, and minimal learning — a company making one major decision yearly collects only ten data points per decade.
  • The Solution: A rehearsal layer embedded in the decision process that simulates multiple pathways using adversarial panels, parameter sweeps, and independent bake-offs — before committing to reality.
  • The Core Mechanisms: Engineer independence in simulated panels; maintain calibration through prediction logging; surface remarkable insights through a "membrane" that captures human judgement.
  • The Compounding Advantage: Every rehearsal's reasoning, assumptions, and human rulings are retained and feed the next decision. Your tenth decision is sharper and cheaper than your first.
  • The Data Sovereignty Angle: Unlike cloud AI vendors, this rehearsal layer keeps your knowledge yours. The reasoning, memory, and learning stay your asset. You own the loop.
WHY THIS MATTERS

Most organisations learn about a decision months after making it, and only about the path they took. By then the cost is sunk. A rehearsal layer moves that learning to before you commit, while it is still cheap.

And the learning stays yours. Every ruling your people make in rehearsal is captured and feeds the next one. Your tenth decision is sharper than your first, not because you hired better people, but because the process compounds.

01 · One observation per decision

The issue we face today is that an organisation can make a decision, commit to it, and then have to wait months before evaluating it, before the post-mortem that tells it whether or not it took the optimal path. And that has significant drawbacks. There are cognitive biases at work; people judge with the benefit of hindsight; post-mortems can create an environment of fear, so calculated risks are no longer taken. An organisation that makes one market-defining decision a year collects ten data points a decade.

What we really want is to understand, and to gather knowledge and learning. And if we gather that not just on the path taken but on the paths that could have been taken, we have so much more opportunity to gather compounding knowledge and learning. You can model a different choice against last quarter's numbers any time you like. What you can't do is rerun the actual quarter and watch the other choice play out.

What you can't do is rerun the actual quarter and watch the other choice play out.

02 · Rehearsal multiplies the evidence

The rehearsal layer solves that problem. Today we only get to review the decision we made. Simulating decision pathways lets us run against multiple assumptions and scenarios, and test them using various instruments. Panels of simulated participants, each with defined characteristics, objectives and beliefs, standing in for the people who aren't in the room: boards, regulators, readers, buyers, the market itself. Adversarial verification, whose only job is to break a conclusion: to test it, to make sure we've thought about how people might break that decision. Parameter sweeps, which show which assumption is actually carrying the answer. And finally, bake-offs between independently built answers, which tell you whether a conclusion is a fact about the world or an artefact of the method that produced it.

Adversarial verification is a testing method where a dedicated team's only job is to break your conclusion. They don't argue for a position; they try to find every flaw, assumption, and edge case that could make your decision wrong.
Why it matters here:
This isn't criticism — it's stress-testing. It ensures you've thought through how people might challenge your decision before you commit to it. This is what separates a decision that survives scrutiny from one that crumbles.

Throughout this entire process, humans are in the loop. Where decisions are being made, or could be made, they can be validated or challenged by a human, ensuring we are learning about the decisions we're making. Ultimately we want every decision to land as close as possible to where the organisation expected it to.

03 · The obvious objection

The first and most obvious objection is always the simulated panels. Aren't they just consensus machines? The naive answer is yes, and the research tends to agree. Models are trained towards a typical answer, so as time goes by they will converge on the same answers. And simulated panellists who share a base model will end up achieving absolutely none of the objectives we set out to deliver.

So what we need to do is engineer independence. We do this by enforcing stances, beliefs and commitments on every simulated participant, and we reapply them throughout the process, because, believe it or not, agents cave to peer pressure if you merely tell them to hold a view. Our panels use adversarial composition and are heavy on challengers, so when we rerun the same question with different casts, the finding is what survives.

Adversarial composition means deliberately building your simulated panel to include sceptics and challengers alongside believers. You're not stacking the deck; you're engineering disagreement. When you run the same question three times with different panel mixes and get the same conclusion, that result survives because it's robust across fundamentally different viewpoints.
Why it matters here:
This is how we prevent the consensus machine problem. Homogeneous panels will always agree. But if independent panels with conflicting beliefs all arrive at the same decision, you've found something real — not an artefact of the method.

Throughout this process, we also look for remarkable insights. We surface those notable snippets to the appropriate persons, who can comment on them, and we learn from those comments.

The brief that sets off this whole process matters as much as the panel you are forming. If you write it to reach a particular conclusion you have in mind, that conclusion will be reached through almost any panel. So the brief has to be thoroughly audited.

There's a number of studies that validate our thinking. One in particular was a Stanford blind study where machine ideas beat researchers' ideas on novelty, but didn't beat them on feasibility. Which is validating, because this is exactly where we see the humans in the loop coming in, in our architecture. We're focused on delivering structured disagreement. Because ultimately every instrument can be gamed. The safeguard is that we test things against each other, adversarially, and openly publish what breaks.

04 · Calibration keeps it honest

I'm sure you've experienced the sycophancy of the large language models. As Plutarch wrote in about 100 AD, "the flatterer's object is to please in everything he does, whereas a true friend always does what is right." Simulations are a product of large language models, so they can flatter. A rehearsal that always says yes, or is never challenging, is worse than no rehearsal at all, because it gives the person or the team doing the rehearsing an unfounded confidence ahead of a live decision.

A rehearsal that always says yes, or is never challenging, is worse than no rehearsal at all.

So we ensure that everything is built upon a calibration spine. Predictions are logged and checked against what actually happened in reality. Drift is measured and corrected.

Calibration spine is the backbone that turns simulation from guesswork into validated forecasting. Every prediction from every rehearsal is logged with a date and assumptions. Later, when reality happens, you compare prediction to outcome. Over hundreds of predictions, you build a calibration curve.
Why it matters here:
Without calibration, you're just building better-sounding guesses. With it, you know whether your rehearsals actually predict reality. Drift gets caught and corrected. This is what separates a rehearsal layer from a very expensive opinion machine.

We recently engaged with an investment bank. Across the research runs we've published, our own verification has refused four of nine. All four still went to the bank, with the refusals attached and the corrections that were made, so that they could see our reasoning, and make decisions aware of our workings and our refusals. Contextually better decisions, because nothing was hidden from them.

Ultimately, we want to engineer the crisis in a simulation, so you never hold it in reality.

Should verification refusals be published alongside decisions?

05 · What rehearsal cannot know

Rehearsals run on the record: the public record, and the institution's own corpus. What the record can't know is what a shop assistant gleaned on his way to work on a Tuesday, and we do not pretend we can. We try to mitigate this by ensuring that when we think we've discovered something remarkable, we escalate it to the appropriate people to review, because in doing that we may well stimulate ideas.

There are some things that we can't solve for. Things like internal politics and influence. And tacit knowledge (don't ask Hattie to be creative if Spurs lost at the weekend, so best not to ask her at all).

Tacit knowledge is knowledge that isn't written down — the insights, instincts, and understanding that people carry in their heads. It's hard to document because it's often intuitive or context-dependent. A colleague might "just know" something works based on years of experience, but struggle to explain why.
Why it matters here:
Rehearsals run on documented data. They can't access what people know but haven't written down. This is a fundamental limitation — and why the membrane (surfacing insights to humans for judgement) is so critical. That's how tacit knowledge gets captured and made permanent.

We try our best by introducing a rehearsal membrane. When a rehearsal surfaces a notable objection or insight, it's pushed through the membrane to the relevant person, who then rules on it. That contribution is placed on the record and becomes an attribution that is available for every run after it. If we do that hundreds of times, the knowledge store fills with harvestable insights: with the very knowledge it was missing when it started. The trick is ensuring that people care enough, that the effort is rewarding and never routine, because the moment it becomes a reflex, the human has to all intents and purposes left the loop.

The rehearsal membrane is the boundary between simulation and human judgement. When a rehearsal surfaces a notable objection or insight, it pushes through the membrane to the right person. They rule on it. Their decision gets recorded and becomes available to every future rehearsal.
Why it matters here:
This is how tacit knowledge gets captured and made permanent. Over hundreds of rulings, the knowledge store fills with insights your organisation discovered in real time. The membrane ensures humans stay in the loop and the system keeps learning from their judgement.

We have limitations, two in particular. Obviously, we can only capture what we promote up as comment-worthy, so whatever hasn't been challenged, or raised as possibly challenge-worthy, stays outside. We also need to ensure that the right knowledge is being pulled into the decision process: whose knowledge gets pulled in is a design decision, not something an organisation chart can decide.

The final piece, and the one we are working towards solving, is how you audit a decision not to act. Because doing nothing produces no outcomes or paths to score. We absolutely log it. But we can't grade it yet...yet!

The claim isn't that we can predict what audiences think and might do with panels. What we do is question which decision survives rehearsal. And those decisions that survive fail less often. Calibration is how we know.

06 · Memory makes it compound

Every rehearsal we run is retained. With them, the reasoning behind them, the rejection and assumption attributions. Just holding that is an archive, which is OK, but we want it to be much more useful. We want these runs to feed reinforcing loops: each run should add to the platform's knowledge.

The key to success is ensuring what deserves attention gets surfaced to the people who rule on it, and what they do with it. Do they dismiss it, and why? Do they accept it, and what are the consequences? This ensures the next rehearsal starts from a learnt position. The advantage is that people are in the process, not reviewers at the end. This gives us access to hundreds of judgements a year, instead of just a handful of post-mortem anecdotes. The reasoning is captured in the moment, accruing that knowledge to the institution rather than leaving it with a single person. A tenth decision will be sharper and cheaper than the first, and not because anyone worked harder, but because of our learnings along the way.

There are two things that organisations lose: the memory of restraint (what didn't we do, and why?) and disagreement records (what was argued, what was the contention, what was the debate, and which rationale won?). If all of that is captured during the process, you don't just have a decision log; you have more: a disagreement log, which tells you about the thinking that went into those decisions.

It all comes back to the fact that committed decisions only give you one path. Rehearsing gives you a lineage across every future examined. Reality gives you one path. Our rehearsal layer gives you the map you were standing on when you chose it.

There was a map illustrating what was known vs unknown by Columbus posted on X. Which I thought was a good illustration of what we try to provide. We look to give you map 2 as you consider your voyage.

Map 1: the world as Columbus knew it, bordered by unknown territory Map 2: the world with the voyage routes overlaid

x.com/Mcuquerella/status/2094455318319169734

Does the idea of compounding organisational learning feel achievable for your context?

07 · And the loop is yours

One big argument in AI right now is who ends up owning what your company learns. Satya Nadella has called it the Reverse Information Paradox. The old problem for people selling knowledge was that to sell it, you have to show it, and once you've shown it, it loses its value. With AI, that flips, because the person using the AI is the one who is giving something up. As a user of AI, I impart my knowledge, and the AI consumes it; the deal we make is that if I give, I get better results. So ultimately, I end up paying for the AI twice: once monetarily, and then again by gifting the AI vendors my know-how. The vendors want me to impart as much knowledge into the system as possible; as I keep on correcting, I'm fixing their wrong answers, and every time I fix a wrong answer, the machine is contextually becoming better and learning more about my business.

Palantir's point is that companies already have decision systems they suggest that you should own that record, but it's severely limited, it's only the record of action. A single pathway.

Pignataro, the ION founder, says in his essay The Wrong Apocalypse that the real value of knowledge work isn't actually the thinking itself; it's how everything gets pulled together, and how the decisions are coordinated. And ultimately, anyone using the frontier models is teaching those models the tacit knowledge of their industry and company.

I think they're all right. But the thing they're not allowing for, which we do, is that they assume only one path is being taken. We allow a multitude of paths to be taken, so we can examine a multitude of paths and understand the decision process around all of them.

And if you think about it: if, as people are saying, the corrections we make are the most valuable thing that leaks out of a company, that they are its unique IP, then the rulings our people make during rehearsals, pure judgement captured in the moment, are probably the most concentrated version of that there is. Which makes the rehearsal layer we're building hugely valuable, and the last operation you should ever run on someone else's learning loop.

The last operation you should ever run on someone else's learning loop.

So we think we have two answers, really, built into how we work. The first is that we allow people to explore pathways and futures without telling anyone. If you go and ask an outside person a question, you're already beginning to leak it. But if you ask it within your walls, you can test the whole idea and nobody knows what you're looking at. And the second is that we have built this to ensure we are not reliant on one single AI model. That's by design. The memory and the learning stay the client's asset. We run the loop. We never keep the data. When we leave, the loop stays.

08 · A layer, not a workshop

Decision rehearsal isn't new. War-gaming has been around for ages. But it's a point-in-time event. People come together for two or so days, you gather real insights, then the room empties and the learning walks out with it.

Every decision has two moments: the moment you intend to do something and then the moment you commit to doing it. There's a gap between those two positions, and that's where our rehearsal layer lives. Between the intent to do something and the commitment to do it.

Think about it like software. No code ever goes through to production without passing test. Nobody schedules a testing workshop; the release process just flows through test. Just as decisions should flow through the rehearsal layer.

When we've gone out to enterprises to talk about the layer, we assumed it would be a bone of contention. But when we put the question to people around consequential decisions, they don't debate the value. What they immediately jump to is how to ensure success, and how the compounding can be ensured. The same four questions kept coming up. How do we ensure that people capture and integrate the roads not taken, when documenting what wasn't done is so easily dropped? How do you ensure the fidelity of the rehearsals and the integrity of their outcomes? How do you manage notability thresholds: the agreed points where a rehearsal result flips the default from proceed to re-evaluate? And how do you capture a surprise, when a rehearsal produces something, the organisation really didn't see coming?

None of this is technology and it's why the layer has to be built in rather than bolted on when the stakes are so high.

Reality runs once. So when it does, you've already been there.

BRING THE DECISION YOU NEED TO GET RIGHT →

Before you go

References

Nick Hencher, Co-Founder & Managing Director · Contexxt · September 2026