The method

How we work

The same method for research, for client work, and for capital. The only thing that changes between them is how expensive we let it get.

This page is the methodology as published

You have probably noticed that the answer usually matches whoever is being paid. A consultancy’s diagnosis lands on a project it can sell. A fund’s diligence lands on the deal it already wanted. A paper written to support a position finds support.

Nobody has to be dishonest for that to happen. It only takes an incentive and nothing in the process designed to push back.

Below is what pushes back. It is not complicated and none of it is secret. It is just work that most firms have no commercial reason to do.

Seven steps

1

State the claim before gathering evidence.

Hypotheses are written as falsifiable statements with operational definitions, before anything is retrieved. What counts as “failure”, as “production”, as “ROI” is fixed in advance, so the definitions cannot drift toward the conclusion while the work is under way.

#01

2

Write down what would prove you wrong.

Every hypothesis carries kill criteria and a disconfirmation path: which search, in which sources, would surface the strongest evidence against us. It is written before we look, because afterwards it is much harder to want to.

#02

3

Gather the evidence blind.

The people and the models retrieving evidence receive the question and a neutral working title. They never receive the thesis, the preferred angle, or prior findings. A search that already knows what it is looking for will find it, which is why retrieval and argument are kept apart by design rather than by discipline.

#03

4

Pay someone to attack it.

An adversarial pass tries to break every claim. Verdicts use six fixed terms rather than a judgment call: contradicted, unsupported, ambiguous, supported with qualification, not assessable, survived specified tests. A human adjudicates anything load‑bearing.

#04

5

Have an independent person check the load‑bearing numbers.

Independent means having nothing to gain. An author cannot verify his own report. An employee cannot produce an error rate we then describe as independent. Where no external verifier is available for a piece of work, the check is labeled internal and non‑independent in the published text. Honest labeling costs a little credibility. An internal check described as independent costs all of it.

#05

6

Hold a gate where abandoning the work is a real option.

Evidence meets hypothesis, and the outcomes are proceed, reframe, or kill. A killed thesis is a published result, not a buried one.

#06

7

Publish the error rate, and what would change the conclusion.

First‑pass error rates with confidence intervals, material and minor reported separately. Two numbers, never conflated: the observed error rate among the claims we checked is a fact about what we checked, and the estimated prevalence across the whole report is an inference from a sample. Reporting the first as though it were the second is exactly the category error this method exists to prevent, so we report both.

#07

Rigor is proportional. Honesty is not.

Running a three‑day brief through journal‑grade machinery is waste. Running a flagship report through brief‑grade machinery and calling it a review is the one thing we will not do.

So the expensive parts scale with the work: systematic search, independent verification, external review, preregistration. Five things never scale, because they are cheap.

  1. Every claim traces to a source a reader can check.
  2. Nothing is asserted beyond what the evidence supports.
  3. Uncertainty is stated wherever it is material.
  4. Scope and limits are declared: what we searched, what we did not, and what would change the answer.
  5. Errors are corrected in place, dated, and visible. Silent edits are prohibited.

When there is not enough time, we downgrade the claim, never the honesty.

Six kinds of claim, and we know which one we are making.

Most dishonesty in consulting and in finance is the presentation of an assumption as a measurement. It rarely involves a lie. It involves a category error nobody corrects.

Category What it means
Observed factDirectly measured or documented
External evidenceSupported by a cited third‑party source
900 Labs findingDerived from our own research
InferenceOur interpretation of the evidence
ForecastA statement about what we expect to happen
Commercial assumptionRequired for a model, not established as fact

Compare “AI will reduce this function’s cost by 40%” with “our base case assumes a 40% reduction, based primarily on X and Y, at medium confidence.” Same business. Entirely different standard.

The taxonomy applies to a research paper, an investment memo, a sales proposal, and this website.

The Truth Floor

Every other part of our operating standard can be relaxed under pressure, with the deviation logged, reasoned, and reviewed afterwards against what happened. One part cannot. No cash‑flow exception, no founder exception, no strategic exception.

900 Labs will not:

  1. tell a client we expect an outcome we do not actually expect
  2. represent weak evidence as strong evidence
  3. present correlation as causation
  4. hide material uncertainty because it makes a sale harder
  5. conceal a material conflict of interest
  6. knowingly sell work whose expected value to the client is negative
  7. manipulate a research conclusion because a commercial outcome depends on it
  8. continue defending a thesis after the evidence has materially changed

What this costs us

The method is worth describing only if the costs are described with it.

  • A firm whose method includes kill criteria will sometimes tell a prospect not to buy, and will not get paid. That is the proof, and it is expensive.
  • Publishing diligence error rates is genuinely risky. “We were wrong on three of ten theses” is not a comfortable sentence in an investor report, and there is no version of this where it becomes comfortable.
  • Rigor costs time, and clients want speed. Proportional classes handle it, but only if the scope is agreed up front. Otherwise every engagement quietly becomes a research project.
  • It requires independent people. Verifiers, refuters and red‑teamers with no stake in the outcome are a recurring cost across all three applications, not a one‑off.

The differentiation exists because it is costly. If it were free, everyone would already have it.

The recipe is not the chef.

We publish the method. What we do at each stage, why, what has to be true before the work proceeds, how a claim is graded, and what would stop us. That is a methods section, and it is what a reader needs in order to judge whether a finding can be trusted. Without it, “we were rigorous” is just a word we chose.

We do not publish the implementation. The retrieval prompts, which models run which lanes, the extraction schemas, the structure of the claim database, the audit scripts. That is a lab notebook. No journal asks for one and no serious reader expects it.

The distinction is not a hedge. Anyone can make cola. The recipe is not what makes the chef, and the same knife chops a salad or skins an animal.

What cannot be copied

A competitor could adopt every step on this page tomorrow. They would still hold zero published error rates, zero calibration data, and no record of what they predicted against what subsequently happened.

The asset is not the technique. It is the history the technique produces: an accumulating database of claims and the sources they rest on, a public error‑rate record that gets longer every time we publish, and first‑party outcome data from our own engagements and deployments. Those take years, and they cannot be acquired by reading a web page.

Which is exactly why we can afford to be generous with the method.

What we have actually run

An honest inventory, since a page about method is worth nothing without one.

  • 01The AI Execution Gap, March 2026, was produced under an earlier version of the pipeline. It used cross‑model adversarial review and primary‑source verification, and its methodology note says exactly that. It was not preregistered, and it does not carry an error rate.
  • 02The current methodology, version 2.2, has been revised twice following external methods review. Its first full application is The Agentic Enterprise, whose protocol is approved and whose public registration is deliberately on hold until the methodology change lands. Registration is permanent, so it is worth being a few days late.
  • 03The first published error rate will therefore appear on our next report, not the last one.
Read the research this produced
900 Labs · Page revised 2026‑08‑09 · Corrections: none