Certain about
uncertainty

Certezaħ builds world models of AI failure, turning uncertainty in agentic systems into measurable risk. A neutral measurement layer for reinsurers, insurers and brokers.

Scroll
The problem

We are building an economy of agents without a science of how it fails

Software used to answer questions. It now acts: approving payments, screening claims, supporting clinical decisions. And it does not fail like traditional software. An uncertain agent does not raise its hand; it guesses, and hands that guess downstream as fact.

The failures that matter most are the ones seen least often. They live in the tail, invisible to average benchmarks. And they are correlated: thousands of deployments rest on the same few foundation models, so one upstream change shifts behaviour everywhere at once. Not a series of incidents, but one accumulation event.

Today this exposure is assessed through questionnaires. They show that policies exist. They do not show how a workflow crosses from normal operation into loss, or which dependencies it shares with everyone else.

How it works

Two layers: simulation, then prediction

Observe failure under controlled conditions. Learn its structure. Recognise it before it occurs.

01

Simulation

We map a workflow first: which agent holds what authority, where the handoffs are, what the human review step actually catches, which external dependencies sit underneath. Then we stress it under controlled fault conditions and record how it breaks. Because the fault is ours, every run is labelled by construction, and the result is an observed failure rate for a real process rather than an estimate.

02

Prediction

Those runs train a world model of the workflow's dynamics, learned from behavioural telemetry: latency, tool calls, handoff payloads, state transitions. The AI is one component among review steps, tools and dependencies, and the model learns how the whole process moves. It identifies when a live workflow drifts toward the regions where failure begins, with no fault injected.

03

From measurement to risk

Failure rates and drift scores become underwriting inputs: loss exceedance curves, accumulation exposure, and correlation across portfolios of systems sharing the same upstream dependencies.

stable basin failure region separatrix
A workflow traced through the learned state space. Drift is measurable
before the boundary is crossed, and long before output degrades.
Why us

Detect, explain, quantify. Never intervene.

We observe and report. We do not modify outputs or gate decisions, because an instrument that changes what it measures is no longer measuring it. Detection surfaces the risk, explanation makes it auditable, quantification turns it into expected loss.

The method is physics, not another language model. A model asked to judge another model inherits every blind spot it was meant to catch. We learn system dynamics from behavioural telemetry, latency, tool calls, handoffs, state transitions, and treat failure as what it is: a trajectory leaving a stable region.

And we do not underwrite. A company that rates the systems it insures is rating its own book. Because we carry no risk, vendor, broker, carrier and reinsurer work from the same figures, and disagree about the price, not the measurement.

Credit risk models became foundational to finance. Cyber risk models became foundational to the internet. Agentic AI needs the same, and does not yet have it.

Certezaħ
Existing evaluation tools

Measures the whole workflow

Authority, handoffs, dependencies and controls, not the model alone

Scores the model alone

A model judging a model, blind to the process around it

Simulation, then a world model of failure

Faults injected on purpose, then learned well enough to anticipate

Expert judgement

Estimates with no observed loss history behind them

Risk expressed as loss

Severity and expected loss an underwriter can price

No standard for conversion

Quality scores that never become a financial quantity

Between agents and across portfolios

Handoff failures and shared dependencies accumulating together

One agent at a time

Correlated failures counted as independent events

Every score traces back

To the telemetry and the runs that produced it

Black box

Trust us, the model says it is fine

What you gain

Underwrite with evidence, deploy with confidence

Price on loss data

Loss exceedance curves for agentic workflows, derived from observed failure rates rather than expert estimate. Higher confidence in the number, and a faster route to it.

See the accumulation

Portfolio correlation across insured systems sharing the same foundation models, frameworks and tooling. The exposure a single-risk view cannot show.

Monitor in force

Continuous failure-propensity scoring on live systems, so a workflow drifting toward failure is visible while cover is running, not after the loss is reported.

Become insurable

For AI vendors: evidence of how a system behaves under stress. The basis for coverage, for enterprise procurement, and for the assurance obligations arriving with the EU AI Act and the Product Liability Directive.

Who we are

Two founders

Rare-event mathematics and interdisciplinary product research, applied to systems where uncertainty carries consequence.

André de Oliveira Gomes

André de Oliveira Gomes

CO-FOUNDER & CEO

André has followed a single question throughout his work: how do complex systems fail when the rare event finally arrives. From a PhD in stochastic dynamical systems and rare-event mathematics to deploying AI in regulated healthcare, he has worked where trust, uncertainty and consequence meet. As CEO he leads the company and its technical architecture, and builds the scientific and actuarial engine behind certezaħ.

LinkedIn ↗
Maximilian Wichmann

Maximilian Wichmann

CO-FOUNDER & CPO

Maximilian believes new possibilities come from asking the questions nobody has asked yet, the way quantum mechanics came from asking something the old rules could not answer. To find answers for the unknown, you first have to know that you don't know. Trained in medical physics and working across interdisciplinary fields, he turns that instinct into products that deliver real value. He leads product, market strategy and operations at certezaħ, and co-defines its mission with André.

LinkedIn ↗
無為

From a world of uncertainty to certezaħ, where AI and humans coexist in dynamic equilibrium

We aim for wu wei, effortless action. A state where humans and AI work together without friction or doubt, where constant manual correction becomes unnecessary.

It cannot be forced. It emerges only on a foundation of safety, reliability and trust, where both sides know their limits: AI receptive enough to know what it doesn't know, humans wise enough to trust the signal. Certezaħ builds that foundation, so you don't wrestle with the AI. You flow with it.

Request early access

We are working with a small number of reinsurers and brokers as design partners, to build the first loss dataset for agentic-AI failure. If that is you, we would like to talk.