Skip to content
tessel

Record it. Understand it. Automate it.

01 · Record

Know what every decision was based on.

Decisions get made in meetings, notebooks and chat threads. A month later, nobody can say why.

Decision ledger and proof graph

Answer “why did we ship this?” in seconds, with the evidence.

Every claim tied to the runs, data and model versions behind it, from question to decision.

Decision record

Open the held-out set?

Holds on the new devicerun 2291 · results by siterun 2304 · after rebalance
The gain is realrun 2287 · 20 fresh splits
No leakagerun 2290 · near-duplicate search

Automatic capture

Your record fills itself. Nobody has to remember to log anything.

Your runs, notebooks and coding agents, recorded as they happen. Or use our SDK.

Captured automaticallylive
14:02Run 2304 started by your coding agentcaptured
14:02Dataset v14 · new-device scans addedcaptured
14:31Sensitivity 0.91 on the new devicecaptured
14:31Linked to “Open the held-out set?”linked

02 · Understand

Find the signals that answer your questions.

Metrics say a model is good. They don’t say why, where it will break, or whether it’s safe to rely on.

Investigations

From “it broke” to “here’s why, and the fix.”

Bring a question about your model. tessel designs and runs the experiments on your model and data, and returns the signals that answer it.

Investigation3 experiments

Why does it miss more on the new device?

  1. Results by site and device
  2. What the misses have in common
  3. Retrain without the new device

Signal: misses cluster at the edge of the image. Rebalance, then re-check.

Interpretability

Catch the shortcuts that won’t survive new data.

Feature visualization and sparse features show what the model relies on.

What the model relies on
Ultrasound of a normal liver, with the vessels a vision model picks up highlightedfeature 7 · liver vessels

Simulation

Spend your time on the questions that could change the decision.

Before anything runs, simulate how far each answer could shift the decision.

How far each answer could shift the decisionbefore anything runs
New device
Gain is real
Leakage
Bigger backbone

03 · Automate

Autonomy you can justify.

The same decisions come up every week. Before a system makes them on its own, you need to know how far to trust it, and the moment that changes.

Decision benchmark

Know how far you can trust a system to make the call.

Models, processes and agents are scored against your own past decisions before they get any autonomy.

Decision benchmark
Past decisionSystem’s callYour team’s call
Ship v2.3 to the new siteyesyes
Retrain on the winter datayesno
Switch to a larger backbonenono
Agrees with your team on 22 of 24 past decisions

Continuous monitoring

Know the moment new data stops looking like what you decided on.

New data and outcomes are checked against what each decision assumed. Discrepancies are flagged automatically, and a person steps in.

New data vs. what the decision assumedweekly

Anything outside the range goes to a person.

Learning from outcomes

Every decision makes the next one better.

What happened is compared with what was expected. The gap updates which signals matter, for your team and for automated decisions.

Expected vs. what happened
DecisionExpectedTurned out
Open the held-out set0.88–0.920.90
Ship v2.3 to the new siteholdsdrifted at week 6
The gap becomes a new signal: site drift is now checked first.

How it works

Start from the decision, not the metric.

Most evals and benchmarks are built before anyone knows the decision they’ll inform.

1Weigh every question2Learn from past decisions3Ask the few that matter4Audit the decision
The decisionShould we open the held-out set? You make the call.
Your record
Should we open the held-out set?Today · evidence supports opening it, after the rebalance
Ship v2.3 to the new site?March · A new device changed the decision
Retrain on the winter data?June · Leakage was found, and fixed
Switch to a larger backbone?August · It didn’t change the decision
What if the new device is different?12h
Is the gain real?8h
Is there leakage?4h
What about the new hospital?16h
Does it hold on low-quality scans?12h
Does it hold for older patients?12h
What if readers disagree?20h
What about pediatric scans?2d
What about rare findings?12h
We want more data first.3w
Train for longer?2d
Clean up the training code first?1d
Relabel the hard cases?2w
Use a bigger backbone?1w
More augmentation?16h
Try a different loss?2d
What about an ensemble?3d
Tune the learning rate?8h

18 questions on the table: what if, what about, we want this.

Where teams start.

Before a launch, a release or a submission

“Is it ready?”

tessel tests each claim the decision rests on, with the held-out set untouched until the end.

A decision brief, every claim tied to its evidence, ready for whoever signs off.

When a model breaks on data it hasn’t seen

“Why did it fail?”

tessel finds where it breaks, shows what the model relies on, and tests the causes one by one.

The mechanism of failure, and the fix: data to add, a threshold to move, a simpler model.

When there are more ideas than time

“What should we try next?”

tessel simulates how each answer could shift the decision, and runs only the few that could change it.

A ranked plan you approve, and results as they land: experiments in hours, not days.

When the same decisions come up every release

“Can this run without us?”

Your past decisions become a benchmark. Routine decisions run on their own; drift goes to a person.

Time back, consistent decisions, and an early warning when new data stops matching.

Questions about tessel.

How is this different from evals or monitoring?

Evals and monitoring produce numbers: often answers to questions you didn’t need to ask, or partial answers to the ones you did. tessel flips that. It starts from the decision, works out which questions could change it, and answers each one properly, running your evals where they help and building the ones that are missing. Doing that for every call takes infrastructure, and that’s what tessel is: a record of every decision, what it rested on and what happened next.

How do I get started?

Book a call and bring one decision you need to make. We connect tessel to the model and data behind it and work with whoever knows what the decision depends on. You see every result as it lands, and the decision goes into your record.

Who does the work, your team or ours?

Your team, on tessel. Your engineers keep working the way they do, with every decision recorded and the right signals in front of them. When you need more hands, or evaluations you can’t run in-house, our forward-deployed engineers do the work with you, on the same platform and the same record, and more of it runs on its own as tessel grows. Either way, you make the call.

Is this a certification?

No. tessel gives whoever certifies your model the evidence to do it: every claim tied to the runs, data and model versions behind it, including where the evidence runs out. The whole record is queryable, so a regulatory submission, an audit or a release review starts from evidence, not memory.

Start with one decision.

Bring us the decision you’re least sure about. We’ll show you what it takes to make it with confidence.