Skip to concept
Protopilot

Find quality

Which version works best—and why?

Compare credible routes on whether people understand what to do and complete the task. Choose only after every version has been tried the same way.

Direct answer

Record what works, then be selective.

What to compare
Whether people understand the task and complete it for each credible version.
When to decide
Only after every version has both observations. A tie needs a stricter test.
What follows
Use the result to focus the next test, then follow the data from real sessions.
Choose a version to record

Simulated observation

Guided test composer

A distinct interface for the same selected hypothesis, domain revision, and task.

Did the tester understand what to do?
Did the tester complete the task?

Comparison decision

What the comparison says

Record both simulated observations for every candidate.

Local simulation · not recorded.

Next: follow the data

Build-first / open loop

Output is not evidence.

Output
Convincing first version
Confidence
Still an untested belief

Speed shortens production. Only a test against an explicit claim shortens uncertainty.

One conductor / three instruments

Build the smallest test for the belief.

The order matters. Each instrument constrains the next, so the prototype stays attached to a person, a costly situation, and a claim that can lose.

  1. 01

    Context frame

    Problem

    situated

    Specimen record / Protopilot validating Protopilot

    Whose signal
    Product leads using AI-assisted builders.
    Conditions
    A convincing first version feels close enough to ship.
    Cost
    Untested assumptions become polished software too early; the problem is blocking.
  2. 02

    Falsification gauge

    Hypothesis

    exposed

    Assumption / want / screener

    “Product leads using AI-assisted builders want a structured way to expose assumptions before creating a working app.”

    Falsifiable answer rule “Yes” is supporting evidence. “No” is refuting evidence. 0 supporting · 0 refuting at baseline
  3. 03

    Variable aperture

    Solution

    narrowed

    Smallest rig / one ordered prototype journey

    1. 1Frame the problem
    2. 2Make hypotheses falsifiable
    3. 3Compose a tester flow

    Enough interface to produce evidence about the claim. Nothing built for applause.

External sample / evidence return

Let evidence update the belief.

The self-demo asks two yes/no questions before the narrow rig. Each answer is judged by its declared support rule; prototype interactions remain separately recorded session events.

Field tray / T-01 External tester
A

Screen

Qualify context

Does your team currently use AI-assisted tools to explore or build product ideas?

Yes / supports + qualifies No / refutes + screens out
B

Answer / non-qualifying evidence

Ask the claim directly

Would a structured way to surface assumptions before building an app be useful to your team?

Yes / supporting No / refuting
C

Present / session events

Run the narrow flow

Three screens record screen_view, interaction, flow_completed, or flow_abandoned without treating those events as screener evidence.

Self-demo truth: screener answers update linked assumptions; prototype usage is recorded separately.

Measured response / yes-no answer Yes = supporting / No = refuting

Mapped by question ID to the structured-discovery assumption.

Evidence port The response updates the linked evidence tally for the next decision.

Bench ready / question first

Start with one testable belief.

Start with the people and costly situation. Leave with a prototype whose confidence can be changed by something real.

Open Protopilot