Skip to content
Aletheonix

Aletheonix · research prototype v0.2

What is worth discovering next?

Aletheonix observes complex systems, identifies what deserves investigation, designs experiments, verifies the evidence and updates what it believes.

  1. 01Observetraces, evals, costs, deploys
  2. 02Questionwhat deserves investigation
  3. 03Experimentmost informative, governed
  4. 04Verifyindependent, reproduced
  5. 05Learnonly from accepted evidence
  6. ↺ Repeat

01Where the question comes from

Most AI starts with a question. Aletheonix starts before it.

AI begins with imitation. Human progress begins with imitation too. But imitation cannot explain where the first thing worth imitating comes from.

Most AI starts here

“Here is the goal.”

A person decides what matters, writes it down, and the system optimises toward it. Everything upstream of the goal is assumed.

Aletheonix starts earlier

  • What changed?
  • What doesn’t make sense?
  • What uncertainty matters?
  • What is worth knowing?
  1. 01Environment
  2. 02Observations
  3. 03Beliefs
  4. 04Prediction error
  5. 05Opportunity
  6. 06Goal

Aletheonix works here: it builds the goal from what it observes.

Most AI begins here.

02The loop

From unexplained observation to verified finding.

Seven stages, six components, one rule between them: no component grades its own work.

  1. The Belief Graph predicts how a healthy system should behave. Detectors compare those predictions with what actually happened. Where they disagree beyond chance and beyond a practical effect size, a candidate appears.

    PREDICTED BANDDEPLOY

    Prediction error · click the point to open a question

  2. Candidates are merged when they describe one event, ranked by what they are worth, and the top one becomes a question in words and as a typed claim. Nobody supplied it.

    Question · supplied by Aletheonix

    Why has this metric worsened since the change point, and is the change real?

    claim: {
      kind:     "agent_outcome_shift",
      entity:   { type: "agent", id: … },
      metric:   "eval_score",
      window:   { start: change_point, end: now },
      expected: baseline,   observed: current
    }
  3. At least three falsifiable hypotheses, always including chance. Each one states, in advance, what every available experiment should show if it is true.

    ANOMALYNamed lever ANamed lever BWorkload compositionMeasurement (evaluator)Chance / transient
  4. For each feasible design, the engine computes the expected information gain over the current posterior and divides by cost. The design is preregistered and hashed before anything runs.

    Expected information gain ÷ cost

    Stratified re-analysis
    EIG
    Cost
    Fresh-sample replication
    EIG
    Cost
    Paired counterfactual
    EIG
    Cost

    A nearly free re-analysis runs first. The posterior moves; the ranking is recomputed before the next design.

  5. The engine asks governance for permission. Only a signed token for sandbox scope lets the run start. Production is never touched; the sandbox checks the token itself.

    ENGINErequestsGOVERNANCEdecidesSANDBOXchecks tokenALLOWPRODUCTIONuntouched · changes are proposalsTOKEN SCOPE { sandbox: true, production: false }
  6. A separate verifier recomputes the outcome from raw per-unit data with its own code, reproduces the experiment on a fresh sample, checks confounders, measurement validity, power and multiplicity, then signs a verdict.

    • preregistration_hash
    • governance_token
    • sandbox_isolation
    • sample_size
    • pairing
    • composition_stable
    • measurement_quality
    • statistical_power
    • reproduction
    • multiplicity
    Signed by the verifier onlyAccepted
  7. Confidence moves only when a signed verdict arrives. Every change is appended to a hash-chained audit log, with the before and after values and the verdict that justified it.

    Mechanism belief · confidence

    before
    after verdict 1
    after verdict 2

    Audit chain

    1c1aaa822f99…5febd0

03Generation is not verification

An AI can generate an explanation. That doesn’t make it true.

The component that proposes hypotheses cannot sign a verdict. A separate verifier recomputes the result from raw records with its own code, reproduces the experiment on a fresh sample under its own permission, and checks confounders, measurement, power and multiplicity. Only its signed verdicts can change what the system believes.

Generator

proposes

Tool version · crm_lookup

A hypothesis. Not yet evidence of anything.

Experiment

returns raw records

Paired counterfactual · tool version · crm_lookup

  • 150 paired units · common random numbers
  • control effect −0.161
  • treatment − control +0.172
  • paired p = 9.6e-24

Verifier

recomputes, reproduces, signs

ACCEPTED

12/12 checks passed · fresh-sample reproduction agrees

Effect removed by the lever, and removed again on a sample the engine never saw.

Belief graph

moves only on ACCEPTED

53%→92%

Update appended to the audit chain with the verdict signature.

Benchmark S17 · seed 0 · autonomous condition · run_main @ 5c0ccc6 · verdict_0011 · signature 64b0df753755b608…

No discovery without a receipt.

The system does not learn because it convinced itself. It learns because the evidence survived verification. In the benchmark, the verifier signed 705 verdicts. 128 were held back as inconclusive, most often because an independent reproduction disagreed.

04The Discovery Graph

What Aletheonix remembers.

Not a chat history. A provenance-linked history of how a system came to know what it knows, including what it ruled out and what it could not settle.

  1. Observationtemporal detector

    eval_score for agent_beta on refunds moved from 0.670 to 0.491, aligned with a deploy.

05First environment

AI-agent systems.

The first place to test self-directed discovery is a system that is measurable, fast and cheap to experiment on, and that increasingly acts on the world.

Measurable

Every step leaves a trace: tool, version, status, latency, cost, outcome.

High-frequency

Thousands of tasks a week give enough signal to separate effects from noise.

Cheap to experiment on

Replays and sandbox forks cost compute, not reagents or months.

Real failures

Regressions, silent errors and misattributed incidents already cost teams time and trust.

Clear interventions

Pin a version, revert a prompt, change a retry policy, re-grade with a reference.

Increasingly consequential

Agents now act on real systems. Understanding them is not optional.

Inputs

Agent traces · logs · evals · incidents · costs · tool calls · configuration changes · deployments · code snapshots · user outcomes

Real systems can be ingested today through an OpenTelemetry importer in observe-only mode: candidates are ranked and questions written, but nothing is called a discovery without a sandbox to test it in.

What it looks for · and where the benchmark tests it

  • ReliabilityS01 S02 S05 S11 S13
  • CostS06 S10 S25 S29
  • LatencyS01 S03 S14 S20 S27
  • Tool failuresS12 S16 S24 S26
  • Delegationopen
  • Evaluation qualityS04 S18 S30
  • SafetyS19
  • Unexpected interactionsS09 S17 S23

06One example

Knowledge and authority are different things.

A real finding from the benchmark, start to finish. The system was not told to look at this agent, this category or this tool.

Benchmark S17 · seed 0 · autonomous condition · run_main @ 5c0ccc6Benchmark output · closed world

Observation

Agent performance falls after a deployment: eval score for agent_beta on refunds tasks drops from 0.670 to 0.491.

  • metric eval_score
  • entity agent_beta · refunds
  • expected 0.670
  • observed 0.491
  • traffic 7.4% affected
  • context a deploy aligned with the change point

Nobody asked about this agent or this category. The detector raised it because the Belief Graph predicted stable levels and the data disagreed.

Belief44%

The belief bar is the posterior of the leading hypothesis. It stays at its prior until verified evidence arrives.

07Evidence

What the prototype has shown, and what it hasn’t.

30

hidden scenarios

28 planted mechanisms, 2 null worlds

3

seeds

every run reproducible from scenario and seed

0

questions supplied

in the autonomous condition

run_main · 30 scenarios × 3 seeds × 3 conditions
MetricAutonomousNo question suppliedHuman questionSame engine, told where to lookThreshold monitorStatic thresholds, no experiments
Planted mechanism found99%96%43%
False discoveries (total)0030
False discovery rate0%0%42%
Top-ranked conclusion correct92%94%36%
Impact calibration error41%39%n/a

08The harder problem

The real frontier isn’t answering questions. It’s choosing them.

Today

v0 chooses among questions inside a structured space: five detector families and a typed taxonomy of hypotheses and levers. There is no language model in the loop. What it demonstrates is self-directed selection and verification within that space, not the open-ended creation of new kinds of question.

That boundary is deliberate. It keeps the experimental claim about the loop and the separation of roles, not about any one model.

Research ahead · none of this works yet

  • Model-generated questionsbehind the same verifier and governance boundary
  • Open-ended hypothesesbeyond a fixed lever catalogue
  • Better opportunity valuationwhat is worth knowing, not only what is surprising
  • Calibrated intervention impacttoday’s predictions are off by ~40%
  • Cross-system learningmemory that transfers between systems
  • Real-world experimental environmentswhere an experiment costs hours, not milliseconds

From “given a goal, solve it” to “given an environment, determine what is worth understanding.” And eventually: what am I not noticing?

Three independent systems that interoperate

Aletheonix, LucidRail and Terranoux

Aletheonix discovers. LucidRail governs. Terranoux acts.

Discovery. Authority. Reality.

Field validation

Bring a system worth understanding.

We are looking for teams running AI-agent systems with traces and, ideally, a replay or shadow environment. You write down what you already know before the run. We find out what Aletheonix finds that you didn’t ask about, and whether it holds up.