What is worth discovering next?
Aletheonix observes complex systems, identifies what deserves investigation, designs experiments, verifies the evidence and updates what it believes.
- traces, evals, costs, deploys
- what deserves investigation
- most informative, governed
- independent, reproduced
- only from accepted evidence
Most AI starts with a question. Aletheonix starts before it.
AI begins with imitation. Human progress begins with imitation too. But imitation cannot explain where the first thing worth imitating comes from.
“Here is the goal.”
A person decides what matters, writes it down, and the system optimises toward it. Everything upstream of the goal is assumed.
- What changed?
- What doesn’t make sense?
- What uncertainty matters?
- What is worth knowing?
- Environment
- Observations
- Beliefs
- Prediction error
- Opportunity
- Goal
Aletheonix works here: it builds the goal from what it observes.
Most AI begins here.
From unexplained observation to verified finding.
Seven stages, six components, one rule between them: no component grades its own work.
The Belief Graph predicts how a healthy system should behave. Detectors compare those predictions with what actually happened. Where they disagree beyond chance and beyond a practical effect size, a candidate appears.
Candidates are merged when they describe one event, ranked by what they are worth, and the top one becomes a question in words and as a typed claim. Nobody supplied it.
Why has this metric worsened since the change point, and is the change real?
claim: { kind: "agent_outcome_shift", entity: { type: "agent", id: … }, metric: "eval_score", window: { start: change_point, end: now }, expected: baseline, observed: current }At least three falsifiable hypotheses, always including chance. Each one states, in advance, what every available experiment should show if it is true.
For each feasible design, the engine computes the expected information gain over the current posterior and divides by cost. The design is preregistered and hashed before anything runs.
Stratified re-analysisFresh-sample replicationPaired counterfactualA nearly free re-analysis runs first. The posterior moves; the ranking is recomputed before the next design.
The engine asks governance for permission. Only a signed token for sandbox scope lets the run start. Production is never touched; the sandbox checks the token itself.
A separate verifier recomputes the outcome from raw per-unit data with its own code, reproduces the experiment on a fresh sample, checks confounders, measurement validity, power and multiplicity, then signs a verdict.
- preregistration_hash
- governance_token
- sandbox_isolation
- sample_size
- pairing
- composition_stable
- measurement_quality
- statistical_power
- reproduction
- multiplicity
AcceptedConfidence moves only when a signed verdict arrives. Every change is appended to a hash-chained audit log, with the before and after values and the verdict that justified it.
1c1aaa822f99…5febd0
An AI can generate an explanation. That doesn’t make it true.
The component that proposes hypotheses cannot sign a verdict. A separate verifier recomputes the result from raw records with its own code, reproduces the experiment on a fresh sample under its own permission, and checks confounders, measurement, power and multiplicity. Only its signed verdicts can change what the system believes.
Tool version · crm_lookup
A hypothesis. Not yet evidence of anything.
Paired counterfactual · tool version · crm_lookup
- 150 paired units · common random numbers
- control effect −0.161
- treatment − control +0.172
- paired p = 9.6e-24
12/12 checks passed · fresh-sample reproduction agrees
Effect removed by the lever, and removed again on a sample the engine never saw.
53%→92%
Update appended to the audit chain with the verdict signature.
Benchmark S17 · seed 0 · autonomous condition · run_main @ 5c0ccc6 · verdict_0011 · signature 64b0df753755b608…
No discovery without a receipt.
The system does not learn because it convinced itself. It learns because the evidence survived verification. In the benchmark, the verifier signed 705 verdicts. 128 were held back as inconclusive, most often because an independent reproduction disagreed.
What Aletheonix remembers.
Not a chat history. A provenance-linked history of how a system came to know what it knows, including what it ruled out and what it could not settle.
- Observation
eval_score for agent_beta on refunds moved from 0.670 to 0.491, aligned with a deploy.
AI-agent systems.
The first place to test self-directed discovery is a system that is measurable, fast and cheap to experiment on, and that increasingly acts on the world.
Measurable
Every step leaves a trace: tool, version, status, latency, cost, outcome.
High-frequency
Thousands of tasks a week give enough signal to separate effects from noise.
Cheap to experiment on
Replays and sandbox forks cost compute, not reagents or months.
Real failures
Regressions, silent errors and misattributed incidents already cost teams time and trust.
Clear interventions
Pin a version, revert a prompt, change a retry policy, re-grade with a reference.
Increasingly consequential
Agents now act on real systems. Understanding them is not optional.
Agent traces · logs · evals · incidents · costs · tool calls · configuration changes · deployments · code snapshots · user outcomes
Real systems can be ingested today through an OpenTelemetry importer in observe-only mode: candidates are ranked and questions written, but nothing is called a discovery without a sandbox to test it in.
- Reliabilityregressions after deploys, silent failuresS01 S02 S05 S11 S13
- Costredundant context, duplicates, long tailsS06 S10 S25 S29
- Latencyslow clients, retry storms, cold startsS01 S03 S14 S20 S27
- Tool failuresmalformed arguments, cascades, unretried errorsS12 S16 S24 S26
- Delegationhandoffs between agents · not yet in the benchmarkopen
- Evaluation qualitygrader drift, blind spots, length biasS04 S18 S30
- Safetypermission denials only so farS19
- Unexpected interactionsconfounded mixes, misattributed incidents, trade-offsS09 S17 S23
Knowledge and authority are different things.
A real finding from the benchmark, start to finish. The system was not told to look at this agent, this category or this tool.
Agent performance falls after a deployment: eval score for agent_beta on refunds tasks drops from 0.670 to 0.491.
- metric eval_score
- entity agent_beta · refunds
- expected 0.670
- observed 0.491
- traffic 7.4% affected
- context a deploy aligned with the change point
Nobody asked about this agent or this category. The detector raised it because the Belief Graph predicted stable levels and the data disagreed.
The belief bar is the posterior of the leading hypothesis. It stays at its prior until verified evidence arrives.
What the prototype has shown, and what it hasn’t.
30
hidden scenarios
28 planted mechanisms, 2 null worlds
3
seeds
every run reproducible from scenario and seed
0
questions supplied
in the autonomous condition
| Metric | AutonomousNo question supplied | Human questionSame engine, told where to look | Threshold monitorStatic thresholds, no experiments |
|---|---|---|---|
| Planted mechanism found | 99% | 96% | 43% |
| False discoveries (total) | 0 | 0 | 30 |
| False discovery rate | 0% | 0% | 42% |
| Top-ranked conclusion correct | 92% | 94% | 36% |
| Impact calibration error | 41% | 39% | n/a |
The real frontier isn’t answering questions. It’s choosing them.
v0 chooses among questions inside a structured space: five detector families and a typed taxonomy of hypotheses and levers. There is no language model in the loop. What it demonstrates is self-directed selection and verification within that space, not the open-ended creation of new kinds of question.
That boundary is deliberate. It keeps the experimental claim about the loop and the separation of roles, not about any one model.
- Model-generated questionsbehind the same verifier and governance boundary
- Open-ended hypothesesbeyond a fixed lever catalogue
- Better opportunity valuationwhat is worth knowing, not only what is surprising
- Calibrated intervention impacttoday’s predictions are off by ~40%
- Cross-system learningmemory that transfers between systems
- Real-world experimental environmentswhere an experiment costs hours, not milliseconds
From “given a goal, solve it” to “given an environment, determine what is worth understanding.” And eventually: what am I not noticing?
Aletheonix, LucidRail and Terranoux
Discovery
What should we understand?
Finds what is worth investigating, tests competing explanations and verifies the evidence.
You are here
Authority
What are we allowed to do?
Decides whether a consequential action may happen, under which grant, and records it.
lucidrail.ai ›
Reality
How do we test it in the physical world?
Bounded physical environments for experiments: laboratories, robotics, sensors, materials.
terranoux.com ›
Aletheonix discovers. LucidRail governs. Terranoux acts.
Bring a system worth understanding.
We are looking for teams running AI-agent systems with traces and, ideally, a replay or shadow environment. You write down what you already know before the run. We find out what Aletheonix finds that you didn’t ask about, and whether it holds up.