Proof before trust
Verification
When an AI says "done", that is not proof. We check whether the result really happened, using tests the AI does not control.
- Automated tests
- Independent reviewers
- Critical review
- Repeatable experiments
01 / AGENDA
We study what AI needs to stay on a goal, check its own progress, and get better with experience.
01Our core question
This one question covers reliability, learning, and independence together. Benchmark scores measure parts of a system. They don’t show whether an agent can finish a long task in a setting it has never seen.
02Four core programs
Proof before trust
When an AI says "done", that is not proof. We check whether the result really happened, using tests the AI does not control.
Staying on track for weeks, not minutes
Long tasks need a clear record of the goal, the plan, and the progress so far. The agent must know what it has proven, what it only assumes, and what is still open.
Every task teaches the next one
A finished task leaves behind memories, reusable know-how, and a record of what failed. We study how to turn that into better work next time, in a way we can measure and control.
Knowing what an action will cause
A world model lets an agent ask: what happens if I do this instead of that? We want agents that predict consequences and carry their skills from software to simulations and, later, to physical machines.
03Where we start
We start with digital work: software engineering, AI research, data analysis, experiment design, model evaluation, and computational research. These tasks matter, and their results can be inspected and repeated.
The agent hands back a checked result with all the evidence: code, raw data, its reasoning, the checks it passed, evidence against it, how confident it is, and steps to reproduce it.
See how a task runs04Principles
A benchmark only matters if it predicts how AI does on real work.
Every important claim needs evidence that others can reproduce.
Permissions, activity logs, and handing off to a person are part of the design, not add-ons.
↗WORK WITH US
Start with a question where a proven answer would change what your team does next.