01 / AGENDA

Build what the model is missing.

We study what AI needs to stay on a goal, check its own progress, and get better with experience.

How long can an AI work on a real goal, get provably correct results, and learn along the way, without a person stepping in?

This one question covers reliability, learning, and independence together. Benchmark scores measure parts of a system. They don’t show whether an agent can finish a long task in a setting it has never seen.

Each one makes
the others stronger.

01 / RESEARCH PROGRAM

Proof before trust

Verification

When an AI says "done", that is not proof. We check whether the result really happened, using tests the AI does not control.

  • Automated tests
  • Independent reviewers
  • Critical review
  • Repeatable experiments
02 / RESEARCH PROGRAM

Staying on track for weeks, not minutes

Long-running agents

Long tasks need a clear record of the goal, the plan, and the progress so far. The agent must know what it has proven, what it only assumes, and what is still open.

  • Task records
  • Planning
  • Progress tracking
  • Handling uncertainty
03 / RESEARCH PROGRAM

Every task teaches the next one

Memory & learning

A finished task leaves behind memories, reusable know-how, and a record of what failed. We study how to turn that into better work next time, in a way we can measure and control.

  • Memory of past tasks
  • Reusable skills
  • Learning from failure
  • Controlled updates
04 / RESEARCH PROGRAM

Knowing what an action will cause

World models & action

A world model lets an agent ask: what happens if I do this instead of that? We want agents that predict consequences and carry their skills from software to simulations and, later, to physical machines.

  • Environment models
  • Actions and effects
  • Skill transfer
  • Simulation and robotics
OUR FIRST PRODUCT

A research agent for technical work.

We start with digital work: software engineering, AI research, data analysis, experiment design, model evaluation, and computational research. These tasks matter, and their results can be inspected and repeated.

The agent hands back a checked result with all the evidence: code, raw data, its reasoning, the checks it passed, evidence against it, how confident it is, and steps to reproduce it.

See how a task runs

Truth over demos.

01

Real work over benchmarks

A benchmark only matters if it predicts how AI does on real work.

02

Evidence over stories

Every important claim needs evidence that others can reproduce.

03

Control built in

Permissions, activity logs, and handing off to a person are part of the design, not add-ons.

Bring us a hard technical problem.

Start with a question where a proven answer would change what your team does next.