Discover and evaluate the unknown.

For nearly 300,000 years, humanity has been driven by the same mission: to conquer the unknown. We ventured into unmapped forests, sailed across uncharted seas, rose into boundless skies, and reached for the distant Moon — pushing the frontier of the known ever farther.

We only care about the unknowns

They are hard, uncertain and may take years to solve. But they are highly valuable to humanity. And they are the only problems we care about.

Now, Apodex Discovery gives AI the same mission: conquer the unknown — with three defining principles.

  1. Principle 1: Problems

    Scouting the problems that matter most to humanity

    Most benchmarks start with questions.
    Apodex Discovery starts with the real world.

  2. Principle 2: Environments

    Building the world AI needs to solve them

    Real problems do not arrive as prompts.
    They arrive as ambitions.

  3. Principle 3: Verifications

    Evaluating the journey when the answer does not yet exist

    In discovery, the answer is often not hidden.
    It simply does not exist yet.

Principle 1: problems

Scouting the problems that matter most to humanity

Most benchmarks start with questions. Apodex Discovery starts with the real world.

Finding the right problems is itself a human endeavor. Even in the age of AI, the questions that matter most cannot simply be generated at scale and filtered by an algorithm.

That is why Apodex Discovery began with an unusually large manual effort: 10 experts, all with STEM PhDs, spent two months scouting 561 industries across 16 sectors to identify problems with genuine scientific, technological, economic, and societal value. From this effort, we built a registry of 423 high-value problems spanning science, medicine, engineering, energy, materials, computing, finance, healthcare, and industry.

The effort is manual by design — because humanity defines what is worth solving, and AI exists to help us solve it.

For each problem, we turn an open-ended ambition into a well-defined question for AI — defining the objective, decomposing the challenge, specifying what success means, and establishing how a solution can ultimately be verified.

Our first release focuses deeply on 20 selected problems — the beginning of an expanding frontier of AI-solvable challenges.

Explore our problems →

10
experts, all with STEM PhDs
2 months
of manual scouting
561
industries surveyed
16
sectors covered
423
high-value problems in the registry — 20 fully built for the first release

Manual by design Humanity defines what is worth solving. AI exists to help us solve it.

Principle 2: Environments

Building the world AI needs to solve them

The hardest problems are often out of reach for AI — not because AI is not ready, but because the world around the problem was never built for it.

Real problems do not arrive as self-contained prompts. They arrive as ambitions. What turns one into the other is everything a prompt leaves out — the specialized tools and data the work runs on, the workflow and context that decide what happens next, and a way to test whether each step actually worked.

A scientist asked to design a better gene-therapy vector, for example, does not answer in one shot. The journey moves through viability, tissue targeting, structural analysis, candidate design, and experimental feedback — each step drawing on different data, tools, and evidence.

Apodex Discovery brings this same structure to AI.

For every problem, we turn an open-ended ambition into an executable world for AI — defining the path of work, making the right resources accessible, and ensuring that progress can be tested along the way.

At its center, a sophisticated multi-agent system navigates the environment — planning, delegating, using tools, interpreting new evidence, and iterating toward a solution.

We do not just give AI a harder question. We build the world it needs to solve it.

Explore our environments & leaderboards →

What our experts build by hand

  • DataCurating the inputs the solver is given, and nothing more.
  • ToolsSelecting and wiring the instruments it is allowed to call.
  • Task decompositionBreaking the problem into tasks that can each be checked.
  • Hidden verifierGrading the submission against ground truth the solver never sees.

Why by hand A model can only generate what it already knows — solving that proves nothing. Hand-built worlds hold real data, real tools, and real expert practice, so the solver has to discover, not recall.

Principle 3: Verifications

Evaluating the journey when the answer does not yet exist

In discovery, the answer is often not hidden — it simply does not exist yet.

A new drug may take years to prove itself. A new material has yet to be made. A scientific hypothesis may still be waiting for the experiment that reveals whether it is true. There is no answer key to look up — reality itself has yet to reveal the answer.

Humanity has always navigated the unknown this way.

Developed with input from more than 100 experts across diverse domains, Apodex Discovery introduces a novel scoring system covering six capabilities for evaluating the journey toward problems whose answers have yet to exist.

An uncharted ocean

Imagine crossing an uncharted ocean. Long before a new shore appears on the horizon, we judge whether the voyage is sound by how well we read the weather, use our instruments, compare possible routes, stay on course, respond to storms, and recognize the limits of our maps.

Discovery with AI should be judged the same way.

When the destination is still beyond the horizon, Apodex Discovery tells us whether AI is navigating the unknown well.

TTools

Tool Use & Execution State Management

Reaching for the right tool, calling it correctly, and letting the result change what happens next.

RRepair

Self-correction under Verification

Finding the actual error rather than talking around it, fixing it, and checking the fix held.

AAlternatives

Hypothesis Management

Holding more than one explanation, weighing the evidence both ways, and running the test that tells them apart.

CCoherence

Long-horizon State Coherence

Carrying constraints and findings through a long problem, so the final answer traces back to the work.

EEvidence

Evidence Fidelity

Claims that trace back to real, checkable evidence — and a clear line between what is known and what is inferred.

SScope

Boundary & Failure Reasoning

Where the conclusion holds, how it could fail, and how those limits bound the recommendation.

“The system must rely on verifiable reasoning and external feedback loops, rather than ‘sounding reasonable’ to muddle through.”
Mr. Tianqiao Chen, Founder & CEO of Apodex