πŸ€— Spaces / CentificAIResearch / ALE - Trust Layer evaluation
β—† 8 agent configurations read-only Β· ALE score never modified

🧭 ALE Trust Layer β€” Can You Trust What a Benchmark Score Says?

Agents' Last Exam (ALE) asks AI agents to do real professional work and grades each attempt with a single number. That number tells you what the agent scored β€” but not whether the work was real, whether the agent told the truth about it, how far it actually got, or whether the same score would come back if you ran it again. The Trust Layer sits beside ALE and answers those questions. It never changes ALE's score; it adds the context the score leaves out.

πŸ† How each agent configuration holds up

Select a task on the left.

πŸ—‚οΈ How the tasks are classified

How to read this table. One row per task. ALE result is how the eight configurations were graded by the benchmark. Trust verdict is how many of those eight survived all four checks. Run-by-run detail, including which dimension blocked each one, is in Explore tasks.

#TaskDomainTier Tool callsExecution time ALE resultTrust verdict

πŸ“ What ALE doesn't grade

ALE reports a single number per run. It does not say why the task failed, or where.

The score carries no reason. A 0 does not distinguish an agent that got most of the way from one that never started.
The trajectory is not graded. What the agent actually did to reach its answer is recorded, but never assessed.
Trustworthiness is not tested. Whether the answer was genuinely produced, and whether the agent reported its own work honestly, go unchecked.

🧭 What the Trust Layer adds

Agent run sealed environment Deliverables what it produced Trajectory every step taken ALE grader grades the deliverables One score no reason, no context THE TRUST LAYER D1 Competence does the score hold up? D2 Reward hacking was the credit earned? D3 Deception was the report honest? D4 Reliability would it hold tomorrow? Trust verdict with the reason why

ALE grades only the deliverables and returns a number. The Trust Layer reads the same deliverables and the trajectory, and reports what the score alone cannot show.

ALE Trust Layer β€” a verification layer over Agents' Last Exam, built at Centific. Results shown are replayed from saved run artifacts; no agent is executed by this Space.