◆ Claude Sonnet 4.6read-only · ALE score never modified
🧭 ALE Trust Layer — Can You Trust What a Benchmark Score Says?
Agents' Last Exam (ALE) asks AI agents to do real professional work and grades each attempt
with a single number. That number tells you what the agent scored — but not whether the work
was real, whether the agent told the truth about it, how far it actually got, or whether the same
score would come back if you ran it again. The Trust Layer sits beside ALE and answers those
questions. It never changes ALE's score; it adds the context the score leaves out.
This is a task-level evaluation of the twenty tasks that have results on all four dimensions.
The Trust Layer re-uses ALE’s own grader on saved run artifacts — no agent is re-run here.
Start with “Explore tasks”, then “Task classification” for the full inventory,
then “How it works” for the method.
Select an example on the left.
🗂️ How the tasks are classified
These are the tasks featured in Explore — the ones where the Trust Layer changed ALE’s
category. For each, the table shows what ALE judged it, what every Trust Layer dimension found
and — where it did not pass — why, and the overall Trust verdict. A task earns a passing verdict
only if all four dimensions pass. The full pass rate is computed over the twenty-task cohort.
How to read this table. One row per task. Each dimension carries its own pass / partial / fail
verdict and, when it did not pass, a short reason (the failure mode). The Trust verdict on the
right applies the all-must-pass rule: any single dimension failing makes the overall verdict fail.
#
Task
Domain
Tier
ALE verdict
D1 Competence
D1 reason
D2 Reward hacking
D2 reason
D3 Deception
D3 reason
D4 Reliability
D4 reason
Trust verdict
📐 What ALE doesn't grade
ALE reports a single number per task. It does not say why the task failed, or where.
The score carries no reason. A 0 does not distinguish an agent that
got most of the way from one that never started.
The trajectory is not graded. What the agent actually did to reach
its answer is recorded, but never assessed.
Trustworthiness is not tested. Whether the answer was genuinely
produced, and whether the agent reported its own work honestly, go unchecked.
🧭 What the Trust Layer adds
ALE grades only the deliverables and returns a number. The Trust Layer reads the
same deliverables and the trajectory, and reports what the score alone cannot show.
ALE Trust Layer — a verification layer over Agents' Last Exam, built at Centific.
Results shown are replayed from saved run artifacts; no agent is executed by this Space.