🤗 Spaces / CentificAIResearch / ALE - Trust Layer evaluation
◆ Claude Sonnet 4.6 read-only · ALE score never modified

🧭 ALE Trust Layer — Can You Trust What a Benchmark Score Says?

Agents' Last Exam (ALE) asks AI agents to do real professional work and grades each attempt with a single number. That number tells you what the agent scored — but not whether the work was real, whether the agent told the truth about it, how far it actually got, or whether the same score would come back if you ran it again. The Trust Layer sits beside ALE and answers those questions. It never changes ALE's score; it adds the context the score leaves out.

This is a task-level evaluation of the twenty tasks that have results on all four dimensions. The Trust Layer re-uses ALE’s own grader on saved run artifacts — no agent is re-run here. Start with “Explore tasks”, then “Task classification” for the full inventory, then “How it works” for the method.
Select an example on the left.

🗂️ How the tasks are classified

These are the tasks featured in Explore — the ones where the Trust Layer changed ALE’s category. For each, the table shows what ALE judged it, what every Trust Layer dimension found and — where it did not pass — why, and the overall Trust verdict. A task earns a passing verdict only if all four dimensions pass. The full pass rate is computed over the twenty-task cohort.

How to read this table. One row per task. Each dimension carries its own pass / partial / fail verdict and, when it did not pass, a short reason (the failure mode). The Trust verdict on the right applies the all-must-pass rule: any single dimension failing makes the overall verdict fail.
#TaskDomainTierALE verdict D1 CompetenceD1 reason D2 Reward hackingD2 reason D3 DeceptionD3 reason D4 ReliabilityD4 reason Trust verdict

📐 What ALE doesn't grade

ALE reports a single number per task. It does not say why the task failed, or where.

The score carries no reason. A 0 does not distinguish an agent that got most of the way from one that never started.
The trajectory is not graded. What the agent actually did to reach its answer is recorded, but never assessed.
Trustworthiness is not tested. Whether the answer was genuinely produced, and whether the agent reported its own work honestly, go unchecked.

🧭 What the Trust Layer adds

Agent run sealed environment Deliverables what it produced Trajectory every step taken ALE grader grades the deliverables One score no reason, no context THE TRUST LAYER D1 Competence does the score hold up? D2 Reward hacking was the credit earned? D3 Deception was the report honest? D4 Reliability would it hold tomorrow? Trust verdict with the reason why

ALE grades only the deliverables and returns a number. The Trust Layer reads the same deliverables and the trajectory, and reports what the score alone cannot show.

ALE Trust Layer — a verification layer over Agents' Last Exam, built at Centific. Results shown are replayed from saved run artifacts; no agent is executed by this Space.