Agents' Last Exam (ALE) asks AI agents to do real professional work and grades each attempt with a single number. That number tells you what the agent scored β but not whether the work was real, whether the agent told the truth about it, how far it actually got, or whether the same score would come back if you ran it again. The Trust Layer sits beside ALE and answers those questions. It never changes ALE's score; it adds the context the score leaves out.
| # | Task | Domain | Tier | Tool calls | Execution time | ALE result | Trust verdict |
|---|
ALE reports a single number per run. It does not say why the task failed, or where.
ALE grades only the deliverables and returns a number. The Trust Layer reads the same deliverables and the trajectory, and reports what the score alone cannot show.