Scoring Reliability
Independent reviewers score anonymized packets to test whether the rubric produces consistent judgments from the same evidence.
Profile management
Security actions require an active login.
Provider submissions
Provider workspace is ready.
Reviewer
Reviewer access is shown only for reviewer-capable accounts. Admin accounts see both Provider and Reviewer tabs.
Reviewer workspace is ready.
EPOB separates human scoring reliability from full experimental reproducibility. Scoring reviewers evaluate frozen run artifacts with the same rubric; replication reviewers rerun the protocol under the same model, seed, harness, container, and resource profile.
Reviewer packets are stored in the EPOB database and appear in Assignments after claim or assignment. Scoring reviewers work from the database-backed packet table, so there is no packet download or checksum step.
Independent reviewers score anonymized packets to test whether the rubric produces consistent judgments from the same evidence.
Independent replication reruns selected cells to test whether the benchmark protocol and frozen configuration reproduce comparable outcomes.
Scoring reliability does not require model access. Experimental reproducibility requires the execution environment and should preserve fresh run evidence.
[email protected] can answer reviewer questions, confirm conflicts, and coordinate the expected time window.
This example shows how a reviewer should translate packet evidence into scores. It is a teaching example, not a hidden answer key for assigned packets.
The packet contains a dependency-aware plan, role-aligned task assignments, trace events for implementation and review, a final artifact, one unresolved handoff, and two avoidable rework loops.
Score what is recorded in the packet. Do not use framework reputation, expected model ability, or guesses about work that is not visible in the trace or artifact.
Numeric scores use the evidence band first, then the concrete facts in the packet. Low quality is not automatically invalid when the evidence is still comparable.
| Option | Example value | How to use every scoring option | Why this example is scored this way |
|---|---|---|---|
| Plan Quality | 76 | Judge whether the plan decomposes the request, preserves dependencies, and keeps deliverables visible. | The plan is usable and dependency-aware, but it leaves one validation handoff underspecified, so it belongs in the 61-80 band instead of 81-100. |
| Assignment Quality | 82 | Check whether specialized work is routed to roles allowed by the packet constraints. | Implementation, review, and validation are assigned to appropriate roles with no material role violation, so this is in the complete role-aligned band. |
| Coordination | 64 | Follow handoffs, dependency order, review incorporation, and whether later work uses earlier evidence. | The trace shows ordering and review use, but one review handoff is not closed cleanly, keeping coordination in the mostly coherent 61-80 band. |
| Deliverable Quality | 68 | Compare final artifacts against the requested deliverables, constraints, custom rules, and review critique evidence. | The final artifact satisfies the main request and tests, but one optional documentation detail is thin, so it is satisfactory with minor weaknesses. |
| Efficiency | 58 | Look for avoidable retries, redundant tasks, dependency churn, or rework relative to packet limits. | The run finishes, but two avoidable rework loops and repeated validation chatter make the execution only moderately efficient. |
| Confidence | 78 | Rate your confidence in the score from evidence clarity, not from how good the run is. | The trace and artifact are clear enough for a strong judgment, but the unresolved handoff leaves some uncertainty. |
| Primary Failure Label | handoff | Choose the single dominant failure mode from the fixed list, or none if no failure class dominates. | The main weakness is the unresolved reviewer-to-implementer handoff; it matters more than decomposition, assignment, or deliverable gaps. |
| Invalid Run Flag | false | Mark true only when the packet is non-comparable, such as missing required evidence or invalid orchestration execution. | The run is imperfect but comparable because the packet includes trace evidence, artifact evidence, and enough context for numeric scoring. |
| Rationale | The trace shows a usable plan and role-aligned assignments, but one unresolved handoff and two avoidable rework loops lower coordination and efficiency. | Write a short evidence-grounded explanation that names the facts driving scores, invalid flag, and failure label. | The rationale cites packet evidence and explains both the strengths and the lower scores without invoking framework reputation. |
Start with the rubric band: 61-80 means mostly usable with minor or moderate gaps. Use 81-100 only when the packet is complete, cleanly coordinated, and critique-aware.
Do not mark invalid just because a run is weak. Mark invalid only when missing evidence or invalid execution makes the run non-comparable.
Pick the dominant failure class. This example has several weaknesses, but the handoff issue best explains the score pattern.
Select scoring reliability, experimental reproducibility, or both. Do not ask scoring reviewers to rerun models.
Record conflicts, prior participation, and whether the reviewer has seen framework names or paper rankings.
Claim or receive database-backed assignments, score from the Assignments table, and keep local copies out of the workflow unless explicitly approved.
Submit scores, invalid-run flags, failure labels, rationale, adjudication notes, replication logs, hashes, and aggregate deltas before writing any reliability claim.
Admin oversight
Latest account records load automatically for admin users.
Group membership controls which role workspaces are available. Admin accounts inherit provider and reviewer access.
| Group | ProviderTab access | ReviewerTab access | Effective permissions |
|---|---|---|---|
|
Admin group
All groups
|
Allowed | Allowed | Admin inherits provider and reviewer access. |
|
Provider group
Submission workspace
|
Allowed | Not allowed | Can submit framework evaluation requests from the Provider tab. |
|
Reviewer group
Review workspace
|
Not allowed | Allowed | Can claim review packets and score assignments from the Reviewer tab. |
Generate one-time reviewer invitation codes, set an expiration time, and audit invitation history after use.