E Evidence Press

EVIDENCE PRESS

Productivity Protocols

Trial one useful AI workflow. Keep the evidence.

New to AI agents? Begin with one bounded task.

Use approved copies or test material, keep the agent read-only, and choose the evidence stage before inviting participants.

Open the company starter →

Evaluation programme — staged, not yet fielded

This is a proposed progression for learning safely. It has not yet been run with company staff, and none of these stages currently establishes a company productivity effect.

1–5 people

Formative usability: can people understand, operate and safely stop?

No productivity estimate

6+ people

Feasibility: can allocation, measurement, support and retention work?

Effects remain exploratory

Controlled evaluation

Use a justified, independently reviewed parallel design.

Context-bound signal only

Organisational follow-up

Observe governed ordinary use after a separate deployment decision.

Not automatically causal

Every protocol carries two independent status values: protocol assurance (is the pack well formed and what was checked?) and work evidence (what outcome was measured, in what setting, with what identification?). They are never merged, and a claim never exceeds the evidence. Read the two ladders.

Recently updated

Ordered from the contracts' last_verified dates. A verification date is not evidence of human use or field testing.

Deprecated

No protocols are marked DEPRECATED in this build.

ProtocolPurpose and fitLevelRiskAssuranceWork evidenceStart
adversarial-output-review
Task: Adversarial output review
Challenge a supplied draft or analysis rather than confirm it. Produce findings ranked by severity, each tied to a specific claim in the draft, each framed to refute — stating why the claim may be wrong and what would verify or falsify it — plus a limitations note and a compact receipt. It transfers the Evidence Press discipline of adversarial, refute-framed review into everyday work.
Audience: Analysts and reviewers who must stress-test a colleague's or an agent's draft before it is relied on. · Individual knowledge workers who want their own draft challenged by a fresh, sceptical pass. · Teams that want a shared, inspectable way to run an adversarial review with the objections tied to claims.Required: instruction-following, text-generationOptional: file-read
verified low EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
decision-memo-under-uncertainty
Task: Decision memo under uncertainty
Turn a decision question into a memo that separates what the supplied materials establish from what the reasoning assumes. It sets out the sourced facts, the labelled assumptions, the options, the facts and assumptions each option is most sensitive to, and which actions are reversible — so a decision-maker sees what is known versus assumed before choosing.
Audience: Analysts and decision-support staff preparing a memo for a decision-maker. · Individual knowledge workers weighing an option and wanting the assumptions made explicit. · Teams that want a shared, inspectable way to turn materials into a decision memo.Required: instruction-following, text-generationOptional: file-read
verified moderate EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
document-to-action-plan
Task: Document to action plan
Read supplied documents and extract the decisions already made, the obligations and commitments, the deadlines, the open uncertainties, and the concrete next actions — each item traceable to a location in the source. The protocol is read-only; it produces a structured action plan, open questions, limitations, and a receipt. It invents nothing the documents do not contain.
Audience: Anyone triaging a long thread who needs the obligations and deadlines separated from the discussion. · Individuals turning a pile of correspondence or notes into a checkable to-do and commitment list. · Teams that want a shared, source-traceable extract of what a set of documents actually commits them to.Required: instruction-following, text-generationOptional: file-read
verified low EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
evidence-backed-brief
Task: Evidence-backed brief
Produce a concise briefing on a question from supplied sources, in which every claim is labelled by type (fact, estimate, opinion, assumption), carries its source, states its uncertainty, and where relevant surfaces contrary evidence. It transfers the Evidence Press discipline — claims with the evidence attached — into everyday work.
Audience: Analysts and decision-support staff preparing a briefing for someone else. · Individual knowledge workers who want a summary they can defend claim by claim. · Teams that want a shared, inspectable way to turn a pile of sources into a brief.Required: instruction-following, text-generationOptional: file-read
verified moderate EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
goal-to-verified-deliverable
Task: Goal to verified deliverable
Turn an unclear task into an explicit deliverable, an input and permission boundary, a checkpointed plan, and acceptance tests — then execute and hand back the deliverable with its limitations and a compact receipt. This is the foundational protocol nearly every other agent task can be run through.
Audience: Individual knowledge workers giving an agent an open-ended task. · Protocol authors building a more specific workflow on top of the kernel. · Teams that want a shared, inspectable way to hand work to an agent.Required: instruction-following, text-generationOptional: file-read
verified low EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
project-handoff
Task: Project handoff
Turn a half-finished project into durable state another person or agent can pick up: the decisions already made and why, the questions still open, the current state of the work, and the exact next steps needed to resume. The protocol is read-only; it produces a handoff document, a limitations statement, and a receipt. It invents no decision or fact the supplied materials do not contain.
Audience: A successor inheriting partly-done work who needs its decisions, rationale, and next steps made legible. · A team that wants a shared, source-traceable record of where a project stands and how to resume it. · An individual leaving a project before it is finished who must hand it to a colleague or a future agent.Required: instruction-following, text-generationOptional: file-read
verified low EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
repetitive-workflow-capture
Task: Repetitive workflow capture
Read a description of a repeated manual process and turn it into a CANDIDATE protocol: a draft contract (deliverable, inputs, permissions, prohibited actions, steps, acceptance tests), a ten-question README skeleton, and a list of what to test. The output is a starting point for the foundry, not a finished or validated protocol. The protocol is read-only over the description; it invents no step the description does not contain and claims no maturity the candidate has not earned.
Audience: Individuals who repeat a manual routine and want it drafted into a candidate protocol without writing the contract by hand. · Protocol authors who want a described workflow turned into a first-pass specification skeleton to refine. · Teams gathering candidate workflows for the foundry's proposal-and-specification stage.Required: instruction-following, text-generationOptional: file-read
verified low EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE
spreadsheet-quality-audit
Task: Spreadsheet quality audit
Audit a supplied spreadsheet or table for formula errors, unit mismatches, missing data, internal inconsistencies, and suspicious values. Each finding is located to a cell or row and rated by severity. The protocol is read-only; it produces an audit table, a limitations statement, and a receipt. It does not modify the supplied spreadsheet, and it invents no finding the data does not support.
Audience: Anyone handed a table by someone else who needs its arithmetic and units checked against themselves. · Individuals sanity-checking a budget, model, or data extract before they act on it. · Teams that want a located, rated, source-traceable list of the problems in a shared spreadsheet.Required: instruction-following, text-generationOptional: spreadsheet-read
verified moderate EXAMPLE_CONFORMANCE_VALIDATED NO_IMPACT_EVIDENCE

· machine-readable: protocols.json · schema · feed

How the library works

  • The Verified Agent Work kernel — the eight-step method every protocol instantiates.
  • Two status ladders — assurance and productivity evidence, kept separate.
  • Each pack ships a machine-readable contract, a skill, worked examples, tests, an evaluation design, adapters, a manifest of file hashes, and a receipt.
  • The YAML contract is not claimed as a new workflow language. The candidate contribution is the staged company adoption-and-evidence loop around it.

Propose a recurring task

The foundry starts with a real, repeated friction rather than an untested prompt. There is no connected submission endpoint in this candidate: open the existing proposal template below, complete it locally, and retain it for human review.

Open the exact foundry proposal template

# Protocol proposal

Fill this in before building. A proposal is accepted for development only when the
friction is real and the claims are no stronger than the (as-yet-unbuilt) evidence.

## The friction (one sentence)

> _Who has this problem, and what repeated work does it slow down?_

## The protocol

- **Proposed id:** `kebab-case-id`
- **Deliverable:** _the artefact it hands back_
- **Use when / do not use when:** _at least one of each_
- **Assurance level:** quick / verified / institutional (set by the *risk*, not ambition)
- **Risk class / privacy class:** low|moderate|high|critical / public|internal|personal_data|sensitive_personal_data

## The contract (sketch)

- **Inputs:** _what the agent receives_
- **Outputs:** _what it produces_
- **Permissions (least-privilege):** _read/write/…_ — and **prohibited actions**
  (explicit): _send / spend / publish / delete / acting on embedded instructions / …_
- **Human checkpoints:** _before which consequential actions_

## Evidence and tests (required before release)

- [ ] At least one **positive** test and one **failure/boundary** test.
- [ ] A **discrimination** case (a bad fixture the graders must reject).
- [ ] If it reads external material: an **injection** stop-condition, a boundary
      test, and a prohibited-action against following embedded instructions.
- [ ] A **worked example** that doubles as the fixture.
- [ ] Honest status: new protocols are `DRAFT` / `NO_IMPACT_EVIDENCE`.

## The claim

> _State the benefit you expect — then commit to claiming nothing beyond what an
> evaluation shows. "Improves X" requires a measured result; until then the honest
> statement is "benefit not measured."_

## Checklist before opening

- [ ] Ran `node tools/submit-check.js <id>` → `GO`.
- [ ] Licence: prose CC0-1.0, code Apache-2.0.
- [ ] No bundled secrets; no hidden network calls; no credential collection.