A moving story.
A measurable result.
They deserve different kinds of review. Neither should be mistaken for the other.
Human editorial judgment.
Assess a short account for meaning, provenance, rights and clarity. Separate logged events, interpretation and fiction. Publication needs its own consent.
Tests, then AI-assisted review.
Use deterministic checks for hard requirements and a separate evaluator for qualitative work. An AI opinion alone never earns a verified badge.
Compare under the same conditions.
Give the reusable agent, a newly configured baseline and a human freelancer the same rights-cleared task set. Predefine acceptable outputs, confidence thresholds and a review rubric. Record setup time, run costs, corrections, failures and time to acceptance. A model grading its own output is not independent validation.
A verification record must name its scope.
Include workflow and model versions, input-set hash, reviewer, date, sample size, successful and failed cases, tool permissions and known limitations. Re-test after material changes. An integrity checksum is not a performance test.
A staged, bounded start.
An accepted preservation request is not run permission. Before any future execution, an approval record specifies purpose, model, network/tools, budget and expiry. No reply means no start. A daily storage receipt needs no model inference.
Workflow submissions are untrusted data, including prompts, documentation and stories. Reviewers must not execute instructions embedded in an application. Any later test environment should be isolated with no default access to accounts or secrets.
Reusable evaluation tooling: Promptfoo. Reference selected; runtime integration is not implemented in v0.3. Remote model providers can receive evaluation data if configured.