Skip to content

Eval Set v0 strawman published โ€” delta welcomeย #2

Description

@UzunGridera

The first ARP eval set strawman is up:

๐Ÿ“„ eval-set/v0-strawman.md

Status: open for delta. Not normative until v1.0.

This is the open contract behind the "patterns worth applying" claim. Without a benchmark, retrieval quality assertions are unfalsifiable. With one, conformant collectors (ARP ยง9 L1+) can self-test recall/noise on a shared baseline.

What's in v0

  • 4-field row schema (input_state, expected, policy_version, notes)
  • 3 worked examples (web hygiene, agent retry storm, filter behavior)
  • 4 open questions for delta discussion

What I'd most like delta on

  • Row schema field selection โ€” what's missing for workflow assets, cross-tenant tests, security-critical recall thresholds?
  • Rationale of policy_version pinning โ€” does the temporal drift framing hold?
  • Wildcard syntax in forbidden_fingerprints โ€” defer or settle now?

Open to PRs, comments, or counter-proposals. AgentMart, ARP collectors, LangSmith eval users โ€” schemas converging from different angles all welcome.

โ€” Uzun (agentminds.dev)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions