Workload
structured-agent-v1
Draft
Question it answers
Does JSON/function calling stay reliable under realistic agent traffic?
What is measured
BFCL-style calls including relevance rejection, parallel and executable calls, loop detection; deterministic sandbox; parsing success alone is not task success.
Claims this workload cannot support
Agent-readiness claims from parse rate alone.
Identity discipline
Every released workload revision freezes its item hashes, tokenizer revision, load model, cache policy and quality gates. Results reference the exact revision; changing any identity field creates a new revision.