id stringclasses 5
values | task_type stringclasses 5
values | prompt stringclasses 5
values | gold stringclasses 5
values | meta stringclasses 5
values |
|---|---|---|---|---|
exact-01 | exact | What is the capital city of Japan? Reply with the city name only. | Tokyo | geography |
numeric-01 | numeric | A shop sells pens at 7 rupees each. What is the total cost of 23 pens? Give the final answer as a number only. | 161 | arithmetic |
json-01 | json | Return a JSON object with exactly two keys, "name" and "age", describing a person called Asha who is 31. Reply with JSON only. | {"name": "Asha", "age": 31} | structured-output |
freeform-01 | free-form | In two sentences, explain why distributed tracing helps diagnose latency in a microservice system. | Tracing links a single request across services, so the slow hop is visible instead of inferred from per-service averages. | explanation |
truncation-01 | truncation | Count aloud from 1 to 400, writing every number in words, with no abbreviation. Finish with the token END. | END | deliberately exceeds max_new_tokens to exercise the truncated failure tag |
gr.Workflow Eval Arena — sample rows
Five rows, one per scorer branch, for the gr.Workflow eval arena example.
| column | meaning |
|---|---|
id |
stable row identifier |
task_type |
exact, numeric, json, free-form, truncation |
prompt |
the prompt sent to every candidate |
gold |
reference answer |
meta |
free-text note |
Deliberately tiny: it exists to exercise every scoring path, not to rank models. Any conclusion about model quality drawn from five rows is not statistically meaningful.
- Downloads last month
- 34