Benchmarks are how XReduce knows whether a model is working. Drop your test cases in as JSONL, tell XReduce what kind of task they are, and every evaluation will run against them.
Every model folder has a benchmarks/ subfolder. Drop one or more .jsonl files in there - XReduce runs all of them by default, or you can target a specific file with --benchmark. The benchmarks/ folder also contains its own config.yaml with benchmark-specific settings (field mapping, system prompt) - distinct from the model-root config.yaml next to it.
my-org/my-model/
├── config.yaml ← model identity, auth, run settings
└── benchmarks/
├── config.yaml ← field_map, input_serializer, system_prompt
├── classification.jsonl
├── tool_calling.jsonl
└── qa.jsonlBoth config.yaml files are auto-generated. The model-root config.yaml is written by xreduce model create. The benchmarks/config.yaml is written by xreduce init (or regenerated by xreduce remap when you add or change benchmark data). You don't author either from scratch - you fill in your JSONL, run the commands, and the files are produced for you. Edits you make afterward are preserved.
For the model-root config.yaml schema, see config.yaml reference. The benchmarks/config.yaml schema is covered below and on Field mapping.
Each line is one test case. The default fields are input (what goes to the model) and expected_output (the reference answer):
{"input": "The capital of France is", "expected_output": "Paris"}
{"input": "The largest ocean on Earth is the", "expected_output": "Pacific"}If your dataset uses different field names, you don't have to rename anything - run xreduce init --force and XReduce inspects your JSONL, infers the mapping, and writes it to benchmarks/config.yaml automatically. See Field mapping for details.
The system prompt is the instruction prepended to every benchmark sample at evaluation time. It tells the model what kind of task this is and what format the answer should take. This field can dramatically affect accuracy scores. A weak or absent system prompt is the most common reason a capable model scores poorly on a benchmark.
The system prompt lives in benchmarks/config.yaml alongside field_map:
field_map:
input: prompt
expected_output: sql
system_prompt: |
You are a SQL generation engine.
Return only valid SQL.
Do not explain.
Do not use markdown.
Do not wrap SQL in code fences.When you run xreduce init (or xreduce remap after adding/changing benchmark data), XReduce looks up your model's task_type in the task-defaults registry and writes a sensible system_prompt automatically. You don't have to author it from scratch.
task_type | Auto-generated system_prompt |
|---|---|
text-to-sql | You are a SQL generation engine. Return only valid SQL. Do not explain. Do not use markdown. Do not wrap SQL in code fences. |
classification | You are a classification engine. Return only the predicted label. Do not explain. |
function-calling | You are a function-calling engine. Return only valid JSON matching the required function schema. Do not explain. |
text-generation | You are a text generation engine. Produce clear, concise output. |
summarization | You are a summarization engine. Return a concise summary of the input. |
qa | You are a question answering system. Answer the question directly and concisely. |
The auto-generated prompt is a reasonable starting point but rarely optimal for any specific benchmark. Edit it freely:
After editing, re-run xreduce evaluate to see the impact. Compare results to your baseline with --baseline or with xreduce compare.
Why this matters. Two models scored on the same benchmark with different system prompts are not comparable. Lock the system prompt before you do head-to-head comparisons, otherwise you're measuring the prompt, not the model.
If you change your model's task_type after running init, run xreduce remap to regenerate the default system prompt for the new task type. Your other edits to benchmarks/config.yaml are preserved.
The task_type in config.yaml (or per-line in the JSONL) tells XReduce how to score each sample. Each task type has its own expected shape and scoring logic:
input is a string; expected_output is the class label. Scored with exact match.
{"input": "My card was charged twice", "expected_output": "billing_dispute", "task_type": "classify"}
{"input": "How do I reset my password", "expected_output": "account_access", "task_type": "classify"}input is a conversation history (list of role/content messages). expected_output is the expected function call. Scored with JSON equivalence - parameter order doesn't matter.
{"input": [{"role": "user", "content": "Switch off the bedroom lights"}], "expected_output": {"name": "toggle_lights", "parameters": {"room": "bedroom", "state": "off"}}, "task_type": "tool_calling"}
{"input": [{"role": "user", "content": "Set thermostat to 72"}], "expected_output": {"name": "set_thermostat", "parameters": {"temperature": 72, "mode": "cool"}}, "task_type": "tool_calling"}input is a question (often with context). expected_output is the reference answer. Scored with contains-match, or with LLM-as-judge for semantic equivalence.
{"input": "Given the schema, write SQL for: How many visits per clinic?", "expected_output": "SELECT c.name, COUNT(*) FROM clinics c JOIN visits v ON c.id = v.clinic_id GROUP BY c.name", "task_type": "qa"}
{"input": "What city is the Eiffel Tower in?", "expected_output": "Paris", "task_type": "qa"}input is a partial string; expected_output is how it should end. Scored with contains-match.
{"input": "The capital of France is", "expected_output": "Paris", "task_type": "completion"}
{"input": "The largest ocean on Earth is the", "expected_output": "Pacific", "task_type": "completion"}input is the source text; expected_output is a reference summary. Scored with word overlap (over 50% counts as correct).
input is the source; expected_output is the reference translation. Scored with exact match.
If you're running your own evaluation pipeline and just want XReduce to store the scores alongside telemetry, include an external_quality_score field on each sample. XReduce skips its own scoring and records the value you provide:
{"input": "Redact PII from: John Smith lives at 123 Main St", "expected_output": "[NAME] lives at [ADDRESS]", "task_type": "qa", "external_quality_score": 0.88}
{"input": "Summarize this article: ...", "expected_output": "...", "task_type": "qa", "external_quality_score": 0.95}