Skip to main content
An evaluator defines what to measure (an LLM-as-judge template, code, or an agent-as-a-judge harness). A task points that evaluator at data: a project’s live traces, a dataset, or an experiment, and controls how much of it gets scored and how often. The GraphQL API covers both halves, plus the related task types for running an experiment on demand and driving prompt optimization. This guide assumes you already know how to form a GraphQL call with an x-api-key header against https://app.arize.com/graphql.

Find the IDs you need

Evaluator and task mutations need a space ID, a model (project) ID, and often an evaluator ID or a named LLM integration ID. Start from viewer.
Once you have an evaluator, task, or run ID, fetch it directly with node(id: "<ID>") { ... on Evaluator { name } }. See Using global node IDs for how these opaque IDs work.

List evaluators and online tasks in a space

Evaluators live in Eval Hub independently of any task; a single evaluator can be attached to many tasks. This query lists both so you can see what exists before you wire one to the other.

Create an LLM-as-judge evaluator and a new version

createEvaluator creates the evaluator and its first version together. createEvaluatorVersion adds a new version (for example, tightening the rails) without losing history. To just rename or re-describe an evaluator without creating a version, use editEvaluator instead; it takes evaluatorId, name, and description.
Reference: createEvaluator, createEvaluatorVersion.

Create an online eval task bound to an evaluator

createEvalTask points an evaluator at live traces from a project. samplingRate is a fraction between 0 and 1, and queryFilter narrows which spans are admitted before sampling is applied.
Reference: createEvalTask.

Pause, resume, resample, or migrate a task to Eval Hub

patchEvalTask updates any subset of a task’s fields; omit what you are not changing. Older tasks created before Eval Hub store their evaluator config inline on the task; migrateTaskToEvalHub (input: just onlineTaskId) converts one to reference Eval Hub evaluators instead. Once migrated, use patchEvalTaskWithEvalHub rather than patchEvalTask, since it takes evaluators (hub references) instead of inline templateEvaluators/codeEvaluators.
Reference: patchEvalTask, patchEvalTaskWithEvalHub.

Run a task on demand and cancel a run

runOnlineTask backfills a task over a historical time range instead of waiting for live traffic, capped at maxSpans (10,000 by default). Both mutations return a union: a success type on the happy path, or TaskError with a message and code.
Reference: runOnlineTask, cancelOnlineTaskRun.

Run an experiment task

runExperimentTask runs a configured LLM generation (or a template eval, code eval, or agent call) across a dataset or dataset version, one runConfigurations entry per instance you want to compare. Only the required fields are shown; add templateConfig, codeEvalConfig, or agentConfig to a run configuration depending on its experimentType.
Reference: runExperimentTask.

Create and update a prompt optimization task

createPromptOptimizationTask runs prompt learning against a dataset (or an experiment’s output), using feedback columns to steer the rewrite. patchPromptOptimizationTask updates it, most commonly to turn continuous optimization on or off.
Reference: createPromptOptimizationTask, patchPromptOptimizationTask.

Delete an evaluator and online tasks

Deleting an evaluator does not delete the tasks it was attached to; detach or delete those tasks separately. deleteOnlineTask is a batch mutation, up to 50 task IDs per request.
Reference: deleteEvaluator, deleteOnlineTask.

Gotchas and behavior notes

CreateEvalTaskMutationInput/PatchEvalTaskMutationInput.modelId and .datasetId are plain ID, not ID!, and the schema does not enforce that you provide one or the other, so a task created with neither has no data source to scope against. RunOnlineTaskMutationInput.onlineTaskId is similarly typed ID even though there is no meaningful way to run a backfill without it.
Evaluator and column names (TemplateEvaluationConfigInput.name, CodeEvaluationConfigInput.name, and others) use the EvalColumnName scalar. The schema does not expose its validation rules; naming constraints are enforced server-side at request time.
Template evaluators use TemplateEvaluationConfigDirection, harness evaluators use HarnessEvaluationConfigDirection, and System One questions use OptimizationDirection. All three have the identical value set (maximize, minimize, none), but they are distinct enum types, so you cannot share a single constant across evaluator kinds in a strongly typed client.
SystemOneEvaluatorInput.llmIntegrationId is typed String!, while the equivalent field on TemplateEvaluationLlmConfigInput, HarnessEvaluatorInput, and OnlineTaskLLMConfigInput is typed ID!. All of them expect the same relay global ID string; only the GraphQL type differs.
patchEvalTask, patchEvalTaskWithEvalHub, deleteOnlineTask, and patchPromptOptimizationTask carry no top-level description, unlike their sibling create mutations. Field-level descriptions are present, but you have to infer each mutation’s purpose from its name and input shape.

Evaluator and task mutations reference

Full argument and field listing for all 14 evaluator and task mutations.

All GraphQL mutations

Browse mutations for every other domain.

API explorer

Try queries and mutations interactively with autocomplete.