> ## Documentation Index
> Fetch the complete documentation index at: https://arize-ax.mintlify.site/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Annotate Traces

> Add human feedback labels and scores to your spans to build ground truth

export const AskAlyx = ({children}) => {
  const gradientId = `askAlyxGradient-${Math.random().toString(36).slice(2)}`;
  return <div style={{
    display: "flex",
    alignItems: "flex-start",
    gap: "0.625rem",
    margin: "1rem 0",
    padding: "0.75rem 1rem",
    borderRadius: "10px",
    border: "1px solid rgba(120, 115, 245, 0.25)",
    background: "linear-gradient(135deg, rgba(255, 110, 196, 0.08), rgba(120, 115, 245, 0.08))"
  }}>
      <svg width="18" height="18" viewBox="0 0 21 17" xmlns="http://www.w3.org/2000/svg" style={{
    flexShrink: 0,
    marginTop: "0.2rem"
  }}>
        <defs>
          <linearGradient id={gradientId} x1="0%" y1="100%" x2="100%" y2="0%">
            <stop offset="0%" stopColor="#FF3CA8" />
            <stop offset="100%" stopColor="#4827C1" />
          </linearGradient>
        </defs>
        <path d="M6.28906 12.7223C6.28889 11.3831 5.24007 10.3385 3.98926 10.3385C2.73859 10.3387 1.68963 11.3832 1.68945 12.7223C1.68945 14.0616 2.73849 15.1059 3.98926 15.1061C5.24018 15.1061 6.28906 14.0617 6.28906 12.7223ZM7.81152 0.557254C9.70554 -0.563135 12.1081 0.0645388 13.2607 1.91468L13.3691 2.09827V2.09925L20.7266 15.402C20.8713 15.6637 20.8667 15.9823 20.7148 16.2399C20.5629 16.4975 20.2864 16.6559 19.9873 16.6559H14.5459C13.0953 16.6474 11.7648 15.848 11.0469 14.5748V14.5739L6.33301 6.19104C5.22656 4.2273 5.87813 1.706 7.80957 0.558231L7.81152 0.557254ZM11.8906 2.91761C11.2374 1.74047 9.78961 1.34962 8.67188 2.01038L8.67285 2.01136C7.61521 2.64 7.19477 3.99924 7.69336 5.13733L7.80566 5.36194V5.36292L12.5186 13.7448C12.9466 14.5038 13.7274 14.9616 14.5557 14.9664H18.5547L11.8906 2.91663V2.91761ZM7.97949 12.7223C7.97949 14.9527 6.21527 16.7965 3.98926 16.7965C1.7634 16.7963 0 14.9526 0 12.7223C0.000173728 10.4921 1.76351 8.64923 3.98926 8.64905C6.21516 8.64905 7.97932 10.492 7.97949 12.7223Z" fill={`url(#${gradientId})`} />
      </svg>
      <span>{children}</span>
    </div>;
};

Your automated evals say a response is "grounded" — but is it really? Sometimes you need a human to weigh in. **Annotations** let your team add ground-truth labels and scores directly on spans, building a feedback loop between humans and your AI.

## How to do it

1. Open a trace and click into any span
2. Click the **Annotate** toggle in the span toolbar
3. Select an **annotation config** (e.g., "Correctness", "Helpfulness") or create a new one
4. Add your label or score — saves automatically

<AskAlyx>**Ask Alyx** to annotate these spans/traces for you — try *"Mark the ungrounded responses as incorrect."*</AskAlyx>

## Span-level and trace-level annotations

Annotations can apply to either a single span or an entire trace. Use a **span-level annotation** when you are judging one operation, such as whether a retrieved document was relevant. Use a **trace-level annotation** when the judgment concerns the complete request or conversation, such as end-to-end response quality.

## Annotation configs

Configs define the schema for your labels. Shared across the project so everyone uses the same schema.

* **Categorical** — fixed labels (e.g., "correct", "incorrect", "partially correct")
* **Continuous** — numeric scores on a range (e.g., 1–5)

Create new configs on the fly from the annotation panel: click **+ New Config**, choose type, add options, save.

### Annotation notes

In addition to labels and scores, you can attach **free-text notes** to any annotation. Notes are useful for explaining edge cases, providing context for disagreements, or flagging spans for follow-up discussion.

## Measure eval quality with annotations

Use annotations as ground truth to measure how well your automated evals perform:

```sql theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
SELECT
    PRECISION(
        predicted = "eval.Groundedness.label",
        actual = "annotation.Human Groundedness.label",
        pos_class = 'grounded'
    )
FROM model
```

## Annotations vs. evals

|              | Annotations               | Evals                            |
| ------------ | ------------------------- | -------------------------------- |
| **Who**      | Humans                    | Automated (LLM-as-judge or code) |
| **Scale**    | Small samples             | Every span                       |
| **Best for** | Ground truth, calibration | Production monitoring            |

Use both: evals for scale, annotations for accuracy.
