foundry-evaluations
Azure/kars/runtimes/openclaw/skills/foundry-evaluations/SKILL.md
Evaluate agent quality using Foundry OpenAI Evals API. Create evaluations, run them against models, and analyze results.
Skill44 starsChanged 3 months ago
- Sends data out
What's in it
- Foundry Evaluations — OpenAI Evals API
- Endpoint
- Operations
- List evaluations
- Create an evaluation
- Run an evaluation
- List evaluators (built-in + custom)
- List evaluation rules
- When to use
- When NOT to use
---
name: foundry-evaluations
description: Evaluate agent quality using Foundry OpenAI Evals API. Create evaluations, run them against models, and analyze results.
metadata: {"openclaw": {"requires": {"env": ["FOUNDRY_PROJECT_ENDPOINT"]}, "primaryEnv": "FOUNDRY_PROJECT_ENDPOINT"}}
---
# Foundry Evaluations — OpenAI Evals API
You can evaluate agent and model quality using the Foundry OpenAI Evals API. Create evaluation definitions with testing criteria, run them against models, and analyze pass/fail results.
## Endpoint
All requests: `http://localhost:8443` with `?api-version=2025-11-15-preview`. Auth is automatic.
## Operations
### List evaluations
```bash
curl -s 'http://localhost:8443/openai/evals?api-version=2025-11-15-preview'
```
### Create an evaluation
```bash
curl -s -X POST 'http://localhost:8443/openai/evals?api-version=2025-11-15-preview' \
-H 'Content-Type: application/json' \
-d '{"name":"quality-check","data_source_config":{"type":"custom","item_schema":{"type":"object","properties":{"input":{"type":"string"},"expected":{"type":"string"}},"required":["input","expected"]}},"testing_criteria":[{"type":"string_check","name":"exact-match","input":"{{sample.output_text}}","reference":"{{item.expected}}","operation":"eq"}]}'
```
### Run an evaluation
```bash
curl -s -X POST 'http://localhost:8443/openai/evals/eval_abc123/runs?api-version=2025-11-15-preview' \
-H 'Content-Type: application/json' \
-d '{"name":"run-1","data_source":{"type":"jsonl","source":{"type":"file_content","content":[{"item":{"input":"2+2","expected":"4"}}]}}}'
```
### List evaluators (built-in + custom)
```bash
curl -s 'http://localhost:8443/evaluators?api-version=2025-11-15-preview'
```
### List evaluation rules
```bash
curl -s 'http://localhost:8443/evaluationrules?api-version=2025-11-15-preview'
```
## When to use
- Measuring agent quality (accuracy, safety, groundedness)
- A/B testing model versions or prompt changes
- Continuous evaluation of production agent responses
- Red-teaming and safety testing
## When NOT to use
- For runtime inference (use /v1/chat/completions or /openai/responses)
- For conversation management (use foundry-conversations skill)
More agent context in Azure/kars
14 other files this repository gives its agents.
Copilot instructions
llms.txt
Skill
- agt-e2e-encryption.github/skills/agt-e2e-encryption/SKILL.md
- kars-deployment.github/skills/kars-deployment/SKILL.md
- mesh-federationmesh-plugin/skills/mesh-federation/SKILL.md
- agt-governanceruntimes/openclaw/skills/agt-governance/SKILL.md
- foundry-agentsruntimes/openclaw/skills/foundry-agents/SKILL.md
- foundry-coderuntimes/openclaw/skills/foundry-code/SKILL.md
- foundry-conversationsruntimes/openclaw/skills/foundry-conversations/SKILL.md
- foundry-deploymentsruntimes/openclaw/skills/foundry-deployments/SKILL.md
- foundry-knowledgeruntimes/openclaw/skills/foundry-knowledge/SKILL.md
- foundry-memoryruntimes/openclaw/skills/foundry-memory/SKILL.md
- foundry-web-searchruntimes/openclaw/skills/foundry-web-search/SKILL.md
- kars-spawnruntimes/openclaw/skills/kars-spawn/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
No reports yet. Be the first to say whether it worked.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

