Iris — stop shipping agents on vibes

Iris is an open-source, self-hosted MCP agent evaluation server for tracing runs, scoring output quality and safety, and monitoring per-agent costs.

Visit Website
Iris — stop shipping agents on vibes

Introduction

Overview

Iris is an open-source evaluation server for MCP-compatible AI agents. It captures agent runs and evaluates output quality, safety, grounding, task completion, and cost, providing information that ordinary infrastructure monitoring cannot derive from successful HTTP responses alone.

The server runs on the user's own machine and registers through MCP, so compatible clients discover its tools when they connect. Agents can submit traces directly, a host hook can forward turns, or existing trace files can be ingested.

Evaluation and Safety Features

Iris combines deterministic checks with optional model-based evaluation. Its shipped capabilities include 25 built-in evaluation rules and tools for logging, evaluation, querying, rule deployment, trace deletion, citation verification, and LLM-as-judge workflows.

Safety checks cover personal data and credential patterns, prompt-injection patterns, prohibited words, stub outputs, and hallucination markers. The capability documentation distinguishes fully supported judgments, limited judgments, open gaps, and questions that do not apply, rather than presenting every evaluation area as equally complete.

Tracing and Cost Controls

Each received run can be represented as a hierarchical trace with agent, model, and tool spans. Iris records per-tool latency, prompt and completion token usage, total tokens, costs, and custom metadata for attribution.

Teams can inspect costs by trace, agent, and time window, then apply thresholds to flag excessive spending. Repeated evaluations can also be compared across runs to expose regressions, pass rates, and flaky cases within the documented limits of those comparisons.

Deployment and Compatibility

Iris works with MCP-compatible clients without requiring an application SDK. The self-hosted package includes a local dashboard and playground and supports both standard input/output and HTTP transports.

Evaluation data remains on the operator's machine by default. Data leaves the machine only when the operator configures trace export or enables an external LLM judge using their own provider key.

Pricing and Availability

The current self-hosted edition is MIT licensed and free, with no account, metering, quota, or evaluation limit. Custom Zod rules, built-in checks, the dashboard, LLM-as-judge, and citation verification are included; external model calls are billed directly by the selected provider.

Hosted team and enterprise capabilities are described only as possible future offerings and are not currently available or priced. No compliance certifications have been started, and self-hosting is the only supported deployment model at present.