# Jetty > Jetty is where professionals find, share, and run reliable AI workflows. You write a runbook — one markdown file that gives an agent its job, its bar for "done", and its checks — and Jetty runs it in a managed sandbox, records every run as a trajectory, and helps you improve it until it works every time. These docs are written for builders: developers who work in Claude Code, Codex, or Cursor. Connect via the Claude Code plugin, the MCP server, the TypeScript SDK, or the OpenAI-compatible REST API. ## For agents - [agent-instructions.md](https://jetty.io/agent-instructions.md): pasteable onboarding trigger — numbered steps, human-only handoffs, ends in a Return block. Start here. - [MACHINE_CONTEXT.md](https://jetty.io/MACHINE_CONTEXT.md): deep reference — state model, CLI verb contracts, runbook format, anti-patterns. - [auth.md](https://jetty.io/auth.md): how an agent obtains, uses, and revokes a Jetty API credential. Machine-readable discovery: https://jetty.io/.well-known/oauth-protected-resource, https://jetty.io/.well-known/oauth-authorization-server, https://jetty.io/.well-known/api-catalog. - [Agent skills index](https://jetty.io/.well-known/agent-skills/index.json): the jetty, jetty-setup, create-runbook and optimize-runbook skills with SHA-256 digests. MCP Server Card at https://jetty.io/.well-known/mcp/server-card.json; ARD capability manifest at https://jetty.io/.well-known/ai-catalog.json. Every HTML page on this site is also available as Markdown: request it with `Accept: text/markdown`. - [llms-full.txt](https://jetty.io/llms-full.txt): every docs page below as Markdown, in one file. ## Get Started From zero to a runbook you trust, in one sitting. - [Get Started](https://jetty.io/docs/get-started): What Jetty runs, and three ways in: start free, install locally, or call it from your app. - [Try a Project](https://jetty.io/runbooks): Browse the Jetty Runbook Directory and try it yourself on your own account. - [Bring Your Agent](https://jetty.io/docs/integrations/bring-your-own-agent): Install Jetty locally: the plugin, the jetty-setup skill, and your first run. - [Connect Your App](https://jetty.io/docs/get-started/connect-your-app): Every on-ramp: the MCP server, the Client SDK, and the API. ## Product Overview What Jetty does, in two motions: run the workload, then read the run. - [Product Overview](https://jetty.io/docs/product-overview): The two halves of Jetty: run AI workloads in managed sandboxes, and investigate how well the agent performed. - [Running AI Workloads](https://jetty.io/docs/product-overview/running-ai-workloads): How a runbook becomes a running workflow: the author → run → investigate → improve lifecycle, and what Jetty assembles around every run. - [Investigating Agent Quality](https://jetty.io/docs/product-overview/investigating-agent-quality): How to read a run and its verdict, add a check at the failure moment, compare runs honestly, and run the investigate-and-improve loop. ## AI Workloads The job you hand to Jetty, and the pieces that carry it out. - [AI Workloads](https://jetty.io/docs/ai-workloads): What counts as an AI workload, and the two pieces every one is built from: a runbook and a runtime. - [Runbooks](https://jetty.io/docs/ai-workloads/runbooks): The hero artifact: one markdown file with the job, the bar for “done”, and the checks. - [Runtime Configuration](https://jetty.io/docs/ai-workloads/runtime-configuration): The agents Jetty runs inside its own sandboxes, which providers back each, and how to pick one. - [antigravity](https://jetty.io/docs/ai-workloads/runtime-configuration/antigravity): Google's current terminal agent (agy) in Gemini-API-key mode: catalog slugs, effort tiers, MCP support. - [claude-code](https://jetty.io/docs/ai-workloads/runtime-configuration/claude-code): The default runtime: Anthropic's agent with MCP, full usage telemetry, and the loop Jetty's skills are tuned against. - [codex](https://jetty.io/docs/ai-workloads/runtime-configuration/codex): OpenAI's agent on a pinned release: Responses-API only, config.toml routing, the reasoning-effort knob. - [gemini-cli](https://jetty.io/docs/ai-workloads/runtime-configuration/gemini-cli): Google's earlier terminal agent; kept for existing runbooks; start new Gemini work on antigravity. - [goose](https://jetty.io/docs/ai-workloads/runtime-configuration/goose): Block's agent, included for comparison: bare model ids, direct providers, best-effort output parsing. - [hermes](https://jetty.io/docs/ai-workloads/runtime-configuration/hermes): Nous Research's agent; the one runtime that reaches all six providers, including self-hosted vLLM. - [opencode](https://jetty.io/docs/ai-workloads/runtime-configuration/opencode): The open-source SST agent with permissions pre-allowed and Playwright + Firecrawl MCP ready on every run. - [pi](https://jetty.io/docs/ai-workloads/runtime-configuration/pi): A small provider-agnostic agent with no MCP client by design: CLI tools and skills instead. - [Integrations](https://jetty.io/docs/ai-workloads/integrations): How a workload connects to the world: integrations are environment variables injected into every run: Slack, databases, GitHub, CRMs, your own app. - [Runs](https://jetty.io/docs/ai-workloads/runs): Every run is captured end to end: each step, its inputs, outputs, files, and status, plus labels to group and compare runs. - [Agent Evaluations](https://jetty.io/docs/ai-workloads/agent-evaluations): Code checks, checklists, and LLM or agent judges live in the runbook and run in flight. Every run records a verdict you can investigate and improve against. - [REST API](https://jetty.io/docs/reference/api): The Jetty REST API: run runbooks as tasks (deploy, run, monitor — the primary endpoint), the OpenAI-compatible chat-completions compatibility layer, schedules, webhooks, and GitHub-PR APIs. - [Machine Instructions](https://jetty.io/docs/reference/api/machine-instructions): The runbook API surface for agents and programmatic callers: tasks, runs, init_params, trajectories, files, MCP config, sweeps, snapshots, webhooks, schedules. - [Actions](https://jetty.io/docs/reference/api/steps): The pre-built actions that run alongside runbooks — AI models, control flow, data processing, evaluation — and the path expressions that wire them together. ## Jetty and Friends Run your agent on Jetty or anywhere, and evaluate on Jetty. - [Jetty and Friends](https://jetty.io/docs/jetty-and-friends): Keep your agent in any framework; use Jetty as the independent grader, store, and A/B harness. Flue and eve worked examples. - [Bring Your Friends and Framework](https://jetty.io/docs/jetty-and-friends/bring-your-friends-and-framework): Keep your agent in its own framework and let Jetty grade every run: the pattern, with Eve and Flue worked examples. - [Eve + Jetty](https://jetty.io/docs/jetty-and-friends/bring-your-friends-and-framework/eve): Mount the @jetty/eve extension into a Vercel eve agent: live grading, a grade-steered experiment bandit, and an eval reporter. Catch a regression before it ships. - [Flue + Jetty](https://jetty.io/docs/jetty-and-friends/bring-your-friends-and-framework/flue): The complete flue-jetty walkthrough: grade every Flue agent run with an independent rubric and catch a regression before it ships. - [Client SDK](https://jetty.io/docs/integrations/sdk): @jetty/sdk for TypeScript: JettyClient, runAndWait, etc. - [MCP Server](https://jetty.io/docs/integrations/mcp-server): 27 tools for Claude Code, Cursor, VS Code, Windsurf, Zed, Codex, and WebMCP. ## Guides Task-shaped walkthroughs for the work you actually do. - [Guides](https://jetty.io/guides): Every guide in one place: use cases from teams running AI workloads, improving a runbook from its own runs, and scheduling workloads. - [Build vs. Buy Evaluation](https://jetty.io/docs/guides/ai-agent-evaluation-cost): What AI agent evaluation costs to build versus buy: the variables that matter, a framework for making the call, and how to test it against your own workload. - [Building AI Evaluations](https://jetty.io/docs/guides/building-ai-evaluations): A way to know, before you ship, whether your agent is actually doing the job it was built to handle at the level of quality you expect. - [Creating a Benchmark](https://jetty.io/docs/guides/creating-a-benchmark): Pick one axis of comparison, write the runbook that defines the job, and run the task set against every candidate in parallel so the comparison is reproducible. - [Defining Agent Checklists](https://jetty.io/docs/guides/ai-agent-checklist): Keep your agent on task with a checklist: the conditions that define done, ticked and rechecked inside every run, then investigated across runs and kept honest in production. - [Improving Agent Runbooks](https://jetty.io/docs/guides/improving-agent-runbooks): Pull the latest runs with your coding agent, tighten the evals they expose, and let the optimize-runbook skill propose evidence-backed edits. - [Research](https://jetty.io/guides/research): Papers Jetty has contributed to, on agent benchmark standards, reproducible ML evaluations and AI accountability. - [Scheduling AI Workloads](https://jetty.io/docs/guides/scheduling-ai-workloads): Run a runbook on a cron cadence so evals stay fresh as models and providers drift. - [Use Cases](https://jetty.io/guides/use-cases): Stories from teams running AI workloads on Jetty: what they set up, how parallel runs and evals changed the work, and what came out. - [Index](https://jetty.io/docs/overview): What Jetty is, who it is for, and every docs page in one place. ## Source - [jettyio/jettyio-skills](https://github.com/jettyio/jettyio-skills): Claude Code plugin, skills, and the MCP server. - [jettyio/jetty-sdk](https://github.com/jettyio/jetty-sdk): the TypeScript client SDK and the Flue integration example.