> ## Documentation Index
> Fetch the complete documentation index at: https://mcpjam-docs-coming-from-mcp-inspector.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Create an eval suite (author-only, does not run)

> Creates a runnable eval suite — the suite record plus its test cases — and responds `201` **synchronously**, WITHOUT executing anything. Use this to author a suite, then run it later with `POST /eval-runs` (passing the returned `suiteId`).

This is distinct from `POST /eval-runs`, which creates a run and **detaches execution**, responding `202` with a `runId`. There is no concurrency cap here (no run is started).

The body uses an ergonomic authoring shape: a suite-level default `model` (and optional `provider`) applies to every test unless the test overrides it; `provider` is derived from a `provider/model` id when neither is supplied. Each test's case body is an ordered `steps` array (prompt / toolCall / interact / assert).

Guest callers are denied (suite creation is a write).



## OpenAPI

````yaml /reference/openapi.json post /projects/{projectId}/eval-suites
openapi: 3.1.0
info:
  title: MCPJam API
  version: 1.0.0-preview
  description: >-
    Programmatic access to MCP servers saved in your MCPJam projects — live
    diagnostics (validate, inspect, export) and operations: call tools, render
    prompts, run eval suites asynchronously and poll their results, and import
    OAuth tokens.


    **The API is in preview**: the surface may change while we finish the
    design. Error `code` values are stable; error `message` strings are not.
    Write clients that ignore unknown response fields.
  contact:
    name: MCPJam
    url: https://github.com/MCPJam/inspector/issues
servers:
  - url: https://app.mcpjam.com/api/v1
    description: Hosted MCPJam
security:
  - bearerAuth: []
tags:
  - name: Hosts
    description: >-
      Project hosts: named model + capability profiles you run chats and eval
      suites against.
  - name: Computer environments
    description: >-
      Custom Computer images: digest-pinned Dockerfile environments your
      project's computers boot from.
  - name: Server diagnostics
    description: Connect-level health checks against a saved MCP server.
  - name: Primitives
    description: 'The server''s MCP primitives: tools, prompts, and resources.'
  - name: Export
    description: Full-server snapshots for diffing and CI.
  - name: Execution
    description: 'Run the server''s primitives: call tools, render prompts.'
  - name: Eval runs
    description: >-
      Asynchronous eval suite runs: create with 202, poll status, iterations,
      and traces.
  - name: OAuth
    description: 'Bring-your-own OAuth: import externally obtained tokens for a server.'
  - name: Chatboxes
    description: >-
      Read-only access to the chatboxes published from a project: listing,
      settings, attached servers, and share links.
  - name: Catalog
    description: >-
      Discover the resources the other routes operate on: your account,
      projects, servers, eval suites, and chat sessions.
  - name: Tunnels
    description: >-
      Relay tunnels that expose local MCP servers through a public URL,
      registered as first-class project servers (the `mcpjam tunnel` CLI flow).
paths:
  /projects/{projectId}/eval-suites:
    post:
      tags:
        - Eval runs
      summary: Create an eval suite (author-only, does not run)
      description: >-
        Creates a runnable eval suite — the suite record plus its test cases —
        and responds `201` **synchronously**, WITHOUT executing anything. Use
        this to author a suite, then run it later with `POST /eval-runs`
        (passing the returned `suiteId`).


        This is distinct from `POST /eval-runs`, which creates a run and
        **detaches execution**, responding `202` with a `runId`. There is no
        concurrency cap here (no run is started).


        The body uses an ergonomic authoring shape: a suite-level default
        `model` (and optional `provider`) applies to every test unless the test
        overrides it; `provider` is derived from a `provider/model` id when
        neither is supplied. Each test's case body is an ordered `steps` array
        (prompt / toolCall / interact / assert).


        Guest callers are denied (suite creation is a write).
      operationId: createEvalSuite
      parameters:
        - $ref: '#/components/parameters/projectId'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvalSuiteCreateRequest'
            example:
              name: smoke
              serverIds:
                - srv_abc123
              model: anthropic/claude-haiku-4.5
              tests:
                - title: lists files
                  steps:
                    - id: s1
                      kind: prompt
                      prompt: List the files in the project root.
                    - id: s2
                      kind: assert
                      assertion:
                        type: toolCalledWith
                        toolName: list_files
                        args:
                          args: {}
      responses:
        '201':
          description: The suite was created.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EvalSuiteCreated'
        '400':
          $ref: '#/components/responses/ValidationError'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
        '429':
          $ref: '#/components/responses/RateLimited'
        '500':
          $ref: '#/components/responses/InternalError'
        '502':
          $ref: '#/components/responses/ServerUnreachable'
components:
  parameters:
    projectId:
      name: projectId
      in: path
      required: true
      description: ID of the hosted project that contains the server.
      schema:
        type: string
  schemas:
    EvalSuiteCreateRequest:
      type: object
      description: >-
        Author-only suite-create body. A suite-level default `model` (and
        optional `provider`) applies to every test unless the test overrides it.
      required:
        - name
        - serverIds
        - model
        - tests
      properties:
        name:
          type: string
          minLength: 1
          description: Suite name.
        description:
          type: string
        serverIds:
          type: array
          minItems: 1
          description: Servers (by canonical project ID) the suite's cases run against.
          items:
            type: string
        serverNames:
          type: array
          description: Optional display names, parallel to `serverIds`.
          items:
            type: string
        model:
          type: string
          minLength: 1
          description: >-
            Suite-level default model id (e.g. `anthropic/claude-haiku-4.5`).
            Used for any test that omits `model`.
        provider:
          type: string
          description: >-
            Optional suite-level default provider. When omitted, the provider is
            derived from a `provider/model` id.
        passCriteria:
          type: object
          properties:
            minimumPassRate:
              type: number
        tags:
          type: array
          description: Accepted for forward-compat; not persisted today (no-op).
          items:
            type: string
        tests:
          type: array
          minItems: 1
          maxItems: 100
          description: Test cases to create in the suite.
          items:
            type: object
            required:
              - title
              - steps
            properties:
              title:
                type: string
                minLength: 1
              steps:
                type: array
                minItems: 1
                description: >-
                  Ordered test steps (the unified test model). The first
                  `prompt` step is the case query; `toolCalledWith` asserts are
                  the expected tool calls; a single model-free `toolCall` step
                  is a render-check.
                items:
                  $ref: '#/components/schemas/EvalTestStep'
              runs:
                type: integer
                minimum: 1
                maximum: 10
                description: Iterations to execute for this case. Defaults to 1.
              model:
                type: string
                description: Per-case model override. Defaults to the suite-level `model`.
              provider:
                type: string
                description: Per-case provider override.
              expectedOutput:
                type: string
              isNegativeTest:
                type: boolean
                description: When `true`, the case passes if NO tools are called.
              scenario:
                type: string
              advancedConfig:
                type: object
                description: Optional `system`, `temperature`, `toolChoice` overrides.
                additionalProperties: true
            additionalProperties: true
      additionalProperties: false
    EvalSuiteCreated:
      type: object
      required:
        - suiteId
        - name
        - caseUpsert
      properties:
        suiteId:
          type: string
        name:
          type: string
        servers:
          type: array
          description: The servers attached to the suite. `name` is present when supplied.
          items:
            type: object
            required:
              - id
            properties:
              id:
                type: string
              name:
                type: string
        caseUpsert:
          type: object
          description: >-
            Per-case create outcomes. Partial failures don't abort the suite; an
            all-failed new suite is rejected with VALIDATION_ERROR.
          properties:
            committed:
              type: array
              items:
                type: object
                properties:
                  id:
                    type: string
                  name:
                    type: string
            failed:
              type: array
              items:
                type: object
                properties:
                  id:
                    type: string
                  name:
                    type: string
                  error:
                    type: string
    EvalTestStep:
      type: object
      description: >-
        One authored test step (the unified test model). `kind` discriminates:
        `prompt` is a user message (model turn); `toolCall` is a deterministic,
        model-free tool call; `interact` is one pure widget action; `assert` is
        an assertion (a `Predicate` like `toolCalledWith` / `widgetRendered`, or
        a DOM `WidgetAssertion`).
      required:
        - id
        - kind
      properties:
        id:
          type: string
          minLength: 1
        kind:
          type: string
          enum:
            - prompt
            - toolCall
            - interact
            - assert
        prompt:
          type: string
          description: 'User message (`kind: prompt`).'
        serverName:
          type: string
          description: 'Server that owns the tool (`kind: toolCall`).'
        toolName:
          type: string
          description: 'Tool name (`kind: toolCall` / `interact`).'
        arguments:
          type: object
          description: 'Tool-call arguments (`kind: toolCall`).'
          additionalProperties: true
        action:
          type: object
          description: 'Widget action (`kind: interact`).'
          additionalProperties: true
        assertion:
          type: object
          description: 'Predicate or widget assertion (`kind: assert`).'
          additionalProperties: true
      additionalProperties: true
    Error:
      type: object
      required:
        - code
        - message
      properties:
        code:
          type: string
          description: >-
            Stable, machine-readable error code. New codes may be added over
            time; treat unknown codes as non-retryable failures unless the HTTP
            status says otherwise.
          enum:
            - UNAUTHORIZED
            - FORBIDDEN
            - NOT_FOUND
            - VALIDATION_ERROR
            - RATE_LIMITED
            - FEATURE_NOT_SUPPORTED
            - SERVER_UNREACHABLE
            - TIMEOUT
            - OAUTH_REQUIRED
            - INTERNAL_ERROR
        message:
          type: string
          description: >-
            Human-readable description. May change between releases — don't
            match on it.
        details:
          type: object
          description: Optional, unstructured context bag.
          additionalProperties: true
  responses:
    ValidationError:
      description: Malformed body or parameters.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: VALIDATION_ERROR
            message: Invalid JSON body
    Unauthorized:
      description: >-
        Missing, invalid, revoked, or orphaned key (`UNAUTHORIZED`) — or the
        **target MCP server** needs an OAuth grant (`OAUTH_REQUIRED`), which is
        a property of the server, not your key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          examples:
            badKey:
              summary: Invalid or revoked key
              value:
                code: UNAUTHORIZED
                message: Invalid API key
            oauthRequired:
              summary: Target server needs an OAuth grant
              value:
                code: OAUTH_REQUIRED
                message: Server requires OAuth authorization
    Forbidden:
      description: Key is valid but not allowed to do this.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: FORBIDDEN
            message: You do not have access to this project
    NotFound:
      description: Unknown project, server, or resource.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: NOT_FOUND
            message: Server not found
    RateLimited:
      description: >-
        Per-key rate limit exceeded (60 requests/minute sustained, bursts up to
        10). Honor `Retry-After` and back off with jitter.
      headers:
        Retry-After:
          description: Seconds to wait before retrying.
          schema:
            type: integer
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: RATE_LIMITED
            message: API key rate limit exceeded. Slow down and retry.
    InternalError:
      description: Something failed on MCPJam's side.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: INTERNAL_ERROR
            message: Unexpected internal error
    ServerUnreachable:
      description: Could not connect to the target MCP server.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: SERVER_UNREACHABLE
            message: Failed to connect to server
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        MCPJam API key (`sk_…`). Create one at [Settings → API
        keys](https://app.mcpjam.com/settings/api-keys). Guest sessions cannot
        use the API, and API keys cannot manage other API keys.

````