HTTP API
Discover models, call inference endpoints or choose the agent execution lifecycle.
In this topic
- 01Your client→
- 02Model endpoint or agent endpoint→
- 03Configured model and access rules→
- 04Response to your client
Raw model endpoints return tool calls for the client to execute. Authorized agent execution follows its own tool policies.
Mellow's local HTTP server lets an application request model output, call exposed tools, and work with Mellow agents. Choose the contract that matches your application: a compatible chat request supplies model input; an agent run delegates execution to the configured agent; a detached task returns an identifier that you monitor separately.
The routes below are implemented in the current server source. Their presence does not mean every backing model, media service, or workspace is configured on a particular Mac.
Establish the server address
The usual local base is http://127.0.0.1:1337. Confirm the address displayed by your running app, or use mellow status and mellow doctor. A configured port or MELLOW_PORT in the CLI environment can change the port used by your tools.
curl -sS http://127.0.0.1:1337/health
curl -sS http://127.0.0.1:1337/v1/models
A successful health response verifies reachability. Listing models identifies requestable model identifiers. Send a small request to an actual returned identifier before adding streaming, media, or tools.
Choose a compatible protocol
| Protocol | Request route | Typical client |
|---|---|---|
| Chat Completions | POST /v1/chat/completions | OpenAI-compatible SDK |
| Text completions | POST /v1/completions | Completion or fill-in-middle client |
| Responses | POST /v1/responses | Responses-compatible client |
| Messages | POST /v1/messages | Anthropic-compatible SDK |
| Ollama chat | POST /api/chat | Ollama-compatible client |
| Ollama generation | POST /api/generate | Ollama generation client |
Compatibility applies to the implemented request/response surface. Do not assume every feature of a hosted vendor API is available. Start with supported text input and expand one feature at a time. For Messages, use the registered /v1/messages route; do not invent an additional provider prefix.
Make one chat request
curl -sS http://127.0.0.1:1337/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "MODEL_ID",
"messages": [
{"role": "system", "content": "Explain concepts in clear, short paragraphs."},
{"role": "user", "content": "What is a local agent workspace?"}
],
"stream": false
}'
Replace MODEL_ID with an identifier from discovery. This loopback example assumes the app permits the request without an access key; add authentication when required by your server configuration.
| Field | Purpose |
|---|---|
model | Exact model identifier |
messages | Ordered conversation messages |
stream | Request streaming instead of a single JSON response |
max_tokens | Explicit output-token limit |
temperature, top_p | Optional sampling overrides |
tools | Function schemas offered by the client |
tool_choice | Tool selection instruction supported by the protocol |
session_id | Conversation/session bookkeeping identifier |
Omit sampler overrides when you intend to use the effective model and app configuration. Do not fill in arbitrary values from an unrelated model example. A stable session identifier does not automatically preserve the client's message history or force KV-cache reuse.
For a non-streamed Chat Completions response, inspect choices[].message, finish_reason, and available usage. A length-limited response is different from a naturally completed one. Handle an empty or missing content field when the message instead contains tool calls.
Read a stream correctly
Chat Completions uses server-sent events. Accumulate content deltas and tool-call deltas in order. A tool's JSON arguments may span several events; parse them only after the call is complete. Finish the response on the protocol's completion event rather than on the first text fragment.
data: {"choices":[{"index":0,"delta":{"content":"A local"},"finish_reason":null}]}
data: {"choices":[{"index":0,"delta":{"content":" workspace…"},"finish_reason":null}]}
data: [DONE]
This is a schematic excerpt, not a complete server response. Preserve protocol-specific event handling when using Responses or Messages instead of assuming every stream has the same shape. Close the request on cancellation and verify that the server settles its active work.
Decide who executes tools
A client supplying function definitions through a compatible chat API generally owns execution of the returned calls. Validate arguments, obtain any required user approval, execute the function, and append the assistant call and matching tool result before requesting continuation.
For Mellow to own the autonomous loop, use the selected agent's run/dispatch surface. Mellow then resolves that agent's permitted tools, policy, iteration budget, and approvals. Tool availability is scoped; it is not inherited from a client having discovered a tool name elsewhere.
Run an existing agent
| Request | Lifecycle |
|---|---|
GET /agents | Discover custom agents and their metadata |
POST /agents/{id}/run | Execute an agent loop within the request lifecycle |
POST /agents/{id}/dispatch | Start a detached task and return task metadata |
GET /tasks/{task_id} | Inspect a detached task |
DELETE /tasks/{task_id} | Request cancellation |
POST /tasks/{task_id}/clarify | Supply an answer to a pending clarification |
Use an identifier returned by agent discovery. A detached request accepts a prompt and optional title:
{
"prompt": "Read the approved project notes and summarize the open decisions.",
"title": "Open decisions"
}
An accepted dispatch returns an identifier and poll location. Follow the returned poll URL and inspect status until a terminal result or an explicit clarification state. Acceptance is not task completion. Respect concurrency-limit responses and avoid retrying a dispatch blindly if the first outcome is unknown.
A clarification request supplies {"response":"your answer"}. Remote run and dispatch paths require their applicable secure-channel and scope checks. A plain bearer key should not be assumed to unlock owner-level host execution.
Discover and call externally exposed tools
| Request | Purpose |
|---|---|
GET /mcp/health | Probe MCP HTTP availability |
GET /mcp/tools | List enabled tools allowed for external callers |
POST /mcp/call | Invoke an exposed tool |
The HTTP invocation body names the tool and supplies an argument object:
{
"name": "TOOL_NAME_FROM_DISCOVERY",
"arguments": {}
}
Replace both the name and arguments using the discovered schema. The server rejects app-only tools before execution even if a caller supplies their name directly. A tool_not_exposable response is an exposure-policy decision, not a prompt to retry under another name.
Use the discovered schema for arguments. The external list can be smaller than the app's internal registry because enabled state and exposure rules apply. The number of tools is not the number of cloud agents or models.
Inspect the returned tool envelope even when HTTP transport succeeded. For clients using MCP stdio, mellow mcp provides the supported bridge; see CLI.
Work with models and media
| Request | Purpose |
|---|---|
GET /v1/models | OpenAI-style model discovery |
GET /v1/tags | Ollama-style model list |
POST /api/show | Model metadata and capabilities |
POST /v1/embeddings | Embedding request |
POST /api/embed | Ollama-format embedding request |
POST /v1/audio/transcriptions | Speech transcription |
GET /v1/images/models | Installed local image-model capabilities |
POST /v1/images/generations | Image generation |
POST /v1/images/edits | Image editing with a supported model |
POST /v1/images/upscale | Image upscaling with a supported model |
POST /v1/images/cancel | Cancel an image job |
POST /v1/videos/quote | Request a cloud video quote |
POST /v1/videos/generations | Start a quoted video job |
GET /v1/videos/jobs/{id} | Inspect video-job state |
GET /v1/videos/jobs/{id}/content | Retrieve completed media |
Each media route has prerequisites beyond server health: appropriate models, provider access, supported input, and, for quoted cloud jobs, the relevant service and consent. Follow Image generation and Voice for setup before writing a client.
Attribute and ingest memory
Use X-Mellow-Agent-Id to attribute supported requests to a discovered agent. Attribution does not bypass authentication. POST /memory/ingest accepts agent_id, conversation_id, and turns, where each turn contains user and assistant text. Optional session_date supplies historical timing; skip_extraction retains transcript input without distillation.
The memory pipeline can require a configured extraction model. An ingested turn and a recalled fact are separate outcomes. See Memory internals for lifecycle and diagnosis.
Administrative and identity routes
Configuration is managed through loopback-only /admin/config/export, /schema, /plan, and /apply operations under that prefix. Runtime diagnostics include /admin/cache-stats, /admin/generation-settings, and /admin/runtime-settings. Use Configuration for the supported document workflow.
Pairing and secure-channel routes include /pair/challenge, /pair, /pair/hello, /pair/code, /pair/unpair, /pair-invite, /secure/session, and /secure/call. Use the supported pairing client and protocol rather than manufacturing requests from a copied access link. Public reachability is not public authorization.
GET /credits/balance reports the configured cloud-routing service's balance when that service is available. It is not a local-model prerequisite and should not be used to diagnose a missing local agent.
Authentication and failure handling
Mellow-issued access keys use sk-mellow and carry scope and lifecycle information. Supply them using the supported bearer-authentication mechanism. Keep keys in a secret store or process environment, not source code, screenshots, or URLs.
Different routes have different requirements: local administrative trust, agent audience, owner scope, pairing state, or secure transport. A valid credential can still be inappropriate for an operation. Read the response body before deciding whether a failure is authentication, authorization, validation, capacity, or execution.
| Failure class | Recovery |
|---|---|
| Connection refused | Confirm app, server, address, and port |
| Authentication rejected | Check key validity, audience, expiry, and revocation |
| Permission denied | Check route scope and required transport |
| Invalid request | Correct fields using the current schema |
| Capacity limit | Wait or reduce concurrent work; avoid duplicate dispatch |
| Model unavailable | Resolve bundle/provider readiness |
| Partial stream or timeout | Determine whether work is still running before retrying |
Browser clients also need an allowed origin and must not expose a reusable host credential in publicly served JavaScript. A same-origin application backend is often the appropriate place to hold that credential.
Implementation reference: Packages/MellowCore/Networking/HTTPHandler.swift. Client recipes show minimal integrations without implying that every vendor SDK operation is supported.