Skip to content
Models, voice and media

Apple’s on-device model

Check Foundation availability, choose suitable tasks and handle its context limits.

In this topic

Mellow exposes Apple's system language model as Foundation when macOS reports it available. It is an option for short writing, transformation and summarization tasks without configuring a separate hosted provider.

Foundation's availability is controlled by macOS. Mellow can run on a system where this model is unavailable, so app installation and Foundation readiness are separate checks.

Requirements

Mellow's Foundation integration requires macOS 26 or later and an eligible Mac with Apple Intelligence enabled. System model assets must finish preparing. The integration also checks that the Foundation Models framework is present in the build.

The source distinguishes an ineligible device, disabled Apple Intelligence, assets not ready, an older OS and a missing framework. Use the reported reason rather than repeatedly downloading an unrelated MLX model.

Get started

  1. Check the Mac's version in About This Mac.
  2. Open System Settings → Apple Intelligence & Siri and enable Apple Intelligence if available.
  3. Wait for macOS to prepare its assets.
  4. Open Mellow's model picker and select Foundation when it is available.
  5. Test a short request using text you supply.

Try “Rewrite this note as three clear action items” with a small paragraph. Inspect the output before moving to a longer conversation or enabling tools.

Choosing suitable work

Foundation has a more constrained context budget than many larger models. System instructions, tool descriptions, prior messages and the requested answer all need room. Start a shorter conversation or reduce the input if a request exceeds the budget.

Foundation is text-only in this integration. An image-input request is rejected as unsupported; select a vision-capable model to work with pictures.

For tasks requiring more context or a different reasoning capability, select a downloaded model or configured provider. Do not assume an unavailable Foundation request will automatically select the fallback you intended.

Mellow's background helper

Mellow can use a Core Model for supporting operations such as text cleanup and memory-related work. Inspect that selection in General. It can differ from the model chosen for a conversation.

When checking privacy or failures, follow the actual operation: voice recognition, cleanup, chat generation and tool execution can take different paths. Selecting Foundation for one of them does not configure all the others.

Model name and availability

API clients should discover available models and request the returned ID. Foundation's service ID is foundation. Use the base address and access policy shown by your running Mellow server; do not assume a screenshot's port matches your installation.

For example, after replacing the local address and supplying any required authorization:

curl --fail http://127.0.0.1:1337/v1/models

Only request foundation when discovery and system availability support it. Do not substitute a guessed downloadable model ID when it is absent.

Basic chat

A small request body for the Chat Completions route is:

{
  "model": "foundation",
  "messages": [
    {"role": "user", "content": "Turn this note into a checklist: review draft, confirm dates, send final copy."}
  ],
  "max_tokens": 160,
  "stream": false
}

Send it to the configured server's /v1/chat/completions endpoint. See HTTP API for authentication and streaming details.

Function / tool calling

The integration adapts supported tool requests to the system framework. Test the exact tool schema your client uses. Tool support does not grant permission to execute that tool against files or another application; the agent and runtime still determine that access.

Context size

Mellow reads the framework's context size where available. Treat that value as a request budget, not a promise that an entire long document will fit. Keep examples short and handle context errors explicitly.

Privacy

Foundation inference is on-device. Connected tools can still contact a network service, and subsequent operations can use other configured models. Review the whole task when you need an offline workflow.

Troubleshooting

Reported stateAction
OS too oldUse another supported model or update macOS
Device not eligibleChoose a downloaded model or provider compatible with the Mac
Apple Intelligence disabledEnable it in System Settings if desired
Model not readyLet macOS finish preparing assets, then retry
Framework missingCheck the installed build
Context too largeReduce input, start a shorter conversation or choose a larger-context model

For client applications, make unavailability a visible state and let the user choose a replacement. A fallback that silently sends the request to a hosted model changes where the content is processed.

Continue exploring · Models, voice and mediaHosted inference →Distinguish model-service readiness and billing from Cloud workspace sign-in.