Apple’s on-device model
Check Foundation availability, choose suitable tasks and handle its context limits.
In this topic
Mellow exposes Apple's system language model as Foundation when macOS reports it available. It is an option for short writing, transformation and summarization tasks without configuring a separate hosted provider.
Foundation's availability is controlled by macOS. Mellow can run on a system where this model is unavailable, so app installation and Foundation readiness are separate checks.
Requirements
Mellow's Foundation integration requires macOS 26 or later and an eligible Mac with Apple Intelligence enabled. System model assets must finish preparing. The integration also checks that the Foundation Models framework is present in the build.
The source distinguishes an ineligible device, disabled Apple Intelligence, assets not ready, an older OS and a missing framework. Use the reported reason rather than repeatedly downloading an unrelated MLX model.
Get started
- Check the Mac's version in About This Mac.
- Open System Settings → Apple Intelligence & Siri and enable Apple Intelligence if available.
- Wait for macOS to prepare its assets.
- Open Mellow's model picker and select Foundation when it is available.
- Test a short request using text you supply.
Try “Rewrite this note as three clear action items” with a small paragraph. Inspect the output before moving to a longer conversation or enabling tools.
Choosing suitable work
Foundation has a more constrained context budget than many larger models. System instructions, tool descriptions, prior messages and the requested answer all need room. Start a shorter conversation or reduce the input if a request exceeds the budget.
Foundation is text-only in this integration. An image-input request is rejected as unsupported; select a vision-capable model to work with pictures.
For tasks requiring more context or a different reasoning capability, select a downloaded model or configured provider. Do not assume an unavailable Foundation request will automatically select the fallback you intended.
Mellow's background helper
Mellow can use a Core Model for supporting operations such as text cleanup and memory-related work. Inspect that selection in General. It can differ from the model chosen for a conversation.
When checking privacy or failures, follow the actual operation: voice recognition, cleanup, chat generation and tool execution can take different paths. Selecting Foundation for one of them does not configure all the others.
Model name and availability
API clients should discover available models and request the returned ID. Foundation's service ID is foundation. Use the base address and access policy shown by your running Mellow server; do not assume a screenshot's port matches your installation.
For example, after replacing the local address and supplying any required authorization:
curl --fail http://127.0.0.1:1337/v1/models
Only request foundation when discovery and system availability support it. Do not substitute a guessed downloadable model ID when it is absent.
Basic chat
A small request body for the Chat Completions route is:
{
"model": "foundation",
"messages": [
{"role": "user", "content": "Turn this note into a checklist: review draft, confirm dates, send final copy."}
],
"max_tokens": 160,
"stream": false
}
Send it to the configured server's /v1/chat/completions endpoint. See HTTP API for authentication and streaming details.
Function / tool calling
The integration adapts supported tool requests to the system framework. Test the exact tool schema your client uses. Tool support does not grant permission to execute that tool against files or another application; the agent and runtime still determine that access.
Context size
Mellow reads the framework's context size where available. Treat that value as a request budget, not a promise that an entire long document will fit. Keep examples short and handle context errors explicitly.
Privacy
Foundation inference is on-device. Connected tools can still contact a network service, and subsequent operations can use other configured models. Review the whole task when you need an offline workflow.
Troubleshooting
| Reported state | Action |
|---|---|
| OS too old | Use another supported model or update macOS |
| Device not eligible | Choose a downloaded model or provider compatible with the Mac |
| Apple Intelligence disabled | Enable it in System Settings if desired |
| Model not ready | Let macOS finish preparing assets, then retry |
| Framework missing | Check the installed build |
| Context too large | Reduce input, start a shorter conversation or choose a larger-context model |
For client applications, make unavailability a visible state and let the user choose a replacement. A fallback that silently sends the request to a hosted model changes where the content is processed.
Continue exploring · Models, voice and mediaHosted inference →Distinguish model-service readiness and billing from Cloud workspace sign-in.