Skip to content
Models, voice and media

Hosted inference

Distinguish model-service readiness and billing from Cloud workspace sign-in.

In this topic

Mellow can use a hosted inference service when that service is configured and available. This is separate from signing in to a Cloud workspace, which supplies workspace-scoped agents and tasks. A successful workspace sign-in does not establish a hosted model catalog or a credit balance.

Use this distinction when a Cloud workspace is connected but the model browser is empty. The workspace may be ready for agent discovery while inference needs a different configured service.

Choose the connection you need

Your goalConfiguration to inspect
Run a downloaded model on this MacLocal Models
Use your own hosted model accountProviders
Discover published agents in an organizationMellow Cloud, then the authenticated Cloud workspace
Use managed hosted inferenceIts model service, identity and account availability
Run work on another paired MacDevices, then that host's agent and provider

For direct provider accounts, follow Provider connections. For workspace tasks, use Cloud workspaces.

Get started

First inspect the service status in the installed app. The current Router configuration is disabled unless explicitly enabled. Its presence in source code or a settings route is not evidence that a public service has been provisioned.

When a deployment supplies hosted inference, enable it through its account controls, complete any required identity setup, and refresh the model catalog. Select an advertised model and send a small request. Confirm that it produces an answer before relying on it for a longer task.

Do not purchase credits as a remedy for a missing service endpoint or failed model discovery. Resolve availability and account errors first.

Finding models

A usable catalog should identify the model and expose the controls it supports. Choose a model from that catalog rather than typing an example name from a screenshot. Provider model names and availability can change independently of the Mellow app version.

An empty catalog can mean discovery failed, the service is disabled, or your account has no eligible models. It is not sufficient evidence that you have exhausted a balance. Keep the error visible and distinguish retrying discovery from signing in again.

Credits

Where a configured service supports metered inference, review its current balance, prices and billing scope in that service's account controls. This reference does not promise welcome credits, a conversion rate or an included subscription.

A workspace role, provider subscription and hosted-inference balance are different things. Check which account would be charged before starting a large request. Hosted image and video operations can have separate quote and approval steps; see Images and video.

Checking your balance in chat

If the installed build displays a balance or usage control, treat it as information for the service named there. A cached balance may lag a recent operation. Refresh through the service's controls before drawing conclusions about a failed request.

Workspaces and shared credits

Use the account and billing information returned by the configured workspace. Do not assume that a published workspace agent shares a personal inference balance. Workspace permissions govern which actions and agents are available; billing depends on that deployment's contract.

Turning it off

Disabling managed inference removes that route from use; local models and independently configured providers remain separate choices. If a saved conversation selected an unavailable hosted model, select another ready model before retrying.

Your privacy

Hosted inference sends the request content required by the selected model to its service. Review both the service policy and any tools enabled for the conversation. Local storage of chat history does not mean remote inference received no content.

Billing diagnostics and conversation content are different records. Review an export before sharing it, and avoid treating the existence of a local billing ledger as proof of a remote service's retention behavior.

Troubleshooting

StateMeaning to investigateNext action
Router is not configuredBuild/deployment does not supply usable managed inferenceUse a ready local model or your configured provider
Workspace connected, no modelsWorkspace connection and inference are separateCheck the provider/model route you selected
Model catalog failsDiscovery, service or account issueRead the actual error and retry discovery after correcting it
Request reports insufficient balanceThe selected service rejected metered workReview that account's balance and billing scope
Model disappearedCatalog or permissions changedRefresh and select an advertised model
Request finished without an answerGeneration or rendering failedInspect the run error before retrying a billable task

Under the hood

The Router integration has its own enabled state and identity-signed requests. Keyless loopback spending is a separate opt-in and defaults off. A local process should not be given billing authority merely because it can reach the local server.

Release configuration fixes the production service address; debug configuration supports controlled test overrides. These implementation choices do not certify deployment availability. Validate the actual account, catalog and a small request in the environment being used.

Continue exploring · Models, voice and mediaSpeaking and dictation →Configure recognition, chat input, text insertion, wake phrases and stop behavior.