Inside inference
Understand model resolution, loading, streaming and request failures.
In this topic
Mellow's local runtime translates a chat or API request into a model load, a generation plan, a stream of output, and cleanup. Model selection, concurrency, and cache controls affect different stages. This chapter gives developers a way to inspect those stages without treating every model family as interchangeable.
Follow the lifecycle
- Resolve the requested model bundle and its metadata.
- Combine model configuration with explicit request and user settings.
- Check the requested configuration and memory constraints before unsafe allocation.
- Acquire the model engine and retain a lease while generation is active.
- Tokenize the prompt, reuse compatible cached state, and prefill remaining input.
- Decode text and structured model output into the caller's stream.
- Finish or cancel, release the lease, and apply idle residency policy.
A downloaded model, a loaded engine, and a completed answer are three different milestones. Tool calling, vision, long context, and repeated turns need their own execution checks.
Generation defaults belong to the effective plan
Explicit request or agent choices can override applicable user defaults and bundle generation configuration. Leaving an override unset allows the model's own configuration to contribute. Clients should avoid populating every sampler field with arbitrary values simply because their SDK accepts them.
Record the effective temperature, top-p, top-k, min-p, output limit, and repetition settings when comparing runs. Greedy decoding and sampling are different modes; a changed sampler can invalidate a performance or quality comparison even when the model filename is unchanged.
Reasoning and tool boundaries depend on the model bundle, tokenizer, chat template, and runtime. Do not conceal a parser defect by adding forced markers or stripping suspicious output after generation. Capture the raw event and the rendered conversation so the failing boundary can be located.
Concurrency has several owners
| Component | Responsibility |
|---|---|
| Batch engine | Schedule compatible requests within a model engine |
| Model registry | Coalesce engine creation for the same model |
| Model lease | Prevent unloading a model while a request still uses it |
| Residency manager | Decide when an idle model should unload |
| GPU coordination | Coordinate producers that share Metal resources |
| Plugin host limits | Bound inference requests initiated by a plugin |
Continuous batching concerns compatible work inside an engine. It does not make every pair of models fit in memory, and it does not grant unlimited tool or plugin concurrency. Inspect queueing and active slots when a request appears to wait.
Cache reuse is about compatible state
Prefix reuse follows the effective prompt and model state, not a magic conversation identifier. Keep session_id stable for conversation bookkeeping, but do not treat it as a command to reuse incompatible cache contents. An informational prefix hash can help detect a prompt change; sending the hash back is not cache control.
Changes to system instructions, tool schemas, media, or compacted conversation context can change reuse. Context compaction may replace older outbound turns with a summary while leaving the visible transcript intact. The next generation therefore needs a cache identity consistent with the new prompt.
Different model architectures retain different state. Full-attention KV, hybrid recurrent/SSM state, and other pooling or sliding-window structures cannot be validated with the same cache metric. Disk cache availability also depends on its directory, budget, and model support.
Use runtime evidence
The local server provides administrative diagnostics including:
GET /admin/cache-stats
GET /admin/generation-settings
GET /admin/runtime-settings
Use these from the local machine under the app's administrative access rules. Compare effective settings and cache counters before and after a real request. A counter increase is useful evidence, but does not establish a coherent answer or correct tool execution by itself.
For performance, retain first-load time, time to first token, prefill rate, decode rate, peak physical memory, and the selected bundle. Separate a cold request from a warm repeat. Use a second turn that genuinely depends on the first when assessing continuity.
Cancellation and unloading
Test cancellation during initial load, prefill, generation, and tool continuation. The UI should settle, input should unlock, and resources should stop growing. Closing a stream should not leave an unowned model load running in the background.
Idle unloading begins after active leases are released. It is separate from deleting downloaded weights or clearing persistent cache. A model disappearing from resident memory is not evidence that its files were removed.
Troubleshooting matrix
| Symptom | Inspect |
|---|---|
| Load rejected | Bundle support, effective memory limit, and requested context |
| Slow first answer | Download completeness, load, prefill, and cold cache |
| Slow later turns | Changed prefix, disabled disk tier, incompatible cache state |
| Queue never settles | Active requests, cancellation, lease ownership |
| Raw tool or reasoning markers | Template/parser contract and full stream |
| Answer changes after cache restore | Baseline without reuse and architecture-specific state restoration |
Support should be stated per tested model and scenario. A runtime implementation or isolated unit test is not a blanket compatibility claim.
Continue exploring · Build with MellowSandbox execution →Follow runtime preparation, environment boundaries and agent code execution.