Inside memory
Trace memory processing, storage and the diagnostics behind remembered context.
In this topic
Memory is a separate continuity system beside conversation history. A transcript preserves what was said; memory extracts useful context and decides what to bring into a later request. Understanding the write and read paths helps explain why a saved chat may not immediately appear as a recalled fact.
For everyday controls, use Memory. This chapter follows the services that implement those controls.
Capture turns without pausing the conversation
MemoryService.bufferTurn accepts the user message, optional assistant message, agent identifier, conversation identifier, optional session date, and optional project identifier. It records a pending signal and rearms the conversation's debounce. It does not run the extraction model synchronously on every chat turn.
When the debounce expires, the session changes, or a caller explicitly flushes the session, the service can distill the buffered material. One structured extraction combines episode information, entities, identity changes, and candidate pinned facts. The configured memory model must be available; pending signals are not the same thing as successfully distilled memory.
Project membership is captured when the turn is buffered, while the live conversation still knows its project. Reconstructing that relationship later could lose the correct project association after navigation or a session reset.
Delegated work has separate session bookkeeping
bufferDelegatedTurn records work performed on an agent's behalf without replacing that agent's active direct-chat conversation. A delegated task arriving during a chat therefore does not force the direct session to flush simply by claiming to be the new active conversation.
When adding a new caller to the memory pipeline, choose the appropriate entry point. Reusing the direct-chat call for background delegation can change lifecycle behavior even when the stored text looks correct.
Distinguish the stored material
| Material | Purpose |
|---|---|
| Pending signals | Durable input awaiting extraction |
| Transcript | Conversation turns available for historical recall |
| Episodes | Distilled session context such as topics and decisions |
| Pinned facts | Useful individual facts retained across sessions |
| Identity information | Small stable context about the person or agent |
Distillation need not create a pinned fact from every conversation. A short exchange may warrant only an episode, and irrelevant content should not be promoted just to increase a memory count.
Retrieve only what the turn needs
The read path evaluates the incoming message, selects relevant memory material, and constructs a bounded context block. Identity, episode, pinned-fact, and transcript recall serve different requests. Asking for the exact wording of an earlier conversation is different from asking which project the person works on.
A token budget limits the material inserted into a request. Consequently, a fact can exist in storage without appearing in every answer. Diagnose retrieval selection and the final model request separately from storage. Increasing a budget is not a substitute for identifying why the correct record was not selected.
Explicit memory search is also available through the tool surface where enabled. Inspect the tool's current schema and results rather than assuming that a free-form request searches every agent's memory.
Consolidation is background maintenance
MemoryConsolidator performs maintenance outside the request path: reducing salience over time, combining similar episodes, promoting repeatedly supported facts, and pruning material according to retention settings. The scheduled path respects its configured interval and idle conditions; an on-demand action is a separate trigger.
Retention, salience, and relevance have different roles. Retention determines how long material is kept. Salience influences its importance. Retrieval decides whether it helps this request. A troubleshooting report should name which of these behaviors is unexpected.
Configuration fields for the memory pipeline
The current configuration model defines these defaults and normalization bounds. Read the saved setting for your installation before interpreting a run; defaults are not evidence of the effective value.
| Field | Default | Accepted normalized range | Meaning |
|---|---|---|---|
memoryBudgetTokens | 800 | 100–4000 | Maximum retrieved-memory context budget |
summaryDebounceSeconds | 60 | 10–3600 | Quiet interval before session distillation |
consolidationIntervalHours | 24 | 1–168 | Scheduled maintenance interval |
salienceFloor | 0.2 | 0–1 | Salience threshold used by maintenance |
episodeRetentionDays | 365 | 0–3650 | Retention configuration for historical material |
These fields live in MemoryConfiguration.swift. Changing a debounce value is not a way to repair an unavailable extraction model. Check readiness first, then tune timing to the intended workflow.
Search indexes and agent boundaries
MemorySearchService maintains agent-scoped search namespaces and separate entry points for pinned facts, episodes, and transcript. Removing an agent's index, purging its namespace storage, and rebuilding an index are distinct maintenance operations. A rebuild should restore searchable state from the intended underlying records; it should not broaden the caller's agent scope.
When retrieval returns nothing, check both the stored record and indexing state. An extraction failure and an indexing failure can produce the same visible empty search while requiring different recovery. Keep index-failure diagnostics with the affected agent and material type.
Attribute API conversations correctly
Clients can use X-Mellow-Agent-Id with an agent identifier from GET /agents. Keep a stable conversation identifier for turns that belong together. An agent header is attribution, not a way to override authentication or another agent's permissions.
The bulk ingestion endpoint accepts a conversation and turn list:
{
"agent_id": "AGENT_ID",
"conversation_id": "research-notes-01",
"turns": [
{"user": "The launch review is on Tuesday.", "assistant": "I will use Tuesday in this conversation."}
]
}
Send this to POST /memory/ingest only when intentionally adding that material. session_date supplies historical timing; skip_extraction requests transcript insertion without the extraction pass. Successful ingestion is not a guarantee that an unavailable extraction model produced a memory summary.
Trace missing memory in order
- Confirm memory is enabled and the turn is nonempty.
- Check agent and conversation attribution.
- Establish whether the caller reached
bufferTurnor the delegated equivalent. - Inspect pending signals and extraction-model readiness.
- Confirm a distilled record exists.
- Inspect retrieval selection and budget on a later request.
- Check whether the model used the supplied evidence accurately.
The service's buffering telemetry separates no caller activity from early exits, including disabled memory and empty messages. Preserve that distinction in diagnostics. Source landmarks include MemoryService.swift, MemorySearchService.swift, the memory database, and the consolidation service. No stage's source implementation alone proves recall quality for a released build.