0
Phase 5

Working Memory

~3 min read

Concept & How It Works

    Why Does It Exist?

    The context window is the agent's RAM: finite, expensive, and shared across prompts, tools, and memory. Managing it is critical for long tasks.

    Real-World Analogy

    Working memory is your desk surface — only what fits there is immediately usable; everything else is in a filing cabinet (long-term memory).
    Loading diagram...

    Visual Workflows

    What is Working Memory?

    Loading diagram...

    Example

    Scenario

    During a 20-step coding task, the agent's working memory holds the current file contents, recent terminal output, and the last 5 user messages.

    Solution

    In Agent Memory, apply Working Memory to this scenario: During a 20-step coding task, the agent's working memory holds the current file contents, recent terminal output, and the last 5 user messages. Identify the inputs, run the technique, validate the output, and note one thing you would monitor in production.

    Practice Task

    Do this before moving to the next module — reading alone is not enough.

    Open the Code Walkthrough below and run it locally. Change one parameter related to Working Memory (e.g. model, temperature, top_k, or tool name), observe the difference in output, and write 2–3 sentences explaining what changed.

    Code Walkthrough

    Highlighted lines show where Working Memory happens in the code.

    Working Memory
    1# Working Memory — minimal example2from openai import OpenAI3
    4client = OpenAI()  # create API client5
    6# Ask the model to explain this topic7response = client.chat.completions.create(  # core API call for Working Memory8    model="gpt-4o-mini",9    messages=[10        {"role": "system", "content": "You explain working memory clearly."},11        {"role": "user", "content": f"What is working memory?"},12    ],13    temperature=0,14)15print(response.choices[0].message.content)  # show output for debugging

    Commands to Remember

    Commands to Remember

    • pip install chromadb # vector store for long-term memory
    • pip install redis # fast session / working memory
    • pip install tiktoken # count tokens before injecting memory

    Common Mistakes

    • Treating Working Memory as a black box without evaluation
    • Ignoring cost and latency in production
    • Skipping error handling for working memory

    Cheat Sheet

    Quick recap — the most important points from this module.

    Cheat Sheet

    quick ref
    • Working Memory
    • Context Window
    • Token Budget
    • Scratchpad