Memory Compression
~2 min read
Concept & How It Works
Why Does It Exist?
Uncompressed memory hits context limits and increases cost. Compression is mandatory for agents running 50+ step tasks.
Real-World Analogy
Memory compression is packing for a trip — bring essentials, leave the 'just in case' items that never get used.
Visual Workflows
What is Memory Compression?
Example
Scenario
A 50KB API response is compressed to 'Endpoint returned 200 with 3 users: Alice, Bob, Carol (IDs 1,2,3)' before entering context.
Solution
In Agent Memory, apply Memory Compression to this scenario: A 50KB API response is compressed to 'Endpoint returned 200 with 3 users: Alice, Bob, Carol (IDs 1,2,3)' before entering context. Identify the inputs, run the technique, validate the output, and note one thing you would monitor in production.
Practice Task
Do this before moving to the next module — reading alone is not enough.
Open the Code Walkthrough below and run it locally. Change one parameter related to Memory Compression (e.g. model, temperature, top_k, or tool name), observe the difference in output, and write 2–3 sentences explaining what changed.
Code Walkthrough
Highlighted lines show where Memory Compression happens in the code.
1# Memory Compression — minimal example2from openai import OpenAI3
4client = OpenAI() # create API client5
6# Ask the model to explain this topic7response = client.chat.completions.create( # core API call for Memory Compression8 model="gpt-4o-mini",9 messages=[10 {"role": "system", "content": "You explain memory compression clearly."},11 {"role": "user", "content": f"What is memory compression?"},12 ],13 temperature=0,14)15print(response.choices[0].message.content) # show output for debuggingCommands to Remember
Commands to Remember
pip install chromadb # vector store for long-term memorypip install redis # fast session / working memorypip install tiktoken # count tokens before injecting memory
Common Mistakes
- Treating Memory Compression as a black box without evaluation
- Ignoring cost and latency in production
- Skipping error handling for memory compression
Cheat Sheet
Quick recap — the most important points from this module.
Cheat Sheet
quick ref- •Memory Compression
- •Summarization
- •Token Budget
- •Distillation