• itscybernews
  • Posts
  • A trending tool stops AI agents drowning in their own homework. Here is the cool part, and the risky part.

A trending tool stops AI agents drowning in their own homework. Here is the cool part, and the risky part.

Context Mode keeps bulky tool output out of an agent's memory. What it does, what it risks, and how to try it safely.

Picture a chef whose counter is buried. Every time they look up a recipe, someone dumps the entire cookbook on the bench, and by lunchtime there is no room left to chop an onion.

That is what happens inside an AI coding agent all day. Every file it reads, every web page it fetches and every log it opens gets piled into its “context window”, the finite short-term memory it thinks with. A trending open-source project called Context Mode wants to clear the counter.

The problem: agents drown in their own homework

Ask an agent to check twenty GitHub issues and the raw text of all twenty lands in its memory, whether it needs it or not. Fill the window and the agent gets slower, pricier and forgetful.

Context Mode, by a developer going by mksglu on GitHub, is an MCP server (the plug-in standard many agents use to reach outside tools) that sits between the agent and its tools. Its idea is simple: keep the bulky stuff out of the conversation and let the agent ask for just the bit it needs.

How it works, in plain English

According to the project’s README, there are three tricks:

Trick

What it does

Sandboxed output

Commands run in a separate process and only the short printed result goes back to the agent. The README’s example: 315 KB becomes 5.4 KB.

A searchable notebook

Full results are stored in a local SQLite database on your machine and found again with keyword search (BM25), so no extra AI calls are needed.

Memory that survives a reset

Edits, git actions, tasks and errors are logged, so when the conversation is compacted the agent can look up what it was doing.

Instead of reading 47 files into memory, the agent writes a tiny script that does the digging and hands back one line. The README’s own numbers for that case: about 700 KB down to 3.6 KB.

Why people are excited

The project lists support for more than a dozen agents, including Claude Code, Codex CLI, Cursor and Gemini CLI. In Claude Code, the README gives a two-line install through the plugin marketplace, plus a ctx-doctor command to check it is healthy.

The reported savings are eye-catching. A review on andrew.ooo repeats the project’s benchmark figures: a 56.2 KB browser snapshot shrinks to 299 bytes, and twenty GitHub issues go from 58.9 KB to 1.1 KB. Those are the author’s own tests, though, and he says plainly that he has not yet run formal benchmarks on answer quality. Fewer bytes is not the same as a smarter agent.

One quick word from today’s sponsor

Leave Granola and get up to 12 months free of Wispr Flow Notetaker + Dictation

If you have paid time left on an individual Granola plan, we'll match it with a Wispr Flow subscription that includes Notetaker and dictation, and add bonus time, up to 12 months total. Sign in or create a Wispr account and submit proof of your plan to check eligibility.

What can go wrong

Anything that sits between your agent and your tools is powerful, so it deserves a hard look. Five things worth knowing:

  • A notebook of everything your agent saw. Indexed results live in plain databases under ~/.context-mode/ on your machine. If a tool ever printed a token or a customer record, it may now be stored there. OWASP’s MCP Top 10 lists this family of risk as Token Mismanagement & Secret Exposure and Context Injection & Over-Sharing.

  • It is a lighter lock than it sounds. The README says each run uses its own process, but the andrew.ooo reviewer notes the code it runs still inherits the process’s file access. It is not an operating-system-level cage.

  • It lets agents use your logged-in tools. The README says command-line tools such as gh, aws, gcloud, kubectl and docker inherit your environment and config. Handy, and also exactly what a bad instruction would love to borrow.

  • Rough edges are real. The same review reports Windows hangs and stats that overstate savings, and notes the licence moved from MIT to Elastic License 2.0, which is source-available rather than open source in the strict sense.

  • Plug-ins are a supply chain. OWASP also lists Software Supply Chain Attacks and Tool Poisoning for MCP servers. Any add-on you install can change what your agent does.

How to try it without regrets

  1. Start in a throwaway project. Do not point it at a repo full of production secrets on day one.

  2. Look at what it stores. Peek inside ~/.context-mode/. The README says ctx_purge permanently deletes indexed content, and databases older than 14 days are cleaned up on startup.

  3. Keep secrets out of tool output. Do not let your agent print keys. A notebook cannot leak what it never wrote down.

  4. Give the agent only the logins it needs. Use read-only or short-lived credentials for cloud tools.

  5. Read the permissions. The review says it enforces rules from your agent’s settings file. Make those rules strict and test them.

  6. Check your own numbers. Compare token use with and without it on a task you know well before believing any headline percentage.

The takeaway

Context Mode is a neat reminder that a smarter agent is sometimes just a tidier one. Less clutter in memory means more room to think, and that idea will outlive this one project.

If you only remember one thing: a tool that remembers everything your agent saw is only as safe as the secrets you kept out of it.

Facts here come from the Context Mode GitHub README, a review on andrew.ooo, and the OWASP MCP Top 10. Context Mode is a fast-moving project, so details may change.Picture a chef whose counter is buried. Every time they look up a recipe, someone dumps the entire cookbook on the bench, and by lunchtime there is no room left to chop an onion.