12k
All articles

How to Cut Token Usage in AI Coding Agents

Cut token usage in AI coding agents by narrowing file access, writing concise project instructions, planning before coding, searching manually, and clearing sessions.

OpenReplay Team
OpenReplay Team
How to Cut Token Usage in AI Coding Agents

To cut token usage in an AI coding agent, reduce what goes into its context. Name the files it should read, keep a short project instructions file, ask for a plan before code, run searches yourself, restart sessions often, and turn off tools the task doesn’t need.

Most people come to this after a session hit its usage limit halfway through a refactor, or after an invoice that doubled while the work stayed the same size.

These habits work the same way in Claude Code, Codex, Cursor and similar agents. Each section below explains why the habit works, gives a short example you can copy, and says whether the evidence is a measurement or reasoned practice.

Key Takeaways

  • Every file an agent reads and every earlier turn it keeps stays in its context. The cheapest token is one the agent never sees.
  • In a study of 124 pull requests run with OpenAI Codex, adding an AGENTS.md file came with a 16.58% drop in median output tokens and a 28.64% drop in median runtime. Median total tokens barely changed (about 1% higher).
  • A Spotify Engineering post measured about 90% mean savings on bulk file reads in a Java monorepo by giving Claude a cheaper model’s summary instead of the files. That figure counts Claude’s context only.
  • Anthropic’s Claude Code docs say that when spend on an API or cloud-provider plan runs higher than expected, the cause is usually a long session nobody cleared or Opus left as the default model.

Why Does Reducing Token Usage in AI Coding Agents Matter?

Tokens cost money twice: once on your bill or usage cap, and again in the quality of the output. Every file an AI coding agent reads and every earlier turn it keeps stays in its context, so the cheapest token is the one the agent never has to see. A smaller, more relevant context also gives the model less irrelevant material to get distracted by. A lean session is usually cheaper and more accurate than a bloated one.

For example, a debugging session that starts with “look through the repo” might pull in dozens of unrelated files before the agent reaches the actual bug. You pay for every one of them, and the model has to reason past all of them.

Narrow What the Agent Can See

To cut an AI coding agent’s exploratory reads, tell it which files to read instead of pointing it at the whole repository. Without that, the agent has to hunt for them, and every file it opens along the way ends up in context. Once you name the files, the agent’s reading is limited to what the task actually needs.

# Before
Why is the checkout total wrong? Look through the repo.

# After
The total in src/cart/total.ts is wrong when a discount code is applied.
Read src/cart/total.ts and src/cart/discounts.ts only. Do not open other files
without asking.

You can also limit what the agent can reach in the first place. If it runs on a dedicated machine, as in a remote box setup for agentic coding, clone only the repositories that task needs.

Keep a Project Instructions File

A project instructions file, usually named AGENTS.md or CLAUDE.md, saves an AI coding agent from working out your conventions again in every session. Most agents support one. Check your agent’s docs for the filename it reads. Claude Code reads CLAUDE.md, and its memory documentation explains that from v2.1.277 it also reads AGENTS.md directly when a repository has no CLAUDE.md.

The evidence here is specific. In a study of 124 pull requests across 10 repositories using OpenAI Codex, runs with an AGENTS.md file had a 28.64% lower median runtime and a 16.58% lower median output token count. The same results table shows median total tokens were roughly unchanged, and in fact about 1% higher with the file. In practice, the file mostly makes the agent write less and finish sooner. It does not shrink everything it reads. The study also has limits: it used one agent (Codex running gpt-5.2-codex), covered only small merged PRs (at most 100 changed lines and five files each), and did not fully evaluate whether the output was correct.

Keep the file short. Claude Code pulls CLAUDE.md into context when each session starts, so you pay for every line in every conversation. Anthropic suggests staying under 200 lines, and notes that Claude sticks to short, specific instructions more reliably than long or vague ones.

# AGENTS.md

## Project
Web API for order processing. Entry point: src/server.ts.

## Structure
- src/routes/: HTTP handlers, one file per resource
- src/services/: business logic; handlers never touch the DB directly
- tests/: mirrors src/; test files end in .test.ts

## Conventions
- Run `npm test` before proposing changes; run `npm run lint` on touched files
- Use the logger in src/lib/log.ts, never console.log
- Do not edit generated files in src/generated/

Ask for a Plan Before Code

Asking an AI coding agent for a plan before it writes code lets you catch a wrong direction before it has read and rewritten files. Throwing away a plan costs almost nothing. A wrong implementation has already spent tokens on every file it touched before you noticed. The plan also gives you a list of files you can cut down before any work starts.

Before writing any code: list the files you intend to read and change, and
the steps you will take, in under 10 bullets. Wait for my approval.

If the plan includes a file that isn’t relevant, remove it from the list before you approve.

Run Searches Yourself and Hand Over the Results

When you run grep yourself and paste the matching lines, the agent only pays for those lines. Asking the agent to find the same thing can cost it every file it opens along the way. Tools like git grep and git diff are fast and cost no tokens to run, so let them do the searching and give the model only the results.

# Find every call site yourself
git grep -n "applyDiscount(" -- '*.ts'

# See only what changed, not whole files
git diff --stat
git diff -- src/cart/total.ts

Then paste the output: “Here are the 4 call sites of applyDiscount: [pasted lines]. Update them to pass the currency argument.”

Spotify built an automated version of this habit. Its setup sends bulk file reads from Claude Code to a cheaper worker model and gives Claude only the summary. A Spotify engineer tested it on a Java monorepo in four scenarios. For each, he set the tokens Claude would use to read the files itself against the tokens it used to read the worker’s summary, and the bulk-read savings averaged around 90%. That figure applies to that one codebase, counts only Claude’s context and not the worker model’s tokens, and covers the bulk-read scenarios only. The Spotify Engineering post also lists its limits: edits still need direct reads because the summaries lack reliable line numbers, reasoning and debugging stay with Claude, and each delegation adds latency.

Keep Sessions Short and Start Fresh

Start a new session between tasks instead of letting one conversation keep growing. A long conversation carries its whole history into each new message, so continuing a stale thread keeps charging you for context you no longer need. Anthropic’s Claude Code cost guidance says that when spend on an API or cloud-provider plan comes in higher than expected, the usual cause is a session that was never cleared or Opus left as the default model. It lists clearing between unrelated tasks among the habits with the biggest effect.

# End of session 1
Summarize in under 15 lines: what we changed, what we decided, what is left.
Write it to NOTES.md.

# Start of session 2 (new chat/session)
Read NOTES.md, then continue with the first remaining item.

The summary keeps your decisions. The new session drops the history that led to them.

Turn Off Tools and Integrations the Task Doesn’t Need

Each integration you enable adds to what the agent loads and can reach for. If the task doesn’t need an integration, disable it. MCP servers are the usual culprit. Claude Code now holds back full MCP tool definitions until Claude needs one, so an idle server costs less than it used to, but its tool names and instructions still sit in context. Anthropic’s Claude Code cost guidance still tells you to run /mcp and switch off servers you aren’t using. A database server, a browser automation server and a ticketing integration you connected last month all add weight to a session that only touches CSS.

For a one-file refactor, open your agent’s settings and switch off the database, browser and issue-tracker servers. Turn them back on when a task needs them. If you’re not sure what each server exposes, the guide to the MCP ecosystem explains how clients and servers fit together. If you maintain your own server, building an MCP server step by step shows how to expose only the tools you define.

Summary

HabitWhat it removes from contextEvidence
Name the filesExploratory file readsReasoned practice
Instructions fileRediscovering conventions; wasted outputarXiv 2601.20404
Plan before codeReads and writes spent on wrong directionsReasoned practice
Search yourselfWhole files opened to find a few linesSpotify Engineering, one Java monorepo
Short sessionsStale conversation historyAnthropic’s Claude Code cost guidance
Disable unused toolsIntegrations the task never callsAnthropic’s Claude Code cost guidance

Conclusion

Token costs grow with the size of the context, and an agent like Claude Code resends the whole conversation with every request, so the context it reads and then rereads adds up fast. Every habit above works by keeping that context small and on task. Pick one, then paste the prompt and files you’re about to send into an LLM token counter and count them before and after the change. Without a measurement, a saving is only a guess.

FAQs

What is the difference between compacting and clearing an agent session?

Compacting swaps older conversation history for a summary and keeps the session going, while clearing throws the history away and starts a fresh context. In Claude Code, the model has to read the whole conversation to summarize it, so compacting a big context is a big request in its own right, while clearing is free. Compact mid-task when earlier decisions still matter, and clear between unrelated tasks.

Do MCP servers I am not using still cost tokens in Claude Code?

Less than they used to. By default, Claude Code uses tool search: at startup the model sees each server's tool names and instructions, and a tool's full schema is fetched only when a task calls for it. Claude Code falls back to loading everything up front when ANTHROPIC_BASE_URL points to a non-first-party host, on Microsoft Foundry deployments hosted on Azure, and on Google Cloud's Agent Platform models older than the Claude 4.5 generation. Servers marked alwaysLoad also load in full. Tool output from any server still enters context.

Does switching to a cheaper model reduce token usage?

No, a cheaper model changes the price of each token, not the number of tokens. The agent still reads the same files and resends the same history. The two levers stack: Anthropic's Claude Code cost guidance names matching the model to the job and clearing between unrelated tasks as the highest-impact habits for reducing spend.

How can I check token usage from inside Claude Code?

Run /usage for session cost, plan usage limits and activity stats, and /context for a breakdown of what fills the current context window. On Pro, Max, Team and Enterprise plans, /usage also points out any behavior, such as long context, that makes up 10% or more of your recent usage, and suggests how to cut it. The dollar figure is always an estimate worked out at list price. On a subscription it doesn't show what you actually pay.

Understand every bug

Uncover frustrations, understand bugs and fix slowdowns like never before with OpenReplay — self-hosted, with full data ownership.

Star on GitHub

We use cookies to improve your experience. By using our site, you accept cookies.