How Much Does It Actually Cost to Run Claude/ChatGPT for Work? A Realistic Token Math Breakdown

The bill you don't see coming

Most people who use Claude or ChatGPT for work every day have never actually looked at what they're paying for. There's a monthly subscription, or an API key with a balance that quietly drains, and the number moves — sometimes slower than expected, sometimes a lot faster. If you've ever wondered why a "light" week of AI-assisted work cost more than a heavy one, the answer is almost never the AI being "expensive." It's usually about what you're sending, not just what you're asking.

This post walks through the actual math — in plain language, no jargon — so you can see where the cost really comes from, and why the fix isn't "use AI less," it's "stop paying to re-explain yourself every single message."

Tokens, in plain English

Every message you send and every reply you get back is broken into small chunks called tokens. As a rough rule of thumb, one token is about three-quarters of a word — so 1,000 words of text is roughly 1,300–1,400 tokens. Short words, common phrases, and code often compress to fewer tokens per word; unusual words, formatting, and non-English text can use more.

Two things trip people up when they first think about this:

  • Both directions count. The words you type (input) and the words the AI writes back (output) are both billed or metered. A long, detailed answer costs more than a one-line answer, even to the same question.
  • Every message re-sends the whole conversation. This is the part almost nobody realizes until they see it laid out. Chat-based AI models don't have persistent memory between messages the way a human colleague does. To generate a reply that makes sense in context, the system has to re-read the entire conversation so far — your first message, its first reply, your second message, its second reply, and so on — every single time you hit send.

Why "just chatting" gets expensive as a conversation grows

Here's the part that actually matters. Picture a conversation that starts short and grows naturally, the way real work conversations do:

  • Message 1: you send 200 words of context, get back 150 words.
  • Message 2: the system now has to process message 1 and your new message, then reply again.
  • Message 3: it re-processes messages 1 and 2, plus your new one.

By message 10, you're not paying for "one new question." You're paying to re-transmit everything that came before it, every time. The conversation history is cumulative — it doesn't reset or compress on its own. A 20-turn conversation isn't twice as expensive as a 10-turn one; because of the re-sending effect, it can end up costing several times more, even though the actual "new" content per message stayed about the same size.

This is directional math, not a pricing table — providers set and change their own per-token rates, and you should always check the current published pricing on the provider's site (anthropic.com or openai.com) if you want exact numbers for your plan. But the mechanism holds regardless of the rate: even a modest per-token cost, multiplied across a growing conversation history, adds up fast. The rate is not the main lever. The size of what you're re-sending is.

Where the waste actually lives

Now add in how most people actually work. A typical knowledge worker doing AI-assisted tasks — drafting emails, summarizing calls, analyzing a spreadsheet, writing a brief, checking a contract clause — doesn't start each task with a blank slate. They re-explain, every single time:

  • "You are an experienced [role]. Act like a senior [X]. Here's our brand voice / our team's context / our formatting preferences..."
  • Background on the company, the project, the audience, the constraints.
  • The same tone and style instructions they typed yesterday, and the day before.

None of that is "new" information to the task at hand — it's boilerplate. But because it gets typed fresh into the chat window (or worse, built up organically inside one long meandering conversation), it gets counted, re-sent, and re-billed every time.

An illustrative example: 10 tasks a day, two different ways

Let's compare two approaches to the same day of work — 10 AI-assisted tasks, each requiring the AI to act as a specific kind of specialist (say, a marketing analyst reviewing campaign numbers, or an operations manager tightening a process doc).

Approach A: ad hoc re-typed prompts

Each of the 10 tasks starts with the person re-typing (or copy-pasting from an old doc) a paragraph or two of role framing: "You're a senior marketing analyst with expertise in attribution and campaign ROI. Please analyze the following data and give me a structured summary with recommendations..." Then the actual task. Then, often, a couple of follow-up messages inside the same thread to clarify or refine — each of which re-sends everything above it. Multiply that boilerplate-plus-follow-up pattern across 10 separate tasks and you're re-paying for essentially the same instructions, dozens of times over, all day, every day.

Approach B: a lean, skill-file-based setup

The role, tone, standards, and working method are written once, carefully, as a compact system prompt — a "skill file" — and loaded at the start of the session instead of retyped. It's front-loaded once, kept tight (no rambling, no redundant instructions), and reused across every task in that role. The person just states the task. No re-explaining who the AI is supposed to be, how to format answers, or what "good" looks like — that's already locked in, concisely, at the top.

The token math difference isn't about the AI "working harder" in Approach B — it's that the repeated, boilerplate portion of every message shrinks dramatically. Less boilerplate re-sent per message, across every message in every task, across every day — that compounds fast in the same way the conversation-growth problem does, just in the opposite direction.

Analytics, without the re-typing
Rafa — Marketing Analytics Specialist AI Skill
Rafa — Marketing Analytics Specialist AI Skill
$29this skill vs $200/hran analytics consultant

A pre-built system prompt that turns Claude or ChatGPT into a senior marketing analyst — campaign ROI, attribution, and reporting standards locked in once, so you're not re-explaining the role every time you paste in new numbers.

View Rafa — Marketing Analytics Specialist →

Why a well-built skill file beats an ad hoc prompt

Not all "reusable prompts" solve this problem. A bloated, overly generic system prompt — five pages of vague instructions "just in case" — can be just as wasteful as retyping, because it's still boilerplate getting re-sent every message, just pasted once instead of typed fresh each time. The lever that actually matters is conciseness: a skill file that says exactly what the role is, exactly how to format output, and exactly what "done" looks like — nothing more — front-loaded a single time per session.

That's the real difference between a skill file and a loose prompt you keep re-writing: it's not just about convenience, it's about permanently shrinking the repeated part of every message you send, for the entire life of that working session.

Ops work, one clean setup
Owen — Operations Manager AI Skill
Owen — Operations Manager AI Skill
$29this skill vs $200/hran ops consultant

Configures Claude or ChatGPT as a senior operations manager — SOPs, process audits, and prioritization frameworks built in, so every new task starts lean instead of re-explaining your operating context from scratch.

View Owen — Operations Manager →

Practical habits that actually move the needle

Regardless of which tool or plan you're on, a few habits reduce the "hidden re-send tax" more than anything else:

  • Start new tasks in new conversations instead of piling everything into one ever-growing thread. A fresh, focused conversation re-sends far less history per message.
  • Front-load role and context once, concisely, rather than re-explaining it inside the task itself every time.
  • Keep the front-loaded instructions tight. A long, sprawling system prompt is still boilerplate being re-sent — trim it to what actually changes the output.
  • Ask for the length of output you actually need. Output tokens count too, and "give me a structured summary" tends to cost less than an open-ended "tell me everything about this."
  • Check current published pricing directly with your provider if you want to model your own costs precisely — rates and models change, and the exact numbers are best confirmed at the source rather than from any third-party estimate.
Automate the repeat work itself
Arthur — AI Operations Agent
Arthur — AI Operations Agent
$49this skill vs $150/hran operations consultant

A configured operations agent skill built for recurring, structured tasks — designed to keep instructions compact and consistent instead of ballooning with every new request.

View Arthur — AI Operations Agent →
Free tool · Solo or team mode
See your own number: AI Token Cost Calculator

Plug in how you (or your team) actually use AI chats and get a real monthly token-waste and dollar estimate for re-prompting vs. a persistent skill file. No signup, no email — just the math from this post applied to your numbers.

Try the calculator →

FAQ

Do I really pay for the AI's replies, not just my own messages?

Yes. Both what you send and what you receive back are counted. A model that writes long, thorough answers by default will use more output tokens than one that writes tight, focused answers — even for the exact same question.

Why does a long conversation cost so much more than a short one, task for task?

Because most chat-based AI tools don't retain memory between messages on their own — to answer message 20 sensibly, the system has to re-process the first 19 messages (and its own 19 replies) along with it. The conversation history is resent in full, every time, so cost tends to grow much faster than the conversation length itself.

Is a skill file just a fancy name for a saved prompt?

Functionally, yes — it's a system prompt or project instruction you write once and reuse. What makes a good one different from an ad hoc prompt is discipline: it states the role, standards, and output format precisely and concisely, so the "boilerplate" portion re-sent with every message stays as small as possible while still doing its job.

Should I worry about exact dollar amounts, or just the behavior?

Focus on the behavior first — trimming repeated context and starting fresh conversations for new tasks will reduce your usage regardless of what the per-token rate happens to be this month. If you want to model exact costs for your own volume, pull the current published rates from your provider's pricing page and do the math against your own typical message counts.

In summary:

The token cost of daily AI work is driven far more by repeated, boilerplate context than by the per-token rate itself — every message re-sends the whole conversation, so ad hoc re-typed prompts compound fast. A concise, front-loaded skill file like Rafa or Owen minimizes that waste by loading the role once, tightly, instead of re-explaining it every single message.

~/get-started

Skills that work. No fluff.

Browse every skill, prompt pack, and agent in the store.

Browse all skills →Or start with free skills