You know the feeling. Message forty of a chat that started sharp is now giving you answers that ignore half your instructions, contradict something it said ten messages ago, or drift back to a generic default you corrected an hour earlier. Nothing "broke." The model didn't get worse. What changed is the pile of text it now has to read before it can respond to you.
What's actually happening inside a long chat
Every reply an AI model gives you is generated by reading the entire conversation up to that point — your original instructions, every follow-up, every correction, every tangent, every "actually, can you also..." — and predicting what should come next. That whole transcript is called the context window, and it does not get smaller as the conversation grows. It gets bigger, message by message, until it's carrying far more text than it started with.
Early in a chat, your instructions are almost the entire context. The model's job is easy: read a clean brief, follow it. By message thirty, your original brief is a small fraction of a much longer document, buried under dozens of exchanges about specific edits, specific numbers, specific one-off requests. The model still technically "has" your original instructions in its context — but they're now diluted, competing for attention against a much larger volume of more recent, more specific, often contradictory text.
This is why long chats don't fail all at once. They degrade gradually. The model starts weighting recent messages more heavily than early ones (which is often the right call — recency usually signals relevance), but when your foundational rules live in that same aging thread, "recent" and "important" get confused. A throwaway comment from five minutes ago can outrank a standards-setting instruction from an hour ago, simply because of where it sits in the transcript.
There's a second, quieter cost: the model has to re-derive your intent from a messier document every single time. Instead of reading "you are a senior growth strategist who always structures output as X, always avoids Y" as a clean instruction, it's parsing that same idea filtered through paraphrases, exceptions you granted mid-conversation, and edits that only applied to one specific task. More text to reconcile means more room for the model to average things out, hedge, or quietly drop a constraint that never got repeated.
The two fixes people reach for — and why both are lossy
Fix one: scroll up and re-paste the original instructions
This is the most common patch. You scroll back to message one, copy your original brief, paste it back in at message forty with something like "just a reminder, please stick to this." It works for about five more replies. Then the same drift happens again, because you haven't actually changed anything about how the context window works — you've just added one more copy of your instructions to an already bloated transcript. The chat is now even longer, the signal-to-noise ratio hasn't improved, and you're stuck doing this every so often for the rest of the project.
Fix two: start a brand new chat and re-type everything
The other common move is more drastic: abandon the thread and start fresh. This does solve the dilution problem — a new chat has a clean, short context, so instructions are read clearly again. But it throws out something valuable along with the noise: the useful task-specific context that made the AI good at this particular project. The examples you fed it, the edge cases you already walked it through, the decisions you already made together — all of that is gone. You're rebuilding from zero, and you're also relying on your memory (or a scramble through old messages) to reconstruct instructions you may not have saved anywhere outside that chat.
Both fixes treat the symptom. Neither one fixes the underlying problem, which is that your ROLE, FORMAT, and STANDARDS live nowhere except inside a growing pile of back-and-forth.
A compact, reusable skill file that locks in a senior growth strategist's role, output format, and standards — so a new chat starts sharp instead of starting from a blank page.
View Sofia — Growth Marketing Strategist →The better fix: separate the signal from the noise
The actual solution isn't a scrolling trick or a fresh start — it's a structural change to where your instructions live. Instead of your ROLE, FORMAT, and STANDARDS sitting inside message one of an ever-growing thread, they live in a small, standalone file: a skill. You load that file fresh into a new chat or a Claude/ChatGPT project whenever you need it, and the conversation that follows stays focused purely on the task at hand.
Think of it as splitting a conversation into two layers that used to be tangled together:
- Signal — the stuff that should never change and never get diluted: who the AI is acting as, exactly how it should format output, what standards and constraints always apply. This lives in the skill file, outside the chat.
- Noise — the stuff that's specific to this task, this week, this one back-and-forth: "make it shorter," "use this client's numbers," "no, the other version." This lives in the chat, and it's fine if it gets long, because it's not competing with your core instructions for the model's attention.
When those two layers are separate, a long project stops being one long, decaying chat. It becomes a series of clean chats, each one starting with the exact same sharp instruction set loaded from the skill file, each one free to run long on the actual work without ever having to re-litigate what "good" looks like. You get the freshness of fix two (a new chat, clean context) without the cost of fix two (retyping everything from memory), because the instructions were never trapped in the old thread to begin with.
This also solves a problem neither of the bad fixes touches: consistency across sessions. If your standards live in your head and get re-typed slightly differently each time, quality drifts from project to project too, not just within one chat. A skill file is the same file every time — version it once, refine it when you learn something new, and every future chat inherits the improvement instantly.
Analytics threads run long — dashboards, event schemas, reporting cadences. Rafa keeps the analytical standards fixed in a skill file so week six of the project follows the same rigor as day one.
View Rafa — Marketing Analytics Specialist →What this looks like in practice
Say you're running a multi-week engagement — a growth audit, an ongoing analytics build-out, a consulting project with a client. Instead of one marathon chat that slowly loses the plot, the pattern looks like this:
- Your ROLE/FORMAT/STANDARDS live in a skill file (a short, structured document — the kind you'd paste as a system prompt or a Claude/ChatGPT project instruction).
- Each work session, or each new phase of the project, starts a fresh chat with that skill file loaded.
- The chat itself is allowed to get as long and messy as the actual task requires, because the foundational instructions aren't riding along in that same transcript, getting outvoted by more recent text.
- When you learn something that should apply to every future session — a formatting preference, a constraint you keep having to repeat — you edit the skill file once, not every chat going forward.
The net effect is that quality stays flat across a long project instead of sloping downward. You're not fighting the context window; you're just not asking it to do a job it was never well-suited for, which is being your permanent memory for standing instructions.
McKinsey-trained structured thinking, loaded fresh into every session — so a long consulting engagement doesn't inherit the drift of one overloaded chat thread.
View Thomas — Business Consultant →Plug in how you (or your team) actually use AI chats and get a real monthly token-waste and dollar estimate for re-prompting vs. a persistent skill file. No signup, no email — just the math from this post applied to your numbers.
Try the calculator →FAQ
Why does an AI seem to "forget" instructions I gave earlier in the same chat?
It hasn't literally forgotten — the text is still there in the context window. But as the conversation grows, your original instructions become a smaller and smaller share of everything the model has to weigh, and more recent messages tend to get more influence over the response. The result looks like forgetting, but it's really dilution.
Isn't starting a new chat the simplest fix?
It fixes the dilution problem but creates a new one: you lose the task-specific context that made the AI useful for this particular project, and you have to remember (or dig up) your original instructions to retype them. A skill file gives you the clean-context benefit of a new chat without the memory tax, because your standing instructions live outside the chat entirely.
What's the difference between a skill file and just a longer, more detailed prompt?
A longer prompt still lives inside the chat and still ages the same way every other message does — it's subject to the same dilution over a long thread. A skill file is separate from any single conversation: you load it fresh at the start of each new chat or project, so it never competes with recent messages for the model's attention and never accumulates the noise of a specific task.
In summary:
Long chats don't get dumber because the model degrades — they get dumber because your standing instructions get buried under a growing pile of task-specific back-and-forth. Keep the ROLE/FORMAT/STANDARDS in a compact, reusable skill file — like Sofia, Rafa, or Thomas — load it fresh into each new chat, and let the conversation carry only the noise it actually needs to.