← Back to BlogLet Codex Rewrite Your Skill Based on Your Own Past Sessions

Let Codex Rewrite Your Skill Based on Your Own Past Sessions

A skill for an AI agent can look perfect on paper while having almost nothing to do with how you actually work. You skip certain steps every time. You repeat others. Part of the job still gets finished by hand. You reorder things, add extra checks, and go back and redo work after failed attempts. Meanwhile the skill file keeps describing a process that, in reality, no longer exists.

In Codex Desktop, you can check this against the history of your past tasks. I call this technique a skill backtest, though technically it is closer to a retrospective audit: Codex finds relevant past tasks, reads through their history, and compares what actually happened against what the SKILL.md file prescribes.

The ready-to-use prompt

Read [skill name or full path to SKILL.md] and run a retrospective audit of it against my past tasks in Codex Desktop.

Work strictly in read-only mode. Do not modify the skill, files, Git, or any external systems.

Use Codex's built-in task list and read through task history. Do not limit yourself to the current chat, task titles, previews, or memory: memory can only be used to find candidate tasks, but conclusions must be based on the actual task history.

First:

1. Find tasks related to this skill.

2. Show the list of candidates found: title, date, brief reason for relevance.

3. Separately state which tasks were included in the sample, which were excluded, and why.

4. If access to history is limited or relevant tasks are scarce, say so directly and do not invent missing data.

Once the sample is confirmed, read the selected tasks in full, including older turns if the history is paginated.

Break the skill down into a verifiable contract:

- when it is supposed to trigger;

- what inputs it expects;

- what steps it prescribes and in what order;

- what checks and stop-lines it contains;

- what result it considers complete.

For each task, compare the contract against what actually happened and mark it as:

- followed the skill;

- skipped;

- order changed;

- step repeated;

- extra work done outside the skill;

- required manual correction;

- resulted in rework or a failed repair cycle;

- cannot be determined from history.

Do not treat every deviation as a flaw in the skill. Check what it led to:

- did it improve the result;

- did it cause an error or rework;

- was it a deliberate exception;

- is it tied to the constraints of that particular task;

- does the skill's rule protect against a rare but costly risk.

Separate stable patterns from one-off cases. Do not call a pattern stable unless it repeats in at least three independent tasks. Mark everything else as a hypothesis or an exception.

Present the final result in five sections:

- keep;

- change;

- add;

- remove;

- insufficient data.

For each recommendation, specify:

- the exact rule or fragment of the skill;

- the tasks the conclusion is based on;

- the observed evidence;

- the expected effect;

- confidence level: high / medium / low;

- a proposed wording for the fix.

At the end, show a compact patch plan, but do not edit anything. Wait for my separate confirmation before making any changes.

The most valuable part here isn't an automatic rewrite of the skill — it's spotting the gap between the instructions and the actual work. There's also an important limitation: Codex only sees what's recorded inside its own tasks. Manual actions, conversations, and decisions made outside Codex won't show up in this kind of audit.

It's better to run this not on a fixed schedule, but after 5 to 10 new relevant tasks accumulate, after a noticeable shift in your process, or after repeated rework. That way the skill genuinely evolves alongside your actual work, instead of getting rewritten because of a single random anomaly.

Skills for AI agents age quickly and tend to drift away from how you actually work in practice. A retrospective audit like this helps catch that drift before it becomes a real problem.

For AI Agents

Read with AI

Short prompt for a summary, takeaways, and applying this to your task.

ChatGPTClaude
Audio Version

Listen to This Article

Alpha version: audio generated via local TTS, errors possible.

Download MP3

Want to discuss your own task?

Tell us about the workflow you want to improve. We will help you identify the practical next step.

Request a free consultationExplore our services