Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

How I Use AI to Document Legacy Code (Without Trusting It Blindly)

24 Dec 2025Category : Blog

About Post

Every long-lived system has a folder that everyone is slightly afraid of. It works. It's been working for years. Nobody is quite sure why, the person who wrote it moved on, and the only documentation is a comment that says // don't touch this.

Documenting that kind of code used to be the task that sat at the bottom of the backlog forever: important, unglamorous, and slow. AI coding agents have changed that for me. Not because they magically understand old code, but because they're very good at the tedious part (reading everything, mapping it, drafting) and leave me free for the part that needs a human (checking it's true).

Here's the workflow I use with Claude Code, step by step, including the parts where I don't trust the AI at all.

The principle: the agent drafts, the code decides

Before the steps, the rule that shapes all of them. An AI agent reading legacy code will produce explanations that sound completely confident. Most of the time they're right. Sometimes they're a plausible story about what the code probably does, which is worse than no documentation, because people will believe it.

So every claim in the final docs has to be traceable back to code, a test, or a person who knows. The agent's job is to make that verification fast, not to replace it.

Step 1: get a map before reading any code

I don't start with "explain this file". I start with the shape of the whole thing. In Claude Code I use a knowledge-graph skill that extracts the relationships across the codebase (which classes call which, which modules depend on each other) so the agent can answer structural questions without reading every file into its context.

Then I ask for an inventory, with a strict format:

List the modules in app/ that relate to billing.
For each: purpose in one sentence, main entry points,
database tables it reads and writes, and other modules
it depends on. Cite file paths for every claim.
If you are guessing, say "unverified".

That last line matters more than anything else in the prompt. Giving the agent explicit permission to say "I'm not sure" produces far more honest output than asking it to sound authoritative.

Step 2: ask what it can't explain

This is my favourite trick. After the inventory, I ask:

What in this module is surprising, inconsistent or
unexplained? List code whose purpose you cannot
determine from the code alone.

The answers are exactly the things that need documenting: the magic number, the special case for one kind of record, the job that runs twice, the column that's written but never read. Those are also the things a new developer would trip over. The agent finds them quickly because it reads with no assumptions, which is the one advantage of having no memory of the project.

Step 3: trace the important flows end to end

Module lists explain structure. Developers usually need to know flows: what happens, in order, when something occurs. So I pick the few flows that matter most and ask the agent to trace each one:

Trace what happens from the moment a payment is recorded until the receipt is sent. List every class, job, event and table involved, in order, with file and line references.

Then I verify the trace myself by following the references. With file and line numbers in hand, this takes minutes instead of an afternoon, because I'm checking a route rather than discovering one.

The rule I don't break: the agent explains what the code does. Only a person, a commit message or a ticket can explain why. When the docs need a "why" and nobody knows, I write "reason unknown" instead of letting the AI invent one.

Step 4: write the docs where developers will find them

Once a module's map and flows are verified, the agent drafts the actual documentation. I keep it small and close to the code:

  • A short README.md in the module folder: purpose, key classes, main flows, gotchas.
  • A few comments on the genuinely surprising lines, explaining why they exist (if known).
  • A line in the project's CLAUDE.md pointing to the module docs, so future agent sessions start from the verified version instead of re-guessing.

That last point creates a nice loop. Each module I document makes every later AI session in that area faster and more accurate, because the agent reads the verified docs first.

Step 5: turn understanding into tests

Documentation describes behaviour. Tests lock it in. For legacy code with few tests, I ask the agent to write characterization tests: tests that capture what the code currently does, quirks included, without judging whether it's right.

Write PHPUnit tests that capture the current behaviour of
LateFeeCalculator, including edge cases you found.
Do not change the class. If a behaviour looks like a bug,
test the current behaviour and flag it in a comment.

Those tests do two jobs. They prove the documentation's claims are true, and they give us a safety net for the day someone finally refactors that scary folder.

Where it goes wrong (and how I catch it)

  • Trusting old comments. Agents read comments and believe them. A comment from years ago may describe code that has since changed. I tell the agent to treat comments as claims to verify, not facts.
  • Missing the hidden callers. Code can be triggered from places search doesn't find easily: scheduled tasks, queue jobs, event listeners, external webhooks, even a cron entry on a server. I ask specifically about each of these.
  • Documenting dead code. The agent will happily write a lovely explanation of a class nothing uses. Checking for callers first saves the effort, and sometimes leads to a satisfying deletion instead.
  • Too much at once. One module per session. Long sessions crowd out the details, and the quality of explanations drops.

Why this is worth doing

The real cost of undocumented code isn't the missing README. It's that every change takes longer, every new developer depends on one person's memory, and every refactor feels like defusing a bomb. Documentation used to be too slow to justify. With an agent doing the reading and drafting, and a human doing the verifying, it's become one of the most useful things I can spend an afternoon on.

If you try this, start with the module people are most afraid of. It's usually the one that needs it most.

What's your approach to documenting old code: write it as you go, schedule a big push, or wait until something breaks? I'd love to hear what actually works for other teams.

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close