AI Agents and the Refactoring That Never Happens(rosenfeld.page) |
AI Agents and the Refactoring That Never Happens(rosenfeld.page) |
I find it hard to read articles where the agentic writing is this obvious. It's a distraction from the message of the text, which I'm sure is worth my time. Is there no way to stop generated writing from sounding like this?
> Claude, generate a 500 word blog post about how people don't refactor anymore because of AI.
Just send me the prompt!
---
Human contexts are way more limited compared to computers. We can't reason about complex software when they go through many branches with so many implications. It's just too hard for humans to keep track of all interconnected pieces. So humans have historically split the system parts into manageable modules that can be understood in isolation and then we spend some time connecting those parts. That's how we can keep the context reasonable for human understanding.
So, when senior engineers found themselves lost while trying to debug an issue in a complex system they would naturally decide to pause and rewrite or refactoring the confusing piece of the system to make it manageable so developers can easily understand what's going on and review future changes.
Usually a system doesn't start that confusing. But as requirements change developers add additional branches and code until the code is no longer manageable. Sometimes the requirements changed significantly since the code was first written and all we have in the code are exceptions rather than the rule. That's usually when historically senior developers would take the time to rewrite that part of the system so they can reason about it.
But AI agents are not as limited as humans context-wise and they can reason about those confusing (to humans) systems and make sense of it. So they simply keep adding additional branches to the existing mess without ever suggesting a major refactoring like a senior developer would do in those cases, unless there's specific harness to tell agents to act like that.
This article is about bringing this into attention so that developers can policy themselves and keep asking themselves whether it's time for a major refactoring instead of relying on the AI agents and trust them because they no longer understand the code because it's too complex for humans to follow. Can you draft an article focused on this concern?
When they get it right, it works well - full context covered, full business logic recreated completely.
When they fail, they fail fairly hard. They miss some functions in a grep, or return too many and fail to address one part, and they end up with holes.
I scrutinise the work heavily, and find I rely on gut feel a lot for what's right and wrong.
This summer I spent quite a while using a coding agent to help me untangle a deep and complicated data processing pipeline. It had itself been built by agents, in a remarkably short amount of time. But it had also become clear that it was riddled with errors and was producing lots of bad data.
What I quickly discovered was that upwards of half of my questions would receive very confidently wrong answers. And even once I had finally diagnosed whatever problem I was currently working on, it was difficult to trust the agent with any bug fixes. Since it was having an even harder time tracing data flows than I was (I'll take this chance to submit for your consideration that faster is not necessarily better), it was proving to be a bit of a monkey's paw. Yes, it would fix the exact bug I asked it to fix, but typically introduce new defects in the process. And yes, I was having this struggle with all of the latest & greatest models.
I ultimately concluded that, in this codebase, the agent was indeed deeply, hopelessly lost. (edit: And probably this code got so bad in the first place because the agents that were used to build it had been lost for a while, but unable to recognize this problem and call their operators' attention to it.)
I agree with the other reply that you're likely to get better results if it has some kind of test case to run that's more authoritative than its own reasoning.
I agree though that one has to be very careful when trying to "fix" things with agents in a big codebase without it introducing new defects.
I've always _wanted_ my code to be clean and easy to follow but life, deadlines, shifting-priorities, etc have stood in the way of that. Now I can finally realize my personal nirvana.
That said, I've had to steer models away from too-heavy of abstraction or similar because it made the code too hard to follow.
Refactoring used to have a very real cost - it was substantial amounts of time that would have to be carved away from working on new features.
Now I can spot a potential refactor, fire off a prompt in an asynchronous coding agent (or on a worktree or whatever), then come back 20 minutes later and either accept it, poke it a bit, or abandon it. Costs me almost nothing.
Refactoring was never a substantial amounts of time for me before llms. Before I could spot a potential refactor and refactor it in 20 minutes and less. Never abandon it. Just constant improvement to the point the previous tech startup that I was working for just drive from itself (I am still paid a a few hours per months for it)
Not exactly a refactor, but high degrees of consistency are what I strive for in a codebase. LLMs get confused by inconsistencies, as do humans.
Interesting choice of words. Not a fix. Not a refactor. A commit.
But it also doesn’t mean these things aren’t problems, they’re obviously huge problems.
I have some speculation that maybe the people doing the model post-training tend to be younger researchers that haven't worked on large complex software systems, so their taste isn't as developed in this regard about what things are important here.
But this is an area that can be steered with appropriate early instructions ("When choosing tradeoffs of implementation, build for long term maintainability and understandability of code over implementing just the exact code necessary to solve the issues. etc. chain of thought often includes information that would have to be repeated in a future agent session, make sure to persist it to code or external docs so that future sessions and user understanding is respected.")
It can also be done as a post-change step with similar effects. And you can use your agents to build this layer into your general modus operandi for dealing with the crimes of generated code. But one of the things that all AI labs should be doing is looking at AGENTS.md on real project as being hard expressions of what failure modes real projects have noticed in models generally. Don't wait fo the bugs to be raised on these things, use express preferences that show that there's a problem. Go trawl github for these in bulk to use for future post-training.
Your discipline only pays off if you already understand your code and/or established clear baseline for your standards before launching into a feature development mania. And it needs to be enforced every turn, or the firehose of code generation knocks the front door down easily.
Thats… not my experience. Like, not at all. They very regularly get lost
Also this article reads like it was written by Sonnet.
Not only that. Good modularity also:
- improves code reusability, reduces unnecessary code duplication
- helps agents and engineers make better data model, data structure, design pattern, naming, and algorithm choices
- surfaces incorrect irregularities or outdated exceptions to a rule
- enables clean, independent upgrades of parts of a system to improve performance
- reduces stale references in code and comments (and the confusion that results, both from agents and humans)
Bad modularity is basically a summary of what usually constitutes pathological code in general, but AI systems seem particularly disposed to sling lots of it (at least humans are constrained in their output rate). How often have you tried to grok an AI-built project and found trivially unreusable code, unnecessary duplication (everywhere!), bad data structure and algorithm choices, and stale references?
They're not bound by the same limits but they're still bound by some limits, yeah?
I'm not an AI expert, so I don't honestly understand why LLM driven agents are as good as they are. But my impression is "trace every caller", most of the time, is still an approximation. Once the code has gotten convoluted enough, cases are going to get dropped.
Sorry, just thinking about it is reviving my frustration…
Try "build a script to trace callers" and "cite each file and line that calls the function"
Won't fix everything but usually helps in my experience
Although you still run into hard-to-find things especially in dynamic/duck typed languages or codebases with heavy use of reflection or code bases used as libraries
Anything over 600 lines is considered for a refactor when it's touched if appropriate, anything over 800 lines must be refactored when touched. (Combination of AGENTS.md and custom lint warnings/errors).
Anything over 1000 lines forms the refactoring backlog. This way, the bleeding stops immediately and a lot of debt that would otherwise sit in a refactoring backlog forever gets chipped away at over time. The debt that doesn't get chipped is code that isn't touched often anyway and has been battle tested, so might as well be left alone.
As ever, no one is willing to allocate time for this, but (with unlimited work tokens) I can parallel path massive cleanup refactors all the time now.
Often times it's not even that the code is bad but rather that it's overengineered. I see it happen so much that I'm tempted to actually go the other way on a toy project. Like what would Claude or Codex come up with if I told it I wanted an enterprise grade, globally scalable, compliant and auditable tic-tac-toe game.
https://github.com/enterprisequalitycoding/fizzbuzzenterpris...
I've yet to work at a place that bothered with much refactoring over adding the thirtieth conditional to new feature....
Why wouldn't we do that? I think there's a point where this comes down to values instead of facts. If you want it to be human readable, that's fine and there are a bunch of therefores from that point. But if you don't necessarily want that for a particular codebase, why refactor if the LLMs can handle it?
The position in the article is reasonable because what would end up happening otherwise is:
- Agents increase complexity, humans can't read it anymore
- Agents increase complexity, agent can't read it's own code anymore
- Agent unable to keep making updates without looping forever (the complexity of the code exceeds the agent's context length). Human doesn't understand either so can't fix.
This isn't hypothetical either, it's basically what ends up happening to most vibe-coded software if the person doing the vibe coding doesn't know how to review the outputs being produced.
That's a bit of a straw man. State of the art agents are limited to ~3.8Mi (1M tokens). That's usually where I run into issues--an LLM can't possibly hold as much context as a human and it's more of an art than a science getting the most important things squeezed in. It's especially prudent for complex codebases/systems.
An agent only knows what it can see. It doesn't know oldCruftyFunction is still critical to Bob's Excel macro that generates financial reports and yanks the codebase in as a bastardised dependency. A lot of times agents give a fall sense of security by making it seem like something complicated and unsafe is actually safe.
They were trained to do what you ask, but unlike with humans, you need to ask for the refactoring yourself. -- They won't necessarily come up with it on their own.
(They also tend to not be around for long enough to live through the consequences of their tech debt actions.)
I even shared the initial prompt in some comment in this thread. I don't understand what is the matter with also using AI help with generating articles. I still reviewed the article and put all the ideas I wanted to discuss in that article. I don't see why people are often complaining about this. The article content is much more important than how it was generated.
The community is trending strongly against wanting to read generated writing. It's unclear how it will shake out in the long run, but for the near term, this preference is clear*. We can debate the reasons or what the correct position is, but it doesn't much matter when there's such a strong community verdict.
We don't have a rule against genai in articles the way we do about text appearing on HN itself (https://news.ycombinator.com/newsguidelines.html#generated) but the pushback is strong enough that it's in authors' interest to write their articles themselves. Even seemingly less invasive tools like grammar checkers and translators leave strong LLM imprints on text these days.
There is an emerging class distinction in the culture: generated writing, or writing that sounds like it, gets stigmatized and relegated to a lower-status bucket in readers' minds. The converse is also true: writing that isn't generated, and doesn't sound like it, gets boosted into a high-status bucket. It's easy enough to turn this to one's advantage: simply write your own writing.
There are even signs that mistakes (e.g. in grammar or spelling) which would formerly have counted against an author and lowered their status, are turning into markers of authenticity. This is something we point out to non-native English speakers who post to HN: don't worry about your English, just write in your own voice, because that is the high-order bit now. (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...)
I think your prompt was great, btw (https://news.ycombinator.com/item?id=49544844) and something most readers, including myself, would prefer to the generated output. Kudos for sharing it! We often hear readers saying they'd "rather see the prompt", but yours may be the first case I've seen of an author actually posting it.
* We've been keeping an ad hoc list of these reactions to illustrate the point: https://news.ycombinator.com/genai-pushback
But I've personally never worked on code with test coverage that good. Refactors were always risky.
It's like "content creator" vs "writer" or "artist".
Fix more things, solve more problems, and do it to a higher degree of quality in less time using LLMs. What's to complain about?
I also had it first review all tests for superfluous ones that could be deleted.
An actual psychologist could explain it much better than me, but i notice these things.