Hacker News

Favorites Setup
Comment by Buttons840 | original | Grieving the loss of details
[−]Buttons840 · 2026-10-10 Sat 17:13 UTC · link
I think programmers are frustrated in part because the code used to be our place to do our thinking, but now LLMs just buzz through the file changing thousands of lines and we can't keep up.

It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that. We just have to raw-dog it by thinking really really hard and remembering how all the code connects together.

I want a tool that gives programmers a place to record their thoughts. Developers need a place to draw and write, and also interleave blocks of code that automatically update to match the actual state of the code.

The closest thing I know to this is org-babel, part of Emacs, which allows you to push code blocks out of an org file into an actual source files, or pull them in from actual source files. This is mostly done manually by invoking functions called `tangle` and `detangle`.

I intend to investigate this further in Emacs, since I'm an Emacs user, but Emacs is never going to be the friendly UI we need to make this tooling common.

[−]devin · 2026-10-10 Sat 17:20 UTC · link
Yeah, I can relate to this. However, I haven't found it too difficult to adjust. I have found myself creating draft PRs, and then just sitting on them and thinking about it for a day or two before I even consider merging it. This lag time is the time that I used to spend typing it out and thinking as I went along, now that's happening later. I have closed more than a couple of my own PRs once I had time to consider them. I am rarely shocked when I wake up the next day, look at it, and think: "Eh, this change is not sufficient because it doesn't address X".
[−]kaashif · 2026-10-10 Sat 17:21 UTC · link
Yeah, this is the tough part. In order to write code that worked, you had to have some kind of mental model of it. Now that's not true.

Now, when someone sends a working PR in, even high quality and well tested, they may actually have no idea how it works.

[−]GolfPopper · 2026-10-10 Sat 19:05 UTC · link
Which, in my experience, means they sometimes cannot fix the bugs they've introduced.
[−]abalashov · 2026-10-11 Sun 04:46 UTC · link
No, indeed. They rely on LLMs to do that, and that necessarily adds a degree of entropy that will eventually cause the thing to vibrate apart, but much later, in the future. Meanwhile, the incentives are to fix (or "fix") the bug now.
[−]xtajv · 2026-10-11 Sun 07:17 UTC · link
I think that software engineering is finally entering the "regulations written in blood" era.
[−]abalashov · 2026-10-11 Sun 07:52 UTC · link
Oh yeah. The AI hype cultural moment will wear off, but we'll be stuck with the technical debt, cognitive debt and slop code, which has made it into critical production systems, for decades, probably.

The code--as long as you don't try to modify or evolve it, engendering further nondeterministic degradation--will mostly work, but 0.01% of the time it won't, and you'll never be able to predict when and where that'll be.

[−]visarga · 2026-10-10 Sat 17:28 UTC · link
> now LLMs just buzz through the file changing thousands of lines and we can't keep up. It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that.

I do it differently, I focus on better recording what the user wanted, the so-called "user intent". To do this, I record all messages typed by the user since the start of the project, whether 3,000 or 10,000 messages. An LLM can churn through them in 10 minutes and derive a fresh, up-to-date interpretation from the raw data. This can be used to judge whether the implementation has diverged from the intent, or, in other words, to realign the code and tests. The messages the user writes are usually designs or corrections, a very rich, compact signal. If the user struggles with something, it could result in a tool, a skill, updates to the project docs, or new tests.

[−]chrisweekly · 2026-10-10 Sat 17:50 UTC · link
Yeah, "Intent-based UX" is evolving to address this challenge, but it's got some catching up to do AND is far from mainstream....
[−]gnatolf · 2026-10-10 Sat 21:01 UTC · link
But a codebase is a state machine, and given the somewhat random style of LLM outputs, those amplify to an extent that the recorded intent needs to include the outputs too, in a way? Or are you basically doing high detail specs?
[−]MHard · 2026-10-10 Sat 17:30 UTC · link
The workflow I use to still keep up with everything is to start coding by hand and only once I have a good idea how the rest is gonna look like and am bored I had off the rest of the pr to the LLM.

Could be just defining the methods without filling them but depending on the mood I code more by hand or less.

[−]Terr_ · 2026-10-10 Sat 17:30 UTC · link
> It's still valuable to deeply understand parts of a program

Part of woe is that once you've reviewed, validated, and comprehended a piece... Later gets casually mangled by some other LLM-generated urgent change.

[−]kaffekaka · 2026-10-11 Sun 06:55 UTC · link
Yes to this. With the speed of ai code generation, the team has to decide not to rip things apart in every pr even though it is possible, since this destroys any hope of maintaining even a bit of understanding.

Some parties mean that making understanding irrelevant of the whole goal of ai driven development, though.

[−]writeslowly · 2026-10-10 Sat 17:38 UTC · link
I feel like we also need better non-LLM driven ways (like better static analysis tools) to analyze LLM-driven changes. The way changes are presented in modern IDEs was designed around reviewing human-created changes and doesn’t really feel like it’s keeping up with presenting and validating what modern LLMs are doing.
[−]radarsat1 · 2026-10-10 Sat 19:50 UTC · link
I've been playing with often asking Claude to generate visualizations for me of what it's doing in the sense of diagrams of different types. Sometimes it's as simple as "show me what you're doing using an infographic". Othertimes I'll ask it for sequence or class diagrams or bipartite graphs if it's designing some sort of mapping. Bipartite graph is also very helpful for following the plan of a big branch rewrite or squash which I often do before making a PR to get rid of all those confusing in-between commits.

I find forcing it to visualize things immensely helpful. I'm usually studying git diffs but when working of a big feature or refactor that can just be too hard.

I've never been very pro "visual programming" and always hated UML et al, but part of me is starting to wonder if it's time for us to give it another serious go.

[−]yurishimo · 2026-10-11 Sun 09:12 UTC · link
UML is not a terrible idea from a systems architecture point of view but UML is absolutely terrible at describing business logic. And in today’s age of CRUD apps and frameworks, for some devs, business logic is the entire problem to be solved.

What is “valid” input? What is the failure scenario at the boundary for the consumers of your application? How are exceptions in the business handled? All of these are questions that might have a somewhat obvious answer, but if your goal as a business is to do something radically different compared to your competition, the “obvious” answer might be the wrong one.

[−]RugnirViking · 2026-10-11 Sun 11:21 UTC · link
The usefulness of UML is highly reliant on the names you choose. And current LLMs are really, really bad at naming things. "The collector's refresh cycle has a bug: when the cache hydrates, the Cartesian grouping is misaligned" type nonsense - it absolutely will name a class "CartesianMisalignmentHandler" if you let it, good luck understanding what that is on a UML diagram
[−]mejutoco · 2026-10-10 Sat 17:44 UTC · link
> I want a tool that gives programmers a place to record their thoughts. Developers need a place to draw and write, and also interleave blocks of code that automatically update to match the actual state of the code.

Sounds, like you already mention with org-mode or similar ones) like literate programming (https://en.wikipedia.org/wiki/Literate_programming) or jupyter notebook.

I think the solution is still code, just at a much higher level of abstraction. Maybe a start is kind of typed ADR or FSM that guides (constrains) the agents. I believe more type checking guarantees will be more and more important for agents.

[−]Buttons840 · 2026-10-11 Sun 07:20 UTC · link
I don't think the solution is code, because I still want a place where I can write my own chain-of-thought, without it having to rise to the level of "code". I just want to put my chicken-scratch drawings and diagrams somewhere. I just want a place to think that won't change out from under me.
[−]bluefirebrand · 2026-10-10 Sat 18:23 UTC · link
> It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that. We just have to raw-dog it by thinking really really hard and remembering how all the code connects together

Which is frankly exhausting to do when you have to keep up with the rate of LLM changes

[−]__MatrixMan__ · 2026-10-10 Sat 18:40 UTC · link
I can sling code at about 20x speed with an LLM, but I can only understand it well enough to support it at 5x speed and I can only make decisions that won't piss off the rest of the company at about 3x speed. My job as a software engineer is to therefore slow down to working merely 3x faster than before despite the extra headroom that the LLM gives me. Anybody can give into the seduction of new features poorly understood, to be a specialist means to bother spending the extra time.

Or at least that's the current model I'm playing with.

[−]zmmmmm · 2026-10-10 Sat 20:40 UTC · link
This is my dilemma as well currently, because there's no obvious place to draw that line. The boundaries are all subjectively defined.

I could change a whole UI completely in 30 minutes to something fundamentally better but then 30 people would all wake up and be upset they weren't consulted and need training for it. That training and consultation will take hours and hours. And probably generate feedback - some of it correct, some of it misguided - that needs to be human negotiated, taking more hours. The effective maximum rate of change is limited so dramatically more by other factors than the technical implementation that we have to completely redesign process now to cater to those factors.

We are in a weird space now because most of the process is still built around a presumption that technical implementation is a lot of work. The main reason to be upset that you weren't consulted about a change is because there's a presumption that you will be stuck with it - ie: it's a lot of work to change it back. But it isn't a lot of work, it's effectively free. All this is just living in inertia right now.

[−]__MatrixMan__ · 2026-10-10 Sat 21:41 UTC · link
I've been handling it as a sort of voluntary A/B test.

A is what you're used to, B is what I recommend. If I can convince people to start using B instead, I can look at the metrics for A and conclude that it's effectively dead, and then I can remove it.

It's working out for me, but maybe not a fair comparison because I only have something like 15 users.

[−]zmmmmm · 2026-10-11 Sun 00:22 UTC · link
Yeah it's an interesting one - I have had the same thought process. If code is free, and I trust the tests and the review process, why not deploy a different branch to production for every user that wants one? it's only at the point where shared resources such as database schemas conflict that it becomes an issue, but a large slice of user requests don't even touch those. If something gets deployed broken in one branch, someone can flip branches and use the one that works. As long as the system maintains strongly enforced safety boundaries, a thousand roses can bloom outside of those boundaries.

It does beg the question where it all leads however ...

[−]abalashov · 2026-10-11 Sun 04:52 UTC · link
> But it isn't a lot of work, it's effectively free. All this is just living in inertia right now.

... but it isn't free. You're introducing a certain amount of entropy and drift every time you let the coding agent loose on it. There's a hidden cost of loss of cohesion that comes with changing anything and then changing it back, and while that cost exists with human developers, too, the pace at which they work and think limits the damage and the risks. LLMs just compound them, but the consequences are long-term, while the incentives are to do the thing now.

[−]__MatrixMan__ · 2026-10-10 Sat 18:36 UTC · link
I've been having agents build knowledge graphs, they're a tremendous mess to start with, but I take the time to manually drag nodes around or group them in meaningful ways so that it's actually human-browsable. This is boring enough to create space for me to think in. It leads me to go on expeditions into the code which surface the missing details. It's also a nice way to communicate context to agents. Like, I can hide all but the relevant nodes from an agent before suggesting that it query the graph to understand which service references which other service via which api, which database tables are read/written by such an action, etc...
[−]tartoran · 2026-10-11 Sun 00:36 UTC · link
Curious, how do you do this? What tools are you using to build and manipulate the graphs, and how do you feed that context back to the agents? The manual organization part sounds particularly interesting.
[−]torstenvl · 2026-10-10 Sat 18:53 UTC · link
I find that there are domains where LLMs are much faster and more skilled than I am, particularly in extremely well-documented but technical and complicated, but a lot of domains where they cannot do anything at all (mostly novel issues, weird architecture issue resolution, etc.).

Building a basic X11 window manager is almost a one shot prompt.

Modifying a UI toolkit to make it work with MSAA/IA2 is simply not possible.

There's a lot of room for deep work left... for now.

[−]ctoth · 2026-10-10 Sat 19:40 UTC · link
When you say this is simply not possible you sparked my interest. I would be curious to chat about your approach?

If I were trying to accomplish this particular goal I would first consider what the agent could see. In particular does it have an accessibility inspector of some kind? or even NVDA hooked up with NVDA Remote so that it can actually see the implicit a11y tree for the toolkit it is working on? My email is in my profile and I would love to chat about this.

[−]joshuahedlund · 2026-10-10 Sat 20:48 UTC · link
> It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that. We just have to raw-dog it by thinking really really hard and remembering how all the code connects together.

I deeply relate to this. When engineers were writing all the code that meant every part was deeply understood by _someone_ on the team, and they could valuably contribute to maintenance and further development. It wasn’t perfect, people leave, people forget things, etc, but the overall coverage was high and valuable.

Now, every agent-produced MR introduces code that is deeply understood by _no one_. It’s the “original developer left five years ago” problem, but now growing on every single new piece of code. Reviewing doesn’t give you the same depth of understanding, and the continually increasing impulse is to just approve, maybe nudge it about some isolated enum types or something, but don’t take the time to understand it, just keep the train going.

But then what happens when something breaks and the cloud agents are down…

[−]bagacrap · 2026-10-10 Sat 23:05 UTC · link
> It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that.

Doesn't the LLM do that, if you want it to? I do use it to write code, but the more striking ability it grants me is a means of understanding legacy code far faster. I can ask "what actually causes this branch to be taken" and it's usually right. Ok sometimes it's not right, but I'm not always right either even after I spend tens of minutes reading code.

It's also exceptionally good (i.e. fast) at looking through git history to figure out where/when a certain behavior originated, which can be difficult (time consuming) if code is continuously being refactored.

Sure you could use it to vibecode. I don't, rather the opposite, I understand my own changes better. But I still fear that I'm going to be obsolete as soon as it figures out what questions to ask. And I'm unwittingly training it to do that.

[−]abalashov · 2026-10-11 Sun 04:48 UTC · link
> But I still fear that I'm going to be obsolete as soon as it figures out what questions to ask.

Code and business requirements, outside of maybe very pedestrian CRUD apps, is diverse enough that I don't see that to be a realistic concern. A far bigger concern is that it may not matter whether they're asking the right questions, and that sufficient brute force can overcome any objections like 92% of the time--good enough for government work, good enough for the MBA frat boy who lords over you and always saw you as a needlessly expensive typist anyway...

[−]StrangeWill · 2026-10-10 Sat 23:43 UTC · link
The annoyance I have is that it isn't one of the other. I spend a lot of time thinking and fighting with whatever frontier model we're on today. I've seen people let LLMs make absolutely atrocious decisions and just click next and collect checks.

I just posted this morning about this type of engineering has secured us numerous customers and put some projects in our backlog that either need significant rework or at least a very close eye to see if their issues crop up.

IDK, the idea that we shouldn't be thinking is a worrying one. I still have to think a lot.

What I'm _mostly_ worried about is that the path to get to high performing senior is basically a burned bridge with our current training techniques, and I'm not sure we'll adapt before a brain-drain situation in the industry.

[−]soundworlds · 2026-10-11 Sun 00:10 UTC · link
Not exactly what you asked for, but I now keep pencil and paper next to me, and that has been incredibly refreshing for brainstorming / mental processing