For me at least I have to be able to believe I can understand any part of the stack if I want in reasonable effort. I would happily dive into semiconductor physics if I happen to want to (I'm a former physicist so the atom level I pretty much already learnt at school).
Honest question: What is preventing you from understanding any part of the stack now?
You could potentially ask an LLM to explain it, but I would venture a guess that wouldn't feel like the right source to help you. You could use the LLM to help you identify source material written by humans in order to learn anything new. It doesn't have to "do all the work for you".
Modern AI systems are powerful. They are not omnipotent or omniscient. Understanding the details is still valuable, especially if you want to push the frontier of any known field.
In any event, the situation isn't hopeless. At least not yet. :)
> You could potentially ask an LLM to explain it, but I would venture a guess that wouldn't feel like the right source to help you.
I don’t think this works, at least for me. Even on the 5.5 versions of Claude it’s still arduous to read at this point.
> You could use the LLM to help you identify source material written by humans in order to learn anything new. It doesn't have to "do all the work for you".
This probably depends on the shop, but humans aren’t writing docs anymore. That was the first thing to go, sadly.
I hear what you're saying. I often tweak the CLAUDE.md file to improve the quality of its output. It's an ongoing experiment.
As for "humans writing the docs", you raise a good point although it's not quite what I imagined when I wrote that line. I was thinking more about having the LLM/Agent create a list of verified output by humans with reputations as experts in whatever field interests you. This is important when pushing the boundaries of research, although it's overkill for wanting to know whether or not a new technology could be useful in your existing project (or the next one).
What use case did you have in mind when you brought up the docs situation? Sounds like you had some direct experience with that one.
You can ask frontier models with all the bells, whistles, language servers etc to give you every example of X. And it gives you 12 examples and says that's all of them. When you know damn well there are at least 40 (but not exactly how many). So you say no, you know it has missed some, such as X23, and X27. So it goes away and comes back again and says yes, there are 41 X, here they are. How much digging should you do to see if its right?
I have many times gone looking myself to find that it has still missed some. Asked it to go check for those, to look harder it goes oopsie and says now im sure ive gotten all of them (has it?)
All this to say, I have been burned, repeatedly, multiple times a day, for the last year+. I still use these tools, but I truly cannot understand how some people treat them as oracles that know everything about our codebases
Excellent points. I agree that frontier models get things wrong all the time. They are NOT oracles. I believe that we need to encourage our peers to become sophisticated operators of these systems, instead of treating the output like Moses and the 10 Commandments. Perhaps we need more specialists, because the major AI labs lack the incentives to venture down the long tail of knowledge. Agents can be a piece of the learning puzzle, despite their imperfections. Again, thanks for your reply. It gave me lots to think about.
You need to define the goal in a way that you will trust the completion. If you are not sure how you yourself would verify that there are 41, then you are in trouble. Verification must be deterministic, or it's worthless.
What the large models do really well nowadays, is that they won't lie to you if your deterministic verification fails. If you tell an OpenAI or Anthropic model that they need to run a certain `grep` or search or whatever command to verify, then they will do it. I haven’t seen them lie about this for a year, and trust them in this.
Surely you can see that for basically every example of this sort of problem, fully defining a deterministic check is the same as finding them all?
Like you're telling me if I had a script that printed all X, and it's my responsibility to ensure it has no bugs, then the agent could tell me all X and I could trust it?
This is not helpful at all? I am capable of running scripts myself and using the output directly.
No? It's like saying "why bother doing maths, a calculator can figure out the solution to any question you have if you just put the right equation in" there might be more to maths, it turns out. Most of it lying inside of finding "the right equation"
In many case it is the sheer accidental complexity. I tried to get into Linux codebase and gosh it's so much harder than FreeBSD. And Chromium, LLVM... I don't believe it has to be that way. I use Common Lisp software (besides Emacs) written by human for anything I can and these have always been a joy to use or work with.
It seems the way industry is heading towards wrt LLM will make this even much, much worse.
It's the Mad Hatter's tea party, you sit down, you figure out how it works, and then poof a third of its gets overwritten and your past investment is significantly washed away by someone who absolutely didn't put the same amount of thought into what is replacing it.
You could potentially ask an LLM to explain it, but I would venture a guess that wouldn't feel like the right source to help you. You could use the LLM to help you identify source material written by humans in order to learn anything new. It doesn't have to "do all the work for you".
Modern AI systems are powerful. They are not omnipotent or omniscient. Understanding the details is still valuable, especially if you want to push the frontier of any known field.
In any event, the situation isn't hopeless. At least not yet. :)
I don’t think this works, at least for me. Even on the 5.5 versions of Claude it’s still arduous to read at this point.
> You could use the LLM to help you identify source material written by humans in order to learn anything new. It doesn't have to "do all the work for you".
This probably depends on the shop, but humans aren’t writing docs anymore. That was the first thing to go, sadly.
As for "humans writing the docs", you raise a good point although it's not quite what I imagined when I wrote that line. I was thinking more about having the LLM/Agent create a list of verified output by humans with reputations as experts in whatever field interests you. This is important when pushing the boundaries of research, although it's overkill for wanting to know whether or not a new technology could be useful in your existing project (or the next one).
What use case did you have in mind when you brought up the docs situation? Sounds like you had some direct experience with that one.
You can ask frontier models with all the bells, whistles, language servers etc to give you every example of X. And it gives you 12 examples and says that's all of them. When you know damn well there are at least 40 (but not exactly how many). So you say no, you know it has missed some, such as X23, and X27. So it goes away and comes back again and says yes, there are 41 X, here they are. How much digging should you do to see if its right?
I have many times gone looking myself to find that it has still missed some. Asked it to go check for those, to look harder it goes oopsie and says now im sure ive gotten all of them (has it?)
All this to say, I have been burned, repeatedly, multiple times a day, for the last year+. I still use these tools, but I truly cannot understand how some people treat them as oracles that know everything about our codebases
What the large models do really well nowadays, is that they won't lie to you if your deterministic verification fails. If you tell an OpenAI or Anthropic model that they need to run a certain `grep` or search or whatever command to verify, then they will do it. I haven’t seen them lie about this for a year, and trust them in this.
Like you're telling me if I had a script that printed all X, and it's my responsibility to ensure it has no bugs, then the agent could tell me all X and I could trust it?
This is not helpful at all? I am capable of running scripts myself and using the output directly.
That’s like saying mathematics is worthless and we have to resort to finger counting?
It seems the way industry is heading towards wrt LLM will make this even much, much worse.
Y'know, the last AI hypecycle was loads more fun than this one is...
Bringing this question out of the thread and back to the current moment...
the recurring answer now is "futility in the face of churn", I think.