I deal with a lot of non-LLM users at work and quite often the conversation hits a dead end, where I'm like , ok but see, your app has this performance blockage due to your not having an index here - do some performance tests and you'll see (it was a hang caused by a FOR UPDATE locking the entire table due to lack of an index). And you can hear the pause (it's all over Slack) where they just aren't going there, because writing a performance suite for the issue in question would be a lot of effort to do by hand and in the "before times" would be a difficult undertaking to justify. Because they don't consider an LLM, a task they most certainly should be doing becomes a non starter. Never mind my own Claude had a whole plan ready to go to do this whole suite for them in about five minutes so they could study the impact of the index, but I really didn't want to just go ahead and do this all for them. At some point your non-LLM coworkers need to stare directly at work we'd never be willing to do before that's now trivial, and in fact is now part of the job.
Right, but performance testing has always been hard and LLMs are awful at it. So maybe your coworker was right?
I’m definitely on the “just use the clanker” end of the spectrum. But knowing when and is good. Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out. At this point, if I have any idea about what’s going on and someone tells me “Claude says”, I will ignore them.
In this case, it sounds like you could have made a ten line repro that shows the problem. Why not just send that, something human interpretable?
What LLMs are you using? Just simply telling Sol 6.1, Opus 5.5, Fable or Astra to performance test something will get you a solid improvement in poorly optimized code.
Giving it a specific plan will get you a solid test harness.
And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.
Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.
Just like with actual research, the real benefits came from me asking why is it like X and not like Y?
The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…
yes but you can use an llm to do the grunt work which still speeds up the process, you don't need to be like "computah speed this up", you can be like "computah, write me a profiling script using lldb to inspect this one hotpath and look for X, Y, and X" etc etc.
> Right, but performance testing has always been hard and LLMs are awful at it. So maybe your coworker was right?
LOL, you're making the same mistake they did. thinking "index" means "too much time spent fetching the rows". read again - it's a FOR UPDATE so the entire table gets locked and other processes get totally blocked.
> Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out.
It wasn't just claude's idea, it was my idea too, the claude topic was that it would write a suite that proves the problem, in this case, very loud logging messages that were occurring for the customer when this quasi-deadlock situation occurred. it was not subtle.
> In this case, it sounds like you could have made a ten line repro that shows the problem.
no, it involved running a galera server and about four other services with a specific set of data conditions, again, read what I wrote, creating a proof of concept suite was not trivial to do by hand.
But why would you have to make the succinct suite by hand? Not what I was suggesting, I have gotten in the habit of sending extremely concise, human readable repro scripts made by Claude.
Partially because I can read and validate it actually shows what it claims to show, and partially because I expect people want to know what I’m telling them.
I’m definitely on the “just use the clanker” end of the spectrum. But knowing when and is good. Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out. At this point, if I have any idea about what’s going on and someone tells me “Claude says”, I will ignore them.
In this case, it sounds like you could have made a ten line repro that shows the problem. Why not just send that, something human interpretable?
Giving it a specific plan will get you a solid test harness.
And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.
These things excel at performance optimizations.
Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.
Just like with actual research, the real benefits came from me asking why is it like X and not like Y?
The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…
LOL, you're making the same mistake they did. thinking "index" means "too much time spent fetching the rows". read again - it's a FOR UPDATE so the entire table gets locked and other processes get totally blocked.
> Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out.
It wasn't just claude's idea, it was my idea too, the claude topic was that it would write a suite that proves the problem, in this case, very loud logging messages that were occurring for the customer when this quasi-deadlock situation occurred. it was not subtle.
> In this case, it sounds like you could have made a ten line repro that shows the problem.
no, it involved running a galera server and about four other services with a specific set of data conditions, again, read what I wrote, creating a proof of concept suite was not trivial to do by hand.
Partially because I can read and validate it actually shows what it claims to show, and partially because I expect people want to know what I’m telling them.
You’d think they’d have found this issue out before now; if there’s no usable index for the predicate, non-locking reads would’ve also been slow.
Also, I look forward to the next update wherein your colleagues create an index, and then discover the joys of gap locking under REPEATABLE-READ.