Right, but performance testing has always been hard and LLMs are awful at it. So maybe your coworker was right?
I’m definitely on the “just use the clanker” end of the spectrum. But knowing when and is good. Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out. At this point, if I have any idea about what’s going on and someone tells me “Claude says”, I will ignore them.
In this case, it sounds like you could have made a ten line repro that shows the problem. Why not just send that, something human interpretable?
What LLMs are you using? Just simply telling Sol 6.1, Opus 5.5, Fable or Astra to performance test something will get you a solid improvement in poorly optimized code.
Giving it a specific plan will get you a solid test harness.
And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.
Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.
Just like with actual research, the real benefits came from me asking why is it like X and not like Y?
The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…
yes but you can use an llm to do the grunt work which still speeds up the process, you don't need to be like "computah speed this up", you can be like "computah, write me a profiling script using lldb to inspect this one hotpath and look for X, Y, and X" etc etc.
> Right, but performance testing has always been hard and LLMs are awful at it. So maybe your coworker was right?
LOL, you're making the same mistake they did. thinking "index" means "too much time spent fetching the rows". read again - it's a FOR UPDATE so the entire table gets locked and other processes get totally blocked.
> Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out.
It wasn't just claude's idea, it was my idea too, the claude topic was that it would write a suite that proves the problem, in this case, very loud logging messages that were occurring for the customer when this quasi-deadlock situation occurred. it was not subtle.
> In this case, it sounds like you could have made a ten line repro that shows the problem.
no, it involved running a galera server and about four other services with a specific set of data conditions, again, read what I wrote, creating a proof of concept suite was not trivial to do by hand.
But why would you have to make the succinct suite by hand? Not what I was suggesting, I have gotten in the habit of sending extremely concise, human readable repro scripts made by Claude.
Partially because I can read and validate it actually shows what it claims to show, and partially because I expect people want to know what I’m telling them.
I’m definitely on the “just use the clanker” end of the spectrum. But knowing when and is good. Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out. At this point, if I have any idea about what’s going on and someone tells me “Claude says”, I will ignore them.
In this case, it sounds like you could have made a ten line repro that shows the problem. Why not just send that, something human interpretable?
Giving it a specific plan will get you a solid test harness.
And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.
These things excel at performance optimizations.
Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.
Just like with actual research, the real benefits came from me asking why is it like X and not like Y?
The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…
LOL, you're making the same mistake they did. thinking "index" means "too much time spent fetching the rows". read again - it's a FOR UPDATE so the entire table gets locked and other processes get totally blocked.
> Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out.
It wasn't just claude's idea, it was my idea too, the claude topic was that it would write a suite that proves the problem, in this case, very loud logging messages that were occurring for the customer when this quasi-deadlock situation occurred. it was not subtle.
> In this case, it sounds like you could have made a ten line repro that shows the problem.
no, it involved running a galera server and about four other services with a specific set of data conditions, again, read what I wrote, creating a proof of concept suite was not trivial to do by hand.
Partially because I can read and validate it actually shows what it claims to show, and partially because I expect people want to know what I’m telling them.
You’d think they’d have found this issue out before now; if there’s no usable index for the predicate, non-locking reads would’ve also been slow.
Also, I look forward to the next update wherein your colleagues create an index, and then discover the joys of gap locking under REPEATABLE-READ.