What LLMs are you using? Just simply telling Sol 6.1, Opus 5.5, Fable or Astra to performance test something will get you a solid improvement in poorly optimized code.
Giving it a specific plan will get you a solid test harness.
And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.
Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.
Just like with actual research, the real benefits came from me asking why is it like X and not like Y?
The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…
yes but you can use an llm to do the grunt work which still speeds up the process, you don't need to be like "computah speed this up", you can be like "computah, write me a profiling script using lldb to inspect this one hotpath and look for X, Y, and X" etc etc.
Giving it a specific plan will get you a solid test harness.
And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.
These things excel at performance optimizations.
Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.
Just like with actual research, the real benefits came from me asking why is it like X and not like Y?
The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…