I think there's a huge problem with people getting into the machine learning field with the AI boom.
In prior settings, there used to be a clear separation of training, development/validation, testing partitions of any given task benchmark. The reason for this is so that you can tune hyperparameters: during training (e.g. learning rate), or after a training run (e.g. calibration), and then once you evaluate your system (could include the model and other pre/ post processing), that was it. The test set performance is the number you report.
There is a rationale behind this workflow, because when demonstrating a method, if you're adjusting ANY part of your system's performance against the result you finally report, you're overfitting to the test set.
Suppose you report a 90% performance on the test set, someone reading that would reasonably assume that the system works more often than it doesn't. But if you've overfit any part of your system (the prompt, the calibration, etc.), you could be tuning a 10% performance to 90%, shrugging and saying "Hey if it works it works!" and then happily reporting that number. Applying that same system to some other data that doesn't have the same quirks of this test set will fail.
How Jev manages to claim calibrated probabilities is beyond me. Calibrated to what?
In prior settings, there used to be a clear separation of training, development/validation, testing partitions of any given task benchmark. The reason for this is so that you can tune hyperparameters: during training (e.g. learning rate), or after a training run (e.g. calibration), and then once you evaluate your system (could include the model and other pre/ post processing), that was it. The test set performance is the number you report.
There is a rationale behind this workflow, because when demonstrating a method, if you're adjusting ANY part of your system's performance against the result you finally report, you're overfitting to the test set.
Suppose you report a 90% performance on the test set, someone reading that would reasonably assume that the system works more often than it doesn't. But if you've overfit any part of your system (the prompt, the calibration, etc.), you could be tuning a 10% performance to 90%, shrugging and saying "Hey if it works it works!" and then happily reporting that number. Applying that same system to some other data that doesn't have the same quirks of this test set will fail.
How Jev manages to claim calibrated probabilities is beyond me. Calibrated to what?