UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]
热度趋势
百分比基于当前可用热度信号,而非评论数或独立用户人数。
一位开发者正在寻求关于使用机器学习和异常检测进行性能回归检测的紧急帮助。他们特别询问了最佳评估设置,尤其是在只有10个健康样本的情况下,是采用留一法交叉验证还是60/20/20的数据分割更优。其目的是在最终确定方法之前,确保评估方法是正确的。
I’m working on performance regression detection using machine learning/anomaly detection.
My setup is basically:
- Healthy runs are used to learn normal behaviour
- Regression runs are used to see whether the model detects the anomaly
- For each counter group I only have about 10 healthy samples
- I’m currently using leave-one-out on the healthy data to set the detection threshold
- The regression samples are not used during training or threshold selection
I’m confused about a few things:
- Do I still need a normal train/validation/test split for this type of one-class anomaly detection?
- With only 10 healthy samples, is leave-one-out better than splitting them into something like 60/20/20?
- Can the regression samples simply act as the unseen test set?
- Would it be better to collect a second independent healthy dataset and use that as a final test for false positives?
- For evaluation, should I mainly use false-positive rate and detection rate/recall rather than MSE/MAE, since I’m not predicting a continuous value?
Just trying to make sure the evaluation setup is correct before I finalise it.