返回
RCreddit.com
14
·23小时前·开发者社区 · RSS

I tested 6 AI interview assistants and turned the results into a public dataset

查看原文

热度趋势

新上榜
最近 24 小时与此前 24 小时对比 · 7 天曲线

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

这条记录涉及编程工具或代码能力更新,适合开发者评估工作流变化和可复用价值。

AI 摘要

一位开发者测试了六款AI面试助手,并将其结果整理成一个公开数据集,其中不包括他们自己的产品CTRLpotato。测试内容包括对其中四款产品进行接收端屏幕共享功能测试,结果显示两款失败,两款表现不一。该开发者正在征求熟悉评估数据集的人士的反馈意见。

I build CTRLpotato, so obvious disclosure first, all six products I tested are competitors. CTRLpotato isn't included in the results and I'm not using this to declare a winner.

Over June and July I tested Cluely, Interview Coder, LockedIn AI, ULTRACODE, Parakeet AI and Final Round AI on Mac and/or Windows.

I originally did this because I kept seeing claims like invisible, undetectable, real-time, etc. and wanted to see what the actual desktop apps did. I ended up with enough screenshots, recordings and notes that keeping everything in six separate reviews became pretty useless, so I normalized it into a dataset.

It's now 66 assessments across 6 products and 11 criteria.

A few results that surprised me:

- All 6 failed the focus-behavior check in the setup I tested.

- All 6 failed the cursor-behavior check.

- Shortcut isolation failed on all 5 products where I tested it. The sixth wasn't tested.

- I tested receiver-side screen sharing on 4 products. 2 failed and 2 were mixed.

- There were also some pretty strange context failures, where an assistant would answer an old coding task or otherwise lose track of what it was supposed to be answering.

I don't want to oversell those numbers. This wasn't a lab experiment where every product got an identical setup. Different versions, platforms and test flows were involved, which is why every row includes the app version, platform, date, what actually happened, limitations and evidence where I have it.

The five result labels are just passed, mixed, failed, not found and not tested. A failure means it failed in the documented setup, not that the feature can never work.

The whole thing is public here:

https://www.ctrlpotato.com/compare

I also put the actual dataset on GitHub with CSV/JSON/JSONL, schema, codebook and checksums:

https://github.com/ae0j/ctrlpotato-ai-interview-assistant-benchmark

And there's a versioned Zenodo DOI if anyone wants to cite or archive it:

https://doi.org/10.5281/zenodo.21915738

The dataset is free to use, including commercially, with attribution.

One thing I'm still unsure about: whether the passed/mixed/failed column actually makes the dataset better. The more I worked on this, the more I felt that the raw observation + limitations were more useful than trying to compress what happened into one label.

Curious what people who work with evaluation datasets think.

I tested 6 AI interview assistants and turned the results into a public dataset · BuzzRadr