Back
Ssignal
20
·1 days ago·1 signals
Archived topic · source no longer tracked

I benchmarked AutoGen, CrewAI, LangGraph, and MetaGPT against my own Agent OS. The "LLM-as-a-judge" paradigm is completely broken. Here is the local data.

Plans & limitsOn-device

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.