·
Archived topic · 归档话题,来源已停止追踪
Kog is going deeper to squeeze more inference out of GPUs
French startup Kog is aiming to significantly accelerate AI inference on conventional GPUs, despite the market's positive reception to specialized chips like Cerebras. Kog claims its technology can achieve "30x faster LLM inference." A demo showcased an impressive 3,000 per-request tokens per second (TPS) using the open-sourced Laneformer 2B, a purpose-built small model with approximately 2 billion parameters. This demonstrates Kog's approach to extracting more performance from existing GPU hardware.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年8月14日 15:00 UTC
- 收录
- 2026年8月14日 15:00
- 来源类型
- 未分类