跳到正文
HNHacker News·
暂不在当前实时榜单

LensVLM-9B by Apple

AI 摘要

Apple introduced LensVLM, an inference framework and post-training recipe, on May 7. This framework allows Vision Language Models (VLMs) to process text as rendered images, addressing the challenge of accuracy deterioration with increased compression. LensVLM, built on Qwen3.5-9B-Base, maintains accuracy comparable to full-text upper bounds at 4.3x effective compression and outperforms baselines up to 10.1x effective compression across seven text QA benchmarks. It also generalizes to multimodal document and code understanding tasks, with accuracy gains increasing with compression.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月23日 21:01 UTC

收录
2026年9月23日 21:01
来源类型
官方发布

本站未收录正文。

前往源站阅读 →
来源·Hacker News·huggingface.co