HNHacker News·
Not on the current live radar
LensVLM-9B by Apple
Apple introduced LensVLM, an inference framework and post-training recipe, on May 7. This framework allows Vision Language Models (VLMs) to process text as rendered images, addressing the challenge of accuracy deterioration with increased compression. LensVLM, built on Qwen3.5-9B-Base, maintains accuracy comparable to full-text upper bounds at 4.3x effective compression and outperforms baselines up to 10.1x effective compression across seven text QA benchmarks. It also generalizes to multimodal document and code understanding tasks, with accuracy gains increasing with compression.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 23, 2026, 21:01 UTC
- Ingested
- Sep 23, 2026, 21:01
- Source type
- Official
Full text isn't available here.
Read at source →