跳到正文
RCreddit.com·
暂不在当前实时榜单

I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?

AI 摘要

A developer distilled an LLM's extraction of court decisions into a GLiNER model and a small multiple-choice model for document extraction. This setup processes about 2.3 documents/s on one GPU. However, its performance is slightly below the original LLM, achieving 0.90 F1 for finding mentions compared to the LLM's 0.935 F1, and 0.85 for exact grouping versus the LLM's 0.93. The developer is seeking feedback on potential mistakes or improvements before processing 5 million documents.

为什么是这条

This report details a specific distillation attempt that, unlike others, openly shares the F1 scores and processing speed of the distilled model compared to the teacher LLM.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月4日 21:00 UTC

收录
2026年10月4日 21:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com