I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?
A developer distilled an LLM's extraction of court decisions into a GLiNER model and a small multiple-choice model for document extraction. This setup processes about 2.3 documents/s on one GPU. However, its performance is slightly below the original LLM, achieving 0.90 F1 for finding mentions compared to the LLM's 0.935 F1, and 0.85 for exact grouping versus the LLM's 0.93. The developer is seeking feedback on potential mistakes or improvements before processing 5 million documents.
This report details a specific distillation attempt that, unlike others, openly shares the F1 scores and processing speed of the distilled model compared to the teacher LLM.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 4, 2026, 21:00 UTC
- Ingested
- Oct 4, 2026, 21:00
- Source type
- Dev community
Full text isn't available here.
Read at source →