Bart- A vintage llm [R]
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
Unbounded Labs has introduced Bart, a vintage large language model (LLM) with 2.82 billion parameters.…
Unbounded Labs is proud to introduce Bart, our vintage LLM: 2.82B parameters trained from scratch on 20.1B tokens of English written before 1931. You can talk to it right now!
Why even make a vintage llm? As proposed by Demis Hassabis, could LLMs reach the same conclusions that the great scientists of the past did? While General Relativity was out of budget, we believe that advancing this field targets the crux of AI research. Are these models capable of original ideas, or are they just spitting out the next token?
The article is our full account, covering where the corpus came from and how we cleaned it, the benchmarks we had to build because none existed, every ablation, the training runs, the post-training, and the mistakes we made along the way.
"What I cannot create, I do not understand" is a quote I love from Richard Feynman. Building Bart was our attempt to actually understand LLMs rather than read about them.
- Best vintage base model at its scale on Vintage CORE, ahead of GPT-1900 on a smaller token budget
- Released the largest vintage SFT dataset we know of: 416k graded question and answer pairs, grounded in pre-1930s text
- Trained the final model in 5 days on an H100, holding 60% MFU the whole way
I am proud of my team. What we built will move the vintage LLM field forward, and it moved us forward as researchers and as people.
We paid for all of it ourselves, about $807 so far. Money is the main thing standing between us and a much larger run.
So I will ask directly: we are looking for compute grants, funding, and mentors for our future endeavors. If you work on pre-training, post-training, or you have GPUs sitting idle, we would like to talk!
We believe that with careful dataset curation, domain expertise, and highly efficient training, we can achieve state-of-the-art results in crucial domains. This is only the beginning for Unbounded Labs; we see no bounds ahead.