Language models for text classification: From bag-of-words to Jev
The Jev AI model has recently gained significant attention within technical communities. This model is part of a broader evolution in language models for text classification, moving beyond earlier approaches like bag-of-words. Key advancements in recurrent neural networks (RNNs) include Long short-term memory (LSTM) networks, introduced in 1997, and gated recurrent units (GRUs), introduced in 2014, which utilize learned gates for information management. More recently, xLSTM: Extended long short-term memory was introduced in 2024.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 29, 2026, 11:06 UTC
IngestedOffset at this time: UTC+0Sep 30, 2026, 01:00 UTC
- Published
- Sep 29, 2026, 11:06
- Ingested
- Sep 30, 2026, 01:00
- Source type
- Unclassified
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
The recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.
While Jev aims to classify things, it’s easy to dismiss Jev as “just a classifier,” and my own view of Jev has evolved quite a bit over the past few days. In particular, my thoughts went from “classifiers used to be my bread & butter; I can easily build this myself” (more on this later) to “wow, this actually works better than I thought.”
Figure 1: Quick overview of the Jev API; more details on that later.
Sure, the latest state-of-the-art GPT and open-weight LLMs can do the same kinds of classification tasks as Jev, while also being capable of much more general decision-making. But Jev’s advantage is that it can handle those classification tasks much faster and more cheaply.
At the other end of the spectrum, for a narrow, well-defined problem, Jev probably won’t classify anything better, faster, or cheaper than a special-purpose classifier. But its selling point is that it is far more general than those task-specific models.
So, what is the methodology behind Jev (based on an educated guess), what can it do, and why is it so popular? I aim to answer all of these later in this article. However, I thought starting with a brief history of language models for decision-making would be a great way to begin. And it hopefully helps demystify some of the hype and show what Jev does very well (”Jev is essentially a text classifier,” but “Jev is also not ‘just’ a text classifier.”)
PS: I am not affiliated with Jev in any way. Also, I am not offered free access to Jev, and this is also not a product endorsement, just a technical article to offer some insights into the history of text classification to help you make sense of the recent hype.
Since this is a long article, I recommend reading it in your browser, where you can access the table of contents menu on the left side.
For completeness, before we put Jev in context (no pun intended), I thought it made the most sense to start chronologically. In this section, I want to take a brief tour of applied text classification via naive Bayes, logistic regression, and the more classic (deep) neural networks before transformer-based models came along.
Back in the day, when I was a grad student 15 years ago, even though recurrent neural networks already existed (more on that later), text classification was usually done with a bag-of-words representation because it was straightforward and could get good results on moderately sized datasets.
In short, we can think of the bag-of-words representation as a method that makes free-form text input of different lengths compatible with classic classifiers (naive Bayes, Logistic Regression, SVMs, Random Forest, XGBoost, to name a few), which expect a fixed-size input vector.
Popular real-world applications include anything from news article classification to email spam filtering. And yes, allegedly even Gmail’s original spam filter used a Naive Bayes model with a bag-of-words representation.
As a side note, I wrote about this approach exactly 12 years ago. It was one of the first things I shared on arXiv.
Figure 2: An old tutorial of mine from 2014 that explains naive Bayes classifiers using a bag-of-words model.
So, what exactly is this bag-of-words representation? It’s a way to convert free-form texts with different lengths, e.g.,
- Training example 1: “Zentropa is the most original movie I’ve seen in years. If you like unique thrillers that are influenced by film noir, then this is just the right cure for all of those Hollywood summer blockbusters clogging the theaters these days. Von Trier’s follow-ups like Breaking the Waves have gotten more acclaim, but this is really his best work.”
- Training example 2: “This film is just plain horrible. John Ritter doing pratt falls, 75% of the actors delivering their lines as if they were reading them from cue cards, poor editing, horrible sound mixing”
- Training example 3: “Zentropa has much in common with The Third Man, another noir-like film set among the rubble of postwar Europe.”
into a fixed-size representation for the aforementioned “classic” classifiers. (The example above is an excerpt from the popular IMDb movie review classification dataset .)
Figure 3: An illustration of a bag-of-words representation.
A bag-of-words model starts by building the vocabulary, which consists of all unique words in the training set (optionally, one can get rid of so-called stopwords like “a” and “the”, which are words that carry little to no semantic meaning in most contexts).
A bag-of-words representation results in these fixed-size inputs by assigning each word in a vocabulary its own position in a vector. We then count how often each word occurs in a document. For example, if we have a vocabulary of 50,000 unique words, it produces a fixed-size vector with 50,000 entries, regardless of whether the input consists of only ten words or 300k words. Note that most entries are zero because each document contains only a small subset of the vocabulary. (Instead of representing the raw counts, there are also normalization schemes like TF-IDF.)
Then, once we have these word frequency vectors, we can train a classifier on a labeled training set, such as emails labeled as spam or non-spam. For example, a logistic regression model would then learn feature weights that correlate certain words (and word counts) with particular labels. For instance, certain words might increase the predicted spam probability, and others may decrease it.
This approach is computationally cheap and can work well when particular words provide strong clues about the label. In a simple classification task such as spam classification, this is often enough to get quick, reasonably accurate results.