Best AI models for general intelligence and capabilities
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
A discussion on AI models for general intelligence highlights the rapid advancements since 2007, when limited internet datasets and compute power restricted model training to milli…
I tried to ask this to LLMs such as gemini 3.1 pro and 3.7 flash apart from Arena AI but I guess real people can provide more better perspectives.
There are a lot of well known alternatives such as SSMs, Liquid AI which uses differential equations, KAN, JEPA, TTT etc.
Apart from transformer many of those are limited. SSM, Liquid AI destroys information and can't perform multi step deduction tasks the best, RLM and others are hard to scale up, KAN isn't supported by native architecture, JEPA has problem with it's reward model and having massive good dataset, TTT and others is somewhat already integrated into transformer based core models. Transformer variant means adding things to transformer. My task here is to reach the ultimate paradigm or atleast better than transformer and understanding why and what makes something more capable generally.
I think other approaches here, even if their limitations are removed in terms of compute and other things somewhat won't perform better than transformer based variant architectures.
Here is what I understand. As time passes by, our energy, compute, dataset, algorithm and design, knowledge, economic, interest and application capacity all grows simultaneously making newer models easier to train and newer paradigms which can't be unlocked today no matter what possible.
Furthermore even if someone in frontier lab reaches back two decades ago in 2007, wouldn't be able to do much with the knowledge as the internet's dataset in 2007 would be limited, so would compute which would only be able to train millions parameters model architecture, the chip and CUDA and other efficient support bases and IDEs won't be present, neither would it have energy to train massive models and public interest, economic incentives. It won't perform better than statistical ML models as were popular back then give or take. Transformers would be unlocked naturally by 2015-2020 because of increase in compute, energy, dataset etc.
Going with it, there could be things which won't perform better now but can replace and beat transformers seriously at scale when more powerful compute and scaling is unlocked. Furthermore if data would be a limiting factor and compute isn't, we could have powerful reward models, synthetic high quality data, more research and data growth as well as more compute heavy models which perform better with more compute but it can drive inference cost and time up.
For my take and opinions, I don't think intelligence is something which can be done in O(n) time personally. World models and neurosymbolic-transformer architecture which requires heavier compute could be unlocked and much more powerful in the future along with some successors of JEPA which I am unsure about. Based on this transformer based architectures would last one or two decades more and things can really shift in the 2040s. Predicting the next based thing for general level intelligence is a hard task though. Opinions?