Today is Guyu (Grain Rain).
It was meant to be a day for sitting by the window and spacing out. But a project just dropped by the open-source community completely jolted me awake. OpenMythos is here.
It just flipped the table.
A Grand Feast in a Tiny Kitchen
Let’s look at some public data first. With a mere 770M parameters, it strictly matches the downstream task performance of a traditional 1.3B Transformer.
What does this mean?
It’s genuinely fierce.
In the past, building large models was all about competing on parameters. To cook a complex dish, they basically bought an entire supermarket. This is actually quite unreasonable. Ingredients pile up like mountains, but very few are actually used. This approach was bound to hit a wall sooner or later.
OpenMythos doesn’t do such foolish things.
It introduced an architecture called “Recurrent-Depth.” Simply put, it means cooking a few more times in the same small kitchen using the same set of pots and pans. Data loops through the identical weights multiple times.
770M—this means achieving double the performance level using half the parameter count of the past. Memory overhead drops off a cliff.
Edge devices are finally saved.
Running AI on smartphones will no longer feel like an old ox pulling a broken cart. The barrier to entry has been completely smashed.
The Hidden Intentions Inside the Matryoshka
However, while everyone is marveling at the halved parameter count, what caught my eye is its routing mechanism.
Here comes the problem.
If it’s just spinning in place, wouldn’t that become a meaningless infinite loop? This is where MoE (Mixture of Experts) is brought onto the stage.
This move is quite brilliant.
Imagine you wrote an article that needs polishing. If you have the same editor review it ten times, they’ll definitely grow numb to it. But OpenMythos stuffs Mixture of Experts into the loop, which is equivalent to swapping in a different group of reviewers for every iteration—the first pass checks for typos, the second smooths out the logic, and the third examines the core premise.
Everyone uses the same set of underlying weights, but the experts activated each time are different.
This successfully avoids the diminishing marginal returns of simply stacking depth.
According to the code disclosed by the open-source community, this design also allows reasoning to occur in the latent space. It no longer has to spit out text line by line just to think.
This hits a blind spot in my knowledge.
I’m also pondering just how much potential this “implicit reasoning” holds for the future.
The Flip Side of the Slimming Game
Let’s do a horizontal comparison.
Over the past two years, big tech companies on the market have casually used hundreds of billions of parameters to overwhelm everyone. For a single complete inference run, the VRAM consumption alone is absurdly high.
Although OpenMythos is an open-source experiment under 1B, it has clearly validated the scaling laws of “recurrent computation.”
To be brutally honest.
We people from Shandong speak plainly. It’s not that magical either. (o_O)
Parameters have indeed been saved, but the total computational volume hasn’t decreased. Think about it. Letting data loop over a dozen times will definitely impact latency. You think you’ve saved space, but it’s actually trading time for space.
Moreover, looking at its mathematical principles, that matrix injection used to control the loop from exploding can easily cause issues if not tuned perfectly. The model might just outright go on strike.
It’s not that simple.
A Few Random Thoughts
I sometimes wonder, what if this line of thinking is pushed to the absolute extreme?
What if future models are shipped at only 100M in size? As long as you give it enough power, it could run infinite recurrent deductions locally until it solves a complex mathematical conjecture?
Relying purely on a local chip squeezing out intelligence loop by loop. There’s truly a sense of old-school geek romance to this.
Honestly.
I might be overthinking it, too. After all, the ceiling of complexity might ultimately be constrained by the model’s inherent ability to compress the laws of the world. Spin in circles too many times, and you might just end up spinning your wheels.
Idle Chatter
By the way.
Today’s coffee beans were definitely over-roasted. The rain outside the window seems to have stopped.
I wonder if those “alchemist” engineers currently training massive models in big tech companies will be able to sleep well tonight after seeing this project—one that wins with brains rather than GPUs.
References:
- GitHub – kyegomez/OpenMythos
- OpenMythos Recasts Claude Mythos as Looped MoE Transformer
- OpenMythos – AI Agents | SkillsLLM
—— Lyra Celest @ Turbulence τ.
